sentry_alert_rule.notification_target_id could previously only point at
a target created outside Terraform (sentryctl/curl/the web UI) --
without this resource, "manage alert rules as code" was only half true.
Same create/destroy-only shape as sentry_alert_rule and for the same
reason: alerting has no PUT /targets/{id} either, confirmed down to
notifystore.Store (Create/List/Get/Delete, no Update).
client.go's notificationTarget type mirrors notifystore.Target's JSON
shape. headers stays raw JSON bytes end to end -- the client has no
opinion about its shape (neither does alerting's own Target type,
json.RawMessage), and the resource layer round-trips it as a plain
JSON-text string a caller provides via Terraform's jsonencode().
secret is marked Sensitive in the schema, but alerting's own
GET /targets/{id} returns it unredacted (confirmed in
notifystore/store.go -- no redaction at the store or handler layer, an
existing property of alerting's API, not something this provider
introduces). A new client test
(TestGetNotificationTargetReturnsSecretUnredacted) documents that real
behavior so a future change to it would be caught here, not discovered
by surprise. Sensitive keeps the value out of plan/apply console output;
it does not keep it out of Terraform state, the standard caveat for any
sensitive attribute, named explicitly in the schema description and
README rather than left implicit.
Examples updated end to end: sentry_alert_rule's example now creates a
real sentry_notification_target and references its .id, instead of a
placeholder string.
Verified: client tests are real httptest.Server round trips. Schema
validation needs no Terraform binary.
TestAccNotificationTargetResource_basic is a real acceptance test,
skip-gated by TF_ACC same as the other two, including a
plancheck.ExpectResourceAction assertion that a config change actually
plans destroy-then-create, and (since secret really does round-trip
unredacted) a real ImportStateVerify on the secret attribute rather than
one papered over with ImportStateVerifyIgnore. Not run against a live
stack in this environment, same disclosed gap as everything else
Docker-gated in this repo.
386 lines
23 KiB
Markdown
386 lines
23 KiB
Markdown
# Project: Sentry — Distributed Log Aggregation & Observability Platform
|
|
|
|
## Mission
|
|
Build an open-core, Kubernetes-native centralized logging platform that rivals
|
|
Splunk on features but wins on cost-per-GB, modern language stack, and honest
|
|
multi-tenant RBAC. Full architecture spec is in `/docs/architecture.md` — read
|
|
it before touching any component. Do not deviate from the storage/query split
|
|
described there without flagging it to me first.
|
|
|
|
## Non-negotiable constraints
|
|
- Distro-agnostic Linux agent: must run identically on RHEL/Debian/Arch/SUSE
|
|
derivatives via a statically-linked musl binary. No glibc runtime deps.
|
|
- Windows support via native ETW/Event Log API, not a WSL shim.
|
|
- AGPLv3 for core + agents. Enterprise module (SSO/multi-tenancy/compliance)
|
|
lives in a separate `enterprise/` directory under a commercial license stub
|
|
— keep the boundary clean from day one, don't let AGPL code import from it.
|
|
- Schema-on-write with OTel semantic conventions as the default schema, with
|
|
schema-on-read fallback for unstructured text.
|
|
- Every UI action must correspond to a documented REST/gRPC call. No
|
|
UI-only logic. CLI (`sentryctl`) and Terraform provider are first-class,
|
|
not afterthoughts. **Status**: `sentryctl` has been built out phase by
|
|
phase since Phase 3. The Terraform provider (`/terraform`) only exists
|
|
as of this note -- three resources (`sentry_dashboard`, full CRUD;
|
|
`sentry_alert_rule` and `sentry_notification_target`, both create/
|
|
destroy only -- `alerting` has no `PUT /rules/{id}` or
|
|
`PUT /targets/{id}` to update against), built on HashiCorp's
|
|
`terraform-plugin-framework`, reusing the exact same REST contracts
|
|
`sentryctl dashboards apply`/web's dashboard export and
|
|
`sentryctl alerts apply` already use. Tenant/RBAC resources are real,
|
|
disclosed future work -- see `/terraform/README.md` for the full
|
|
accounting of what is and isn't built, and the same
|
|
"written but not run against a live stack" verification caveat as
|
|
everything else Docker-gated in this repo.
|
|
|
|
## Tech stack (pinned — do not substitute without discussion)
|
|
| Component | Language/Tool |
|
|
|-------------------|------------------------|
|
|
| Edge agent | Rust, musl target |
|
|
| Transport | Redpanda (Kafka API) |
|
|
| Ingest/parse | Go |
|
|
| Analytical store | ClickHouse |
|
|
| Full-text index | Tantivy (Rust) |
|
|
| Control plane/API | Go, gRPC + REST gateway |
|
|
| Frontend | SvelteKit + TypeScript |
|
|
| Deployment | Kubernetes Operator (Go, kubebuilder), Helm, docker-compose for local/homelab |
|
|
|
|
## Repo conventions
|
|
- Monorepo, one top-level dir per component (see structure below).
|
|
- Rust: workspace-based, `cargo clippy --all-targets -- -D warnings` must pass.
|
|
- Go: standard `go vet` + `golangci-lint`, no globals for shared state.
|
|
- Every component ships with: unit tests, a `README.md`, and a Dockerfile
|
|
using distroless or scratch base images where feasible.
|
|
- Conventional commits. Every PR-sized change should be a logically complete,
|
|
independently revertible unit.
|
|
- Prefer boring, well-understood dependencies over novel ones. This is
|
|
infrastructure software; operators need to trust it.
|
|
|
|
## What "done" looks like for Phase 0 (MVP)
|
|
|
|
**Status: shipped.** A single log line, generated on a Linux host by the
|
|
Rust agent, flows: agent → Redpanda → Go ingest service → ClickHouse, and
|
|
is queryable via a minimal SQL endpoint and visible in a bare-bones
|
|
SvelteKit table view. Verified end-to-end on real hardware, not just in
|
|
CI — see `/docs/phase-0-runbook.md`. No alerting, no multi-tenancy, no
|
|
dashboards — that discipline held for the whole phase.
|
|
|
|
## What "done" looks like for Phase 1
|
|
|
|
**Status: shipped.** A Windows Event Log entry and a Linux journald entry
|
|
are both queryable via SQL (the ClickHouse path) and via free-text search
|
|
(the Tantivy path), from the same UI, within a few seconds of being
|
|
generated. Verified end-to-end on the live stack, including the same
|
|
`record_id` coming back from both query paths for the same record — see
|
|
`/docs/phase-1-runbook.md`.
|
|
|
|
ETW and WEF (Windows Event Forwarding) were *designed* in this phase but
|
|
not required to be running for "done": ETW ships behind a feature flag
|
|
most environments won't enable (it needs elevated privileges), and WEF's
|
|
receiver-side was explicitly deferred rather than built. Only the Event
|
|
Log source needed to actually be running end-to-end, and did. The
|
|
Windows-specific agent code itself (`EvtSubscribe`, ETW, service
|
|
registration) remains unverified on real Windows — no Windows toolchain
|
|
existed anywhere in the environment this was built in; flagged
|
|
prominently in `/agent/README.md` and the runbook.
|
|
|
|
## What "done" looks like for Phase 2
|
|
|
|
A single query bar in the web UI and a single `sentryctl query` command
|
|
can express filter + free-text + stats in one query (e.g. `service=api |
|
|
where status>=500 | stats count by host | sort -count`, or
|
|
`message:"connection refused" | stats count by host`), execute correctly
|
|
against both ClickHouse and Tantivy in one compiled plan, and return in
|
|
well under a second for a 1M-row fixture dataset (rough benchmark, not a
|
|
formal SLA — see `/docs/phase-2-runbook.md` for the actual measurement).
|
|
Raw ClickHouse SQL remains available as an escape hatch, compiling to the
|
|
same execution plan/IR as the pipe syntax so performance doesn't depend
|
|
on which syntax a query uses.
|
|
|
|
Non-goals for this phase (same "resist scope creep" discipline as every
|
|
phase so far): no alerting, no dashboards, no multi-tenancy — this phase
|
|
is the query layer only. The two separate placeholder pages/endpoints
|
|
from Phase 0/1 (`/query` raw-SQL-only, `/search` free-text-only) are
|
|
retired, replaced by one `/query` endpoint and one query page.
|
|
|
|
See `/docs/query-language-design.md` for the grammar, IR, and
|
|
ClickHouse/Tantivy routing strategy, and
|
|
`/docs/query-language-reference.md` for the user-facing syntax reference
|
|
once built.
|
|
|
|
## What "done" looks like for Phase 3
|
|
|
|
**Status: shipped.** A user can build a multi-panel dashboard from saved Phase 2 queries (at
|
|
least a line chart panel and a table panel, working end-to-end against
|
|
live data), save an alert rule that fires a Slack webhook when a
|
|
condition is met (threshold comparison, or "absence" — the query returned
|
|
zero rows in its own time window), and see the delivery attempt logged —
|
|
all from the web UI, without touching the API directly. See
|
|
`/docs/phase-3-dashboard-design.md` and `/docs/phase-3-alerting-design.md`
|
|
for the data models and the alerting evaluator's firing/resolved state
|
|
machine, and `/docs/phase-3-runbook.md` for the live-stack verification,
|
|
including a load test of the alert evaluator against ~500 concurrent
|
|
rules.
|
|
|
|
This phase adds PostgreSQL as a new pinned-stack component (see the
|
|
dashboard design doc for why ClickHouse can't do this job — dashboards
|
|
and alert state need real row-level locking and transactional
|
|
read-modify-write, which ClickHouse's MergeTree family doesn't provide),
|
|
scoped strictly to control-plane config: dashboards, panels, notification
|
|
targets, alert rules, alert state, delivery log. Log data itself stays on
|
|
ClickHouse/Tantivy only, unchanged.
|
|
|
|
Non-goals for this phase (same discipline as every phase so far):
|
|
- No multi-tenancy enforcement and no `enterprise/` module work — single
|
|
tenant/org assumed. Most new tables (`dashboards`, `alert_rules`,
|
|
`notification_targets`) carry a `tenant_id` column so part of Phase 4's
|
|
retrofit doesn't require a migration + backfill — but `alert_state` and
|
|
`delivery_log` do not (an inconsistency found during Phase 4 planning,
|
|
not caught at the time); Phase 4 adds `tenant_id` to those two and
|
|
backfills via a join through `alert_rules.id`, and — per
|
|
`/docs/phase-4-isolation-design.md` — tenant isolation itself turned
|
|
out to live at the ClickHouse/Tantivy connection layer, not via these
|
|
columns at all, since Phase 2's raw-SQL escape hatch can never be
|
|
covered by a row filter regardless of which tables carry one.
|
|
- No raw-SQL dashboard panels (time-range injection isn't reliable
|
|
against arbitrary SQL) — pipe-syntax queries only.
|
|
- No per-group/multi-row threshold alerting (e.g. "alert separately per
|
|
host") — a threshold rule's query must resolve to a single row.
|
|
- No debounce on the way down — a firing alert resolves on the first
|
|
false evaluation, no symmetric "stay firing for N more minutes" hold.
|
|
- No Kubernetes Operator/Helm deployment work — still docker-compose,
|
|
`/deploy` remains stubbed.
|
|
|
|
## What "done" looks like for Phase 4
|
|
|
|
**Status: in progress, not shipped.** RBAC enforcement (`api/authz`), the
|
|
`alerting`↔`api` service-identity credential, tenant-scoped dashboards,
|
|
append-only audit logging, and — since the second pass on this phase —
|
|
real per-tenant ClickHouse provisioning and query routing
|
|
(`enterprise/internal/tenantprovision`, `enterprise/internal/chrunner`,
|
|
wired into a new `enterprise/cmd/enterprise-api` binary alongside plain
|
|
`api/cmd/api`) are all built and tested — real integration tests exist
|
|
for the ClickHouse pieces, but this environment lost Docker/database
|
|
access partway through the phase, so only the audit-logging guarantees
|
|
were actually confirmed against a live database; the rest is untested
|
|
beyond "compiles, and skips cleanly when no live database is
|
|
configured" (see `/docs/phase-4-runbook.md`'s verification-status
|
|
section). Human SSO login is now built for both protocols
|
|
(`enterprise/internal/loginhandler`: `GET /auth/oidc/login` +
|
|
`GET /auth/oidc/callback`, and `GET /auth/saml/login` +
|
|
`POST /auth/saml/acs` via `enterprise/internal/saml`'s `crewjam/saml`
|
|
wiring, both issuing a real session cookie after resolving tenant/role
|
|
from `tenant_memberships`) — genuinely verified, unlike the ClickHouse
|
|
pieces, via a real fake IdP for each protocol that performs actual
|
|
cryptographic signing and verification (`coreos/go-oidc`'s `oidctest`
|
|
for OIDC, `crewjam/saml/samlidp` for SAML — `loginhandler_test.go` and
|
|
`saml_test.go`, all passing, including the full login round trip and
|
|
negative paths for both), though never tried against a real external IdP
|
|
or through a running `enterprise-auth` container. Writing the SAML test
|
|
caught and fixed two real bugs in `internal/saml.ParseResponse`: a
|
|
missing `r.ParseForm()` call that would have silently broken every real
|
|
ACS POST, and email-attribute matching that missed the standard LDAP
|
|
"mail" OID IdPs send by default. Tantivy per-tenant index routing is now
|
|
built too
|
|
(`search/src/registry.rs` + `enterprise/internal/searchclient`) —
|
|
**genuinely verified**, like the OIDC login flow: Tantivy is an embedded
|
|
library, not a networked service, so the isolation probe (three tenants,
|
|
same search term, scoped search returns only that tenant's document)
|
|
actually ran in this environment, no Docker needed. That same
|
|
Docker-free advantage is what caught a real bug while closing the last
|
|
of Phase 4 task 8's four adversarial probes (a mid-provisioning tenant
|
|
must be refused, not served): `search/src/registry.rs`'s `IndexRegistry`
|
|
opened-or-created an index for any syntactically-valid `tenant_id`,
|
|
meaning a query against a tenant that exists in `rbacstore` but isn't
|
|
active yet would have silently returned zero results from a
|
|
freshly-created empty index instead of being refused --
|
|
`chrunner`'s ClickHouse routing had the equivalent guarantee for free
|
|
(a mid-provisioning tenant simply isn't in its startup-built connection
|
|
map) but Tantivy, a separate process with no Postgres access, had no
|
|
way to know. Fixed with a new `enterprise/internal/searchclient.
|
|
TenantChecker` (backed by `rbacstore.TenantIsActive`); both halves of
|
|
the fix verified Docker-free (`chrunner_test.go`'s and
|
|
`searchclient_test.go`'s `TestSearchRefusesMidProvisioningTenant`-shaped
|
|
tests) — see `api/queryapi/tenant_isolation_gap_test.go` for the full
|
|
accounting of all four probes, now all closed. The deployment-
|
|
topology gap that briefly was the largest one is now closed for both
|
|
Helm and docker-compose: `deploy/helm/sentry/templates/api.yaml`/
|
|
`enterprise-api.yaml` are mutually exclusive on the same
|
|
`enterprise.enabled` flag that turns on RBAC/audit/SSO, rendering to the
|
|
same Service name/port either way — a Helm-deployed cluster can't
|
|
accidentally run the wrong one. `docker-compose.yml`'s `api`/
|
|
`enterprise-api` services are now the same mutually-exclusive choice,
|
|
gated behind `COMPOSE_PROFILES` (`.env` checks in `single-tenant` as the
|
|
zero-config default) and sharing a host port/network-alias trick so
|
|
`alerting`/`web` need no conditional logic either way — verified via
|
|
`docker compose config` (renders/validates without a daemon, confirms
|
|
the two never both appear for one profile selection), not an actual
|
|
`docker compose up` in this environment. Per-resource dashboard grants
|
|
(the RBAC matrix's "(own/granted)"
|
|
qualifier) are now enforced too: `api/dashboards.PermissionStore` (core
|
|
interface) implemented by `enterprise/internal/rbacstore.
|
|
DashboardPermissions`, wired in only by `enterprise-api` — an Editor can
|
|
now only edit/delete a dashboard they created or were granted access to,
|
|
not every dashboard in their tenant; managing grants themselves is
|
|
stricter still (creator/Admin/Owner only, closing a self-escalation
|
|
path). Verified against a fake store (`api/dashboards/handler_test.go`);
|
|
real integration tests exist but haven't run against a live Postgres,
|
|
same disclosed gap as the rest of this phase's Postgres-backed pieces.
|
|
`sentryctl dashboards permissions list|grant|revoke` is now the CLI
|
|
surface for this — `PUT`/`DELETE /dashboards/{id}/permissions/{userId}`
|
|
previously had no caller but Go tests and curl.
|
|
`deploy/operator`'s `Tenant` CRD and `enterprise-api -provision-tenant`
|
|
are now unified too, deliberately lightweight rather than making the
|
|
K8s controller a second real actor: `-provision-tenant` stays the sole
|
|
caller of ClickHouse/`rbacstore`, and (via a new
|
|
`enterprise/internal/tenantcrd`, gated on `TENANT_CRD_NAMESPACE`) syncs
|
|
its real result into the CRD — a real credential Secret, not the
|
|
previous placeholder that authenticated against nothing, and status
|
|
fields the reconciler derives `Phase`/`Ready` from instead of
|
|
independently guessing "Active" the moment a Tenant object exists. The
|
|
tenant-picker is now fully built, backend and frontend: an identity with
|
|
more than one `tenant_memberships` row gets a real `GET
|
|
/auth/memberships`/`POST /auth/select-tenant` round trip (a short-lived
|
|
pending-login token, distinct from a real session by both Go type and
|
|
JWT claim name — a real token-confusion bug this design's own tests
|
|
caught before it shipped) instead of the flat refusal Phase 4 shipped
|
|
with earlier, and `web/src/routes/select-tenant` is the page that
|
|
actually calls it — see the Phase 4 exit-criteria paragraph below for
|
|
what changed to make that verifiable in this environment. Ingest
|
|
tenant-awareness — the gap this section used to
|
|
call "undesigned" — now has a real, if intentionally partial, design:
|
|
`ingest` (AGPL core) gained an optional `TenantResolver`
|
|
(`ingest/internal/grpcserver`), a per-tenant bearer credential an agent
|
|
presents (minted via `enterprise-auth
|
|
-create-ingest-credential-tenant=<id>`, validated over the network via a
|
|
new `POST /internal/authorize-ingest` endpoint — never an `enterprise/`
|
|
import, same boundary shape as `api/authz.Authorizer`), and the
|
|
resolved tenant ID is attached to every record as a `tenant_id` Kafka
|
|
message header before it's produced. **ClickHouse write-routing is now
|
|
built too**: `enterprise/cmd/enterprise-ingest` (another "second binary,"
|
|
mirroring `enterprise-api`) reuses `ingest/consumer`'s own flush loop
|
|
with `enterprise/internal/chwriter.Registry` swapped in as the writer —
|
|
one dedicated ClickHouse connection per tenant, routing each batch's
|
|
records by their `tenant_id` tag, fail-closed on an untagged or
|
|
unprovisioned tenant. Building it found and fixed a real bug:
|
|
`tenantprovision.ProvisionClickHouse` originally granted a tenant's
|
|
ClickHouse user `SELECT` only, which would have made every real
|
|
per-tenant write fail with a permission error — fixed by granting
|
|
`SELECT, INSERT` (one credential, both directions; no cross-tenant
|
|
boundary is crossed by also allowing INSERT within a tenant's own
|
|
database). **Tantivy write-routing is now built too** —
|
|
`search/src/consumer.rs` (a completely independent Redpanda consumer,
|
|
not called through `ingest` or `enterprise-ingest` at all: a different
|
|
codebase and process) now resolves each record's `tenant_id` header
|
|
through the *same* `IndexRegistry` the read side already used, and
|
|
routes the write there instead of always into the default index. Unlike
|
|
the ClickHouse side, this needed no "second binary": `IndexRegistry`
|
|
already lives in this AGPL-core binary (Tantivy has no grant system to
|
|
gate a commercially-licensed credential behind, so there was never an
|
|
import-boundary reason to split it out), so read and write share one
|
|
registry directly. The periodic Tantivy commit now commits every tenant
|
|
index that's seen a write, not just the default one
|
|
(`IndexRegistry::commit_all`). **The active-tenant gap this same change
|
|
originally disclosed is now closed too**: `search/src/tenants.rs`'s
|
|
`ActiveTenantTracker` polls a new `GET /internal/active-tenants`
|
|
endpoint on `enterprise-auth` (`search` has no Postgres access, unlike
|
|
`chwriter.Registry`'s direct `rbacstore` query or the read side's
|
|
`searchclient.TenantChecker`, so this needed a network call — the same
|
|
"network boundary, not import boundary" shape `ingest`'s
|
|
`TenantResolver` already uses against the same service, authenticated
|
|
with a RoleService credential the same way `alerting` authenticates to
|
|
`api`) and `consumer.rs` refuses any tagged record whose tenant isn't in
|
|
the polled allowlist. Off unless both `ENTERPRISE_AUTH_URL` and
|
|
`ENTERPRISE_AUTH_SERVICE_TOKEN` are set (same "off unless configured"
|
|
default as everything else optional in this codebase); when they are,
|
|
startup blocks on the first fetch succeeding and later refresh failures
|
|
keep serving the last-known-good set rather than clearing it. Verified
|
|
with real HTTP round trips against a hand-rolled TCP test server in this
|
|
environment, no live enterprise-auth needed. **This closing move exposed
|
|
the ClickHouse side's own gap by comparison** — `chwriter.Registry`'s
|
|
writer map was still a startup-only snapshot with no refresh at all, a
|
|
real asymmetry once Tantivy's tracker refreshed every minute and
|
|
ClickHouse's didn't — so `Registry.StartRefreshing` (new) closes that
|
|
too: same one-minute interval, same last-known-good posture on a failed
|
|
refresh, opening connections for newly-active tenants and closing ones
|
|
no longer active. Both engines now share the same active-tenant
|
|
staleness bound instead of one being materially staler than the other.
|
|
**The tenant-picker page is now built too**:
|
|
`web/src/routes/select-tenant` calls `GET /auth/memberships`/
|
|
`POST /auth/select-tenant` via `fetch(..., {credentials: 'include'})`
|
|
(new `$lib/api.ts` functions), which needed a second CORS posture
|
|
alongside the wildcard-friendly one `enterprise-api` already had —
|
|
`api/httpserver.WithCredentialedCORS`, set to a literal origin via a new
|
|
`CORS_ALLOWED_ORIGIN` on `enterprise-auth` — since browsers refuse to
|
|
honor a wildcard `Access-Control-Allow-Origin` on a credentialed
|
|
request. **Genuinely verified in a real browser in this environment**:
|
|
a throwaway Node server standing in for `enterprise-auth`'s exact wire
|
|
contract (including its plain-text `http.Error` bodies, not JSON) on a
|
|
different origin than `web`'s dev server, driven through the full
|
|
cross-origin pending-login-cookie round trip, a real click choosing a
|
|
tenant, and the post-selection redirect — plus the missing/expired-
|
|
pending-login error path — with no Docker or live Postgres/IdP needed,
|
|
since the point was exercising `web`'s own fetch/CORS/cookie wiring, not
|
|
`enterprise-auth`'s internals (already covered by that package's own
|
|
tests). See `/web/README.md`'s "Tenant picker" section for the exact
|
|
setup. What's left in this phase now is entirely the caveats already
|
|
disclosed above, not an unbuilt feature: the ClickHouse/Postgres-backed
|
|
pieces have never run against a real database in this environment, and
|
|
nothing here has been tried against a real external IdP or a real
|
|
running multi-container deployment. Full accounting:
|
|
`/docs/security/threat-model.md`; step-by-step verification procedure
|
|
(not yet run against a live cluster in this environment):
|
|
`/docs/phase-4-runbook.md`. The rest of this section describes the exit
|
|
bar this phase is aiming at, not a completed state.
|
|
|
|
Two tenants can be provisioned with SSO (OIDC or SAML), each with their
|
|
own users, roles, dashboards, and alert rules, fully isolated at the
|
|
ClickHouse/Tantivy connection layer — not by a row filter — with
|
|
adversarial integration tests proving no cross-tenant data leakage,
|
|
including via the raw-SQL escape hatch and ClickHouse's own `system.*`
|
|
tables. A tenant admin can see a query audit trail for their tenant,
|
|
backed by append-only storage a compromised application credential
|
|
cannot alter (enforced by database grants, not just convention) and
|
|
periodically anchored outside the database so tampering is detectable
|
|
even against a privileged attacker. See `/docs/phase-4-isolation-design.md`
|
|
for the tenant isolation model and why it lives at the connection layer,
|
|
`/docs/phase-4-rbac-design.md` for the role/permission model, and
|
|
`/docs/security/threat-model.md` for the auth flows and audit-log
|
|
integrity guarantees, written for a prospective enterprise customer's
|
|
security team.
|
|
|
|
The tenant-isolation, provisioning, SSO, and RBAC-enforcement mechanisms
|
|
live entirely in `enterprise/` (commercial license), confirmed
|
|
explicitly rather than assumed: AGPL core (`/api`, `/alerting`, `/web`)
|
|
stays genuinely single-tenant, with no multi-tenant mechanism present at
|
|
all — `enterprise/` supplies tenant-scoped implementations of core's
|
|
already-shipped `querylang/executor.SQLRunner`/`SearchClient` interfaces
|
|
rather than core growing tenant awareness. Query-compiler-level "compile
|
|
time" enforcement, as originally proposed, turned out not to be
|
|
achievable in any module once Phase 2's opaque raw-SQL passthrough is
|
|
accounted for — the honest, implemented guarantee is that every code
|
|
path (compiled query or raw SQL) is forced through a tenant-scoped
|
|
database connection/index that the database's own access control
|
|
enforces, not a compiler-injected filter.
|
|
|
|
Non-goals for this phase (same discipline as every phase so far):
|
|
- No deny-override permissions — per-resource grants (e.g. a specific
|
|
user getting edit access to one dashboard) are additive only; a full
|
|
allow/deny ACL system is future work.
|
|
- No data retention/deletion policy design for tenant deprovisioning —
|
|
the provisioning state machine includes a `deprovisioning` state, but
|
|
what actually happens to a deprovisioned tenant's data is a separate,
|
|
not-yet-designed compliance question.
|
|
- No general multi-cluster orchestration in `/deploy` — scoped to
|
|
proving the per-tenant ClickHouse/Tantivy isolation model works, not a
|
|
fully general multi-cluster system.
|
|
- No protection against a privileged ClickHouse/Postgres administrator —
|
|
the isolation and audit-log guarantees in this phase are structural
|
|
defenses against application-layer bugs and injection, not against
|
|
someone with database superuser access; that's an operational control,
|
|
out of scope here and named explicitly, not silently assumed away.
|
|
|
|
## When in doubt
|
|
Ask before: changing the pinned stack, adding a new external dependency
|
|
that pulls in a large transitive tree, or making an architectural decision
|
|
that isn't already specified in `/docs/architecture.md`.
|