Files
cairnobs/CLAUDE.md
T
jcoffey-dev 278b24cf67 Add sentry_dashboard_panel, resolving panels as their own resource
The open design question named in the last three commits' README --
"a sentry_dashboard_panel resource (or a panels list block on this one)"
-- is resolved: separate resource, matching api/dashboards.Handler's own
shape (a panel is created/updated/deleted independently of its parent
dashboard via its own endpoints, never by rewriting the dashboard's
whole panel list). A nested list block would have forced every panel to
be rewritten on any single panel's change, hiding fine-grained diffs a
separate resource shows naturally -- the more idiomatic Terraform
pattern for independently-lifecycled child resources, and the one that
matches what the API actually does.

Unlike sentry_alert_rule/sentry_notification_target, this resource
supports a real in-place Update -- api/dashboards.Handler actually has a
PUT /dashboards/{id}/panels/{panelId}. Only dashboard_id forces
RequiresReplace: UpdatePanel's SQL matches WHERE id = $panelID AND
dashboard_id = $dashboardID, so changing dashboard_id through the
existing panel's URL wouldn't move it, it would just fail to match --
there's no API operation for "move a panel to a different dashboard."

Panels have no standalone GET endpoint -- only GET /dashboards/{id},
which includes the full panels array. client.go's new getPanel fetches
the parent dashboard and finds the panel by ID within it, returning the
same *apiError{StatusCode: 404} shape a direct GET would whether the
dashboard itself or just the panel within it is gone, so isNotFound
works identically either way. This also means a bare panel ID isn't
enough to import from -- ImportState takes "dashboard_id/panel_id" and
splits on the last "/", the one resource here with a composite import
identifier.

query_language never accepts "sql" for panels specifically -- confirmed
in api/dashboards's own validatePanel ("dashboards only support
pipe-syntax queries, since the dashboard time-range picker is injected
as leading query terms"), a real constraint from the API this client
doesn't re-validate client-side (same "let the API be the one source of
truth for validation" posture the other resources already take), but
documented in the schema so it's not a surprise 400 from Create.

sentry_dashboard_panel gets a matching data source too
(dashboard_id + id both Required, unlike the other three data sources'
single Required id, since getPanel itself needs both).

Verified: client tests are real httptest.Server round trips, including
getPanel finding the right panel within a real dashboard response and
returning a recognizable not-found both when the panel is missing and
when the parent dashboard itself is gone. Schema validation needs no
Terraform binary. TestAccDashboardPanelResource_basic and
TestAccDashboardPanelDataSource_basic are real acceptance tests,
skip-gated by TF_ACC same as the other six -- the resource test proves a
genuine in-place update (a title change, no plancheck needed since
in-place update is the default expectation here, unlike the
create/destroy-only resources). Not run against a live stack in this
environment, same disclosed gap as everything else Docker-gated in this
repo.
2026-08-15 10:40:27 -07:00

390 lines
24 KiB
Markdown

# Project: Sentry — Distributed Log Aggregation & Observability Platform
## Mission
Build an open-core, Kubernetes-native centralized logging platform that rivals
Splunk on features but wins on cost-per-GB, modern language stack, and honest
multi-tenant RBAC. Full architecture spec is in `/docs/architecture.md` — read
it before touching any component. Do not deviate from the storage/query split
described there without flagging it to me first.
## Non-negotiable constraints
- Distro-agnostic Linux agent: must run identically on RHEL/Debian/Arch/SUSE
derivatives via a statically-linked musl binary. No glibc runtime deps.
- Windows support via native ETW/Event Log API, not a WSL shim.
- AGPLv3 for core + agents. Enterprise module (SSO/multi-tenancy/compliance)
lives in a separate `enterprise/` directory under a commercial license stub
— keep the boundary clean from day one, don't let AGPL code import from it.
- Schema-on-write with OTel semantic conventions as the default schema, with
schema-on-read fallback for unstructured text.
- Every UI action must correspond to a documented REST/gRPC call. No
UI-only logic. CLI (`sentryctl`) and Terraform provider are first-class,
not afterthoughts. **Status**: `sentryctl` has been built out phase by
phase since Phase 3. The Terraform provider (`/terraform`) only exists
as of this note -- four resources (`sentry_dashboard` and
`sentry_dashboard_panel`, both full CRUD, panels as their own resource
rather than a nested block since the API manages them independently
of their parent dashboard; `sentry_alert_rule` and
`sentry_notification_target`, both create/destroy only -- `alerting`
has no `PUT /rules/{id}` or `PUT /targets/{id}` to update against),
each paired with a read-only data source, built on HashiCorp's
`terraform-plugin-framework`, reusing the exact same REST contracts
`sentryctl dashboards apply`/web's dashboard export and
`sentryctl alerts apply` already use. Tenant/RBAC resources are real,
disclosed future work -- see
`/terraform/README.md` for the full accounting of what is and isn't
built, and the same
"written but not run against a live stack" verification caveat as
everything else Docker-gated in this repo.
## Tech stack (pinned — do not substitute without discussion)
| Component | Language/Tool |
|-------------------|------------------------|
| Edge agent | Rust, musl target |
| Transport | Redpanda (Kafka API) |
| Ingest/parse | Go |
| Analytical store | ClickHouse |
| Full-text index | Tantivy (Rust) |
| Control plane/API | Go, gRPC + REST gateway |
| Frontend | SvelteKit + TypeScript |
| Deployment | Kubernetes Operator (Go, kubebuilder), Helm, docker-compose for local/homelab |
## Repo conventions
- Monorepo, one top-level dir per component (see structure below).
- Rust: workspace-based, `cargo clippy --all-targets -- -D warnings` must pass.
- Go: standard `go vet` + `golangci-lint`, no globals for shared state.
- Every component ships with: unit tests, a `README.md`, and a Dockerfile
using distroless or scratch base images where feasible.
- Conventional commits. Every PR-sized change should be a logically complete,
independently revertible unit.
- Prefer boring, well-understood dependencies over novel ones. This is
infrastructure software; operators need to trust it.
## What "done" looks like for Phase 0 (MVP)
**Status: shipped.** A single log line, generated on a Linux host by the
Rust agent, flows: agent → Redpanda → Go ingest service → ClickHouse, and
is queryable via a minimal SQL endpoint and visible in a bare-bones
SvelteKit table view. Verified end-to-end on real hardware, not just in
CI — see `/docs/phase-0-runbook.md`. No alerting, no multi-tenancy, no
dashboards — that discipline held for the whole phase.
## What "done" looks like for Phase 1
**Status: shipped.** A Windows Event Log entry and a Linux journald entry
are both queryable via SQL (the ClickHouse path) and via free-text search
(the Tantivy path), from the same UI, within a few seconds of being
generated. Verified end-to-end on the live stack, including the same
`record_id` coming back from both query paths for the same record — see
`/docs/phase-1-runbook.md`.
ETW and WEF (Windows Event Forwarding) were *designed* in this phase but
not required to be running for "done": ETW ships behind a feature flag
most environments won't enable (it needs elevated privileges), and WEF's
receiver-side was explicitly deferred rather than built. Only the Event
Log source needed to actually be running end-to-end, and did. The
Windows-specific agent code itself (`EvtSubscribe`, ETW, service
registration) remains unverified on real Windows — no Windows toolchain
existed anywhere in the environment this was built in; flagged
prominently in `/agent/README.md` and the runbook.
## What "done" looks like for Phase 2
A single query bar in the web UI and a single `sentryctl query` command
can express filter + free-text + stats in one query (e.g. `service=api |
where status>=500 | stats count by host | sort -count`, or
`message:"connection refused" | stats count by host`), execute correctly
against both ClickHouse and Tantivy in one compiled plan, and return in
well under a second for a 1M-row fixture dataset (rough benchmark, not a
formal SLA — see `/docs/phase-2-runbook.md` for the actual measurement).
Raw ClickHouse SQL remains available as an escape hatch, compiling to the
same execution plan/IR as the pipe syntax so performance doesn't depend
on which syntax a query uses.
Non-goals for this phase (same "resist scope creep" discipline as every
phase so far): no alerting, no dashboards, no multi-tenancy — this phase
is the query layer only. The two separate placeholder pages/endpoints
from Phase 0/1 (`/query` raw-SQL-only, `/search` free-text-only) are
retired, replaced by one `/query` endpoint and one query page.
See `/docs/query-language-design.md` for the grammar, IR, and
ClickHouse/Tantivy routing strategy, and
`/docs/query-language-reference.md` for the user-facing syntax reference
once built.
## What "done" looks like for Phase 3
**Status: shipped.** A user can build a multi-panel dashboard from saved Phase 2 queries (at
least a line chart panel and a table panel, working end-to-end against
live data), save an alert rule that fires a Slack webhook when a
condition is met (threshold comparison, or "absence" — the query returned
zero rows in its own time window), and see the delivery attempt logged —
all from the web UI, without touching the API directly. See
`/docs/phase-3-dashboard-design.md` and `/docs/phase-3-alerting-design.md`
for the data models and the alerting evaluator's firing/resolved state
machine, and `/docs/phase-3-runbook.md` for the live-stack verification,
including a load test of the alert evaluator against ~500 concurrent
rules.
This phase adds PostgreSQL as a new pinned-stack component (see the
dashboard design doc for why ClickHouse can't do this job — dashboards
and alert state need real row-level locking and transactional
read-modify-write, which ClickHouse's MergeTree family doesn't provide),
scoped strictly to control-plane config: dashboards, panels, notification
targets, alert rules, alert state, delivery log. Log data itself stays on
ClickHouse/Tantivy only, unchanged.
Non-goals for this phase (same discipline as every phase so far):
- No multi-tenancy enforcement and no `enterprise/` module work — single
tenant/org assumed. Most new tables (`dashboards`, `alert_rules`,
`notification_targets`) carry a `tenant_id` column so part of Phase 4's
retrofit doesn't require a migration + backfill — but `alert_state` and
`delivery_log` do not (an inconsistency found during Phase 4 planning,
not caught at the time); Phase 4 adds `tenant_id` to those two and
backfills via a join through `alert_rules.id`, and — per
`/docs/phase-4-isolation-design.md` — tenant isolation itself turned
out to live at the ClickHouse/Tantivy connection layer, not via these
columns at all, since Phase 2's raw-SQL escape hatch can never be
covered by a row filter regardless of which tables carry one.
- No raw-SQL dashboard panels (time-range injection isn't reliable
against arbitrary SQL) — pipe-syntax queries only.
- No per-group/multi-row threshold alerting (e.g. "alert separately per
host") — a threshold rule's query must resolve to a single row.
- No debounce on the way down — a firing alert resolves on the first
false evaluation, no symmetric "stay firing for N more minutes" hold.
- No Kubernetes Operator/Helm deployment work — still docker-compose,
`/deploy` remains stubbed.
## What "done" looks like for Phase 4
**Status: in progress, not shipped.** RBAC enforcement (`api/authz`), the
`alerting``api` service-identity credential, tenant-scoped dashboards,
append-only audit logging, and — since the second pass on this phase —
real per-tenant ClickHouse provisioning and query routing
(`enterprise/internal/tenantprovision`, `enterprise/internal/chrunner`,
wired into a new `enterprise/cmd/enterprise-api` binary alongside plain
`api/cmd/api`) are all built and tested — real integration tests exist
for the ClickHouse pieces, but this environment lost Docker/database
access partway through the phase, so only the audit-logging guarantees
were actually confirmed against a live database; the rest is untested
beyond "compiles, and skips cleanly when no live database is
configured" (see `/docs/phase-4-runbook.md`'s verification-status
section). Human SSO login is now built for both protocols
(`enterprise/internal/loginhandler`: `GET /auth/oidc/login` +
`GET /auth/oidc/callback`, and `GET /auth/saml/login` +
`POST /auth/saml/acs` via `enterprise/internal/saml`'s `crewjam/saml`
wiring, both issuing a real session cookie after resolving tenant/role
from `tenant_memberships`) — genuinely verified, unlike the ClickHouse
pieces, via a real fake IdP for each protocol that performs actual
cryptographic signing and verification (`coreos/go-oidc`'s `oidctest`
for OIDC, `crewjam/saml/samlidp` for SAML — `loginhandler_test.go` and
`saml_test.go`, all passing, including the full login round trip and
negative paths for both), though never tried against a real external IdP
or through a running `enterprise-auth` container. Writing the SAML test
caught and fixed two real bugs in `internal/saml.ParseResponse`: a
missing `r.ParseForm()` call that would have silently broken every real
ACS POST, and email-attribute matching that missed the standard LDAP
"mail" OID IdPs send by default. Tantivy per-tenant index routing is now
built too
(`search/src/registry.rs` + `enterprise/internal/searchclient`) —
**genuinely verified**, like the OIDC login flow: Tantivy is an embedded
library, not a networked service, so the isolation probe (three tenants,
same search term, scoped search returns only that tenant's document)
actually ran in this environment, no Docker needed. That same
Docker-free advantage is what caught a real bug while closing the last
of Phase 4 task 8's four adversarial probes (a mid-provisioning tenant
must be refused, not served): `search/src/registry.rs`'s `IndexRegistry`
opened-or-created an index for any syntactically-valid `tenant_id`,
meaning a query against a tenant that exists in `rbacstore` but isn't
active yet would have silently returned zero results from a
freshly-created empty index instead of being refused --
`chrunner`'s ClickHouse routing had the equivalent guarantee for free
(a mid-provisioning tenant simply isn't in its startup-built connection
map) but Tantivy, a separate process with no Postgres access, had no
way to know. Fixed with a new `enterprise/internal/searchclient.
TenantChecker` (backed by `rbacstore.TenantIsActive`); both halves of
the fix verified Docker-free (`chrunner_test.go`'s and
`searchclient_test.go`'s `TestSearchRefusesMidProvisioningTenant`-shaped
tests) — see `api/queryapi/tenant_isolation_gap_test.go` for the full
accounting of all four probes, now all closed. The deployment-
topology gap that briefly was the largest one is now closed for both
Helm and docker-compose: `deploy/helm/sentry/templates/api.yaml`/
`enterprise-api.yaml` are mutually exclusive on the same
`enterprise.enabled` flag that turns on RBAC/audit/SSO, rendering to the
same Service name/port either way — a Helm-deployed cluster can't
accidentally run the wrong one. `docker-compose.yml`'s `api`/
`enterprise-api` services are now the same mutually-exclusive choice,
gated behind `COMPOSE_PROFILES` (`.env` checks in `single-tenant` as the
zero-config default) and sharing a host port/network-alias trick so
`alerting`/`web` need no conditional logic either way — verified via
`docker compose config` (renders/validates without a daemon, confirms
the two never both appear for one profile selection), not an actual
`docker compose up` in this environment. Per-resource dashboard grants
(the RBAC matrix's "(own/granted)"
qualifier) are now enforced too: `api/dashboards.PermissionStore` (core
interface) implemented by `enterprise/internal/rbacstore.
DashboardPermissions`, wired in only by `enterprise-api` — an Editor can
now only edit/delete a dashboard they created or were granted access to,
not every dashboard in their tenant; managing grants themselves is
stricter still (creator/Admin/Owner only, closing a self-escalation
path). Verified against a fake store (`api/dashboards/handler_test.go`);
real integration tests exist but haven't run against a live Postgres,
same disclosed gap as the rest of this phase's Postgres-backed pieces.
`sentryctl dashboards permissions list|grant|revoke` is now the CLI
surface for this — `PUT`/`DELETE /dashboards/{id}/permissions/{userId}`
previously had no caller but Go tests and curl.
`deploy/operator`'s `Tenant` CRD and `enterprise-api -provision-tenant`
are now unified too, deliberately lightweight rather than making the
K8s controller a second real actor: `-provision-tenant` stays the sole
caller of ClickHouse/`rbacstore`, and (via a new
`enterprise/internal/tenantcrd`, gated on `TENANT_CRD_NAMESPACE`) syncs
its real result into the CRD — a real credential Secret, not the
previous placeholder that authenticated against nothing, and status
fields the reconciler derives `Phase`/`Ready` from instead of
independently guessing "Active" the moment a Tenant object exists. The
tenant-picker is now fully built, backend and frontend: an identity with
more than one `tenant_memberships` row gets a real `GET
/auth/memberships`/`POST /auth/select-tenant` round trip (a short-lived
pending-login token, distinct from a real session by both Go type and
JWT claim name — a real token-confusion bug this design's own tests
caught before it shipped) instead of the flat refusal Phase 4 shipped
with earlier, and `web/src/routes/select-tenant` is the page that
actually calls it — see the Phase 4 exit-criteria paragraph below for
what changed to make that verifiable in this environment. Ingest
tenant-awareness — the gap this section used to
call "undesigned" — now has a real, if intentionally partial, design:
`ingest` (AGPL core) gained an optional `TenantResolver`
(`ingest/internal/grpcserver`), a per-tenant bearer credential an agent
presents (minted via `enterprise-auth
-create-ingest-credential-tenant=<id>`, validated over the network via a
new `POST /internal/authorize-ingest` endpoint — never an `enterprise/`
import, same boundary shape as `api/authz.Authorizer`), and the
resolved tenant ID is attached to every record as a `tenant_id` Kafka
message header before it's produced. **ClickHouse write-routing is now
built too**: `enterprise/cmd/enterprise-ingest` (another "second binary,"
mirroring `enterprise-api`) reuses `ingest/consumer`'s own flush loop
with `enterprise/internal/chwriter.Registry` swapped in as the writer —
one dedicated ClickHouse connection per tenant, routing each batch's
records by their `tenant_id` tag, fail-closed on an untagged or
unprovisioned tenant. Building it found and fixed a real bug:
`tenantprovision.ProvisionClickHouse` originally granted a tenant's
ClickHouse user `SELECT` only, which would have made every real
per-tenant write fail with a permission error — fixed by granting
`SELECT, INSERT` (one credential, both directions; no cross-tenant
boundary is crossed by also allowing INSERT within a tenant's own
database). **Tantivy write-routing is now built too**
`search/src/consumer.rs` (a completely independent Redpanda consumer,
not called through `ingest` or `enterprise-ingest` at all: a different
codebase and process) now resolves each record's `tenant_id` header
through the *same* `IndexRegistry` the read side already used, and
routes the write there instead of always into the default index. Unlike
the ClickHouse side, this needed no "second binary": `IndexRegistry`
already lives in this AGPL-core binary (Tantivy has no grant system to
gate a commercially-licensed credential behind, so there was never an
import-boundary reason to split it out), so read and write share one
registry directly. The periodic Tantivy commit now commits every tenant
index that's seen a write, not just the default one
(`IndexRegistry::commit_all`). **The active-tenant gap this same change
originally disclosed is now closed too**: `search/src/tenants.rs`'s
`ActiveTenantTracker` polls a new `GET /internal/active-tenants`
endpoint on `enterprise-auth` (`search` has no Postgres access, unlike
`chwriter.Registry`'s direct `rbacstore` query or the read side's
`searchclient.TenantChecker`, so this needed a network call — the same
"network boundary, not import boundary" shape `ingest`'s
`TenantResolver` already uses against the same service, authenticated
with a RoleService credential the same way `alerting` authenticates to
`api`) and `consumer.rs` refuses any tagged record whose tenant isn't in
the polled allowlist. Off unless both `ENTERPRISE_AUTH_URL` and
`ENTERPRISE_AUTH_SERVICE_TOKEN` are set (same "off unless configured"
default as everything else optional in this codebase); when they are,
startup blocks on the first fetch succeeding and later refresh failures
keep serving the last-known-good set rather than clearing it. Verified
with real HTTP round trips against a hand-rolled TCP test server in this
environment, no live enterprise-auth needed. **This closing move exposed
the ClickHouse side's own gap by comparison** — `chwriter.Registry`'s
writer map was still a startup-only snapshot with no refresh at all, a
real asymmetry once Tantivy's tracker refreshed every minute and
ClickHouse's didn't — so `Registry.StartRefreshing` (new) closes that
too: same one-minute interval, same last-known-good posture on a failed
refresh, opening connections for newly-active tenants and closing ones
no longer active. Both engines now share the same active-tenant
staleness bound instead of one being materially staler than the other.
**The tenant-picker page is now built too**:
`web/src/routes/select-tenant` calls `GET /auth/memberships`/
`POST /auth/select-tenant` via `fetch(..., {credentials: 'include'})`
(new `$lib/api.ts` functions), which needed a second CORS posture
alongside the wildcard-friendly one `enterprise-api` already had —
`api/httpserver.WithCredentialedCORS`, set to a literal origin via a new
`CORS_ALLOWED_ORIGIN` on `enterprise-auth` — since browsers refuse to
honor a wildcard `Access-Control-Allow-Origin` on a credentialed
request. **Genuinely verified in a real browser in this environment**:
a throwaway Node server standing in for `enterprise-auth`'s exact wire
contract (including its plain-text `http.Error` bodies, not JSON) on a
different origin than `web`'s dev server, driven through the full
cross-origin pending-login-cookie round trip, a real click choosing a
tenant, and the post-selection redirect — plus the missing/expired-
pending-login error path — with no Docker or live Postgres/IdP needed,
since the point was exercising `web`'s own fetch/CORS/cookie wiring, not
`enterprise-auth`'s internals (already covered by that package's own
tests). See `/web/README.md`'s "Tenant picker" section for the exact
setup. What's left in this phase now is entirely the caveats already
disclosed above, not an unbuilt feature: the ClickHouse/Postgres-backed
pieces have never run against a real database in this environment, and
nothing here has been tried against a real external IdP or a real
running multi-container deployment. Full accounting:
`/docs/security/threat-model.md`; step-by-step verification procedure
(not yet run against a live cluster in this environment):
`/docs/phase-4-runbook.md`. The rest of this section describes the exit
bar this phase is aiming at, not a completed state.
Two tenants can be provisioned with SSO (OIDC or SAML), each with their
own users, roles, dashboards, and alert rules, fully isolated at the
ClickHouse/Tantivy connection layer — not by a row filter — with
adversarial integration tests proving no cross-tenant data leakage,
including via the raw-SQL escape hatch and ClickHouse's own `system.*`
tables. A tenant admin can see a query audit trail for their tenant,
backed by append-only storage a compromised application credential
cannot alter (enforced by database grants, not just convention) and
periodically anchored outside the database so tampering is detectable
even against a privileged attacker. See `/docs/phase-4-isolation-design.md`
for the tenant isolation model and why it lives at the connection layer,
`/docs/phase-4-rbac-design.md` for the role/permission model, and
`/docs/security/threat-model.md` for the auth flows and audit-log
integrity guarantees, written for a prospective enterprise customer's
security team.
The tenant-isolation, provisioning, SSO, and RBAC-enforcement mechanisms
live entirely in `enterprise/` (commercial license), confirmed
explicitly rather than assumed: AGPL core (`/api`, `/alerting`, `/web`)
stays genuinely single-tenant, with no multi-tenant mechanism present at
all — `enterprise/` supplies tenant-scoped implementations of core's
already-shipped `querylang/executor.SQLRunner`/`SearchClient` interfaces
rather than core growing tenant awareness. Query-compiler-level "compile
time" enforcement, as originally proposed, turned out not to be
achievable in any module once Phase 2's opaque raw-SQL passthrough is
accounted for — the honest, implemented guarantee is that every code
path (compiled query or raw SQL) is forced through a tenant-scoped
database connection/index that the database's own access control
enforces, not a compiler-injected filter.
Non-goals for this phase (same discipline as every phase so far):
- No deny-override permissions — per-resource grants (e.g. a specific
user getting edit access to one dashboard) are additive only; a full
allow/deny ACL system is future work.
- No data retention/deletion policy design for tenant deprovisioning —
the provisioning state machine includes a `deprovisioning` state, but
what actually happens to a deprovisioned tenant's data is a separate,
not-yet-designed compliance question.
- No general multi-cluster orchestration in `/deploy` — scoped to
proving the per-tenant ClickHouse/Tantivy isolation model works, not a
fully general multi-cluster system.
- No protection against a privileged ClickHouse/Postgres administrator —
the isolation and audit-log guarantees in this phase are structural
defenses against application-layer bugs and injection, not against
someone with database superuser access; that's an operational control,
out of scope here and named explicitly, not silently assumed away.
## When in doubt
Ask before: changing the pinned stack, adding a new external dependency
that pulls in a large transitive tree, or making an architectural decision
that isn't already specified in `/docs/architecture.md`.