Ingest tenant-awareness was named "undesigned, not just unbuilt" across CLAUDE.md/threat-model.md/the runbook since early Phase 4 -- the last major standing gap. Scoping was agreed via AskUserQuestion: a config-supplied tenant_id + shared-secret token ingest validates (smaller real implementation, no new PKI), over per-tenant mTLS certs. This change builds that identity mechanism end to end and attaches it to every record at the point it enters the system; it deliberately does NOT build per-tenant write-routing for ClickHouse or Tantivy -- that's real, separately-scoped follow-up work, disclosed explicitly everywhere this was previously called undesigned, not silently left half-done. New pieces: - metadata/migrations/0034 + enterprise/internal/rbacstore/ ingest_credentials.go: a per-tenant bearer credential, only its SHA-256 hash ever persisted (same reasoning a password gets hashed, not stored raw) -- CreateIngestCredential returns the plaintext exactly once, ValidateIngestCredential/RevokeIngestCredential/ ListIngestCredentialsForTenant round it out. - enterprise-auth gains -create-ingest-credential-tenant/ -list-ingest-credentials-tenant/-revoke-ingest-credential (same offline-operator-flag shape as every other credential-minting flag in this binary) and a new POST /internal/authorize-ingest endpoint (internal/authhandler) validating a presented token and resolving its tenant -- a genuinely different credential type from session-backed /internal/authorize, so it doesn't touch session.Manager at all. - ingest (AGPL core) gains an optional TenantResolver (internal/grpcserver, nil by default) and its HTTP client implementation (internal/tenantresolver.HTTPResolver) -- a plain HTTP call to enterprise-auth's new endpoint, never an enterprise/ import, same "network boundary, not import boundary" shape api/authz.HTTPAuthorizer already uses for the query path. PushBatch now requires an `authorization: Bearer <token>` gRPC metadata entry once a resolver is configured, fails the whole batch closed on a missing/invalid credential (never falls back to "no tenant"), and attaches the resolved tenant ID to every record as a `tenant_id` Kafka message header before producing it. Verified with real round trips at every layer, no Docker needed: rbacstore's credential CRUD (skip-gated on live Postgres, same as every other rbacstore integration test this phase), authhandler's new endpoint (real HTTP via httptest, including the regression test that a session token must not validate as an ingest credential), tenantresolver (real HTTP client against httptest, same pattern as authz.HTTPAuthorizer's own tests), and grpcserver's PushBatch (fake resolver/producer -- no resolver leaves messages unchanged, a configured resolver attaches the right header or fails closed on a bad/missing token). Helm: ingest.requireTenantCredential (default false) is a deliberate, separate opt-in from enterprise.enabled -- turning ENTERPRISE_AUTH_URL on for ingest requires every agent to already hold a credential or be refused outright, so it must not default on just because enterprise.enabled does (same reasoning api.yaml's ENTERPRISE_AUTH_URL isn't tied to enterprise.enabled directly either). docker-compose.yml leaves it unset, same as ever. Docs updated everywhere this was called "undesigned": CLAUDE.md, docs/architecture.md, docs/security/threat-model.md (including its summary table, now split into "identity: built" vs "write-routing: not yet"), docs/phase-4-runbook.md (new §13), enterprise/README.md.
ingest
Go service sitting between the Rust agent and ClickHouse. Two halves in one
binary, selected with --mode:
- server — mTLS gRPC front end (
LogIngest.PushBatch) that agents connect to. Assigns each record a server-siderecord_id(a UUID, overwriting whatever the agent sent — agents always send it empty) and otherwise forwards records proto-encoded onto Redpanda unchanged. Still kept thin — one field assignment, no real normalization — so agent- facing latency isn't coupled to ClickHouse write performance.record_idhas to be assigned exactly once, here, rather than independently by each downstream consumer: Phase 1's Tantivy indexer and the ClickHouse writer both read the same Redpanda messages and need to agree on the same ID for the same record to join search hits back to rows — two consumers generating their own IDs would produce mismatched ones for what's supposed to be the same record. - consumer — reads back off Redpanda, normalizes into the ClickHouse row
shape (
internal/normalize), and batch-writes via the native protocol driver. Commits Redpanda offsets only after a successful ClickHouse write, so a ClickHouse outage causes redelivery on restart rather than data loss. - all (default) — both, in one process. This is what docker-compose
runs. Splitting into two deployments later (e.g. to scale them
independently in k8s) is a manifest change, not a code change — see
--mode.
Why Redpanda stays in the path
Confirmed with the project owner during Phase 0 planning: the gRPC front
end produces to Redpanda rather than writing ClickHouse directly. This
exercises the pinned transport layer from day one and keeps agents from
ever needing Kafka credentials — mTLS to ingest is the only network
egress an agent has. See /docs/architecture.md.
Dependencies worth knowing about
- github.com/segmentio/kafka-go — pure Go, no cgo, chosen over franz-go/confluent-kafka-go specifically to keep the distroless build simple (confirmed with the project owner; see git history / PR discussion for the tradeoffs considered).
- github.com/ClickHouse/clickhouse-go/v2 — official client, native protocol, pure Go (no cgo).
- golang.org/x/sync/errgroup — used in
cmd/ingest/main.goto run the server and consumer halves concurrently and propagate the first error. - github.com/google/uuid — was already in the dependency graph
transitively (via clickhouse-go); promoted to a direct dependency for
record_idgeneration ininternal/grpcserver, so not a new addition to the transitive tree.
Configuration
All via environment variables (see internal/config/config.go for the
full list and defaults) — no config file format for Phase 0:
| Var | Default | Purpose |
|---|---|---|
GRPC_LISTEN_ADDR |
:4317 |
Agent-facing gRPC listen address |
TLS_CERT_FILE / TLS_KEY_FILE |
/etc/sentry-ingest/server{,-key}.pem |
ingest's own mTLS identity |
TLS_CLIENT_CA_FILE |
/etc/sentry-ingest/ca.pem |
CA used to verify agent client certs |
REDPANDA_BROKERS |
localhost:9092 |
Comma-separated broker list |
REDPANDA_TOPIC |
sentry.logs.raw |
Must match the topic provisioned in /transport |
REDPANDA_CONSUMER_GROUP |
sentry-ingest |
Consumer group id |
CLICKHOUSE_ADDR |
localhost:9000 |
Native protocol port, not HTTP |
CLICKHOUSE_DATABASE / _USERNAME / _PASSWORD |
sentry / default / `` |
|
CONSUMER_BATCH_MAX_SIZE |
500 |
Records per ClickHouse batch insert |
CONSUMER_BATCH_FLUSH_INTERVAL_MS |
2000 |
Max time a partial batch waits before flushing |
Building & testing
go build ./...
go vet ./...
go test ./...
Requires google.golang.org/protobuf/cmd/protoc-gen-go and
google.golang.org/grpc/cmd/protoc-gen-go-grpc only if you're
regenerating /proto's Go bindings — ingest itself just imports the
already-generated github.com/sentry/sentry/proto module (see the
replace directive in go.mod, pointing at ../proto).
# from the repo root, not ingest/
docker build -f ingest/Dockerfile -t sentry-ingest .
Testing notes
internal/consumer and internal/grpcserver depend on Redpanda and
ClickHouse only through small interfaces (reader/chWriter in consumer,
batchProducer in grpcserver), so the flush/commit/error-handling logic is
unit-tested against fakes — no embedded broker or database needed. What's
not covered by these tests: the real kafka.Reader/kafka.Writer
wiring and the ClickHouse native-protocol driver itself. Those are only
exercised by the docker-compose end-to-end flow described in
/docs/phase-0-runbook.md.