Phase 4: Helm chart enforces api vs enterprise-api, closing the deployment-topology gap
deploy/helm/sentry/templates/api.yaml and the new enterprise-api.yaml
are mutually exclusive, gated on opposite sides of the same
enterprise.enabled flag -- exactly one renders, both as a Deployment+
Service named {{ .Release.Name }}-api on port 8080, so every consumer
(alerting's API_QUERY_URL, web's build args) needs zero conditional
logic of its own. This is the concrete fix for what the threat model
named as the single largest remaining gap once both storage engines'
isolation mechanisms were built: previously nothing forced or flagged
whether a deployment ran the tenant-isolated binary. Now the same flag
that turns on RBAC/audit/SSO also chooses the query binary.
Verified by parsing (not eyeballing) helm template's rendered output
under both value sets: exactly one sentry-api Deployment/Service either
way, with the right image, and kubeconform -strict clean against the
real Kubernetes 1.31 schema. Not applied to a live cluster (still no
cluster in this environment) -- docker-compose.yml also still runs
plain api unconditionally, so this enforcement is Helm-only for now.
Updated the threat model, architecture doc, CLAUDE.md, and deploy/
READMEs to reflect this and to name what's left: ingest has no tenant
concept for either storage engine (undesigned), and the Tenant CRD
(deploy/operator) and enterprise-api -provision-tenant are still two
separate, unreconciled provisioning mechanisms.
This commit is contained in:
+29
-19
@@ -13,28 +13,34 @@ non-goals). Two pieces:
|
||||
## What "multi-tenant-aware" means here, precisely
|
||||
|
||||
Per `/docs/phase-4-isolation-design.md`, tenant isolation itself lives at
|
||||
the **application layer** inside `enterprise/` (one `api` process holds a
|
||||
map of per-tenant ClickHouse connection pools; one `search` process holds
|
||||
a map of per-tenant Tantivy indices) -- not at the Kubernetes layer. This
|
||||
directory is **not** "one Deployment per tenant" or a general
|
||||
multi-cluster system; that's an explicit Phase 4 non-goal (see
|
||||
`/CLAUDE.md`). What it *does* add, matching that same document's exit
|
||||
criteria ("real per-tenant secret management, replacing today's single
|
||||
shared `CLICKHOUSE_PASSWORD`"):
|
||||
the **application layer** inside `enterprise/` (`enterprise-api` holds a
|
||||
map of per-tenant ClickHouse connection pools via `internal/chrunner`;
|
||||
`search` holds a map of per-tenant Tantivy indices via
|
||||
`src/registry.rs`) -- not at the Kubernetes layer. This directory is
|
||||
**not** "one Deployment per tenant" or a general multi-cluster system;
|
||||
that's an explicit Phase 4 non-goal (see `/CLAUDE.md`). What it *does*
|
||||
add:
|
||||
|
||||
- A `Tenant` CRD + controller that generates and manages one dedicated
|
||||
ClickHouse credential Secret per tenant (`operator/internal/controller`).
|
||||
- A Helm chart that can install zero-or-more `Tenant` CRs
|
||||
(`values.tenants`) alongside the rest of the stack.
|
||||
(`values.tenants`) alongside the rest of the stack, and — the newer
|
||||
piece — swaps `api`'s Deployment for `enterprise-api`'s whenever
|
||||
`enterprise.enabled` is true, so which query binary actually serves
|
||||
traffic is no longer a separately-forgettable decision (see
|
||||
`helm/sentry/README.md`'s "`api` vs `enterprise-api`" section).
|
||||
|
||||
The Operator does **not** call ClickHouse (no `CREATE DATABASE`/`CREATE
|
||||
USER`/`GRANT`) and does not touch the Tantivy filesystem or
|
||||
`enterprise/internal/rbacstore` -- that's `enterprise/internal/
|
||||
tenantprovision`, still unbuilt (see the Phase 4 task 5 summary). A
|
||||
`Tenant` reaching `status.phase: Active` here means "this tenant has a
|
||||
K8s Secret," not "this tenant's ClickHouse database/grants exist" --
|
||||
those are two different systems' state machines that aren't reconciled
|
||||
together yet, named explicitly rather than implied.
|
||||
**Two still-separate mechanisms, not yet unified**: the Operator's
|
||||
`Tenant` CRD manages only the K8s-side credential Secret — it does not
|
||||
call ClickHouse (no `CREATE DATABASE`/`CREATE USER`/`GRANT`) or touch
|
||||
the Tantivy filesystem. `enterprise-api -provision-tenant=<id>` is what
|
||||
actually does that (`enterprise/internal/tenantprovision`, built and
|
||||
tested — see `/enterprise/README.md`), driven independently via
|
||||
`rbacstore`, not from the `Tenant` CRD's reconcile loop. A `Tenant`
|
||||
reaching `status.phase: Active` here means "this tenant has a K8s
|
||||
Secret," not "this tenant's ClickHouse database/grants exist" — running
|
||||
both mechanisms for the same tenant ID today requires two separate
|
||||
operator actions, named explicitly rather than implied to be one.
|
||||
|
||||
## Verification status -- read before trusting this against a real cluster
|
||||
|
||||
@@ -62,11 +68,15 @@ access was available to fetch these tools, but no cluster):
|
||||
cleanly under both default values and a `enterprise.enabled: true` +
|
||||
two-tenant override; the rendered output was checked with `kubeconform
|
||||
-strict` against the real Kubernetes 1.31 OpenAPI schema for every
|
||||
built-in resource kind (22-29 resources depending on values, 0
|
||||
built-in resource kind (22-31 resources depending on values, 0
|
||||
invalid) -- this catches schema mistakes (wrong field names, wrong
|
||||
types) but not whether the resources actually reconcile correctly
|
||||
together on a live cluster (Job/StatefulSet startup ordering, PVC
|
||||
provisioning, actual pod scheduling).
|
||||
provisioning, actual pod scheduling). Specifically confirmed by
|
||||
parsing the rendered YAML (not just eyeballing it): exactly one
|
||||
`Deployment`/`Service` named `sentry-api` renders in each mode, with
|
||||
the `enterprise.enabled: true` render using the `enterprise-api` image
|
||||
and the default render using plain `api`'s.
|
||||
- Docker image builds (`operator/Dockerfile` and every other
|
||||
`Dockerfile` this chart references) were **not** verified in this
|
||||
session -- Docker's daemon wasn't reachable here either (see the
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# deploy/helm/sentry
|
||||
|
||||
A Helm chart covering every `docker-compose.yml` service (Redpanda,
|
||||
ClickHouse, Postgres, ingest, search, api, alerting, web) plus, when
|
||||
ClickHouse, Postgres, ingest, search, alerting, web) plus, when
|
||||
`enterprise.enabled: true`: enterprise-auth, the `deploy/operator`
|
||||
tenant-operator, and `Tenant` CRs from `values.tenants`. See
|
||||
`/deploy/README.md` for what "multi-tenant-aware" does and doesn't mean
|
||||
@@ -12,6 +12,27 @@ This chart never builds images -- push every image its `values.yaml`
|
||||
references to a registry the cluster can pull from first, same division
|
||||
of labor as `docker compose build` vs. `docker compose up`.
|
||||
|
||||
## `api` vs `enterprise-api`: one Deployment, chosen by `enterprise.enabled`
|
||||
|
||||
`templates/api.yaml` and `templates/enterprise-api.yaml` are mutually
|
||||
exclusive, gated on opposite sides of the same `enterprise.enabled` flag
|
||||
-- exactly one of them ever renders, both under the same
|
||||
`{{ .Release.Name }}-api` Service name and port 8080. This is the fix
|
||||
for what `/docs/security/threat-model.md` named as Phase 4's single
|
||||
largest remaining gap once both storage engines' isolation mechanisms
|
||||
were built: previously nothing forced or even flagged whether a
|
||||
deployment ran the tenant-isolated binary. Now it's not a second knob to
|
||||
remember -- the same flag that turns on RBAC/audit/SSO also swaps which
|
||||
query binary actually serves `/query` and `/dashboards` traffic. Every
|
||||
consumer (`alerting`'s `API_QUERY_URL`, `web`'s build args) needs zero
|
||||
conditional logic of its own, since both variants answer on the same
|
||||
name/port.
|
||||
|
||||
`enterprise-api` starts with an empty tenant set until
|
||||
`-provision-tenant` has been run for at least one tenant (see
|
||||
`/enterprise/README.md`) -- until then it's up and healthy, but every
|
||||
`/query` request correctly fails closed with no tenant to route to.
|
||||
|
||||
## Startup ordering
|
||||
|
||||
`docker-compose.yml` uses `depends_on: condition: service_healthy` /
|
||||
@@ -56,10 +77,15 @@ kubectl get secret sentry-tenant-acme-clickhouse sentry-tenant-globex-clickhouse
|
||||
|
||||
This proves the K8s-side half of Phase 4's "two tenants... with their
|
||||
own users, roles, dashboards" exit criteria (`/CLAUDE.md`) -- a real
|
||||
per-tenant credential Secret exists for each. It does **not** by itself
|
||||
give either tenant a working login, dashboard, or ClickHouse database:
|
||||
those need the OIDC/SAML login handlers, `internal/tenantprovision`, and
|
||||
`internal/rbacstore` wiring the Phase 4 task 5 summary names as deferred.
|
||||
per-tenant credential Secret exists for each, generated by
|
||||
`deploy/operator`'s `Tenant` controller. It does **not** by itself give
|
||||
either tenant a working ClickHouse database or Tantivy index -- that
|
||||
needs `enterprise-api -provision-tenant=<id>` (a separate, deliberately
|
||||
manual operator action; the `Tenant` CRD and `-provision-tenant` are two
|
||||
independent mechanisms today, not yet unified -- see
|
||||
`/enterprise/README.md`), and OIDC login (built, but still needs a
|
||||
manual `tenant_memberships` row -- see `/docs/phase-4-runbook.md` §3a)
|
||||
before a human can actually query as that tenant.
|
||||
|
||||
## `web`'s image needs rebuilding per environment
|
||||
|
||||
|
||||
@@ -1,3 +1,13 @@
|
||||
{{/*
|
||||
Mutually exclusive with enterprise-api.yaml's Deployment+Service, gated
|
||||
the opposite way -- see that file's doc comment for why: "does a
|
||||
deployment run the tenant-isolated binary or not" should be a single
|
||||
values.yaml decision (enterprise.enabled), not two independently
|
||||
driftable ones. Both render a Service named {{ .Release.Name }}-api on
|
||||
port 8080, so every consumer (alerting's API_QUERY_URL, web's build
|
||||
args) needs zero conditional logic of its own.
|
||||
*/}}
|
||||
{{- if not .Values.enterprise.enabled }}
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
@@ -44,15 +54,11 @@ spec:
|
||||
secretKeyRef:
|
||||
name: {{ .Release.Name }}-postgres
|
||||
key: password
|
||||
{{- if .Values.enterprise.enabled }}
|
||||
# Turns on authz.RequireRole*/RequireRoleOrService enforcement
|
||||
# on /query and /dashboards -- see api/authz and
|
||||
# /docs/phase-4-rbac-design.md. Off (unset) when
|
||||
# enterprise.enabled is false, matching every nil-authorizer
|
||||
# no-op default in this codebase.
|
||||
- name: ENTERPRISE_AUTH_URL
|
||||
value: "http://{{ .Release.Name }}-enterprise-auth:8082"
|
||||
{{- end }}
|
||||
# No ENTERPRISE_AUTH_URL here -- this file only renders when
|
||||
# enterprise.enabled is false (see the top of this file), so
|
||||
# authz.RequireRole*/RequireRoleOrService stay a permanent
|
||||
# no-op for this Deployment. enterprise-api.yaml is where
|
||||
# that enforcement actually turns on.
|
||||
ports:
|
||||
- name: http
|
||||
containerPort: 8080
|
||||
@@ -77,3 +83,4 @@ spec:
|
||||
ports:
|
||||
- name: http
|
||||
port: 8080
|
||||
{{- end }}
|
||||
|
||||
@@ -0,0 +1,111 @@
|
||||
{{/*
|
||||
Mutually exclusive with api.yaml's Deployment+Service -- see that file's
|
||||
doc comment. This is the concrete fix for the deployment-topology gap
|
||||
/docs/security/threat-model.md names as the single largest remaining
|
||||
Phase 4 issue once both storage engines' isolation mechanisms were
|
||||
built: "nothing forces or flags whether a deployment runs the isolated
|
||||
binary." With this file, it's not a separate knob to forget -- the same
|
||||
enterprise.enabled that turns on RBAC/audit/SSO also swaps which query
|
||||
binary actually serves traffic. Uses the "api" selector label (not
|
||||
"enterprise-api") deliberately, so the shared Service name+port below
|
||||
routes to whichever Deployment is actually rendered, with zero
|
||||
conditional logic needed in any consumer (alerting, web).
|
||||
*/}}
|
||||
{{- if .Values.enterprise.enabled }}
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: {{ .Release.Name }}-api
|
||||
labels:
|
||||
{{- include "sentry.labels" . | nindent 4 }}
|
||||
{{- include "sentry.selectorLabels" (list $ "api") | nindent 4 }}
|
||||
app.kubernetes.io/component: enterprise-api
|
||||
spec:
|
||||
replicas: {{ .Values.api.replicas }}
|
||||
selector:
|
||||
matchLabels:
|
||||
{{- include "sentry.selectorLabels" (list $ "api") | nindent 6 }}
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
{{- include "sentry.selectorLabels" (list $ "api") | nindent 8 }}
|
||||
spec:
|
||||
initContainers:
|
||||
{{- include "sentry.waitForTCP" (list "clickhouse" (printf "%s-clickhouse" .Release.Name) "9000") | nindent 8 }}
|
||||
{{- include "sentry.waitForTCP" (list "postgres" (printf "%s-postgres" .Release.Name) "5432") | nindent 8 }}
|
||||
{{- include "sentry.waitForTCP" (list "search" (printf "%s-search" .Release.Name) "50052") | nindent 8 }}
|
||||
{{- include "sentry.waitForTCP" (list "enterprise-auth" (printf "%s-enterprise-auth" .Release.Name) "8082") | nindent 8 }}
|
||||
containers:
|
||||
- name: enterprise-api
|
||||
image: "{{ .Values.enterprise.apiImage.repository }}:{{ .Values.enterprise.apiImage.tag }}"
|
||||
imagePullPolicy: {{ .Values.global.imagePullPolicy }}
|
||||
env:
|
||||
# :8080, not enterprise-api's own :8083 default -- this
|
||||
# container occupies the same Service/port every consumer
|
||||
# (alerting's API_QUERY_URL, web's build args) already
|
||||
# expects "-api:8080" to mean. See this file's doc comment.
|
||||
- name: HTTP_LISTEN_ADDR
|
||||
value: ":8080"
|
||||
- name: CLICKHOUSE_ADDR
|
||||
value: "{{ .Release.Name }}-clickhouse:9000"
|
||||
# tenantprovision's admin connection -- the same credential
|
||||
# clickhouse-migrate uses, needs access_management, never a
|
||||
# tenant-scoped grant. See enterprise/internal/tenantprovision's
|
||||
# doc comment.
|
||||
- name: CLICKHOUSE_ADMIN_USERNAME
|
||||
value: "default"
|
||||
- name: CLICKHOUSE_ADMIN_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: {{ .Release.Name }}-clickhouse
|
||||
key: password
|
||||
- name: SEARCH_GRPC_ADDR
|
||||
value: "{{ .Release.Name }}-search:50052"
|
||||
- name: POSTGRES_ADDR
|
||||
value: "{{ .Release.Name }}-postgres:5432"
|
||||
- name: POSTGRES_DATABASE
|
||||
value: sentry_metadata
|
||||
- name: POSTGRES_USERNAME
|
||||
value: sentry
|
||||
- name: POSTGRES_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: {{ .Release.Name }}-postgres
|
||||
key: password
|
||||
# Restricted audit_writer Postgres role (Phase 4 task 4) --
|
||||
# its own pool, never the shared "sentry" credential above.
|
||||
# See enterprise/internal/audit's doc comment.
|
||||
- name: AUDIT_WRITER_USERNAME
|
||||
value: "audit_writer"
|
||||
- name: AUDIT_WRITER_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: {{ .Release.Name }}-postgres
|
||||
key: auditWriterPassword
|
||||
- name: ENTERPRISE_AUTH_URL
|
||||
value: "http://{{ .Release.Name }}-enterprise-auth:8082"
|
||||
ports:
|
||||
- name: http
|
||||
containerPort: 8080
|
||||
readinessProbe:
|
||||
exec:
|
||||
command: ["/enterprise-api", "-healthcheck"]
|
||||
initialDelaySeconds: 5
|
||||
periodSeconds: 5
|
||||
resources:
|
||||
{{- toYaml .Values.api.resources | nindent 12 }}
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: {{ .Release.Name }}-api
|
||||
labels:
|
||||
{{- include "sentry.labels" . | nindent 4 }}
|
||||
{{- include "sentry.selectorLabels" (list $ "api") | nindent 4 }}
|
||||
spec:
|
||||
selector:
|
||||
{{- include "sentry.selectorLabels" (list $ "api") | nindent 4 }}
|
||||
ports:
|
||||
- name: http
|
||||
port: 8080
|
||||
{{- end }}
|
||||
@@ -140,6 +140,15 @@ enterprise:
|
||||
image:
|
||||
repository: sentry-enterprise-auth
|
||||
tag: latest
|
||||
# enterprise-api (templates/enterprise-api.yaml) -- swaps in for
|
||||
# api.yaml's plain api Deployment when enterprise.enabled is true, on
|
||||
# the same Service/port every consumer already expects. Built from the
|
||||
# repo root (needs api/ and proto/, not just enterprise/), unlike
|
||||
# enterprise-auth's image above -- see
|
||||
# enterprise/cmd/enterprise-api/Dockerfile.
|
||||
apiImage:
|
||||
repository: sentry-enterprise-api
|
||||
tag: latest
|
||||
replicas: 1
|
||||
resources: {}
|
||||
# Leave empty to auto-generate (>= 32 bytes) and persist across
|
||||
|
||||
Reference in New Issue
Block a user