Phase 4: Helm chart enforces api vs enterprise-api, closing the deployment-topology gap
deploy/helm/sentry/templates/api.yaml and the new enterprise-api.yaml
are mutually exclusive, gated on opposite sides of the same
enterprise.enabled flag -- exactly one renders, both as a Deployment+
Service named {{ .Release.Name }}-api on port 8080, so every consumer
(alerting's API_QUERY_URL, web's build args) needs zero conditional
logic of its own. This is the concrete fix for what the threat model
named as the single largest remaining gap once both storage engines'
isolation mechanisms were built: previously nothing forced or flagged
whether a deployment ran the tenant-isolated binary. Now the same flag
that turns on RBAC/audit/SSO also chooses the query binary.
Verified by parsing (not eyeballing) helm template's rendered output
under both value sets: exactly one sentry-api Deployment/Service either
way, with the right image, and kubeconform -strict clean against the
real Kubernetes 1.31 schema. Not applied to a live cluster (still no
cluster in this environment) -- docker-compose.yml also still runs
plain api unconditionally, so this enforcement is Helm-only for now.
Updated the threat model, architecture doc, CLAUDE.md, and deploy/
READMEs to reflect this and to name what's left: ingest has no tenant
concept for either storage engine (undesigned), and the Tenant CRD
(deploy/operator) and enterprise-api -provision-tenant are still two
separate, unreconciled provisioning mechanisms.
This commit is contained in:
@@ -164,16 +164,23 @@ container. Tantivy per-tenant index routing is now built too
|
||||
**genuinely verified**, like the OIDC login flow: Tantivy is an embedded
|
||||
library, not a networked service, so the isolation probe (three tenants,
|
||||
same search term, scoped search returns only that tenant's document)
|
||||
actually ran in this environment, no Docker needed. What still keeps
|
||||
this phase from being done: SAML login (protocol wiring exists, no ACS
|
||||
handler calls it, following OIDC's now-built pattern), ingest itself has
|
||||
no tenant concept for either storage engine (every record lands in the
|
||||
one shared ClickHouse database and Tantivy index no matter what —
|
||||
undesigned, not just unbuilt), and a deployment-topology gap that's now
|
||||
the single largest one: nothing yet forces or even flags whether a given
|
||||
deployment is actually running the isolated binary (`enterprise-api`)
|
||||
versus the plain single-tenant one (`api`); both still exist and nothing
|
||||
currently prevents mixing them up. Full accounting:
|
||||
actually ran in this environment, no Docker needed. The deployment-
|
||||
topology gap that briefly was the largest one is now closed for Helm:
|
||||
`deploy/helm/sentry/templates/api.yaml`/`enterprise-api.yaml` are
|
||||
mutually exclusive on the same `enterprise.enabled` flag that turns on
|
||||
RBAC/audit/SSO, rendering to the same Service name/port either way — a
|
||||
Helm-deployed cluster can't accidentally run the wrong one.
|
||||
`docker-compose.yml` still runs plain `api` unconditionally, though
|
||||
(local/dev parity with the Helm chart's enforcement is real remaining
|
||||
work). What still keeps this phase from being done: SAML login (protocol
|
||||
wiring exists, no ACS handler calls it, following OIDC's now-built
|
||||
pattern), ingest itself has no tenant concept for either storage engine
|
||||
(every record lands in the one shared ClickHouse database and Tantivy
|
||||
index no matter what — undesigned, not just unbuilt), and the two
|
||||
provisioning mechanisms (`deploy/operator`'s `Tenant` CRD and
|
||||
`enterprise-api -provision-tenant`) still aren't unified — running both
|
||||
for the same tenant ID is two separate operator actions today. Full
|
||||
accounting:
|
||||
`/docs/security/threat-model.md`; step-by-step verification procedure
|
||||
(not yet run against a live cluster in this environment):
|
||||
`/docs/phase-4-runbook.md`. The rest of this section describes the exit
|
||||
|
||||
+29
-19
@@ -13,28 +13,34 @@ non-goals). Two pieces:
|
||||
## What "multi-tenant-aware" means here, precisely
|
||||
|
||||
Per `/docs/phase-4-isolation-design.md`, tenant isolation itself lives at
|
||||
the **application layer** inside `enterprise/` (one `api` process holds a
|
||||
map of per-tenant ClickHouse connection pools; one `search` process holds
|
||||
a map of per-tenant Tantivy indices) -- not at the Kubernetes layer. This
|
||||
directory is **not** "one Deployment per tenant" or a general
|
||||
multi-cluster system; that's an explicit Phase 4 non-goal (see
|
||||
`/CLAUDE.md`). What it *does* add, matching that same document's exit
|
||||
criteria ("real per-tenant secret management, replacing today's single
|
||||
shared `CLICKHOUSE_PASSWORD`"):
|
||||
the **application layer** inside `enterprise/` (`enterprise-api` holds a
|
||||
map of per-tenant ClickHouse connection pools via `internal/chrunner`;
|
||||
`search` holds a map of per-tenant Tantivy indices via
|
||||
`src/registry.rs`) -- not at the Kubernetes layer. This directory is
|
||||
**not** "one Deployment per tenant" or a general multi-cluster system;
|
||||
that's an explicit Phase 4 non-goal (see `/CLAUDE.md`). What it *does*
|
||||
add:
|
||||
|
||||
- A `Tenant` CRD + controller that generates and manages one dedicated
|
||||
ClickHouse credential Secret per tenant (`operator/internal/controller`).
|
||||
- A Helm chart that can install zero-or-more `Tenant` CRs
|
||||
(`values.tenants`) alongside the rest of the stack.
|
||||
(`values.tenants`) alongside the rest of the stack, and — the newer
|
||||
piece — swaps `api`'s Deployment for `enterprise-api`'s whenever
|
||||
`enterprise.enabled` is true, so which query binary actually serves
|
||||
traffic is no longer a separately-forgettable decision (see
|
||||
`helm/sentry/README.md`'s "`api` vs `enterprise-api`" section).
|
||||
|
||||
The Operator does **not** call ClickHouse (no `CREATE DATABASE`/`CREATE
|
||||
USER`/`GRANT`) and does not touch the Tantivy filesystem or
|
||||
`enterprise/internal/rbacstore` -- that's `enterprise/internal/
|
||||
tenantprovision`, still unbuilt (see the Phase 4 task 5 summary). A
|
||||
`Tenant` reaching `status.phase: Active` here means "this tenant has a
|
||||
K8s Secret," not "this tenant's ClickHouse database/grants exist" --
|
||||
those are two different systems' state machines that aren't reconciled
|
||||
together yet, named explicitly rather than implied.
|
||||
**Two still-separate mechanisms, not yet unified**: the Operator's
|
||||
`Tenant` CRD manages only the K8s-side credential Secret — it does not
|
||||
call ClickHouse (no `CREATE DATABASE`/`CREATE USER`/`GRANT`) or touch
|
||||
the Tantivy filesystem. `enterprise-api -provision-tenant=<id>` is what
|
||||
actually does that (`enterprise/internal/tenantprovision`, built and
|
||||
tested — see `/enterprise/README.md`), driven independently via
|
||||
`rbacstore`, not from the `Tenant` CRD's reconcile loop. A `Tenant`
|
||||
reaching `status.phase: Active` here means "this tenant has a K8s
|
||||
Secret," not "this tenant's ClickHouse database/grants exist" — running
|
||||
both mechanisms for the same tenant ID today requires two separate
|
||||
operator actions, named explicitly rather than implied to be one.
|
||||
|
||||
## Verification status -- read before trusting this against a real cluster
|
||||
|
||||
@@ -62,11 +68,15 @@ access was available to fetch these tools, but no cluster):
|
||||
cleanly under both default values and a `enterprise.enabled: true` +
|
||||
two-tenant override; the rendered output was checked with `kubeconform
|
||||
-strict` against the real Kubernetes 1.31 OpenAPI schema for every
|
||||
built-in resource kind (22-29 resources depending on values, 0
|
||||
built-in resource kind (22-31 resources depending on values, 0
|
||||
invalid) -- this catches schema mistakes (wrong field names, wrong
|
||||
types) but not whether the resources actually reconcile correctly
|
||||
together on a live cluster (Job/StatefulSet startup ordering, PVC
|
||||
provisioning, actual pod scheduling).
|
||||
provisioning, actual pod scheduling). Specifically confirmed by
|
||||
parsing the rendered YAML (not just eyeballing it): exactly one
|
||||
`Deployment`/`Service` named `sentry-api` renders in each mode, with
|
||||
the `enterprise.enabled: true` render using the `enterprise-api` image
|
||||
and the default render using plain `api`'s.
|
||||
- Docker image builds (`operator/Dockerfile` and every other
|
||||
`Dockerfile` this chart references) were **not** verified in this
|
||||
session -- Docker's daemon wasn't reachable here either (see the
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# deploy/helm/sentry
|
||||
|
||||
A Helm chart covering every `docker-compose.yml` service (Redpanda,
|
||||
ClickHouse, Postgres, ingest, search, api, alerting, web) plus, when
|
||||
ClickHouse, Postgres, ingest, search, alerting, web) plus, when
|
||||
`enterprise.enabled: true`: enterprise-auth, the `deploy/operator`
|
||||
tenant-operator, and `Tenant` CRs from `values.tenants`. See
|
||||
`/deploy/README.md` for what "multi-tenant-aware" does and doesn't mean
|
||||
@@ -12,6 +12,27 @@ This chart never builds images -- push every image its `values.yaml`
|
||||
references to a registry the cluster can pull from first, same division
|
||||
of labor as `docker compose build` vs. `docker compose up`.
|
||||
|
||||
## `api` vs `enterprise-api`: one Deployment, chosen by `enterprise.enabled`
|
||||
|
||||
`templates/api.yaml` and `templates/enterprise-api.yaml` are mutually
|
||||
exclusive, gated on opposite sides of the same `enterprise.enabled` flag
|
||||
-- exactly one of them ever renders, both under the same
|
||||
`{{ .Release.Name }}-api` Service name and port 8080. This is the fix
|
||||
for what `/docs/security/threat-model.md` named as Phase 4's single
|
||||
largest remaining gap once both storage engines' isolation mechanisms
|
||||
were built: previously nothing forced or even flagged whether a
|
||||
deployment ran the tenant-isolated binary. Now it's not a second knob to
|
||||
remember -- the same flag that turns on RBAC/audit/SSO also swaps which
|
||||
query binary actually serves `/query` and `/dashboards` traffic. Every
|
||||
consumer (`alerting`'s `API_QUERY_URL`, `web`'s build args) needs zero
|
||||
conditional logic of its own, since both variants answer on the same
|
||||
name/port.
|
||||
|
||||
`enterprise-api` starts with an empty tenant set until
|
||||
`-provision-tenant` has been run for at least one tenant (see
|
||||
`/enterprise/README.md`) -- until then it's up and healthy, but every
|
||||
`/query` request correctly fails closed with no tenant to route to.
|
||||
|
||||
## Startup ordering
|
||||
|
||||
`docker-compose.yml` uses `depends_on: condition: service_healthy` /
|
||||
@@ -56,10 +77,15 @@ kubectl get secret sentry-tenant-acme-clickhouse sentry-tenant-globex-clickhouse
|
||||
|
||||
This proves the K8s-side half of Phase 4's "two tenants... with their
|
||||
own users, roles, dashboards" exit criteria (`/CLAUDE.md`) -- a real
|
||||
per-tenant credential Secret exists for each. It does **not** by itself
|
||||
give either tenant a working login, dashboard, or ClickHouse database:
|
||||
those need the OIDC/SAML login handlers, `internal/tenantprovision`, and
|
||||
`internal/rbacstore` wiring the Phase 4 task 5 summary names as deferred.
|
||||
per-tenant credential Secret exists for each, generated by
|
||||
`deploy/operator`'s `Tenant` controller. It does **not** by itself give
|
||||
either tenant a working ClickHouse database or Tantivy index -- that
|
||||
needs `enterprise-api -provision-tenant=<id>` (a separate, deliberately
|
||||
manual operator action; the `Tenant` CRD and `-provision-tenant` are two
|
||||
independent mechanisms today, not yet unified -- see
|
||||
`/enterprise/README.md`), and OIDC login (built, but still needs a
|
||||
manual `tenant_memberships` row -- see `/docs/phase-4-runbook.md` §3a)
|
||||
before a human can actually query as that tenant.
|
||||
|
||||
## `web`'s image needs rebuilding per environment
|
||||
|
||||
|
||||
@@ -1,3 +1,13 @@
|
||||
{{/*
|
||||
Mutually exclusive with enterprise-api.yaml's Deployment+Service, gated
|
||||
the opposite way -- see that file's doc comment for why: "does a
|
||||
deployment run the tenant-isolated binary or not" should be a single
|
||||
values.yaml decision (enterprise.enabled), not two independently
|
||||
driftable ones. Both render a Service named {{ .Release.Name }}-api on
|
||||
port 8080, so every consumer (alerting's API_QUERY_URL, web's build
|
||||
args) needs zero conditional logic of its own.
|
||||
*/}}
|
||||
{{- if not .Values.enterprise.enabled }}
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
@@ -44,15 +54,11 @@ spec:
|
||||
secretKeyRef:
|
||||
name: {{ .Release.Name }}-postgres
|
||||
key: password
|
||||
{{- if .Values.enterprise.enabled }}
|
||||
# Turns on authz.RequireRole*/RequireRoleOrService enforcement
|
||||
# on /query and /dashboards -- see api/authz and
|
||||
# /docs/phase-4-rbac-design.md. Off (unset) when
|
||||
# enterprise.enabled is false, matching every nil-authorizer
|
||||
# no-op default in this codebase.
|
||||
- name: ENTERPRISE_AUTH_URL
|
||||
value: "http://{{ .Release.Name }}-enterprise-auth:8082"
|
||||
{{- end }}
|
||||
# No ENTERPRISE_AUTH_URL here -- this file only renders when
|
||||
# enterprise.enabled is false (see the top of this file), so
|
||||
# authz.RequireRole*/RequireRoleOrService stay a permanent
|
||||
# no-op for this Deployment. enterprise-api.yaml is where
|
||||
# that enforcement actually turns on.
|
||||
ports:
|
||||
- name: http
|
||||
containerPort: 8080
|
||||
@@ -77,3 +83,4 @@ spec:
|
||||
ports:
|
||||
- name: http
|
||||
port: 8080
|
||||
{{- end }}
|
||||
|
||||
@@ -0,0 +1,111 @@
|
||||
{{/*
|
||||
Mutually exclusive with api.yaml's Deployment+Service -- see that file's
|
||||
doc comment. This is the concrete fix for the deployment-topology gap
|
||||
/docs/security/threat-model.md names as the single largest remaining
|
||||
Phase 4 issue once both storage engines' isolation mechanisms were
|
||||
built: "nothing forces or flags whether a deployment runs the isolated
|
||||
binary." With this file, it's not a separate knob to forget -- the same
|
||||
enterprise.enabled that turns on RBAC/audit/SSO also swaps which query
|
||||
binary actually serves traffic. Uses the "api" selector label (not
|
||||
"enterprise-api") deliberately, so the shared Service name+port below
|
||||
routes to whichever Deployment is actually rendered, with zero
|
||||
conditional logic needed in any consumer (alerting, web).
|
||||
*/}}
|
||||
{{- if .Values.enterprise.enabled }}
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: {{ .Release.Name }}-api
|
||||
labels:
|
||||
{{- include "sentry.labels" . | nindent 4 }}
|
||||
{{- include "sentry.selectorLabels" (list $ "api") | nindent 4 }}
|
||||
app.kubernetes.io/component: enterprise-api
|
||||
spec:
|
||||
replicas: {{ .Values.api.replicas }}
|
||||
selector:
|
||||
matchLabels:
|
||||
{{- include "sentry.selectorLabels" (list $ "api") | nindent 6 }}
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
{{- include "sentry.selectorLabels" (list $ "api") | nindent 8 }}
|
||||
spec:
|
||||
initContainers:
|
||||
{{- include "sentry.waitForTCP" (list "clickhouse" (printf "%s-clickhouse" .Release.Name) "9000") | nindent 8 }}
|
||||
{{- include "sentry.waitForTCP" (list "postgres" (printf "%s-postgres" .Release.Name) "5432") | nindent 8 }}
|
||||
{{- include "sentry.waitForTCP" (list "search" (printf "%s-search" .Release.Name) "50052") | nindent 8 }}
|
||||
{{- include "sentry.waitForTCP" (list "enterprise-auth" (printf "%s-enterprise-auth" .Release.Name) "8082") | nindent 8 }}
|
||||
containers:
|
||||
- name: enterprise-api
|
||||
image: "{{ .Values.enterprise.apiImage.repository }}:{{ .Values.enterprise.apiImage.tag }}"
|
||||
imagePullPolicy: {{ .Values.global.imagePullPolicy }}
|
||||
env:
|
||||
# :8080, not enterprise-api's own :8083 default -- this
|
||||
# container occupies the same Service/port every consumer
|
||||
# (alerting's API_QUERY_URL, web's build args) already
|
||||
# expects "-api:8080" to mean. See this file's doc comment.
|
||||
- name: HTTP_LISTEN_ADDR
|
||||
value: ":8080"
|
||||
- name: CLICKHOUSE_ADDR
|
||||
value: "{{ .Release.Name }}-clickhouse:9000"
|
||||
# tenantprovision's admin connection -- the same credential
|
||||
# clickhouse-migrate uses, needs access_management, never a
|
||||
# tenant-scoped grant. See enterprise/internal/tenantprovision's
|
||||
# doc comment.
|
||||
- name: CLICKHOUSE_ADMIN_USERNAME
|
||||
value: "default"
|
||||
- name: CLICKHOUSE_ADMIN_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: {{ .Release.Name }}-clickhouse
|
||||
key: password
|
||||
- name: SEARCH_GRPC_ADDR
|
||||
value: "{{ .Release.Name }}-search:50052"
|
||||
- name: POSTGRES_ADDR
|
||||
value: "{{ .Release.Name }}-postgres:5432"
|
||||
- name: POSTGRES_DATABASE
|
||||
value: sentry_metadata
|
||||
- name: POSTGRES_USERNAME
|
||||
value: sentry
|
||||
- name: POSTGRES_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: {{ .Release.Name }}-postgres
|
||||
key: password
|
||||
# Restricted audit_writer Postgres role (Phase 4 task 4) --
|
||||
# its own pool, never the shared "sentry" credential above.
|
||||
# See enterprise/internal/audit's doc comment.
|
||||
- name: AUDIT_WRITER_USERNAME
|
||||
value: "audit_writer"
|
||||
- name: AUDIT_WRITER_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: {{ .Release.Name }}-postgres
|
||||
key: auditWriterPassword
|
||||
- name: ENTERPRISE_AUTH_URL
|
||||
value: "http://{{ .Release.Name }}-enterprise-auth:8082"
|
||||
ports:
|
||||
- name: http
|
||||
containerPort: 8080
|
||||
readinessProbe:
|
||||
exec:
|
||||
command: ["/enterprise-api", "-healthcheck"]
|
||||
initialDelaySeconds: 5
|
||||
periodSeconds: 5
|
||||
resources:
|
||||
{{- toYaml .Values.api.resources | nindent 12 }}
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: {{ .Release.Name }}-api
|
||||
labels:
|
||||
{{- include "sentry.labels" . | nindent 4 }}
|
||||
{{- include "sentry.selectorLabels" (list $ "api") | nindent 4 }}
|
||||
spec:
|
||||
selector:
|
||||
{{- include "sentry.selectorLabels" (list $ "api") | nindent 4 }}
|
||||
ports:
|
||||
- name: http
|
||||
port: 8080
|
||||
{{- end }}
|
||||
@@ -140,6 +140,15 @@ enterprise:
|
||||
image:
|
||||
repository: sentry-enterprise-auth
|
||||
tag: latest
|
||||
# enterprise-api (templates/enterprise-api.yaml) -- swaps in for
|
||||
# api.yaml's plain api Deployment when enterprise.enabled is true, on
|
||||
# the same Service/port every consumer already expects. Built from the
|
||||
# repo root (needs api/ and proto/, not just enterprise/), unlike
|
||||
# enterprise-auth's image above -- see
|
||||
# enterprise/cmd/enterprise-api/Dockerfile.
|
||||
apiImage:
|
||||
repository: sentry-enterprise-api
|
||||
tag: latest
|
||||
replicas: 1
|
||||
resources: {}
|
||||
# Leave empty to auto-generate (>= 32 bytes) and persist across
|
||||
|
||||
+15
-8
@@ -147,15 +147,22 @@ escape hatch is opaque to any compiler-injected filter.
|
||||
`chrunner`/`searchclient` becomes tenant-aware on the write side,
|
||||
which is undesigned, not merely unbuilt.
|
||||
- `deploy/operator`'s `Tenant` CRD still manages only the K8s-side
|
||||
artifact (a credential Secret); the Helm chart has no service
|
||||
definition for `enterprise-api` yet.
|
||||
artifact (a credential Secret); it doesn't call
|
||||
`enterprise-api -provision-tenant` or otherwise trigger ClickHouse-side
|
||||
provisioning. The two mechanisms are independent today, not reconciled
|
||||
into one state machine.
|
||||
|
||||
The deployment-topology gap — giving the system an actual way to route
|
||||
traffic to `enterprise-api` instead of `api` (a Helm service, or at
|
||||
minimum a documented, enforced convention) — is now the single largest
|
||||
remaining gap between this system and the isolation model it was
|
||||
designed to have; both storage engines' connection/index-layer
|
||||
mechanisms themselves are built.
|
||||
**The deployment-topology gap is closed for the Helm chart**: `deploy/
|
||||
helm/sentry/templates/api.yaml`/`enterprise-api.yaml` are mutually
|
||||
exclusive on `enterprise.enabled`, rendering to the same Service name
|
||||
and port either way, so a Helm-deployed cluster can't accidentally run
|
||||
the wrong binary — the same flag that turns on RBAC/audit/SSO now also
|
||||
chooses the query binary. `docker-compose.yml` still runs plain `api`
|
||||
unconditionally, so this enforcement doesn't yet extend to local/dev.
|
||||
With both storage engines' connection/index-layer mechanisms built and
|
||||
deployment topology enforced at the Helm layer, the largest remaining
|
||||
gaps are ingest's lack of tenant-awareness (undesigned) and unifying the
|
||||
`Tenant` CRD with `-provision-tenant` into one provisioning flow.
|
||||
|
||||
## Licensing boundary
|
||||
|
||||
|
||||
+40
-7
@@ -314,17 +314,50 @@ Docker access, alongside `enterprise/internal/loginhandler`'s OIDC
|
||||
tests (§3a) — both are unusually strong evidence precisely because they
|
||||
needed no infrastructure this environment lacked.
|
||||
|
||||
## 10. Confirm the Helm chart actually enforces the binary swap
|
||||
|
||||
No live cluster needed for this either — `helm template`'s output is
|
||||
plain YAML, parseable without a cluster:
|
||||
|
||||
```sh
|
||||
cd deploy/helm/sentry
|
||||
helm template sentry . --include-crds > /tmp/default.yaml
|
||||
helm template sentry . --include-crds --set enterprise.enabled=true \
|
||||
--set 'tenants[0].name=acme' --set 'tenants[0].displayName=Acme Corp' \
|
||||
> /tmp/enterprise.yaml
|
||||
|
||||
python3 -c "
|
||||
import yaml
|
||||
for f in ['/tmp/default.yaml', '/tmp/enterprise.yaml']:
|
||||
docs = list(yaml.safe_load_all(open(f)))
|
||||
deploys = [d for d in docs if d and d.get('kind')=='Deployment' and d.get('metadata',{}).get('name')=='sentry-api']
|
||||
print(f, '->', [d['spec']['template']['spec']['containers'][0]['image'] for d in deploys])
|
||||
"
|
||||
# expect: default.yaml -> ['sentry-api:latest'], enterprise.yaml -> ['sentry-enterprise-api:latest']
|
||||
# and exactly one Deployment named sentry-api in each file.
|
||||
```
|
||||
|
||||
Real cluster (`kind create cluster`, or similar): `helm install` with
|
||||
each set of values and confirm `kubectl get deploy sentry-api -o
|
||||
jsonpath='{.spec.template.spec.containers[0].image}'` matches, and that
|
||||
`kubectl get svc sentry-api` routes to whichever one is actually running.
|
||||
|
||||
## Known gaps (do not treat this phase as done without reading these)
|
||||
|
||||
Full accounting: `/docs/security/threat-model.md`. Headline items:
|
||||
|
||||
- **Both storage engines' isolation exists but is opt-in.**
|
||||
`enterprise-api` (§8, §9) gives real per-tenant ClickHouse *and*
|
||||
Tantivy isolation, but plain `api` (still the default in
|
||||
`docker-compose.yml`/`web`'s base URL) has neither, and nothing flags
|
||||
which one a given deployment is actually running. This is now the
|
||||
single largest gap — not a missing mechanism, a missing enforcement/
|
||||
default.
|
||||
- **Both storage engines' isolation exists, and the Helm chart now
|
||||
enforces which binary runs.** `deploy/helm/sentry/templates/api.yaml`/
|
||||
`enterprise-api.yaml` are mutually exclusive on `enterprise.enabled`
|
||||
(§10) -- a Helm-deployed cluster can't accidentally run the
|
||||
non-isolated binary once that flag is set. `docker-compose.yml` still
|
||||
runs plain `api` unconditionally alongside a separately-started
|
||||
`enterprise-api` (§8), so this enforcement doesn't extend to local/dev
|
||||
yet.
|
||||
- The `Tenant` CRD (`deploy/operator`) and `enterprise-api
|
||||
-provision-tenant` are still two independent provisioning mechanisms
|
||||
-- running both for the same tenant ID today takes two separate
|
||||
operator actions, not one.
|
||||
- **Ingest has no tenant concept for either storage engine.** Every
|
||||
record `ingest` produces lands in the one shared ClickHouse database
|
||||
and the one shared Tantivy index no matter what. A newly-provisioned
|
||||
|
||||
@@ -44,16 +44,23 @@ searchclient`'s tests run a real in-process gRPC server and confirm the
|
||||
wire-level `SearchRequest` carries the right `tenant_id`. All pass, for
|
||||
real, no disclaimer needed for this specific claim.
|
||||
|
||||
**But plain `api/cmd/api` still runs with one shared ClickHouse
|
||||
connection and no tenant-scoped search client**, and nothing in this
|
||||
repo automatically routes traffic to `enterprise-api` instead —
|
||||
`docker-compose.yml` includes it "available, not defaulted into the
|
||||
traffic path" (same shape as `enterprise-auth`'s own addition), and the
|
||||
Helm chart has no service for it at all yet. **A deployment is only as
|
||||
isolated as which binary is actually serving traffic** — this is an
|
||||
operational decision nothing currently enforces or even surfaces as a
|
||||
warning. This is now the single largest gap in the isolation story, not
|
||||
a missing mechanism.
|
||||
**The Helm chart now closes this for K8s deployments; `docker-compose.yml`
|
||||
still doesn't.** `deploy/helm/sentry/templates/api.yaml` and
|
||||
`enterprise-api.yaml` are mutually exclusive, gated on opposite sides of
|
||||
the same `enterprise.enabled` flag, rendering to the same Service
|
||||
name/port — so a Helm-deployed cluster runs exactly one of the two
|
||||
binaries, chosen by the same flag that turns on RBAC/audit/SSO, not a
|
||||
second independently-forgettable decision. Verified by parsing (not
|
||||
eyeballing) the rendered YAML under both values: exactly one `sentry-api`
|
||||
Deployment either way, with the right image. **`docker-compose.yml`
|
||||
still runs plain `api` unconditionally** and includes `enterprise-api`
|
||||
as an extra, separately-started service — local/dev parity with the Helm
|
||||
chart's enforcement is real remaining work. And this only constrains
|
||||
*deployment*, not *operation*: nothing stops an operator from manually
|
||||
running plain `api`'s image against a cluster that has tenants
|
||||
provisioned, pointing at the same ClickHouse/Postgres. The Helm chart
|
||||
makes the *default*, chart-managed path correct; it isn't a runtime
|
||||
guard against misconfiguration.
|
||||
|
||||
**Ingest is not tenant-aware for either storage engine**, and this is
|
||||
more load-bearing than it sounds: `chrunner`/`searchclient` prove *read*
|
||||
@@ -360,7 +367,8 @@ terms:
|
||||
| Tantivy per-tenant index routing (`search/src/registry.rs`) | **Enforced, verified live** — real Tantivy indices, real cross-tenant probe, all passing |
|
||||
| Tantivy tenant_id resolution (`enterprise/internal/searchclient`) | **Enforced, verified live** — real gRPC wire-level test |
|
||||
| Ingest tenant-awareness (ClickHouse and Tantivy both) | **Not implemented, undesigned** — every ingested record lands in the single shared database/index regardless of tenant |
|
||||
| Deployment actually routing traffic to `enterprise-api` | **Not implemented** — no Helm service, no default wiring; now the largest gap in the isolation story |
|
||||
| Deployment actually routing traffic to `enterprise-api` (Helm) | **Enforced** — `api`/`enterprise-api` are mutually exclusive, same flag as RBAC/audit/SSO |
|
||||
| Deployment actually routing traffic to `enterprise-api` (docker-compose) | **Not implemented** — `docker-compose.yml` runs plain `api` unconditionally |
|
||||
| Human SSO login — OIDC | **Built, verified with a real fake IdP** (not yet tried against a real external IdP) |
|
||||
| Human SSO login — SAML | **Not implemented** |
|
||||
| Multi-tenant-membership login (tenant picker) | **Not implemented** — refused with a clear error, not guessed |
|
||||
|
||||
Reference in New Issue
Block a user