diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 4b27fa5..8a94f07 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -1,6 +1,6 @@ # Contributing to Cairn OBS -Thanks for your interest in contributing to **Cairn OBS** — open-core, Kubernetes-native log aggregation and observability. Contributions of all kinds are welcome: bug reports, feature requests, code, documentation, and testing. +Thanks for your interest in contributing to **Cairn OBS** — self-hosted, Kubernetes-native log aggregation and observability. Contributions of all kinds are welcome: bug reports, feature requests, code, documentation, and testing. ## Code of Conduct @@ -77,7 +77,7 @@ Three workflows run on a pull request: docker compose up ``` The web UI is on and the API on `:8080`; alerting is on `:8081` and enterprise auth on `:8082`. The agent connects to ingest over mTLS gRPC on `:4317`. `search` publishes no host port — it's reachable only on the compose network. -3. `COMPOSE_PROFILES` in `.env` selects the query-serving binary: `single-tenant` (default) or `enterprise` for the multi-tenant path. They're mutually exclusive, the same choice Helm's `enterprise.enabled` flag makes for a real cluster. Override per invocation with `COMPOSE_PROFILES=enterprise docker compose up`. +3. `COMPOSE_PROFILES` in `.env` selects the query-serving binary. Report against `single-tenant` unless the bug is in `enterprise/` itself, which is off the roadmap (see the README): `single-tenant` (default) or `enterprise` for the multi-tenant path. They're mutually exclusive, the same choice Helm's `enterprise.enabled` flag makes for a real cluster. Override per invocation with `COMPOSE_PROFILES=enterprise docker compose up`. 4. Exercise the path you changed end to end — for ingest or query work that means getting a real log line in and querying it back, not just a passing unit test. ## Review Process diff --git a/PROJECT-SPEC.md b/PROJECT-SPEC.md index 3d86a60..19075c1 100644 --- a/PROJECT-SPEC.md +++ b/PROJECT-SPEC.md @@ -1,9 +1,9 @@ # Project: Cairn OBS — Distributed Log Aggregation & Observability Platform ## Mission -Build an open-core, Kubernetes-native centralized logging platform that rivals -Splunk on features but wins on cost-per-GB, modern language stack, and honest -multi-tenant RBAC. Full architecture spec is in `/docs/architecture.md` — read +Build a self-hosted, Kubernetes-native centralized logging platform that +rivals Splunk on features but wins on cost-per-GB and a modern language +stack. Full architecture spec is in `/docs/architecture.md` — read it before touching any component. Do not deviate from the storage/query split described there without flagging it to me first. diff --git a/README.md b/README.md index eb2af44..7c8c252 100644 --- a/README.md +++ b/README.md @@ -8,9 +8,9 @@

- Open-core, Kubernetes-native log aggregation and observability.
+ Self-hosted, Kubernetes-native log aggregation and observability.
Built to match Splunk on capability while winning on cost-per-GB,
- with honest multi-tenant RBAC and a modern language stack.
+ on a modern language stack.
Positioned against Cribl too — see positioning for why that is a different claim, and what it means we still have to build.

@@ -80,7 +80,7 @@ web/ SvelteKit frontend cli/ cairnobsctl proto/ gRPC service definitions terraform/ Terraform provider -enterprise/ SSO, multi-tenancy, per-tenant provisioning +enterprise/ SSO, multi-tenancy, per-tenant provisioning (off the roadmap) deploy/ Helm charts and Kubernetes operator docs/ Architecture, design docs, per-phase runbooks hack/ Development scripts @@ -102,14 +102,11 @@ alerting is on `:8081` and enterprise auth on `:8082`. The agent connects to ingest over mTLS gRPC on `:4317`. `search` is reachable only on the compose network — it publishes no host port. -`COMPOSE_PROFILES` in `.env` selects the query-serving binary — `single-tenant` -(default) or `enterprise` for the multi-tenant path. They're mutually -exclusive, the same choice Helm's `enterprise.enabled` flag makes for a real -cluster. Override per invocation: - -```sh -COMPOSE_PROFILES=enterprise docker compose up -``` +`COMPOSE_PROFILES` in `.env` selects the query-serving binary. Leave it at +`single-tenant`: that is the supported shape, and the one every runbook, +the demo and the docs assume. `enterprise` swaps in the multi-tenant +binaries and still works, but it is no longer a direction this project is +taking — see [Multi-tenancy is not the plan](#multi-tenancy-is-not-the-plan). ### Signing in @@ -187,25 +184,59 @@ privileges. Details in [`agent/README.md`](agent/README.md). Terraform provider coverage is partial by necessity: dashboards and panels have full CRUD, while alert rules and notification targets are create/destroy only, because `alerting` exposes no `PUT /rules/{id}` or `PUT /targets/{id}` to -update against. Tenant and RBAC resources are disclosed future work — -[`terraform/README.md`](terraform/README.md) accounts for exactly what exists. +update against. Tenant and RBAC resources are not planned at all, for the +reason below — [`terraform/README.md`](terraform/README.md) accounts for +exactly what exists. + +### Multi-tenancy is not the plan + +Phase 4 works, and it is not the direction. Cairn OBS is a self-hosted +system, and the way to separate two environments is to run two +installations, not two tenants inside one. + +Tenancy turned out to be the wrong boundary for that job on three counts, +each of which is visible in this repository rather than hypothetical: + +- **Shared fate on ingest.** `chwriter.WriteBatch` is all-or-nothing across + tenants by design — one tenant's failed write refuses the whole batch and + stops offset progress for every other tenant. A mistake in one environment + stalls ingestion in the rest. +- **A shared superuser.** `CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT` makes the + admin credential a ClickHouse superuser over every tenant's data, as + `docker-compose.yml`'s own comment spells out. Two environments behind one + credential are not separated in any sense that matters. +- **One upgrade cadence.** A single binary and a single set of migrations + move every tenant together, which is the opposite of what separate + environments are for. + +A full installation idles at roughly 1.3 GB — ClickHouse and Redpanda are +almost all of it, and every Go service is 11–21 MB — so running one per +environment is cheap enough that sharing a control plane buys nothing. That +is the same argument `positioning.md` already makes about storage, applied +to compute. + +So `enterprise/` stays in the tree, AGPLv3 and working, and is not on the +roadmap. It remains the answer for serving other people's data — a question +this project is not asking. Nothing in Phases 8-11 depends on it. ### What would close the gap to production-ready Named here so the list above reads as a plan rather than an apology, and so anyone evaluating this knows what they would be waiting for: -1. **A second identity provider.** SSO works against Auth0 for both OIDC and - SAML; one vendor is an implementation, two is a standard. -2. **A real cluster.** The Helm chart has been installed against a local +1. **A real cluster.** The Helm chart has been installed against a local `kind` cluster, which proves the manifests and nothing about scheduling, storage classes or node failure. -3. **The Windows agent on Windows.** The code is written and reviewed; no +2. **The Windows agent on Windows.** The code is written and reviewed; no Windows toolchain has ever compiled it, let alone run it. -4. **Sustained load.** Every phase was verified functionally. Nothing here has +3. **Sustained load.** Every phase was verified functionally. Nothing here has been run at volume for long enough to find the failures that only show up after a week. -5. **Somebody else's data.** Every deployment so far has been ours. +4. **Somebody else's data.** Every deployment so far has been ours. + +A second identity provider used to head this list. It has come off it: SSO +belongs to `enterprise/`, and with multi-tenancy off the roadmap a second +IdP is no longer something a self-hosted deployment is waiting on. None of that is research; it is time on real infrastructure. It is also exactly the list a pilot deployment would work through, which is the honest diff --git a/docs/status.md b/docs/status.md index fa545a7..41023ba 100644 --- a/docs/status.md +++ b/docs/status.md @@ -30,7 +30,20 @@ Phases 8-11 are a second axis rather than a continuation of the first: 0-7 built the destination, and those four build the road to it. The argument for taking that on, including the part where cheap storage removes the usual reason to buy a pipeline at all, is -[`positioning.md`](positioning.md). Nothing in them is started. +[`positioning.md`](positioning.md). Nothing in them is started. None of +them depends on Phase 4. + +**Decision, 2026-09-05: tenancy is for other people's data; installations +are for environments.** Phase 4 stays shipped and stays in the tree, and +comes off the roadmap. Separating two environments means running two +installations, not two tenants in one, for three reasons this repository +demonstrates rather than assumes: `chwriter.WriteBatch` is all-or-nothing +across tenants, so one tenant's failure stalls offset progress for all of +them; `CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT` puts every tenant's data +behind one superuser credential; and one binary with one set of +migrations moves every tenant together. A whole installation idles at +about 1.3 GB, so the sharing buys nothing worth those three. See the +README's "Multi-tenancy is not the plan". **Known verification gaps**, carried forward rather than buried: @@ -48,7 +61,8 @@ removes the usual reason to buy a pipeline at all, is `COMPOSE_PROFILES=single-tenant` — so it exercises the OSS path and says nothing about RBAC, tenant isolation or per-tenant ClickHouse. Do not read a healthy demo as evidence for Phase 4. -- **Phase 4's SSO has been tried against one IdP, not two.** OIDC and SAML +- **Phase 4's SSO has been tried against one IdP, not two.** Recorded as + fact rather than as pending work — see the decision above. OIDC and SAML were both verified end to end against a real Auth0 developer tenant, browser round trips included. A second, independent IdP has never been tried, and no production-grade cluster has run this — the Kubernetes diff --git a/terraform/README.md b/terraform/README.md index afadb59..c880b75 100644 --- a/terraform/README.md +++ b/terraform/README.md @@ -188,7 +188,10 @@ Terraform state; see "`cairnobs_notification_target`'s `secret` attribute" above. **Also not built, all real and disclosed, not attempted here:** -- Tenant/RBAC resources (`enterprise-auth`'s tenant/membership/grant +- Tenant/RBAC resources -- **not planned**, since multi-tenancy came off + the roadmap (see the README's "Multi-tenancy is not the plan"). The + original accounting stands for anyone who picks it up anyway: + (`enterprise-auth`'s tenant/membership/grant surface) -- meaningfully different auth model (offline operator flags today, not a stable REST API a provider could safely drive idempotently -- see `/enterprise/README.md`'s "Bootstrapping a tenant"