Merge pull request #19 from Coffey-Labs/docs/tenancy-off-the-roadmap

Take multi-tenancy off the roadmap
This commit is contained in:
Coffey Labs
2026-09-04 19:01:57 -07:00
committed by GitHub
5 changed files with 75 additions and 27 deletions
+2 -2
View File
@@ -1,6 +1,6 @@
# Contributing to Cairn OBS # Contributing to Cairn OBS
Thanks for your interest in contributing to **Cairn OBS**open-core, Kubernetes-native log aggregation and observability. Contributions of all kinds are welcome: bug reports, feature requests, code, documentation, and testing. Thanks for your interest in contributing to **Cairn OBS**self-hosted, Kubernetes-native log aggregation and observability. Contributions of all kinds are welcome: bug reports, feature requests, code, documentation, and testing.
## Code of Conduct ## Code of Conduct
@@ -77,7 +77,7 @@ Three workflows run on a pull request:
docker compose up docker compose up
``` ```
The web UI is on <http://localhost:3000> and the API on `:8080`; alerting is on `:8081` and enterprise auth on `:8082`. The agent connects to ingest over mTLS gRPC on `:4317`. `search` publishes no host port — it's reachable only on the compose network. The web UI is on <http://localhost:3000> and the API on `:8080`; alerting is on `:8081` and enterprise auth on `:8082`. The agent connects to ingest over mTLS gRPC on `:4317`. `search` publishes no host port — it's reachable only on the compose network.
3. `COMPOSE_PROFILES` in `.env` selects the query-serving binary: `single-tenant` (default) or `enterprise` for the multi-tenant path. They're mutually exclusive, the same choice Helm's `enterprise.enabled` flag makes for a real cluster. Override per invocation with `COMPOSE_PROFILES=enterprise docker compose up`. 3. `COMPOSE_PROFILES` in `.env` selects the query-serving binary. Report against `single-tenant` unless the bug is in `enterprise/` itself, which is off the roadmap (see the README): `single-tenant` (default) or `enterprise` for the multi-tenant path. They're mutually exclusive, the same choice Helm's `enterprise.enabled` flag makes for a real cluster. Override per invocation with `COMPOSE_PROFILES=enterprise docker compose up`.
4. Exercise the path you changed end to end — for ingest or query work that means getting a real log line in and querying it back, not just a passing unit test. 4. Exercise the path you changed end to end — for ingest or query work that means getting a real log line in and querying it back, not just a passing unit test.
## Review Process ## Review Process
+3 -3
View File
@@ -1,9 +1,9 @@
# Project: Cairn OBS — Distributed Log Aggregation & Observability Platform # Project: Cairn OBS — Distributed Log Aggregation & Observability Platform
## Mission ## Mission
Build an open-core, Kubernetes-native centralized logging platform that rivals Build a self-hosted, Kubernetes-native centralized logging platform that
Splunk on features but wins on cost-per-GB, modern language stack, and honest rivals Splunk on features but wins on cost-per-GB and a modern language
multi-tenant RBAC. Full architecture spec is in `/docs/architecture.md` — read stack. Full architecture spec is in `/docs/architecture.md` — read
it before touching any component. Do not deviate from the storage/query split it before touching any component. Do not deviate from the storage/query split
described there without flagging it to me first. described there without flagging it to me first.
+50 -19
View File
@@ -8,9 +8,9 @@
</p> </p>
<p align="center"> <p align="center">
Open-core, Kubernetes-native log aggregation and observability.<br> Self-hosted, Kubernetes-native log aggregation and observability.<br>
Built to match Splunk on capability while winning on cost-per-GB,<br> Built to match Splunk on capability while winning on cost-per-GB,<br>
with honest multi-tenant RBAC and a modern language stack.<br> on a modern language stack.<br>
Positioned against Cribl too — see <a href="docs/positioning.md">positioning</a> Positioned against Cribl too — see <a href="docs/positioning.md">positioning</a>
for why that is a different claim, and what it means we still have to build. for why that is a different claim, and what it means we still have to build.
</p> </p>
@@ -80,7 +80,7 @@ web/ SvelteKit frontend
cli/ cairnobsctl cli/ cairnobsctl
proto/ gRPC service definitions proto/ gRPC service definitions
terraform/ Terraform provider terraform/ Terraform provider
enterprise/ SSO, multi-tenancy, per-tenant provisioning enterprise/ SSO, multi-tenancy, per-tenant provisioning (off the roadmap)
deploy/ Helm charts and Kubernetes operator deploy/ Helm charts and Kubernetes operator
docs/ Architecture, design docs, per-phase runbooks docs/ Architecture, design docs, per-phase runbooks
hack/ Development scripts hack/ Development scripts
@@ -102,14 +102,11 @@ alerting is on `:8081` and enterprise auth on `:8082`. The agent connects to
ingest over mTLS gRPC on `:4317`. `search` is reachable only on the compose ingest over mTLS gRPC on `:4317`. `search` is reachable only on the compose
network — it publishes no host port. network — it publishes no host port.
`COMPOSE_PROFILES` in `.env` selects the query-serving binary`single-tenant` `COMPOSE_PROFILES` in `.env` selects the query-serving binary. Leave it at
(default) or `enterprise` for the multi-tenant path. They're mutually `single-tenant`: that is the supported shape, and the one every runbook,
exclusive, the same choice Helm's `enterprise.enabled` flag makes for a real the demo and the docs assume. `enterprise` swaps in the multi-tenant
cluster. Override per invocation: binaries and still works, but it is no longer a direction this project is
taking — see [Multi-tenancy is not the plan](#multi-tenancy-is-not-the-plan).
```sh
COMPOSE_PROFILES=enterprise docker compose up
```
### Signing in ### Signing in
@@ -187,25 +184,59 @@ privileges. Details in [`agent/README.md`](agent/README.md).
Terraform provider coverage is partial by necessity: dashboards and panels have Terraform provider coverage is partial by necessity: dashboards and panels have
full CRUD, while alert rules and notification targets are create/destroy only, full CRUD, while alert rules and notification targets are create/destroy only,
because `alerting` exposes no `PUT /rules/{id}` or `PUT /targets/{id}` to because `alerting` exposes no `PUT /rules/{id}` or `PUT /targets/{id}` to
update against. Tenant and RBAC resources are disclosed future work — update against. Tenant and RBAC resources are not planned at all, for the
[`terraform/README.md`](terraform/README.md) accounts for exactly what exists. reason below — [`terraform/README.md`](terraform/README.md) accounts for
exactly what exists.
### Multi-tenancy is not the plan
Phase 4 works, and it is not the direction. Cairn OBS is a self-hosted
system, and the way to separate two environments is to run two
installations, not two tenants inside one.
Tenancy turned out to be the wrong boundary for that job on three counts,
each of which is visible in this repository rather than hypothetical:
- **Shared fate on ingest.** `chwriter.WriteBatch` is all-or-nothing across
tenants by design — one tenant's failed write refuses the whole batch and
stops offset progress for every other tenant. A mistake in one environment
stalls ingestion in the rest.
- **A shared superuser.** `CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT` makes the
admin credential a ClickHouse superuser over every tenant's data, as
`docker-compose.yml`'s own comment spells out. Two environments behind one
credential are not separated in any sense that matters.
- **One upgrade cadence.** A single binary and a single set of migrations
move every tenant together, which is the opposite of what separate
environments are for.
A full installation idles at roughly 1.3 GB — ClickHouse and Redpanda are
almost all of it, and every Go service is 1121 MB — so running one per
environment is cheap enough that sharing a control plane buys nothing. That
is the same argument `positioning.md` already makes about storage, applied
to compute.
So `enterprise/` stays in the tree, AGPLv3 and working, and is not on the
roadmap. It remains the answer for serving other people's data — a question
this project is not asking. Nothing in Phases 8-11 depends on it.
### What would close the gap to production-ready ### What would close the gap to production-ready
Named here so the list above reads as a plan rather than an apology, and so Named here so the list above reads as a plan rather than an apology, and so
anyone evaluating this knows what they would be waiting for: anyone evaluating this knows what they would be waiting for:
1. **A second identity provider.** SSO works against Auth0 for both OIDC and 1. **A real cluster.** The Helm chart has been installed against a local
SAML; one vendor is an implementation, two is a standard.
2. **A real cluster.** The Helm chart has been installed against a local
`kind` cluster, which proves the manifests and nothing about scheduling, `kind` cluster, which proves the manifests and nothing about scheduling,
storage classes or node failure. storage classes or node failure.
3. **The Windows agent on Windows.** The code is written and reviewed; no 2. **The Windows agent on Windows.** The code is written and reviewed; no
Windows toolchain has ever compiled it, let alone run it. Windows toolchain has ever compiled it, let alone run it.
4. **Sustained load.** Every phase was verified functionally. Nothing here has 3. **Sustained load.** Every phase was verified functionally. Nothing here has
been run at volume for long enough to find the failures that only show up been run at volume for long enough to find the failures that only show up
after a week. after a week.
5. **Somebody else's data.** Every deployment so far has been ours. 4. **Somebody else's data.** Every deployment so far has been ours.
A second identity provider used to head this list. It has come off it: SSO
belongs to `enterprise/`, and with multi-tenancy off the roadmap a second
IdP is no longer something a self-hosted deployment is waiting on.
None of that is research; it is time on real infrastructure. It is also None of that is research; it is time on real infrastructure. It is also
exactly the list a pilot deployment would work through, which is the honest exactly the list a pilot deployment would work through, which is the honest
+16 -2
View File
@@ -30,7 +30,20 @@ Phases 8-11 are a second axis rather than a continuation of the first:
0-7 built the destination, and those four build the road to it. The 0-7 built the destination, and those four build the road to it. The
argument for taking that on, including the part where cheap storage argument for taking that on, including the part where cheap storage
removes the usual reason to buy a pipeline at all, is removes the usual reason to buy a pipeline at all, is
[`positioning.md`](positioning.md). Nothing in them is started. [`positioning.md`](positioning.md). Nothing in them is started. None of
them depends on Phase 4.
**Decision, 2026-09-05: tenancy is for other people's data; installations
are for environments.** Phase 4 stays shipped and stays in the tree, and
comes off the roadmap. Separating two environments means running two
installations, not two tenants in one, for three reasons this repository
demonstrates rather than assumes: `chwriter.WriteBatch` is all-or-nothing
across tenants, so one tenant's failure stalls offset progress for all of
them; `CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT` puts every tenant's data
behind one superuser credential; and one binary with one set of
migrations moves every tenant together. A whole installation idles at
about 1.3 GB, so the sharing buys nothing worth those three. See the
README's "Multi-tenancy is not the plan".
**Known verification gaps**, carried forward rather than buried: **Known verification gaps**, carried forward rather than buried:
@@ -48,7 +61,8 @@ removes the usual reason to buy a pipeline at all, is
`COMPOSE_PROFILES=single-tenant` — so it exercises the OSS path and says `COMPOSE_PROFILES=single-tenant` — so it exercises the OSS path and says
nothing about RBAC, tenant isolation or per-tenant ClickHouse. Do not read nothing about RBAC, tenant isolation or per-tenant ClickHouse. Do not read
a healthy demo as evidence for Phase 4. a healthy demo as evidence for Phase 4.
- **Phase 4's SSO has been tried against one IdP, not two.** OIDC and SAML - **Phase 4's SSO has been tried against one IdP, not two.** Recorded as
fact rather than as pending work — see the decision above. OIDC and SAML
were both verified end to end against a real Auth0 developer tenant, were both verified end to end against a real Auth0 developer tenant,
browser round trips included. A second, independent IdP has never been browser round trips included. A second, independent IdP has never been
tried, and no production-grade cluster has run this — the Kubernetes tried, and no production-grade cluster has run this — the Kubernetes
+4 -1
View File
@@ -188,7 +188,10 @@ Terraform state; see "`cairnobs_notification_target`'s `secret`
attribute" above. attribute" above.
**Also not built, all real and disclosed, not attempted here:** **Also not built, all real and disclosed, not attempted here:**
- Tenant/RBAC resources (`enterprise-auth`'s tenant/membership/grant - Tenant/RBAC resources -- **not planned**, since multi-tenancy came off
the roadmap (see the README's "Multi-tenancy is not the plan"). The
original accounting stands for anyone who picks it up anyway:
(`enterprise-auth`'s tenant/membership/grant
surface) -- meaningfully different auth model (offline operator flags surface) -- meaningfully different auth model (offline operator flags
today, not a stable REST API a provider could safely drive today, not a stable REST API a provider could safely drive
idempotently -- see `/enterprise/README.md`'s "Bootstrapping a tenant" idempotently -- see `/enterprise/README.md`'s "Bootstrapping a tenant"