Splunk and Cribl are not the same competitor and the claim is not the same claim twice. Splunk is the destination and Cairn OBS replaces it, which is what phases 0-7 were for. Cribl is the road: routing, reduction, enrichment, redaction and replay on the way to wherever data is going. Cairn OBS is already a road in shape -- agent, Redpanda, ingest -- and exposes none of a pipeline's controls. The agent cannot filter, drop, sample, mask or re-route anything; ingest normalises a schema and writes it; there is exactly one destination and it is us. Reconciling the two turns up something a cost-led project has to face rather than paper over: most people buy Cribl because Splunk is expensive per gigabyte, so being genuinely cheap per gigabyte removes the main reason to buy Cribl in front of us. That makes the strongest pitch "one system where there were two" rather than "we are also a pipeline vendor" -- but that pitch only survives a buyer if we also do the four things people buy a pipeline for that are not about spend: routing to several destinations, redacting before data leaves the network, archive and replay, and not being locked to one analytics vendor. Those are about control, which is better ground anyway: cost advantages get matched and architectural ones do not. The consequence is uncomfortable and is written down as a decision rather than left to be discovered: competing with Cribl means being able to send data to S3, Splunk HEC, Elastic, OTLP and Kafka -- building features whose purpose is to help data leave this platform. A project that refuses lock-in in its licence and then builds it into its egress would be lying about itself. Four phases follow, ordered so each pays for itself: processing (8), routing (9), archive and replay (10), fleet (11). Processing without routing still shrinks what is stored; routing without processing forwards everything and helps nobody. One design decision is called out now because it collides with a non-negotiable constraint. Cribl's rule language is JavaScript, and embedding a JS engine in a statically-linked musl agent would end "no glibc runtime deps" as a claim. The recommendation is a declarative rule DSL -- matchers and typed actions, no arbitrary code -- deliberately less expressive, small enough to audit and safe to push to ten thousand hosts. The retention/TTL question in architecture.md is no longer deferrable and now says so: Phase 10 asks it from the other side.
180 lines
7.4 KiB
Markdown
180 lines
7.4 KiB
Markdown
<p align="center">
|
|
<picture>
|
|
<source media="(prefers-color-scheme: dark)"
|
|
srcset="web/src/lib/assets/logo-horizontal-dark.svg">
|
|
<img src="web/src/lib/assets/logo-horizontal-light.svg"
|
|
alt="Cairn OBS" width="420">
|
|
</picture>
|
|
</p>
|
|
|
|
<p align="center">
|
|
Open-core, Kubernetes-native log aggregation and observability.<br>
|
|
Built to match Splunk on capability while winning on cost-per-GB,<br>
|
|
with honest multi-tenant RBAC and a modern language stack.<br>
|
|
Positioned against Cribl too — see <a href="docs/positioning.md">positioning</a>
|
|
for why that is a different claim, and what it means we still have to build.
|
|
</p>
|
|
|
|
<p align="center">
|
|
Licensed <strong>AGPLv3 in its entirety</strong> — including
|
|
<code>enterprise/</code>. See <a href="#licensing">Licensing</a>.
|
|
</p>
|
|
|
|
## What it does
|
|
|
|
Logs flow from a statically-linked Rust edge agent through Redpanda into a Go
|
|
ingest pipeline, landing in ClickHouse for analytics and Tantivy for full-text
|
|
search. One query language spans both stores, compiling to a single execution
|
|
plan:
|
|
|
|
```
|
|
service=api | where status>=500 | stats count by host | sort -count
|
|
message:"connection refused" | stats count by host
|
|
```
|
|
|
|
Raw ClickHouse SQL stays available as an escape hatch and compiles to the same
|
|
IR, so performance doesn't depend on which syntax you write.
|
|
|
|
On top of that sit dashboards, an alerting evaluator with threshold and
|
|
absence rules, a CLI (`cairnobsctl`), a Terraform provider, and AI-assisted
|
|
query authoring that runs against a self-hosted Ollama model by default — no
|
|
cloud dependency.
|
|
|
|
## Architecture
|
|
|
|
| Component | Stack |
|
|
|---|---|
|
|
| Edge agent | Rust, musl static target |
|
|
| Transport | Redpanda (Kafka API) |
|
|
| Ingest / parse | Go |
|
|
| Analytical store | ClickHouse |
|
|
| Full-text index | Tantivy (Rust) |
|
|
| Control plane / API | Go, gRPC + REST gateway |
|
|
| Control-plane metadata | PostgreSQL |
|
|
| Frontend | SvelteKit + TypeScript |
|
|
| Deployment | Kubernetes Operator (kubebuilder), Helm, docker-compose |
|
|
|
|
PostgreSQL is scoped strictly to control-plane config — dashboards, panels,
|
|
alert rules and state, notification targets, delivery log — because those need
|
|
row-level locking and transactional read-modify-write that ClickHouse's
|
|
MergeTree family doesn't provide. Log data itself never touches it.
|
|
|
|
Full spec: [`docs/architecture.md`](docs/architecture.md). Read it before
|
|
changing any component; the storage/query split in particular is deliberate.
|
|
|
|
## Repository layout
|
|
|
|
Monorepo, one top-level directory per component, each with its own `README.md`,
|
|
unit tests, and Dockerfile.
|
|
|
|
```
|
|
agent/ Rust edge agent (Linux + Windows)
|
|
transport/ Redpanda topics and schemas
|
|
ingest/ Go ingest and parse pipeline
|
|
storage/ ClickHouse schema and migrations
|
|
search/ Tantivy full-text index service
|
|
api/ Control plane, query compiler, RBAC
|
|
alerting/ Rule evaluator and notification delivery
|
|
metadata/ PostgreSQL schema and migrations
|
|
web/ SvelteKit frontend
|
|
cli/ cairnobsctl
|
|
proto/ gRPC service definitions
|
|
terraform/ Terraform provider
|
|
enterprise/ SSO, multi-tenancy, per-tenant provisioning
|
|
deploy/ Helm charts and Kubernetes operator
|
|
docs/ Architecture, design docs, per-phase runbooks
|
|
hack/ Development scripts
|
|
```
|
|
|
|
`enterprise/` stays a separate module that core never imports from. Since the
|
|
Phase 6 relicensing that boundary is architectural rather than legal — it keeps
|
|
core buildable and deployable standalone, and keeps tenant resolution
|
|
server-side.
|
|
|
|
## Running locally
|
|
|
|
```sh
|
|
docker compose up
|
|
```
|
|
|
|
The web UI comes up on <http://localhost:3000> and the API on `:8080`;
|
|
alerting is on `:8081` and enterprise auth on `:8082`. The agent connects to
|
|
ingest over mTLS gRPC on `:4317`. `search` is reachable only on the compose
|
|
network — it publishes no host port.
|
|
|
|
`COMPOSE_PROFILES` in `.env` selects the query-serving binary — `single-tenant`
|
|
(default) or `enterprise` for the multi-tenant path. They're mutually
|
|
exclusive, the same choice Helm's `enterprise.enabled` flag makes for a real
|
|
cluster. Override per invocation:
|
|
|
|
```sh
|
|
COMPOSE_PROFILES=enterprise docker compose up
|
|
```
|
|
|
|
Kubernetes deployment via the Helm chart in [`deploy/`](deploy/README.md).
|
|
|
|
## Status
|
|
|
|
Built in phases; each has a runbook in `docs/` recording how it was verified.
|
|
Full per-phase detail is in [`docs/status.md`](docs/status.md).
|
|
|
|
| Phase | Scope | Status |
|
|
|---|---|---|
|
|
| 0 | Agent → Redpanda → ingest → ClickHouse, queryable end-to-end | Shipped |
|
|
| 1 | Windows Event Log + journald, SQL and full-text paths | Shipped |
|
|
| 2 | Unified query language across both stores | Shipped |
|
|
| 3 | Dashboards, alert rules, notification delivery | Shipped |
|
|
| 4 | RBAC, tenant isolation, audit logging, per-tenant ClickHouse | Shipped |
|
|
| 5 | Frontend redesign and design system | Shipped |
|
|
| 6 | License compliance audit and remediation | Shipped |
|
|
| 7 | AI-assisted query authoring | Shipped |
|
|
|
|
**Phase 4 is shipped, and the environment that proved it is gone.** Every
|
|
control the phase defines was verified against real infrastructure at least
|
|
once — a docker-compose stack with real ClickHouse and Postgres, a local `kind`
|
|
cluster, and both SSO protocols against a real Auth0 tenant — finding eight
|
|
bugs that no amount of Docker-free testing could have caught. The prototype
|
|
VPS was retired on 2026-09-04, so that verification is a record rather than
|
|
something you can re-run: see
|
|
[`docs/phase-4-runbook.md`](docs/phase-4-runbook.md).
|
|
|
|
Two limits worth stating plainly. SSO has been tried against one IdP, not two,
|
|
and no production-grade cluster has run this. And `demo.cairnobs.org` is **not**
|
|
evidence for any of it — the demo runs the single-tenant profile, so it
|
|
exercises the OSS path and says nothing about RBAC or tenant isolation.
|
|
|
|
The Windows agent code (`EvtSubscribe`, ETW, service registration) has never
|
|
run on real Windows — no Windows toolchain existed in the build environment.
|
|
ETW additionally sits behind a feature flag, since it needs elevated
|
|
privileges. Details in [`agent/README.md`](agent/README.md).
|
|
|
|
Terraform provider coverage is partial by necessity: dashboards and panels have
|
|
full CRUD, while alert rules and notification targets are create/destroy only,
|
|
because `alerting` exposes no `PUT /rules/{id}` or `PUT /targets/{id}` to
|
|
update against. Tenant and RBAC resources are disclosed future work —
|
|
[`terraform/README.md`](terraform/README.md) accounts for exactly what exists.
|
|
|
|
## Contributing
|
|
|
|
- Conventional commits. Every change should be a logically complete,
|
|
independently revertible unit.
|
|
- Rust: `cargo clippy --all-targets -- -D warnings` must pass.
|
|
- Go: `go vet` and `golangci-lint`, no globals for shared state.
|
|
- Every UI action must map to a documented REST/gRPC call — no UI-only logic.
|
|
The CLI and Terraform provider are first-class, not afterthoughts.
|
|
- Prefer boring, well-understood dependencies. This is infrastructure software;
|
|
operators need to trust it.
|
|
|
|
## Licensing
|
|
|
|
Copyright (C) 2026 Coffey Labs.
|
|
|
|
AGPLv3, no exceptions — see [`LICENSE`](LICENSE). `enterprise/` was under a
|
|
commercial-license stub from Phase 4 through Phase 5; Phase 6 relicensed it to
|
|
match core. The full record and its business-model consequences are in
|
|
[`docs/compliance/license-audit-report.md`](docs/compliance/license-audit-report.md).
|
|
|
|
The default AI deployment uses `qwen2.5-coder` (Apache-2.0) via Ollama,
|
|
chosen specifically to keep that license purity intact. The cloud adapter is
|
|
pluggable, opt-in, and off by default.
|