Commit Graph
8 Commits
Author SHA1 Message Date
jcoffey-dev 7cc2fd8c78 Draft the Phase 8 processing design
Starts from the distribution channel and the safety invariant it
protects, and derives the language from them, rather than designing a
rule language and asking later how to ship it.

Four decisions proposed. A rule is a matcher plus ordered typed actions,
with no expressions and nothing resembling eval -- less expressive than
Cribl on purpose, and the only shape that can be pushed to ten thousand
hosts and audited by reading it. Total evaluation is the primary safety
guarantee, with apply-then-verify as a backstop. One spec with two
implementations means a language-neutral conformance suite is the
specification and should be built first, not last. Distribution reuses
DesiredOverride, following extra_file_paths as the precedent for a
repeated field.

It also corrects something #21 got wrong. That change said a fatal rule
would strand an agent the way a corrupted ingest endpoint does. Reading
apply_override's actual semantics, overrides live only in the running
process's memory and are never written to disk, so a restarted agent
boots clean and re-syncs -- a fatal rule set crash-loops rather than
strands, and the agent keeps checking in, so it stays correctable. That
is a much better failure mode, and it was acquired by accident: the
"don't persist" choice was made to avoid filesystem writes on read-only
images, not for safety. This design promotes it to a constraint, since
persisting overrides later would silently convert every crash-loop into
a strand. positioning.md is corrected to match rather than left
disagreeing.

Five open questions are left open rather than answered to look decisive,
the sharpest being that there is no staged rollout today: an edit
reaches every matching agent at once, which for executable rules is the
difference between breaking one host and breaking all of them.

The v1 acceptance test is real data, not a fixture: two processes on the
maintainer's own workstation account for 308 of 325 journal entries in
five minutes, and a suppress rule should remove about 60% of that host's
volume.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 20:09:00 -07:00
jcoffey-dev 18a55ccd5a Realign the roadmap with what is already built
Three of the roadmap's claims were contradicted by the repository
itself.

Fleet management was Phase 11, "Planned", and positioning.md said config
still flowed to the agent from the host rather than from the platform.
agent-management-design.md has recorded the opposite for some time:
punch list complete, verified live, with central authoring, versioning,
rollout on the next check-in, observation, and a restart command a real
agent picks up and acts on. There is no Phase 11 now. Its remainder is
either already named there -- stop/uninstall, per-host multi-row
alerting, a rule-per-host generator -- or belongs to Phase 8, since
distributing rules is the one genuinely new thing the mechanism has to
carry, and a rule language nobody can push to a fleet is not worth
having.

That inverts the old ordering argument, which put fleet last on the
grounds that it manages configuration the earlier phases define. Sound
reasoning; the world went the other way and built the mechanism first.
Recorded rather than quietly dropped, because the instinct behind it is
a good one that happened not to apply.

Retention was listed as a question Phase 10 would finally have to
answer. Half of it is answered: api/logretention serves operator-driven
preview and delete with an owner-only per-agent floor. What is missing
is an automatic TTL, so Phase 10 owns tiering and automatic TTL rather
than retention from nothing.

The rule-language recommendation is now settled rather than proposed,
and not on its own authority: DesiredOverride is already a closed typed
shape that cannot carry arbitrary code, so no channel exists that would
deliver JavaScript to an agent even if the language argument had gone
the other way.

That surfaced a requirement nothing had written down. Agent management
rests on an invariant it states outright -- every editable field
degrades behaviour without cutting off the agent's ability to receive
the next correction. Processing rules break it: a rule that panics or
loops strands the agent exactly the way a corrupted ingest endpoint
would, across every host it reached first. Phase 8 now owes either total
evaluation or apply-then-verify with rollback, chosen deliberately
rather than discovered mid-rollout.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 19:16:01 -07:00
jcoffey-dev 0ee2e9183b Take multi-tenancy off the roadmap
Cairn OBS is self-hosted, and the way to separate two environments is to
run two installations rather than two tenants inside one. Tenancy is the
wrong boundary for that, on three counts this repository demonstrates
rather than assumes: chwriter.WriteBatch is all-or-nothing across
tenants, so one tenant's failure stalls offset progress for every other;
CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT puts every tenant's data behind a
single superuser credential, as docker-compose.yml's own comment says;
and one binary with one set of migrations moves every tenant together,
which is the opposite of what separate environments are for. A whole
installation idles at about 1.3 GB, so the sharing buys nothing.

The project led with multi-tenant RBAC in the README banner and in
PROJECT-SPEC's goal statement. Both now say what it is instead:
self-hosted. "Open-core" goes with them -- it was already inaccurate,
since CONTRIBUTING states there is no feature gate and no paid tier, and
with enterprise/ off the roadmap there will not be one.

A second identity provider comes off the list of things standing between
this and production-ready. SSO belongs to enterprise/, and a self-hosted
deployment is not waiting on it. Terraform's tenant/RBAC resources move
from "disclosed future work" to not planned.

Nothing is scrubbed from the record. Phase 4 stays shipped, its runbook
stays, and its known gaps stay stated -- rewriting that history would
contradict the candour the Status section is built on. enterprise/ stays
in the tree, AGPLv3 and working, as the answer to a question this
project is not asking.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 18:31:23 -07:00
jcoffey-dev d49ddb943e Add the AI axis: plain English as an option, analysis as the end state
Cost is the argument against Splunk and control is the argument against
Cribl. AI is the third, and the difference there is not a feature
comparison, it is where the model runs.

Plain-English querying shipped in Phase 7 and stays an option rather than
a replacement for writing a query: every generated query compiles through
the same IR and executor as a hand-written one, with the same tenant
scoping, cost guardrails and audit logging. The model suggests, it does
not get a private path to the data.

AI-assisted analysis and explanation is the end state and is not built.
Authoring answers "how do I ask this"; the valuable question is "what
does this mean" -- what changed in a result set, why an alert fired and
what preceded it, summarising an incident from the records around it.
Recorded on the status page as an end-state goal rather than a numbered
phase, because it is a property the product keeps rather than a thing to
finish and tick off.

Local is the non-negotiable part, and it is worth stating as position
rather than as a bullet: the default runs qwen2.5-coder through Ollama on
the customer's own hardware, Apache-2.0 weights chosen so Phase 6's
licence work survives contact with the model, and the cloud adapter is
opt-in and off by default. Logs are the most sensitive unstructured data
most organisations hold -- credentials in stack traces, customer
identifiers, internal topology -- so an assistant that reads them is
either running where the data already is, or it is a data-egress decision
wearing a helpful interface.

The constraint it imposes is stated too, because it bounds what can be
promised: a 7B model on a customer's hardware will not match a frontier
model, and the honest claim is not that it is as clever but that it is
good enough at a bounded task and runs somewhere you control. Analysis
features have to be designed to that budget rather than assuming an API
is one call away.
2026-09-04 16:35:44 -07:00
jcoffey-dev e57486add5 Position against Cribl as well as Splunk, and say what that costs us
Splunk and Cribl are not the same competitor and the claim is not the
same claim twice. Splunk is the destination and Cairn OBS replaces it,
which is what phases 0-7 were for. Cribl is the road: routing, reduction,
enrichment, redaction and replay on the way to wherever data is going.
Cairn OBS is already a road in shape -- agent, Redpanda, ingest -- and
exposes none of a pipeline's controls. The agent cannot filter, drop,
sample, mask or re-route anything; ingest normalises a schema and writes
it; there is exactly one destination and it is us.

Reconciling the two turns up something a cost-led project has to face
rather than paper over: most people buy Cribl because Splunk is expensive
per gigabyte, so being genuinely cheap per gigabyte removes the main
reason to buy Cribl in front of us. That makes the strongest pitch "one
system where there were two" rather than "we are also a pipeline vendor"
-- but that pitch only survives a buyer if we also do the four things
people buy a pipeline for that are not about spend: routing to several
destinations, redacting before data leaves the network, archive and
replay, and not being locked to one analytics vendor. Those are about
control, which is better ground anyway: cost advantages get matched and
architectural ones do not.

The consequence is uncomfortable and is written down as a decision rather
than left to be discovered: competing with Cribl means being able to send
data to S3, Splunk HEC, Elastic, OTLP and Kafka -- building features whose
purpose is to help data leave this platform. A project that refuses
lock-in in its licence and then builds it into its egress would be lying
about itself.

Four phases follow, ordered so each pays for itself: processing (8),
routing (9), archive and replay (10), fleet (11). Processing without
routing still shrinks what is stored; routing without processing forwards
everything and helps nobody.

One design decision is called out now because it collides with a
non-negotiable constraint. Cribl's rule language is JavaScript, and
embedding a JS engine in a statically-linked musl agent would end "no
glibc runtime deps" as a claim. The recommendation is a declarative rule
DSL -- matchers and typed actions, no arbitrary code -- deliberately less
expressive, small enough to audit and safe to push to ten thousand hosts.

The retention/TTL question in architecture.md is no longer deferrable and
now says so: Phase 10 asks it from the other side.
2026-09-04 16:29:29 -07:00
Coffey Labs 252c207ccf Say that Phase 4 shipped, and that the environment proving it is gone (#7)
The status file and the README both still said Phase 4 was not shipped
because the environment had lost Docker and database access partway
through, and that only the audit-logging guarantees had been confirmed
against a live database. That stopped being true some time ago.
phase-4-runbook.md records the opposite in detail: Docker access came
back, a real docker-compose stack ran with real ClickHouse and Postgres
and two provisioned tenants, a local kind cluster ran the Helm chart end
to end, and both SSO protocols were verified against a real Auth0 tenant
with full browser round trips. Eight real bugs came out of that, six from
compose and two from the chart's first real install -- none of them
findable without the infrastructure.

PROJECT-SPEC.md sends readers to status.md and tells them to read it
before assuming a capability works end to end, so the one file that is
meant to be authoritative was the one understating the project by the
widest margin.

Correcting it matters more now than it would have last week, because the
evidence cannot be regenerated: proto.cairnobs.org and the VPS under it
were retired on 2026-09-04, taking the mTLS CA, the server certificate
and six enrolled agents with them. The runbooks are what is left.

Three gaps are now stated rather than implied:

The prototype is gone, so none of this can be re-run today without
building one. The DNS was kept for that; the certificates deliberately
were not.

demo.cairnobs.org is live and is not evidence for Phase 4. It runs
COMPOSE_PROFILES=single-tenant, so it exercises the OSS path and says
nothing about RBAC, tenant isolation or per-tenant ClickHouse. A healthy
demo proving multi-tenancy is exactly the wrong inference to leave
available.

SSO has been tried against one IdP and one local kind cluster, not two
IdPs and not a production-grade cluster.

The Terraform entry now names its cause instead of pointing at another
file: alerting exposes no PUT for rules or targets, and neither
rulestore.Store nor notifystore.Store has an Update method to wire one
to, so Terraform destroys and recreates -- which resets alert_state and
delivery-log continuity.
2026-09-04 15:29:46 -07:00
jcoffey-dev f756a9d4f6 Rename the project spec and update every reference to it
The charter file carried a tool-specific name while being the repository's own
document: mission, non-negotiable constraints, the pinned stack, repo
conventions and phase status, cited as authority by thirty files across the
agent, api, deploy, docs, search and terraform trees.

PROJECT-SPEC.md says what it is. All 42 references are updated in the same
commit, including the relative link in docs/status.md, so nothing points at a
filename that no longer exists.
2026-08-28 15:57:12 -07:00
jcoffey-dev 914c0af467 docs: split project status out of CLAUDE.md
CLAUDE.md was doing two jobs: durable repo conventions, and a ~500-line
phase-by-phase status narrative that duplicates the per-phase runbooks
and goes stale the moment a phase ships.

Keep mission, constraints, pinned stack, conventions, and "when in doubt"
in CLAUDE.md (572 -> 78 lines). Move the phase record verbatim to
docs/status.md, prefaced with a summary table and the known verification
gaps. Content is byte-identical; nothing was reworded or dropped.

Link both directions, and point the README's status section at the new
file.
2026-08-22 18:28:43 -07:00