ec4b860ba8b3054c20929eb57162a30986abff5c
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ec4b860ba8 |
Specify what aggregate_count emits
The last unanswered action, and the only one whose output is not the input with edits -- it emits a record that never existed, which is why it was deferred twice. It emits the window's first record unchanged, tagged with cairnobs.aggregated, cairnobs.count, and the observed window bounds. That follows the convention the agent already uses for heartbeat and host-metrics records rather than inventing a second synthetic-record mechanism, and keeping the first record intact means a reader sees a real example of what was collapsed instead of an invented summary. It tags even when the count is one. Emitting a bare record there would be tidier and would make cairnobs.count present only sometimes, so summing it silently breaks on quiet windows. window_last is the last record that actually contributed, never window_start + window_ms, because a window flushed early must not claim an end that never happened. Specifying it surfaced a problem the other nine actions do not have. Windows are measured on record time, so a window can only be closed by a later record arriving. suppress_duplicates never has anything pending; aggregate_count holds state, so a matching stream that goes quiet leaves its aggregate unemitted indefinitely -- data loss dressed as latency. Emission therefore has a second trigger, end of stream, which the corpus defines as an implicit flush after the last input and which production gets from the batch flush. The cost is stated rather than hidden: window_ms becomes a maximum, not a guarantee, and one burst can produce more than one aggregate. And it has a consequence nobody should meet in production first: stats count undercounts aggregated data silently, so every panel and alert counting rows changes meaning the moment a rule aggregates the data behind it. Nothing here fixes that. The correct idiom is summing cairnobs.count; teaching the query layer to do it automatically is a Phase 2 change to the IR, recorded as the open question this decision leaves in its place rather than quietly inherited. Six cases added, corpus at 45. The validator's unspecified-action guard stays in place with an empty set, still rejecting anything added to it. Signed-off-by: John Coffey <[email protected]> |
||
|
|
c2f74b4e1a |
Put apply-then-verify in v1, and specify it
It was a defensible cut while a canary might have caught a bad rollout partway through. With no canary it is the only thing that recovers a host without an operator noticing, and "total evaluation should make it unreachable" is what every crash-loop was before it happened. Specified rather than named. A two-field file is written and fsynced before rules are applied: the version being attempted, and whether it is trying, good or quarantined. The rule set itself is never written, which is what keeps the crash-loop-not-strand property -- an agent that loses the file still boots clean and re-syncs. A version found still "trying" at boot is quarantined, the agent starts with no rules, and it reports the quarantined version so the failure is visible rather than merely survived. A different version clears the quarantine, because pushing new rules is the correction. Three failure modes decided instead of discovered. An agent that cannot write the file logs once and runs with total evaluation alone: degrading to "no rules" would punish every read-only deployment for a failure that has not happened, and the backstop is best-effort by construction rather than by accident. An agent killed for an unrelated reason quarantines a blameless rule set, which is a deliberate false positive -- the alternative is claiming to distinguish "died because of the rules" from "died while they happened to be loaded", which it cannot do honestly. And a rule set fatal on only some hosts quarantines per host, which is the closest thing to a canary this design has: the first host to hit it reports while the rest carry on. Also states that the conformance corpus cannot test any of this, so nobody tries. The corpus pins rule semantics -- records in, records out. Process death and file state across restarts are not expressible that way and need agent-side tests driving a real process through crash and restart. Signed-off-by: John Coffey <[email protected]> |
||
|
|
02ebbc86a1 |
Settle Phase 8's regex and rollout questions
regex-lite on the agent, full regex at ingest, conformance corpus limited to the syntax both accept. The agent had no regex dependency at all and the full crate is a megabyte-plus against a binary whose pitch is that it is small and static. The rejected alternative, regex ingest-side only, would have given up redacting PII before it leaves the host -- the one capability that most needs to be on the agent, and most of the reason on-host processing exists at all. Checked rather than assumed: every pattern the corpus uses compiles and behaves under regex-lite, including named captures and replace_all replacing every occurrence, which is what mask needs and what a redaction stopping at the first hit would get wrong. Both engines also resolve alternation leftmost-first, confirmed by running the same pattern through regex-lite and Go's regexp and getting the same output. A case now pins that, since it is the kind of semantic two independent implementations can differ on silently. No canary gate. An edit reaches every matching agent at once, and this design had suggested that might have to block the feature. It does not, but the risk is accepted rather than waved away: total evaluation is the real mitigation, a fatal rule set crash-loops rather than strands because overrides are never persisted, apply-then-verify makes that self-healing, and applied_override_version already shows the blast radius. What is genuinely given up is the ability to stop a bad rollout partway through -- everything else shortens the outage without preventing the rule reaching every host first. That raises the stakes on apply-then-verify, which is still open. Without a canary, total evaluation stops being a nice property of a well-designed DSL and becomes the only thing between a bad rule and every host at once, so cutting the backstop from v1 is a harder call than it looked when it was written down. Signed-off-by: John Coffey <[email protected]> |
||
|
|
7fabc6a067 |
Build the Phase 8 conformance corpus
The design argues the conformance suite is the specification and should be built before either implementation, since two hand-written implementations of one language diverge unless something shared pins them. This is that suite: 38 cases in /processing, a language-neutral top-level directory for the same reason /proto is one -- the Rust agent and the Go ingest tier both consume the definition and neither owns it. Nothing executes the cases, because neither implementation exists. A stdlib-only validator checks the corpus stays well-formed and runs in CI: known actions, addressable fields, compilable patterns, names matching filenames, and no case depending on record_id, which is withheld so the question of whether an ingest-side rule can see one stays open. The validator was checked against seven deliberately broken cases before being trusted, since "38/38 valid" means nothing from a validator that cannot fail. Writing the cases first has already paid for itself twice. It forced two determinism decisions the prose had left vague, both of which a conformance suite cannot avoid answering. Sampling is counter-based rather than random: random is statistically nicer and impossible to assert on. Windows are measured on record timestamps rather than wall-clock, which makes replay deterministic and, not incidentally, makes backfill behave correctly where a wall-clock window would not. And it made the missing aggregate_count answer concrete. The design does not say what that action emits, or what a query not expecting a synthetic record sees. Rather than invent one by writing cases, the validator rejects any case using it, so the design question has to be answered before the behaviour can be frozen by accident. The corpus includes the acceptance case from real measured data: the two processes that account for roughly 60% of a real workstation's journal volume, and the one kernel message worth keeping. Signed-off-by: John Coffey <[email protected]> |
||
|
|
7cc2fd8c78 |
Draft the Phase 8 processing design
Starts from the distribution channel and the safety invariant it protects, and derives the language from them, rather than designing a rule language and asking later how to ship it. Four decisions proposed. A rule is a matcher plus ordered typed actions, with no expressions and nothing resembling eval -- less expressive than Cribl on purpose, and the only shape that can be pushed to ten thousand hosts and audited by reading it. Total evaluation is the primary safety guarantee, with apply-then-verify as a backstop. One spec with two implementations means a language-neutral conformance suite is the specification and should be built first, not last. Distribution reuses DesiredOverride, following extra_file_paths as the precedent for a repeated field. It also corrects something #21 got wrong. That change said a fatal rule would strand an agent the way a corrupted ingest endpoint does. Reading apply_override's actual semantics, overrides live only in the running process's memory and are never written to disk, so a restarted agent boots clean and re-syncs -- a fatal rule set crash-loops rather than strands, and the agent keeps checking in, so it stays correctable. That is a much better failure mode, and it was acquired by accident: the "don't persist" choice was made to avoid filesystem writes on read-only images, not for safety. This design promotes it to a constraint, since persisting overrides later would silently convert every crash-loop into a strand. positioning.md is corrected to match rather than left disagreeing. Five open questions are left open rather than answered to look decisive, the sharpest being that there is no staged rollout today: an edit reaches every matching agent at once, which for executable rules is the difference between breaking one host and breaking all of them. The v1 acceptance test is real data, not a fixture: two processes on the maintainer's own workstation account for 308 of 325 journal entries in five minutes, and a suppress rule should remove about 60% of that host's volume. Signed-off-by: John Coffey <[email protected]> |