Phase 1: Windows log collection + full-text search
Extends the agent, ingest, storage, api, and web with Windows Event Log/ETW sourcing and Tantivy-backed free-text search, per the approved Phase 1 plan. - CLAUDE.md: materialized on disk (never existed as a file before) with a new Phase 1 "done looks like" section. - agent: Windows Event Log (EvtSubscribe) and ETW sources, Windows service wrapper (install/uninstall/run-service), both feature- and target_os-gated so Linux builds/tests/clippy stay unaffected. Also fixed two pre-existing Phase 0 clippy gaps (dead-code on default-features-only builds, a type-inference edge case) found while testing every feature combination properly for the first time. UNVERIFIED on real Windows -- no Windows toolchain existed anywhere in the build environment; flagged prominently in three places. - proto/ingest: new record_id field, assigned once server-side in ingest's gRPC front end so ClickHouse and Tantivy agree on the same ID for the same record. - storage: record_id column + bloom filter index, verified against a live ClickHouse. - search: new service, Tantivy index, rskafka consumer as an independent second consumer group on the same Redpanda topic ingest already reads. - api/web: new /search endpoint and page, sharing the query page's result-table shape and component. - hack/windows-fixture: sends realistic Windows-shaped data straight to ingest, so the pipeline's handling of it is verifiable without a Windows host. Verified end-to-end on the live docker-compose stack: the same record_id comes back from both /query and /search for the same log line, including for windows-fixture's synthetic Windows Event Log data. Real bugs found and fixed along the way: api/Dockerfile missing proto/ in its build context, search's logs being completely silent (RUST_LOG gap), and search/target/ missing from .gitignore/.dockerignore.
This commit is contained in:
+15
-3
@@ -4,9 +4,17 @@ Go service sitting between the Rust agent and ClickHouse. Two halves in one
|
||||
binary, selected with `--mode`:
|
||||
|
||||
- **server** — mTLS gRPC front end (`LogIngest.PushBatch`) that agents
|
||||
connect to. Forwards each record, proto-encoded and unchanged, onto
|
||||
Redpanda. Does no normalization — kept thin so agent-facing latency isn't
|
||||
coupled to ClickHouse write performance.
|
||||
connect to. Assigns each record a server-side `record_id` (a UUID,
|
||||
overwriting whatever the agent sent — agents always send it empty) and
|
||||
otherwise forwards records proto-encoded onto Redpanda unchanged. Still
|
||||
kept thin — one field assignment, no real normalization — so agent-
|
||||
facing latency isn't coupled to ClickHouse write performance.
|
||||
`record_id` has to be assigned exactly once, here, rather than
|
||||
independently by each downstream consumer: Phase 1's Tantivy indexer
|
||||
and the ClickHouse writer both read the same Redpanda messages and need
|
||||
to agree on the same ID for the same record to join search hits back to
|
||||
rows — two consumers generating their own IDs would produce mismatched
|
||||
ones for what's supposed to be the same record.
|
||||
- **consumer** — reads back off Redpanda, normalizes into the ClickHouse row
|
||||
shape (`internal/normalize`), and batch-writes via the native protocol
|
||||
driver. Commits Redpanda offsets only after a successful ClickHouse
|
||||
@@ -35,6 +43,10 @@ egress an agent has. See `/docs/architecture.md`.
|
||||
protocol, pure Go (no cgo).
|
||||
- **golang.org/x/sync/errgroup** — used in `cmd/ingest/main.go` to run the
|
||||
server and consumer halves concurrently and propagate the first error.
|
||||
- **github.com/google/uuid** — was already in the dependency graph
|
||||
transitively (via clickhouse-go); promoted to a direct dependency for
|
||||
`record_id` generation in `internal/grpcserver`, so not a new addition
|
||||
to the transitive tree.
|
||||
|
||||
## Configuration
|
||||
|
||||
|
||||
Reference in New Issue
Block a user