Files
jcoffey-dev 7a86008062 Complete the low-risk half of the Sentry -> Cairn OBS rebrand
Sweeps the references that carry no runtime coupling, and fixes one that
turned out to be a real bug rather than stale branding.

Docker network: sentry_default -> cairnobs_default across 23 runbook and
test-header `docker run` commands. Compose derives the network from the
directory name, so this lands together with renaming the working copy to
cairnobs/ -- the two are only correct as one change.

Stale references corrected: four Dockerfile "repo root (sentry/)"
headers; .env pointing at the long-renamed deploy/helm/sentry/ chart;
five Helm comments describing the topic as sentry.logs.raw when all four
code paths have defaulted to cairnobs.logs.raw for some time; an
absolute /home/john/Projects/sentry/ path in the operator's package doc,
now repo-relative; the hand-written Tenant CRD description in both of
its identical copies, whose Go source already said Cairn OBS.

Migration 0043 repoints the default tenant's data source. 0026 seeded it
with ('sentry', '/var/lib/sentry-search') to match what
api/internal/config then defaulted to; the rebrand later moved those
defaults to "cairnobs" and /var/lib/cairnobs-search without moving the
already-applied row, leaving the default tenant naming a ClickHouse
database nothing writes to. Scoped to the exact stale values so it is a
no-op on any deployment that set them deliberately. 0026's comment is
annotated as superseded; its applied SQL is untouched.

Deliberately not included: the gRPC wire packages (sentry.logs.v1,
sentry.agent.v1) and proto/sentry/ import paths, which cannot change
without a lockstep agent/server upgrade; the Helm chart's
sentry_metadata database and sentry role, which need a real Postgres
migration on existing deployments; and the compliance audit records in
docs/compliance/, which are a dated historical record.

go build, go vet, and go test pass for ingest and deploy/operator.
2026-08-22 18:39:25 -07:00

1.9 KiB

alert-load-test

Seeds a realistic number of concurrent alert rules via /alerting's real create API (not a direct DB insert) and measures whether the evaluator's claim scheduling keeps up under load. See /docs/phase-3-alerting-design.md's "Load-testing plan" and /docs/phase-3-runbook.md for the methodology and real measured results.

# 1. Push real data so rule queries have real work to do (reuses
#    hack/benchmark-fixture):
cd ../benchmark-fixture
go run . --count 500000

# 2. Run a webhook-sink so the (never-firing, by design) rules have a
#    valid notification target to point at:
docker run -d --name cairnobs-webhook-sink --network cairnobs_default \
  -p 9099:9099 -v $(pwd)/../webhook-sink:/src -w /src golang:1.25-alpine go run .

# 3. Run the load test:
cd ../alert-load-test
go run . --rule-count 500 --eval-interval-seconds 60 --duration 3m30s

Each rule queries a different host's count over the last minute (earliest=-1m host="host-01" | stats count) against real ClickHouse data, with threshold_value set unreachably high so rules stay ok -- this isolates evaluator/ClickHouse scheduling throughput from delivery-worker load (a query that never fires still exercises the exact same claim → /query → evaluate → ApplyTransition path every tick).

The report shows, per rule, the observed intervals between consecutive last_evaluated_at changes (polled at --poll-interval), compared against the configured eval_interval_seconds. A real, significant finding from actually running this: the evaluator's claim batch size and worker-pool concurrency limit defaulted to the same number (20), so 500 rules all due at once took 125s to cycle through instead of the configured 60s. Fixed by separating EVALUATOR_CLAIM_BATCH_SIZE from EVALUATOR_WORKER_POOL_SIZE (see alerting/internal/config/config.go).

Cleans up the seeded rules and notification target on exit unless --no-cleanup is passed.