Files
cairnobs/hack/alert-load-test/README.md
T
jcoffey-dev 7a86008062 Complete the low-risk half of the Sentry -> Cairn OBS rebrand
Sweeps the references that carry no runtime coupling, and fixes one that
turned out to be a real bug rather than stale branding.

Docker network: sentry_default -> cairnobs_default across 23 runbook and
test-header `docker run` commands. Compose derives the network from the
directory name, so this lands together with renaming the working copy to
cairnobs/ -- the two are only correct as one change.

Stale references corrected: four Dockerfile "repo root (sentry/)"
headers; .env pointing at the long-renamed deploy/helm/sentry/ chart;
five Helm comments describing the topic as sentry.logs.raw when all four
code paths have defaulted to cairnobs.logs.raw for some time; an
absolute /home/john/Projects/sentry/ path in the operator's package doc,
now repo-relative; the hand-written Tenant CRD description in both of
its identical copies, whose Go source already said Cairn OBS.

Migration 0043 repoints the default tenant's data source. 0026 seeded it
with ('sentry', '/var/lib/sentry-search') to match what
api/internal/config then defaulted to; the rebrand later moved those
defaults to "cairnobs" and /var/lib/cairnobs-search without moving the
already-applied row, leaving the default tenant naming a ClickHouse
database nothing writes to. Scoped to the exact stale values so it is a
no-op on any deployment that set them deliberately. 0026's comment is
annotated as superseded; its applied SQL is untouched.

Deliberately not included: the gRPC wire packages (sentry.logs.v1,
sentry.agent.v1) and proto/sentry/ import paths, which cannot change
without a lockstep agent/server upgrade; the Helm chart's
sentry_metadata database and sentry role, which need a real Postgres
migration on existing deployments; and the compliance audit records in
docs/compliance/, which are a dated historical record.

go build, go vet, and go test pass for ingest and deploy/operator.
2026-08-22 18:39:25 -07:00

43 lines
1.9 KiB
Markdown

# alert-load-test
Seeds a realistic number of concurrent alert rules via `/alerting`'s real
create API (not a direct DB insert) and measures whether the evaluator's
claim scheduling keeps up under load. See
`/docs/phase-3-alerting-design.md`'s "Load-testing plan" and
`/docs/phase-3-runbook.md` for the methodology and real measured results.
```sh
# 1. Push real data so rule queries have real work to do (reuses
# hack/benchmark-fixture):
cd ../benchmark-fixture
go run . --count 500000
# 2. Run a webhook-sink so the (never-firing, by design) rules have a
# valid notification target to point at:
docker run -d --name cairnobs-webhook-sink --network cairnobs_default \
-p 9099:9099 -v $(pwd)/../webhook-sink:/src -w /src golang:1.25-alpine go run .
# 3. Run the load test:
cd ../alert-load-test
go run . --rule-count 500 --eval-interval-seconds 60 --duration 3m30s
```
Each rule queries a different host's count over the last minute
(`earliest=-1m host="host-01" | stats count`) against real ClickHouse
data, with `threshold_value` set unreachably high so rules stay `ok` --
this isolates evaluator/ClickHouse scheduling throughput from
delivery-worker load (a query that never fires still exercises the exact
same claim → `/query` → evaluate → `ApplyTransition` path every tick).
The report shows, per rule, the observed intervals between consecutive
`last_evaluated_at` changes (polled at `--poll-interval`), compared
against the configured `eval_interval_seconds`. A real, significant
finding from actually running this: the evaluator's claim batch size and
worker-pool concurrency limit defaulted to the same number (20), so 500
rules all due at once took 125s to cycle through instead of the
configured 60s. Fixed by separating `EVALUATOR_CLAIM_BATCH_SIZE` from
`EVALUATOR_WORKER_POOL_SIZE` (see `alerting/internal/config/config.go`).
Cleans up the seeded rules and notification target on exit unless
`--no-cleanup` is passed.