Files
cairnobs/hack/alert-load-test/README.md
T
jcoffey-dev c920e0f2c4 Finish the Cairn OBS rename through services, docs, and assets
The rename commit before this one covered module paths and the obvious
user-facing strings; this is the rest of it -- the places where "sentry"
was a default value, a filename, or a picture rather than a word in a
sentence.

Defaults that changed: CLICKHOUSE_DATABASE (sentry -> cairnobs),
POSTGRES_DATABASE (sentry_metadata -> cairnobs_metadata), and
POSTGRES_USERNAME (sentry -> cairnobs), across api/alerting/ingest and
the enterprise binaries, plus the compose files and migrate scripts that
create those objects. These are *defaults*, so a deployment that sets
them explicitly is unaffected -- but any deployment relying on the old
defaults must have its environment updated before it picks this up, or
it will come up pointing at a database that doesn't exist.

Also: the light-mode logo variants (the dark ones existed alone, so the
landing page and sidebar rendered a dark mark on a light background),
regenerated favicons, and the docs/README/threat-model prose that still
said Sentry.
2026-08-22 16:12:08 -07:00

43 lines
1.9 KiB
Markdown

# alert-load-test
Seeds a realistic number of concurrent alert rules via `/alerting`'s real
create API (not a direct DB insert) and measures whether the evaluator's
claim scheduling keeps up under load. See
`/docs/phase-3-alerting-design.md`'s "Load-testing plan" and
`/docs/phase-3-runbook.md` for the methodology and real measured results.
```sh
# 1. Push real data so rule queries have real work to do (reuses
# hack/benchmark-fixture):
cd ../benchmark-fixture
go run . --count 500000
# 2. Run a webhook-sink so the (never-firing, by design) rules have a
# valid notification target to point at:
docker run -d --name cairnobs-webhook-sink --network sentry_default \
-p 9099:9099 -v $(pwd)/../webhook-sink:/src -w /src golang:1.25-alpine go run .
# 3. Run the load test:
cd ../alert-load-test
go run . --rule-count 500 --eval-interval-seconds 60 --duration 3m30s
```
Each rule queries a different host's count over the last minute
(`earliest=-1m host="host-01" | stats count`) against real ClickHouse
data, with `threshold_value` set unreachably high so rules stay `ok` --
this isolates evaluator/ClickHouse scheduling throughput from
delivery-worker load (a query that never fires still exercises the exact
same claim → `/query` → evaluate → `ApplyTransition` path every tick).
The report shows, per rule, the observed intervals between consecutive
`last_evaluated_at` changes (polled at `--poll-interval`), compared
against the configured `eval_interval_seconds`. A real, significant
finding from actually running this: the evaluator's claim batch size and
worker-pool concurrency limit defaulted to the same number (20), so 500
rules all due at once took 125s to cycle through instead of the
configured 60s. Fixed by separating `EVALUATOR_CLAIM_BATCH_SIZE` from
`EVALUATOR_WORKER_POOL_SIZE` (see `alerting/internal/config/config.go`).
Cleans up the seeded rules and notification target on exit unless
`--no-cleanup` is passed.