Files
cairnobs/hack/alert-load-test/README.md
T
jcoffey-dev c920e0f2c4 Finish the Cairn OBS rename through services, docs, and assets
The rename commit before this one covered module paths and the obvious
user-facing strings; this is the rest of it -- the places where "sentry"
was a default value, a filename, or a picture rather than a word in a
sentence.

Defaults that changed: CLICKHOUSE_DATABASE (sentry -> cairnobs),
POSTGRES_DATABASE (sentry_metadata -> cairnobs_metadata), and
POSTGRES_USERNAME (sentry -> cairnobs), across api/alerting/ingest and
the enterprise binaries, plus the compose files and migrate scripts that
create those objects. These are *defaults*, so a deployment that sets
them explicitly is unaffected -- but any deployment relying on the old
defaults must have its environment updated before it picks this up, or
it will come up pointing at a database that doesn't exist.

Also: the light-mode logo variants (the dark ones existed alone, so the
landing page and sidebar rendered a dark mark on a light background),
regenerated favicons, and the docs/README/threat-model prose that still
said Sentry.
2026-08-22 16:12:08 -07:00

1.9 KiB

alert-load-test

Seeds a realistic number of concurrent alert rules via /alerting's real create API (not a direct DB insert) and measures whether the evaluator's claim scheduling keeps up under load. See /docs/phase-3-alerting-design.md's "Load-testing plan" and /docs/phase-3-runbook.md for the methodology and real measured results.

# 1. Push real data so rule queries have real work to do (reuses
#    hack/benchmark-fixture):
cd ../benchmark-fixture
go run . --count 500000

# 2. Run a webhook-sink so the (never-firing, by design) rules have a
#    valid notification target to point at:
docker run -d --name cairnobs-webhook-sink --network sentry_default \
  -p 9099:9099 -v $(pwd)/../webhook-sink:/src -w /src golang:1.25-alpine go run .

# 3. Run the load test:
cd ../alert-load-test
go run . --rule-count 500 --eval-interval-seconds 60 --duration 3m30s

Each rule queries a different host's count over the last minute (earliest=-1m host="host-01" | stats count) against real ClickHouse data, with threshold_value set unreachably high so rules stay ok -- this isolates evaluator/ClickHouse scheduling throughput from delivery-worker load (a query that never fires still exercises the exact same claim → /query → evaluate → ApplyTransition path every tick).

The report shows, per rule, the observed intervals between consecutive last_evaluated_at changes (polled at --poll-interval), compared against the configured eval_interval_seconds. A real, significant finding from actually running this: the evaluator's claim batch size and worker-pool concurrency limit defaulted to the same number (20), so 500 rules all due at once took 125s to cycle through instead of the configured 60s. Fixed by separating EVALUATOR_CLAIM_BATCH_SIZE from EVALUATOR_WORKER_POOL_SIZE (see alerting/internal/config/config.go).

Cleans up the seeded rules and notification target on exit unless --no-cleanup is passed.