Files
cairnobs/hack/alert-load-test
jcoffey-dev 9435115ab7 Phase 3: dashboards and alerting
Saved, shareable multi-panel dashboards (table/line/bar/single-stat
panels via gridstack + uPlot, global + per-panel time range, JSON
export/import) and threshold/absence alert rules with an
ok/pending/firing evaluator and webhook/Slack/PagerDuty delivery.

- New /metadata component: Postgres control-plane store for dashboards,
  panels, notification targets, alert rules/state, and delivery log --
  see docs/phase-3-dashboard-design.md for why ClickHouse's MergeTree
  family isn't a fit for this access pattern (needs real row-level
  locking and read-your-writes consistency).
- api/internal/dashboards: dashboard/panel CRUD, pure -- panel query
  execution stays client-side, reusing the existing /query endpoint.
- New /alerting service: rule/target CRUD, a ticker-driven evaluator
  (claim-then-evaluate concurrency control, transactional-outbox
  delivery, query errors and threshold zero-rows never coerced into a
  false transition) and webhook/Slack/PagerDuty delivery with
  retry/backoff. See docs/phase-3-alerting-design.md for the full
  state-machine design and the four correctness properties it
  implements.
- web: /dashboards and /alerts UIs; cli: sentryctl dashboards/alerts
  list/get/apply, seeding a future Terraform provider's JSON contract.
- hack/alert-load-test: 500 rules against real ClickHouse data, real
  measured results in docs/phase-3-runbook.md.

Five real bugs found by actually running this against a live stack
(documented in the runbook, not just fixed silently): a latent Phase 2
bug where ClickHouse rejected the timestamp format used for
earliest=/latest= queries; a "now" literal token injected into query
text; a GridStack/uPlot layout-timing race; JS's Date.parse being too
lenient to use as a timestamp-detection heuristic; a rule's "enabled"
field silently defaulting to false when omitted; and the evaluator's
claim-batch-size and worker-pool-concurrency defaulting to the same
value, causing 500 concurrently-due rules to take 125s to cycle through
instead of the configured 60s.
2026-08-13 17:29:38 -07:00
..
2026-08-13 17:29:38 -07:00
2026-08-13 17:29:38 -07:00
2026-08-13 17:29:38 -07:00

alert-load-test

Seeds a realistic number of concurrent alert rules via /alerting's real create API (not a direct DB insert) and measures whether the evaluator's claim scheduling keeps up under load. See /docs/phase-3-alerting-design.md's "Load-testing plan" and /docs/phase-3-runbook.md for the methodology and real measured results.

# 1. Push real data so rule queries have real work to do (reuses
#    hack/benchmark-fixture):
cd ../benchmark-fixture
go run . --count 500000

# 2. Run a webhook-sink so the (never-firing, by design) rules have a
#    valid notification target to point at:
docker run -d --name sentry-webhook-sink --network sentry_default \
  -p 9099:9099 -v $(pwd)/../webhook-sink:/src -w /src golang:1.25-alpine go run .

# 3. Run the load test:
cd ../alert-load-test
go run . --rule-count 500 --eval-interval-seconds 60 --duration 3m30s

Each rule queries a different host's count over the last minute (earliest=-1m host="host-01" | stats count) against real ClickHouse data, with threshold_value set unreachably high so rules stay ok -- this isolates evaluator/ClickHouse scheduling throughput from delivery-worker load (a query that never fires still exercises the exact same claim → /query → evaluate → ApplyTransition path every tick).

The report shows, per rule, the observed intervals between consecutive last_evaluated_at changes (polled at --poll-interval), compared against the configured eval_interval_seconds. A real, significant finding from actually running this: the evaluator's claim batch size and worker-pool concurrency limit defaulted to the same number (20), so 500 rules all due at once took 125s to cycle through instead of the configured 60s. Fixed by separating EVALUATOR_CLAIM_BATCH_SIZE from EVALUATOR_WORKER_POOL_SIZE (see alerting/internal/config/config.go).

Cleans up the seeded rules and notification target on exit unless --no-cleanup is passed.