Saved, shareable multi-panel dashboards (table/line/bar/single-stat panels via gridstack + uPlot, global + per-panel time range, JSON export/import) and threshold/absence alert rules with an ok/pending/firing evaluator and webhook/Slack/PagerDuty delivery. - New /metadata component: Postgres control-plane store for dashboards, panels, notification targets, alert rules/state, and delivery log -- see docs/phase-3-dashboard-design.md for why ClickHouse's MergeTree family isn't a fit for this access pattern (needs real row-level locking and read-your-writes consistency). - api/internal/dashboards: dashboard/panel CRUD, pure -- panel query execution stays client-side, reusing the existing /query endpoint. - New /alerting service: rule/target CRUD, a ticker-driven evaluator (claim-then-evaluate concurrency control, transactional-outbox delivery, query errors and threshold zero-rows never coerced into a false transition) and webhook/Slack/PagerDuty delivery with retry/backoff. See docs/phase-3-alerting-design.md for the full state-machine design and the four correctness properties it implements. - web: /dashboards and /alerts UIs; cli: sentryctl dashboards/alerts list/get/apply, seeding a future Terraform provider's JSON contract. - hack/alert-load-test: 500 rules against real ClickHouse data, real measured results in docs/phase-3-runbook.md. Five real bugs found by actually running this against a live stack (documented in the runbook, not just fixed silently): a latent Phase 2 bug where ClickHouse rejected the timestamp format used for earliest=/latest= queries; a "now" literal token injected into query text; a GridStack/uPlot layout-timing race; JS's Date.parse being too lenient to use as a timestamp-detection heuristic; a rule's "enabled" field silently defaulting to false when omitted; and the evaluator's claim-batch-size and worker-pool-concurrency defaulting to the same value, causing 500 concurrently-due rules to take 125s to cycle through instead of the configured 60s.
43 lines
1.9 KiB
Markdown
43 lines
1.9 KiB
Markdown
# alert-load-test
|
|
|
|
Seeds a realistic number of concurrent alert rules via `/alerting`'s real
|
|
create API (not a direct DB insert) and measures whether the evaluator's
|
|
claim scheduling keeps up under load. See
|
|
`/docs/phase-3-alerting-design.md`'s "Load-testing plan" and
|
|
`/docs/phase-3-runbook.md` for the methodology and real measured results.
|
|
|
|
```sh
|
|
# 1. Push real data so rule queries have real work to do (reuses
|
|
# hack/benchmark-fixture):
|
|
cd ../benchmark-fixture
|
|
go run . --count 500000
|
|
|
|
# 2. Run a webhook-sink so the (never-firing, by design) rules have a
|
|
# valid notification target to point at:
|
|
docker run -d --name sentry-webhook-sink --network sentry_default \
|
|
-p 9099:9099 -v $(pwd)/../webhook-sink:/src -w /src golang:1.25-alpine go run .
|
|
|
|
# 3. Run the load test:
|
|
cd ../alert-load-test
|
|
go run . --rule-count 500 --eval-interval-seconds 60 --duration 3m30s
|
|
```
|
|
|
|
Each rule queries a different host's count over the last minute
|
|
(`earliest=-1m host="host-01" | stats count`) against real ClickHouse
|
|
data, with `threshold_value` set unreachably high so rules stay `ok` --
|
|
this isolates evaluator/ClickHouse scheduling throughput from
|
|
delivery-worker load (a query that never fires still exercises the exact
|
|
same claim → `/query` → evaluate → `ApplyTransition` path every tick).
|
|
|
|
The report shows, per rule, the observed intervals between consecutive
|
|
`last_evaluated_at` changes (polled at `--poll-interval`), compared
|
|
against the configured `eval_interval_seconds`. A real, significant
|
|
finding from actually running this: the evaluator's claim batch size and
|
|
worker-pool concurrency limit defaulted to the same number (20), so 500
|
|
rules all due at once took 125s to cycle through instead of the
|
|
configured 60s. Fixed by separating `EVALUATOR_CLAIM_BATCH_SIZE` from
|
|
`EVALUATOR_WORKER_POOL_SIZE` (see `alerting/internal/config/config.go`).
|
|
|
|
Cleans up the seeded rules and notification target on exit unless
|
|
`--no-cleanup` is passed.
|