Phase 3: dashboards and alerting
Saved, shareable multi-panel dashboards (table/line/bar/single-stat panels via gridstack + uPlot, global + per-panel time range, JSON export/import) and threshold/absence alert rules with an ok/pending/firing evaluator and webhook/Slack/PagerDuty delivery. - New /metadata component: Postgres control-plane store for dashboards, panels, notification targets, alert rules/state, and delivery log -- see docs/phase-3-dashboard-design.md for why ClickHouse's MergeTree family isn't a fit for this access pattern (needs real row-level locking and read-your-writes consistency). - api/internal/dashboards: dashboard/panel CRUD, pure -- panel query execution stays client-side, reusing the existing /query endpoint. - New /alerting service: rule/target CRUD, a ticker-driven evaluator (claim-then-evaluate concurrency control, transactional-outbox delivery, query errors and threshold zero-rows never coerced into a false transition) and webhook/Slack/PagerDuty delivery with retry/backoff. See docs/phase-3-alerting-design.md for the full state-machine design and the four correctness properties it implements. - web: /dashboards and /alerts UIs; cli: sentryctl dashboards/alerts list/get/apply, seeding a future Terraform provider's JSON contract. - hack/alert-load-test: 500 rules against real ClickHouse data, real measured results in docs/phase-3-runbook.md. Five real bugs found by actually running this against a live stack (documented in the runbook, not just fixed silently): a latent Phase 2 bug where ClickHouse rejected the timestamp format used for earliest=/latest= queries; a "now" literal token injected into query text; a GridStack/uPlot layout-timing race; JS's Date.parse being too lenient to use as a timestamp-detection heuristic; a rule's "enabled" field silently defaulting to false when omitted; and the evaluator's claim-batch-size and worker-pool-concurrency defaulting to the same value, causing 500 concurrently-due rules to take 125s to cycle through instead of the configured 60s.
This commit is contained in:
@@ -0,0 +1,98 @@
|
||||
# alerting
|
||||
|
||||
Alert rule CRUD, the ticker-driven evaluator, and webhook/Slack/PagerDuty
|
||||
delivery. See `/docs/phase-3-alerting-design.md` for the full design
|
||||
(data model, the ok/pending/firing state machine, and the four
|
||||
correctness properties this implementation follows exactly).
|
||||
|
||||
## Running
|
||||
|
||||
```sh
|
||||
POSTGRES_PASSWORD=sentry-dev-only API_QUERY_URL=http://localhost:8080 go run ./cmd/alerting
|
||||
```
|
||||
|
||||
Talks to the same `sentry_metadata` Postgres database as `/api`
|
||||
(different tables — see `/metadata/README.md`), and to `/api`'s
|
||||
`POST /query` over plain HTTP for rule evaluation. Never connects to
|
||||
ClickHouse or Tantivy directly.
|
||||
|
||||
## HTTP API
|
||||
|
||||
```
|
||||
POST /rules create a rule
|
||||
GET /rules list rules (with current state)
|
||||
GET /rules/{id} get a rule (with current state)
|
||||
DELETE /rules/{id}
|
||||
GET /rules/{id}/deliveries delivery log for a rule, most recent first
|
||||
|
||||
POST /targets create a notification target
|
||||
GET /targets
|
||||
GET /targets/{id}
|
||||
DELETE /targets/{id}
|
||||
|
||||
GET /healthz
|
||||
```
|
||||
|
||||
A rule's `condition_type` is `"threshold"` (requires `comparator` +
|
||||
`threshold_value`, and the query must resolve to exactly one row) or
|
||||
`"absence"` (fires when the query returns zero rows in its own
|
||||
`earliest=`/`latest=` window — no separate window field). A notification
|
||||
target's `kind` is `"webhook"`, `"slack"`, or `"pagerduty"` — all three
|
||||
deliver via the same HTTP POST + retry/backoff mechanism
|
||||
(`internal/delivery/webhook.go`); slack/pagerduty are payload formatters
|
||||
only, not separate delivery paths.
|
||||
|
||||
## Environment variables
|
||||
|
||||
| Var | Default |
|
||||
|---|---|
|
||||
| `HTTP_LISTEN_ADDR` | `:8081` |
|
||||
| `POSTGRES_ADDR` | `localhost:5432` |
|
||||
| `POSTGRES_DATABASE` | `sentry_metadata` |
|
||||
| `POSTGRES_USERNAME` | `sentry` |
|
||||
| `POSTGRES_PASSWORD` | (empty — must be set) |
|
||||
| `API_QUERY_URL` | `http://localhost:8080` |
|
||||
| `CORS_ALLOWED_ORIGIN` | `*` |
|
||||
| `EVALUATOR_TICK_SECONDS` | `5` — how often the scheduler checks for due rules |
|
||||
| `EVALUATOR_CLAIM_BATCH_SIZE` | `1000` — how many due rules one tick can pull off the queue |
|
||||
| `EVALUATOR_WORKER_POOL_SIZE` | `20` — bounded concurrency for `/query` calls within a claimed batch |
|
||||
| `EVALUATOR_QUERY_TIMEOUT_SECONDS` | `30` — per-evaluation `POST /query` timeout |
|
||||
|
||||
`EVALUATOR_CLAIM_BATCH_SIZE` and `EVALUATOR_WORKER_POOL_SIZE` are
|
||||
deliberately separate knobs, not the same number — see
|
||||
`internal/config/config.go`'s doc comment for the real bug this
|
||||
separation fixes (found by `hack/alert-load-test`, see
|
||||
`/docs/phase-3-runbook.md`): with both capped at 20, 500 rules due at
|
||||
once took 125s to cycle through instead of the configured 60s.
|
||||
|
||||
## Package layout
|
||||
|
||||
```
|
||||
cmd/alerting/ wires config, Postgres pool, api client; runs the
|
||||
HTTP server + evaluator + delivery worker concurrently (errgroup)
|
||||
internal/httpapi/ REST handlers -- Handler/RegisterRoutes, same shape as api/internal/dashboards
|
||||
internal/rulestore/ pgx CRUD for alert_rules + alert_state; ClaimDueRules
|
||||
(fix 1's atomic claim) and ApplyTransition (fix 2's transactional outbox)
|
||||
internal/notifystore/ pgx CRUD for notification_targets
|
||||
internal/queryclient/ thin HTTP client to api's POST /query -- no querylang import here
|
||||
internal/evaluator/ the ticker + worker pool; transitions.go is the pure,
|
||||
exhaustively-tested ok/pending/firing state machine;
|
||||
condition.go implements fixes 3/4 (errors never
|
||||
coerced to "condition false"; threshold zero-rows
|
||||
is an error, not a 0)
|
||||
internal/delivery/ webhook.go is the claim-and-send worker (all three
|
||||
kinds go through it); slack.go/pagerduty.go are
|
||||
payload formatters only
|
||||
```
|
||||
|
||||
## Building & testing
|
||||
|
||||
```sh
|
||||
go build ./...
|
||||
go vet ./...
|
||||
go test ./...
|
||||
```
|
||||
|
||||
```sh
|
||||
docker build -f Dockerfile -t sentry-alerting . # context is alerting/, not the repo root -- no /proto needed
|
||||
```
|
||||
Reference in New Issue
Block a user