Phase 3: dashboards and alerting
Saved, shareable multi-panel dashboards (table/line/bar/single-stat panels via gridstack + uPlot, global + per-panel time range, JSON export/import) and threshold/absence alert rules with an ok/pending/firing evaluator and webhook/Slack/PagerDuty delivery. - New /metadata component: Postgres control-plane store for dashboards, panels, notification targets, alert rules/state, and delivery log -- see docs/phase-3-dashboard-design.md for why ClickHouse's MergeTree family isn't a fit for this access pattern (needs real row-level locking and read-your-writes consistency). - api/internal/dashboards: dashboard/panel CRUD, pure -- panel query execution stays client-side, reusing the existing /query endpoint. - New /alerting service: rule/target CRUD, a ticker-driven evaluator (claim-then-evaluate concurrency control, transactional-outbox delivery, query errors and threshold zero-rows never coerced into a false transition) and webhook/Slack/PagerDuty delivery with retry/backoff. See docs/phase-3-alerting-design.md for the full state-machine design and the four correctness properties it implements. - web: /dashboards and /alerts UIs; cli: sentryctl dashboards/alerts list/get/apply, seeding a future Terraform provider's JSON contract. - hack/alert-load-test: 500 rules against real ClickHouse data, real measured results in docs/phase-3-runbook.md. Five real bugs found by actually running this against a live stack (documented in the runbook, not just fixed silently): a latent Phase 2 bug where ClickHouse rejected the timestamp format used for earliest=/latest= queries; a "now" literal token injected into query text; a GridStack/uPlot layout-timing race; JS's Date.parse being too lenient to use as a timestamp-detection heuristic; a rule's "enabled" field silently defaulting to false when omitted; and the evaluator's claim-batch-size and worker-pool-concurrency defaulting to the same value, causing 500 concurrently-due rules to take 125s to cycle through instead of the configured 60s.
This commit is contained in:
@@ -0,0 +1,69 @@
|
||||
# metadata
|
||||
|
||||
PostgreSQL schema and migration tooling for Sentry's control-plane
|
||||
config: dashboards, alert rules, and everything else that isn't log data.
|
||||
See `/docs/phase-3-dashboard-design.md` and
|
||||
`/docs/phase-3-alerting-design.md` for why this is a separate database
|
||||
from `/storage` (ClickHouse) rather than new ClickHouse tables — short
|
||||
version: dashboards and alert state need real row-level locking and
|
||||
transactional read-modify-write, which ClickHouse's MergeTree family
|
||||
doesn't provide.
|
||||
|
||||
## Schema
|
||||
|
||||
Six tables across two features, one shared database (`sentry_metadata`):
|
||||
|
||||
- `dashboards`, `dashboard_panels` — owned by `/api` (`api/internal/dashboards`)
|
||||
- `notification_targets`, `alert_rules`, `alert_state`, `delivery_log` —
|
||||
owned by `/alerting`
|
||||
|
||||
"Owned" here is a documentation convention, not a technical boundary —
|
||||
both services connect to the same Postgres instance/database, each with
|
||||
its own hand-written SQL for the tables it's responsible for. Nothing is
|
||||
shared across service `internal/` trees for this, matching the existing
|
||||
repo convention that only `/proto` is shared code (and even that isn't
|
||||
shared logic, just generated bindings).
|
||||
|
||||
## Migration tooling: mirrors `/storage/migrate.sh`, not a framework
|
||||
|
||||
Same reasoning as `/storage/README.md`: pulling in `golang-migrate` for
|
||||
what's currently six `CREATE TABLE` statements is premature machinery.
|
||||
`migrate.sh` applies `migrations/*.sql` in filename order over `psql`,
|
||||
tracking what's applied in a `schema_migrations` table, one DDL object
|
||||
per file (kept for repo-wide consistency of what a migration "version"
|
||||
means, even though Postgres itself supports multi-statement transactions
|
||||
unlike ClickHouse's HTTP interface).
|
||||
|
||||
## Running
|
||||
|
||||
```sh
|
||||
docker compose up -d # starts a standalone Postgres for local work
|
||||
POSTGRES_PASSWORD=sentry-dev-only ./migrate.sh # applies migrations/*.sql
|
||||
```
|
||||
|
||||
Environment variables `migrate.sh` reads (all optional except
|
||||
`POSTGRES_PASSWORD`, matching the root docker-compose.yml's
|
||||
`metadata-postgres` service):
|
||||
|
||||
| Var | Default |
|
||||
|---|---|
|
||||
| `POSTGRES_HOST` | `localhost` |
|
||||
| `POSTGRES_PORT` | `5432` |
|
||||
| `POSTGRES_USER` | `sentry` |
|
||||
| `POSTGRES_PASSWORD` | (empty — must be set) |
|
||||
| `POSTGRES_DATABASE` | `sentry_metadata` |
|
||||
|
||||
The database itself isn't created by `migrate.sh` — the `postgres:16-alpine`
|
||||
image auto-creates `POSTGRES_DB` on first startup, unlike ClickHouse where
|
||||
`migrate.sh` has to issue `CREATE DATABASE IF NOT EXISTS` itself.
|
||||
|
||||
There's also a `Dockerfile` (bash + the `postgresql16-client` package
|
||||
baked in, `migrations/` copied in at build time) used by the root-level
|
||||
`docker-compose.yml` as a one-shot init service (`metadata-migrate`) —
|
||||
no runtime package install, no host volume mount needed.
|
||||
|
||||
## Adding a migration
|
||||
|
||||
Add `migrations/000N_description.sql` with the next sequential number and
|
||||
a single DDL statement. `migrate.sh` picks it up automatically — no
|
||||
registration step.
|
||||
Reference in New Issue
Block a user