Phase 3: dashboards and alerting

Saved, shareable multi-panel dashboards (table/line/bar/single-stat
panels via gridstack + uPlot, global + per-panel time range, JSON
export/import) and threshold/absence alert rules with an
ok/pending/firing evaluator and webhook/Slack/PagerDuty delivery.

- New /metadata component: Postgres control-plane store for dashboards,
  panels, notification targets, alert rules/state, and delivery log --
  see docs/phase-3-dashboard-design.md for why ClickHouse's MergeTree
  family isn't a fit for this access pattern (needs real row-level
  locking and read-your-writes consistency).
- api/internal/dashboards: dashboard/panel CRUD, pure -- panel query
  execution stays client-side, reusing the existing /query endpoint.
- New /alerting service: rule/target CRUD, a ticker-driven evaluator
  (claim-then-evaluate concurrency control, transactional-outbox
  delivery, query errors and threshold zero-rows never coerced into a
  false transition) and webhook/Slack/PagerDuty delivery with
  retry/backoff. See docs/phase-3-alerting-design.md for the full
  state-machine design and the four correctness properties it
  implements.
- web: /dashboards and /alerts UIs; cli: sentryctl dashboards/alerts
  list/get/apply, seeding a future Terraform provider's JSON contract.
- hack/alert-load-test: 500 rules against real ClickHouse data, real
  measured results in docs/phase-3-runbook.md.

Five real bugs found by actually running this against a live stack
(documented in the runbook, not just fixed silently): a latent Phase 2
bug where ClickHouse rejected the timestamp format used for
earliest=/latest= queries; a "now" literal token injected into query
text; a GridStack/uPlot layout-timing race; JS's Date.parse being too
lenient to use as a timestamp-detection heuristic; a rule's "enabled"
field silently defaulting to false when omitted; and the evaluator's
claim-batch-size and worker-pool-concurrency defaulting to the same
value, causing 500 concurrently-due rules to take 125s to cycle through
instead of the configured 60s.
This commit is contained in:
2026-08-13 17:29:38 -07:00
parent fb5049a747
commit 9435115ab7
88 changed files with 7463 additions and 298 deletions
+82 -3
View File
@@ -97,6 +97,41 @@ services:
CLICKHOUSE_HTTP: "http://clickhouse:8123"
CLICKHOUSE_PASSWORD: "sentry-dev-only"
# Control-plane metadata store (dashboards, alert rules -- see
# /docs/phase-3-dashboard-design.md for why this is Postgres rather
# than new ClickHouse tables). Log data stays on ClickHouse/Tantivy
# only, unaffected.
metadata-postgres:
image: postgres:16-alpine
container_name: sentry-metadata-postgres
environment:
POSTGRES_DB: sentry_metadata
POSTGRES_USER: sentry
POSTGRES_PASSWORD: "sentry-dev-only" # not a real secret, same framing as CLICKHOUSE_PASSWORD above
volumes:
- metadata-postgres-data:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U sentry -d sentry_metadata"]
interval: 5s
timeout: 5s
retries: 30
# One-shot: applies /metadata/migrations/*.sql, then exits 0. api waits
# on this completing successfully, same shape as clickhouse-migrate.
metadata-migrate:
build:
context: ./metadata
container_name: sentry-metadata-migrate
depends_on:
metadata-postgres:
condition: service_healthy
environment:
POSTGRES_HOST: "metadata-postgres"
POSTGRES_PORT: "5432"
POSTGRES_USER: "sentry"
POSTGRES_PASSWORD: "sentry-dev-only"
POSTGRES_DATABASE: "sentry_metadata"
ingest:
build:
context: . # needs both ingest/ and proto/
@@ -148,25 +183,68 @@ services:
depends_on:
clickhouse-migrate:
condition: service_completed_successfully
metadata-migrate:
condition: service_completed_successfully
ports:
- "8080:8080"
environment:
CLICKHOUSE_ADDR: "clickhouse:9000"
CLICKHOUSE_PASSWORD: "sentry-dev-only"
SEARCH_GRPC_ADDR: "search:50052"
POSTGRES_ADDR: "metadata-postgres:5432"
POSTGRES_DATABASE: "sentry_metadata"
POSTGRES_USERNAME: "sentry"
POSTGRES_PASSWORD: "sentry-dev-only"
healthcheck:
# alerting (Phase 3 task 5) depends_on api -- without this, that
# dependency can only mean "container started," not "actually
# listening," and would hammer a not-yet-ready api with errors on
# every evaluator tick during stack startup. api's image is
# distroless (no shell, no wget) so this execs the api binary's own
# -healthcheck self-check mode instead of an external tool.
test: ["CMD", "/api", "-healthcheck"]
interval: 5s
timeout: 5s
retries: 30
alerting:
build:
context: alerting # self-contained, no /proto needed -- see alerting/Dockerfile
dockerfile: Dockerfile
container_name: sentry-alerting
depends_on:
metadata-migrate:
condition: service_completed_successfully
api:
condition: service_healthy
ports:
- "8081:8081"
environment:
POSTGRES_ADDR: "metadata-postgres:5432"
POSTGRES_DATABASE: "sentry_metadata"
POSTGRES_USERNAME: "sentry"
POSTGRES_PASSWORD: "sentry-dev-only"
API_QUERY_URL: "http://api:8080"
healthcheck:
test: ["CMD", "/alerting", "-healthcheck"]
interval: 5s
timeout: 5s
retries: 30
web:
build:
context: web
args:
# Baked in at build time (static site, not a server) as
# localhost:8080 -- this is fetched from the *browser*, which
# resolves against the host's mapped port, not the compose
# network's service DNS name.
# localhost:8080/8081 -- fetched from the *browser*, which
# resolves against the host's mapped ports, not the compose
# network's service DNS names.
VITE_API_BASE_URL: "http://localhost:8080"
VITE_ALERTING_API_BASE_URL: "http://localhost:8081"
container_name: sentry-web
depends_on:
- api
- alerting
ports:
- "3000:3000"
@@ -174,3 +252,4 @@ volumes:
redpanda-data:
clickhouse-data:
search-index-data:
metadata-postgres-data: