Phase 3: dashboards and alerting
Saved, shareable multi-panel dashboards (table/line/bar/single-stat panels via gridstack + uPlot, global + per-panel time range, JSON export/import) and threshold/absence alert rules with an ok/pending/firing evaluator and webhook/Slack/PagerDuty delivery. - New /metadata component: Postgres control-plane store for dashboards, panels, notification targets, alert rules/state, and delivery log -- see docs/phase-3-dashboard-design.md for why ClickHouse's MergeTree family isn't a fit for this access pattern (needs real row-level locking and read-your-writes consistency). - api/internal/dashboards: dashboard/panel CRUD, pure -- panel query execution stays client-side, reusing the existing /query endpoint. - New /alerting service: rule/target CRUD, a ticker-driven evaluator (claim-then-evaluate concurrency control, transactional-outbox delivery, query errors and threshold zero-rows never coerced into a false transition) and webhook/Slack/PagerDuty delivery with retry/backoff. See docs/phase-3-alerting-design.md for the full state-machine design and the four correctness properties it implements. - web: /dashboards and /alerts UIs; cli: sentryctl dashboards/alerts list/get/apply, seeding a future Terraform provider's JSON contract. - hack/alert-load-test: 500 rules against real ClickHouse data, real measured results in docs/phase-3-runbook.md. Five real bugs found by actually running this against a live stack (documented in the runbook, not just fixed silently): a latent Phase 2 bug where ClickHouse rejected the timestamp format used for earliest=/latest= queries; a "now" literal token injected into query text; a GridStack/uPlot layout-timing race; JS's Date.parse being too lenient to use as a timestamp-detection heuristic; a rule's "enabled" field silently defaulting to false when omitted; and the evaluator's claim-batch-size and worker-pool-concurrency defaulting to the same value, causing 500 concurrently-due rules to take 125s to cycle through instead of the configured 60s.
This commit is contained in:
@@ -209,6 +209,27 @@ ClickHouse-side text indexing, a different join strategy, or simply
|
||||
raising `max_query_size` server-side with matching memory sizing) is
|
||||
explicitly future work, out of scope for Phase 2.
|
||||
|
||||
### Post-Phase-2 fix: `earliest=`/`latest=` never actually worked against live ClickHouse
|
||||
|
||||
Found during Phase 3's dashboard time-range picker work (the first thing
|
||||
to run a relative `earliest=`/`latest=` query against real ClickHouse
|
||||
end-to-end — none of Phase 2's own runbook queries or unit tests
|
||||
happened to exercise it): `executor/sql.go` formatted `TimeRange` bounds
|
||||
with `time.RFC3339Nano` (e.g. `2026-08-12T20:17:40.223505479Z`), which
|
||||
ClickHouse's *implicit* string→`DateTime64` cast for a column-vs-literal
|
||||
comparison rejects outright — `code: 53, Cannot convert string ... to
|
||||
type DateTime64(9, 'UTC')`. ClickHouse's implicit cast is strict and
|
||||
wants `'YYYY-MM-DD HH:MM:SS[.fractional]'` (space-separated, no `T`/`Z`);
|
||||
the lenient ISO-8601-accepting `parseDateTimeBestEffort` is a different,
|
||||
explicitly-invoked function, not what a plain `WHERE timestamp >= '...'`
|
||||
comparison uses. Fixed by `formatClickHouseDateTime64` in `sql.go`. The
|
||||
Phase 2 unit test that covered this (`TestBuildSQLTimeRange`) only
|
||||
asserted the generated SQL *string*, against a fake `SQLRunner` — it
|
||||
never caught this because nothing in that test actually asked real
|
||||
ClickHouse whether the SQL was valid. Left here as a pointed reminder of
|
||||
why this project's "actually run it" discipline exists: a passing test
|
||||
suite and a working feature are not the same claim.
|
||||
|
||||
## Where this lives: `api/internal/querylang/`
|
||||
|
||||
Not a new top-level component. This subsystem always executes in-process
|
||||
|
||||
Reference in New Issue
Block a user