Commit Graph
4 Commits
Author SHA1 Message Date
jcoffey-dev 966241070e Name Coffey Labs LLC as the copyright holder
ci / fork-checks (pull_request) Skipped
ci / build (pull_request) Skipped
github/ci (branch) GitHub Actions
ci / github (pull_request) Successful in 6m29s
Coffey Labs is now Coffey Labs LLC, an Arizona limited liability company. Copyright lines and SPDX-FileCopyrightText headers naming Coffey Labs or John Coffey now name Coffey Labs LLC. Upstream copyright notices are unchanged.
2026-10-05 23:33:00 -07:00
jcoffey-dev 30d4cef0e7 Metric history: only the calculating node stores cluster-wide gauges
ci / build (pull_request) Skipped
ci / fork-checks (pull_request) Skipped
github/ci (branch) GitHub Actions
ci / github (pull_request) Successful in 7m5s
queue.count, user.count and domain.count count the whole cluster, and
only the node with the metrics-calculation role works them out. Every
node still stored them. On the others the queue gauge only moves with
local queue events, so it had drifted below zero (production: node 0 at
18,446,744,073,709,551,596, node 1 at ...613, i.e. -20 and -3), and
the account and domain counts stayed at 0. A reader taking the latest
reading got whichever node wrote last.

sample() now takes whether the node calculates them and leaves them out
otherwise. A unit test covers both cases.
2026-09-30 18:51:22 -07:00
jcoffey-dev 20abf69d31 x:Metric: say which node wrote each sample
ci / fork-checks (pull_request) Skipped
ci / build (pull_request) Skipped
github/ci (branch) GitHub Actions
ci / github (pull_request) Successful in 7m5s
Each node stores histograms as running totals since it started. A sample
didn't say which node wrote it (the node was only in the id's low bits),
so a reader couldn't diff totals per node, and the console diffed across
nodes: on the three-node production cluster the delivery attempt time
read 14.7 s over the last hour against 0.7 s from the nodes' own figures.

x:Metric/get now returns nodeId alongside timestamp, both from the id.
The telemetry suite checks every sample carries it.
2026-09-30 11:43:13 -07:00
jcoffey-dev 9f7035588f Monitoring: metric history, one edition of metrics, and x:Metric over the stored samples (MON-4 to MON-7, MON-9, MON-17, MON-39)
Every node writes a sample per metric on metricsCollectionInterval:
counters as the increase since its last sample, gauges always, histograms
as totals when changed. Samples are x:Metric in the registry's encoding
under the telemetry key class, ids time-ordered. x:Metric/get and /query
read them with metric and timestamp filters and full paging, hide what's
past holdMetricsFor, and the data purge deletes it. The is_enterprise
split is gone, so every gauge and histogram is collected and exported, and
queue.count is set from the queue itself. The shared metrics suite runs.
2026-09-19 08:21:45 -07:00