Metric history: only the calculating node stores cluster-wide gauges #139

Merged
jcoffey-dev merged 1 commits from fix/cluster-gauges-one-node into main 2026-10-01 01:59:11 +00:00
Owner

queue.count, user.count and domain.count count the whole cluster, and only the node with the metrics-calculation role works them out. Every node still stored them.

The latest stored readings in production show the result:

gauge node 0 node 1 node 2
queue.count 18,446,744,073,709,551,596 (−20) 18,446,744,073,709,551,613 (−3) 8
user.count 0 0 7
domain.count 0 0 11

On the non-calculating nodes, the queue gauge only moves with local queue events and drifts below zero. A reader taking the latest reading got whichever node wrote last.

  • sample(calculates) leaves the three out on a node without the role. store_metrics passes roles.metrics_calculate.
  • Unit test: only_the_calculating_node_stores_cluster_gauges.
  • The Prometheus and OTel exports are unchanged.

The console side (per-node gauge readings, and skipping the junk already stored) is in a separate inbuxa-admin PR.

No new strings.

queue.count, user.count and domain.count count the whole cluster, and only the node with the metrics-calculation role works them out. Every node still stored them. The latest stored readings in production show the result: | gauge | node 0 | node 1 | node 2 | |---|---|---|---| | queue.count | 18,446,744,073,709,551,596 (−20) | 18,446,744,073,709,551,613 (−3) | 8 | | user.count | 0 | 0 | 7 | | domain.count | 0 | 0 | 11 | On the non-calculating nodes, the queue gauge only moves with local queue events and drifts below zero. A reader taking the latest reading got whichever node wrote last. - sample(calculates) leaves the three out on a node without the role. store_metrics passes roles.metrics_calculate. - Unit test: only_the_calculating_node_stores_cluster_gauges. - The Prometheus and OTel exports are unchanged. The console side (per-node gauge readings, and skipping the junk already stored) is in a separate inbuxa-admin PR. No new strings.
jcoffey-dev added 1 commit 2026-10-01 01:51:34 +00:00
Metric history: only the calculating node stores cluster-wide gauges
ci / build (pull_request) Skipped
ci / fork-checks (pull_request) Skipped
github/ci (branch) GitHub Actions
ci / github (pull_request) Successful in 7m5s
30d4cef0e7
queue.count, user.count and domain.count count the whole cluster, and
only the node with the metrics-calculation role works them out. Every
node still stored them. On the others the queue gauge only moves with
local queue events, so it had drifted below zero (production: node 0 at
18,446,744,073,709,551,596, node 1 at ...613, i.e. -20 and -3), and
the account and domain counts stayed at 0. A reader taking the latest
reading got whichever node wrote last.

sample() now takes whether the node calculates them and leaves them out
otherwise. A unit test covers both cases.
jcoffey-dev merged commit cd7a0f4163 into main 2026-10-01 01:59:11 +00:00
jcoffey-dev deleted branch fix/cluster-gauges-one-node 2026-10-01 01:59:11 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: inbuxa/inbuxa-server#139