Twelve hosts running six services read as somebody's side project. Fifty hosts across twenty-one services read as an estate, which is what a visitor is trying to see themselves in. Thirty-one Linux, eighteen Windows, one Linux host whose agent is gone. The proportions are the point: Windows now carries Active Directory, IIS, SQL Server, Exchange, file shares, Remote Desktop, print, WSUS and SCCM rather than appearing as a Security channel on one box. Linux gains a load-balancer tier, an outbound proxy, MySQL beside Postgres, RabbitMQ, Elasticsearch, three Kubernetes nodes, CI, Vault, OpenLDAP, BIND and a backup server. Fifteen new generators, each writing what the real daemon writes -- HAProxy's timing quintuple, MySQL slow-query blocks, W3C extended format for IIS, kubelet PLEG lines, BIND query logging with NXDOMAIN, Squid's TCP_DENIED, SQL Server deadlock and I/O-stall messages -- with the structured fields carried in attributes so both halves of the query language have something to work on. Three things found while doing it, each of which would have shipped as a quiet wrongness: linuxHosts() decided Windows by `service == "eventlog"`. That held while eventlog was the only Windows role; with IIS and SQL Server on Windows it would have given every one of them a journald system stream -- sshd and UFW lines on a Windows box. It now decides from `os`. worker-02 lost its filling disk when the fleet was rewritten, which is the story worker-disk-filling thresholds on. Restored at its original rate, with a comment saying why it cannot move. Five dashboard panels were written as `timechart`, which this query language does not have -- its stages are where/stats/sort/fields/head/ tail. They are raw ClickHouse SQL now, which the reference recommends for exactly this, and a `pie` panel was dropped before it shipped because the web's VizType union has no such member even though the API accepts one. Five new dashboards: platform/Kubernetes, directory and DNS, messaging and search, the Windows server estate, and edge/proxy. Every panel was checked against the language's real stage list and the web's real viz types. Volume roughly quadruples: about 316 records/minute at rate-scale 1, and 1.9M per nightly reset at the demo's own settings against about 0.5M before. ClickHouse will not notice; reset time and disk on the demo box might, so the README says so and names RATE_SCALE as the lever.
110 lines
3.0 KiB
JSON
110 lines
3.0 KiB
JSON
{
|
|
"name": "Messaging and search",
|
|
"description": "Queue depth, consumer health and the search cluster beside them",
|
|
"default_earliest": "-24h",
|
|
"default_latest": "now",
|
|
"panels": [
|
|
{
|
|
"title": "Queue events",
|
|
"query": "service=rabbitmq | stats count",
|
|
"viz_type": "single_stat",
|
|
"position_x": 0,
|
|
"position_y": 0,
|
|
"width": 3,
|
|
"height": 3,
|
|
"query_language": "spl",
|
|
"sort_order": 0
|
|
},
|
|
{
|
|
"title": "Memory alarms",
|
|
"query": "service=rabbitmq event_kind=alarm | stats count",
|
|
"viz_type": "single_stat",
|
|
"position_x": 3,
|
|
"position_y": 0,
|
|
"width": 3,
|
|
"height": 3,
|
|
"query_language": "spl",
|
|
"sort_order": 1
|
|
},
|
|
{
|
|
"title": "Searches",
|
|
"query": "service=elasticsearch event_kind=search | stats count",
|
|
"viz_type": "single_stat",
|
|
"position_x": 6,
|
|
"position_y": 0,
|
|
"width": 3,
|
|
"height": 3,
|
|
"query_language": "spl",
|
|
"sort_order": 2
|
|
},
|
|
{
|
|
"title": "Long GC pauses",
|
|
"query": "service=elasticsearch event_kind=gc | stats count",
|
|
"viz_type": "single_stat",
|
|
"position_x": 9,
|
|
"position_y": 0,
|
|
"width": 3,
|
|
"height": 3,
|
|
"query_language": "spl",
|
|
"sort_order": 3
|
|
},
|
|
{
|
|
"title": "Queue depth over time",
|
|
"query": "SELECT toStartOfInterval(timestamp, INTERVAL 30 MINUTE) AS bucket, attributes['queue'] AS queue, max(toFloat64OrZero(attributes['queue_depth'])) AS depth FROM logs WHERE service = 'rabbitmq' GROUP BY bucket, queue ORDER BY bucket",
|
|
"viz_type": "line",
|
|
"position_x": 0,
|
|
"position_y": 3,
|
|
"width": 12,
|
|
"height": 5,
|
|
"viz_config": {
|
|
"x_column": "bucket",
|
|
"value_column": "depth",
|
|
"series_column": "queue"
|
|
},
|
|
"query_language": "sql",
|
|
"sort_order": 4
|
|
},
|
|
{
|
|
"title": "Deepest queues",
|
|
"query": "service=rabbitmq | stats max(queue_depth) as max_depth by queue | sort -max_depth",
|
|
"viz_type": "bar",
|
|
"position_x": 0,
|
|
"position_y": 8,
|
|
"width": 6,
|
|
"height": 5,
|
|
"viz_config": {
|
|
"x_column": "queue",
|
|
"value_column": "max_depth"
|
|
},
|
|
"query_language": "spl",
|
|
"sort_order": 5
|
|
},
|
|
{
|
|
"title": "Search latency by index",
|
|
"query": "service=elasticsearch event_kind=search | stats avg(duration_ms) as avg_ms by index | sort -avg_ms",
|
|
"viz_type": "bar",
|
|
"position_x": 6,
|
|
"position_y": 8,
|
|
"width": 6,
|
|
"height": 5,
|
|
"viz_config": {
|
|
"x_column": "index",
|
|
"value_column": "avg_ms"
|
|
},
|
|
"query_language": "spl",
|
|
"sort_order": 6
|
|
},
|
|
{
|
|
"title": "Indexing throttles and GC",
|
|
"query": "service=elasticsearch event_kind!=search | sort -timestamp | head 25 | fields timestamp, host, index, event_kind, message",
|
|
"viz_type": "table",
|
|
"position_x": 0,
|
|
"position_y": 13,
|
|
"width": 12,
|
|
"height": 5,
|
|
"query_language": "spl",
|
|
"sort_order": 7
|
|
}
|
|
]
|
|
}
|