Fix two bugs found by actually running the Phase 0 pipeline end-to-end

Both surfaced only by running docker compose up for real, not from review:

- ClickHouse's official image silently disables network access for the
  default user unless CLICKHOUSE_USER or CLICKHOUSE_PASSWORD is set to a
  genuinely non-empty value (an explicit empty password still triggers
  it). Set a dev-only password across clickhouse, clickhouse-migrate,
  ingest, and api in both docker-compose.yml files.

- rpk cluster health and rpk topic ... don't accept --brokers; health
  checks need -X admin.hosts=... (port 9644), topic commands need
  -X brokers=... (port 9092). The old script's retry loop silently
  swallowed the resulting "unknown flag" error and retried forever,
  which blocked ingest from ever starting.

Verified: agent -> ingest -> Redpanda -> ClickHouse -> api round-trip
confirmed with a real log line on a real host.
This commit is contained in:
2026-08-13 09:31:44 -07:00
parent b6b092c912
commit fe854b1091
5 changed files with 62 additions and 10 deletions
+9 -1
View File
@@ -15,9 +15,17 @@ both are invoked.
```sh
docker compose up -d
REDPANDA_BROKERS=localhost:9092 ./provision-topics.sh
./provision-topics.sh # defaults (localhost:9092 / localhost:9644) match this compose file
```
Two separate addresses matter here, confirmed by actually running this
against a live Redpanda container: `rpk cluster health` talks to the
**Admin API** (`REDPANDA_ADMIN_HOSTS`, port 9644), while `rpk topic ...`
talks to the **Kafka API** (`REDPANDA_BROKERS`, port 9092) — and neither
accepts a `--brokers` flag directly, both need `-X admin.hosts=...` /
`-X brokers=...`. Get this wrong and it doesn't error loudly: it just
retries the health check forever without ever reporting why.
## In the full stack
The root-level `docker-compose.yml` builds this directory's `Dockerfile`
+14 -4
View File
@@ -6,18 +6,28 @@
# sibling container on the root compose's network, or in CI.
set -euo pipefail
# NOTE: confirmed by actually running this against a live Redpanda
# container -- neither `rpk cluster health` nor `rpk topic ...` accept a
# `--brokers` flag in this rpk version. Health checks hit the Admin API
# (-X admin.hosts=..., port 9644); topic commands hit the Kafka API
# (-X brokers=..., port 9092). Getting this wrong doesn't error loudly:
# `rpk cluster health --brokers ...` fails with "unknown flag" but that
# failure was swallowed by this script's own `> /dev/null 2>&1` retry
# loop, which just silently retried the malformed command forever instead
# of ever becoming healthy.
BROKERS="${REDPANDA_BROKERS:-localhost:9092}"
ADMIN_HOSTS="${REDPANDA_ADMIN_HOSTS:-localhost:9644}"
TOPIC="${REDPANDA_TOPIC:-sentry.logs.raw}"
PARTITIONS="${REDPANDA_TOPIC_PARTITIONS:-6}"
echo "Waiting for Redpanda at ${BROKERS}..."
until rpk cluster health --brokers "${BROKERS}" --exit-when-healthy > /dev/null 2>&1; do
echo "Waiting for Redpanda admin API at ${ADMIN_HOSTS}..."
until rpk cluster health -X "admin.hosts=${ADMIN_HOSTS}" --exit-when-healthy > /dev/null 2>&1; do
sleep 1
done
if rpk topic list --brokers "${BROKERS}" | awk 'NR>1{print $1}' | grep -qx "${TOPIC}"; then
if rpk topic list -X "brokers=${BROKERS}" | awk 'NR>1{print $1}' | grep -qx "${TOPIC}"; then
echo "Topic '${TOPIC}' already exists, skipping."
else
echo "Creating topic '${TOPIC}' (${PARTITIONS} partitions)..."
rpk topic create "${TOPIC}" --brokers "${BROKERS}" --partitions "${PARTITIONS}" --replicas 1
rpk topic create "${TOPIC}" -X "brokers=${BROKERS}" --partitions "${PARTITIONS}" --replicas 1
fi