Fix two bugs found by actually running the Phase 0 pipeline end-to-end
Both surfaced only by running docker compose up for real, not from review: - ClickHouse's official image silently disables network access for the default user unless CLICKHOUSE_USER or CLICKHOUSE_PASSWORD is set to a genuinely non-empty value (an explicit empty password still triggers it). Set a dev-only password across clickhouse, clickhouse-migrate, ingest, and api in both docker-compose.yml files. - rpk cluster health and rpk topic ... don't accept --brokers; health checks need -X admin.hosts=... (port 9644), topic commands need -X brokers=... (port 9092). The old script's retry loop silently swallowed the resulting "unknown flag" error and retried forever, which blocked ingest from ever starting. Verified: agent -> ingest -> Redpanda -> ClickHouse -> api round-trip confirmed with a real log line on a real host.
This commit is contained in:
+9
-1
@@ -15,9 +15,17 @@ both are invoked.
|
||||
|
||||
```sh
|
||||
docker compose up -d
|
||||
REDPANDA_BROKERS=localhost:9092 ./provision-topics.sh
|
||||
./provision-topics.sh # defaults (localhost:9092 / localhost:9644) match this compose file
|
||||
```
|
||||
|
||||
Two separate addresses matter here, confirmed by actually running this
|
||||
against a live Redpanda container: `rpk cluster health` talks to the
|
||||
**Admin API** (`REDPANDA_ADMIN_HOSTS`, port 9644), while `rpk topic ...`
|
||||
talks to the **Kafka API** (`REDPANDA_BROKERS`, port 9092) — and neither
|
||||
accepts a `--brokers` flag directly, both need `-X admin.hosts=...` /
|
||||
`-X brokers=...`. Get this wrong and it doesn't error loudly: it just
|
||||
retries the health check forever without ever reporting why.
|
||||
|
||||
## In the full stack
|
||||
|
||||
The root-level `docker-compose.yml` builds this directory's `Dockerfile`
|
||||
|
||||
@@ -6,18 +6,28 @@
|
||||
# sibling container on the root compose's network, or in CI.
|
||||
set -euo pipefail
|
||||
|
||||
# NOTE: confirmed by actually running this against a live Redpanda
|
||||
# container -- neither `rpk cluster health` nor `rpk topic ...` accept a
|
||||
# `--brokers` flag in this rpk version. Health checks hit the Admin API
|
||||
# (-X admin.hosts=..., port 9644); topic commands hit the Kafka API
|
||||
# (-X brokers=..., port 9092). Getting this wrong doesn't error loudly:
|
||||
# `rpk cluster health --brokers ...` fails with "unknown flag" but that
|
||||
# failure was swallowed by this script's own `> /dev/null 2>&1` retry
|
||||
# loop, which just silently retried the malformed command forever instead
|
||||
# of ever becoming healthy.
|
||||
BROKERS="${REDPANDA_BROKERS:-localhost:9092}"
|
||||
ADMIN_HOSTS="${REDPANDA_ADMIN_HOSTS:-localhost:9644}"
|
||||
TOPIC="${REDPANDA_TOPIC:-sentry.logs.raw}"
|
||||
PARTITIONS="${REDPANDA_TOPIC_PARTITIONS:-6}"
|
||||
|
||||
echo "Waiting for Redpanda at ${BROKERS}..."
|
||||
until rpk cluster health --brokers "${BROKERS}" --exit-when-healthy > /dev/null 2>&1; do
|
||||
echo "Waiting for Redpanda admin API at ${ADMIN_HOSTS}..."
|
||||
until rpk cluster health -X "admin.hosts=${ADMIN_HOSTS}" --exit-when-healthy > /dev/null 2>&1; do
|
||||
sleep 1
|
||||
done
|
||||
|
||||
if rpk topic list --brokers "${BROKERS}" | awk 'NR>1{print $1}' | grep -qx "${TOPIC}"; then
|
||||
if rpk topic list -X "brokers=${BROKERS}" | awk 'NR>1{print $1}' | grep -qx "${TOPIC}"; then
|
||||
echo "Topic '${TOPIC}' already exists, skipping."
|
||||
else
|
||||
echo "Creating topic '${TOPIC}' (${PARTITIONS} partitions)..."
|
||||
rpk topic create "${TOPIC}" --brokers "${BROKERS}" --partitions "${PARTITIONS}" --replicas 1
|
||||
rpk topic create "${TOPIC}" -X "brokers=${BROKERS}" --partitions "${PARTITIONS}" --replicas 1
|
||||
fi
|
||||
|
||||
Reference in New Issue
Block a user