Fix two preflight defects a live production rehearsal found

Ran the rehearsal read-only against a live production instance. It
completed, and two preflight checks were wrong in ways a test instance
could not have shown.

1. Store-backend detection missed the config entirely. A config generated
   by `stalwart --init` declares `type = "rocksdb"` inside a
   `[store.rocksdb]` section, which is what every test fixture here used.
   The production config has no section headers at all and declares
   `store.rocksdb.type = "rocksdb"` flat. Detection only matched a bare
   `type` key, so it reported "no known store backend type found in config"
   and left topology.store_backend empty.

   That mattered more than a warning suggests: backup.Run treated an
   unrecognized backend as a *skip*, so a real run would have continued
   with no filesystem or database backup at all - the single artifact that
   phase exists to produce, quietly absent. Flat dotted keys are now
   detected, and an unrecognized backend is a hard failure rather than a
   skip.

2. The cluster warning didn't say where it matched. It fires on any
   occurrence of "cluster" anywhere in the config, which is the correct
   bias - a missed cluster corrupts a shared store - but on the production
   config the only match was inside the value of an unrelated setting,
   leaving a whole config to search to establish that. It now names the
   location and distinguishes a match in the setting name from one in its
   value.

Both verified against the real config: store-backend now reports
"rocksdb (store.rocksdb)", and the cluster warning names the setting,
making a false positive dismissible at a glance.

No production data is in this commit: the fixtures use example.com and the
flat-key shape only. Coverage numbers from that run matched the earlier
scrubbed-corpus measurement exactly.
This commit is contained in:
2026-08-23 21:58:53 -07:00
parent 1684c88877
commit a69f0bbff1
7 changed files with 199 additions and 13 deletions
+12 -6
View File
@@ -166,15 +166,21 @@ func (c *Checker) Run(ctx context.Context, store *checkpoint.Store, rs *checkpoi
}
if _, err := runCheck("cluster-config", func() (CheckResult, string) {
clustered, err := LooksClustered(c.opts.ConfigPath)
mentions, err := ClusterMentions(c.opts.ConfigPath)
if err != nil {
return CheckResult{Status: StatusFail, Detail: err.Error()}, ""
}
if clustered {
return CheckResult{
Status: StatusWarn,
Detail: "config mentions clustering - confirm every peer node is stopped before this run proceeds; the tool does not verify this for you",
}, ""
if len(mentions) > 0 {
shown := mentions
if len(shown) > 5 {
shown = shown[:5]
}
detail := fmt.Sprintf("config mentions clustering at %s", strings.Join(shown, ", "))
if len(mentions) > len(shown) {
detail += fmt.Sprintf(" (and %d more)", len(mentions)-len(shown))
}
detail += " - confirm every peer node is stopped before this run proceeds; the tool does not verify this for you"
return CheckResult{Status: StatusWarn, Detail: detail}, ""
}
return CheckResult{Status: StatusOK, Detail: "no cluster configuration detected"}, ""
}); err != nil {