Fix two defects a production-clone dress rehearsal exposed

Streamed a clone of a production store into the smoke VM - 3.6 GB, 12,361
settings, 6 accounts across 9 domains - and migrated it 0.15.5 -> 0.16.14
with the tool's own phases. The migration succeeded. Two defects surfaced
that no smaller instance could have shown, plus one finding worth recording.

1. Account roles broke on production-shaped names. v0.16 stores an account
   as a local part plus a domain reference: a v0.15 account named
   "[email protected]" becomes name "john" with a domainId. The generator
   passed the full address and the server rejected it outright ("Invalid
   email local part"), failing the apply. The smoke instance used bare
   usernames - alice, bob - and never exercised this.

   Fixed to use the local part. And because local parts are unique only
   within a domain - [email protected] and [email protected] both become
   "postmaster" - an ambiguous one is now refused with a warning rather
   than risking an upsert that grants Admin to the wrong account. Verified
   on the clone: the one admin came out with roles {"@type": "Admin"} and
   the other five accounts untouched.

2. Cutover's health check conflated liveness with credentials. A config
   fallback-admin does not survive the migration - v0.16's config is a
   store pointer, so the old [authentication.fallback-admin] block simply
   ceases to exist - so the credentials supplied for the pre-migration
   instance came back 401 on the migrated one, and the check reported the
   service as never having answered. It had answered; it was up and serving
   on all ten ports. Liveness and credentials are now separate: any
   response proves the service is up, and credentials that stopped working
   are a warning that names this cause.

Also recorded: a failed apply leaves the store in bootstrap mode, where
only Bootstrap objects are accessible. A half-applied plan is not a
partially configured server but an unusable one.

Timing, which is the other reason to rehearse: the recovery-mode conversion
of that 3.6 GB store took 2 seconds. A migration window is dominated by
waiting and verification, not data volume.

No production data in this commit; fixtures use example.net and the shapes
involved.
This commit is contained in:
2026-08-23 22:32:17 -07:00
parent 0d83283caa
commit 2ca9522f9a
6 changed files with 245 additions and 21 deletions
+40
View File
@@ -7,6 +7,7 @@ import (
"context"
"encoding/json"
"fmt"
"net/http"
"sort"
"strings"
"time"
@@ -287,3 +288,42 @@ func (c *Client) WaitForPing(ctx context.Context, timeout time.Duration) error {
}
}
}
// WaitForResponse polls until the instance answers an HTTP request at all,
// whatever the status, or timeout elapses.
//
// This is the liveness question, and it is deliberately separate from
// WaitForPing's "and my credentials work". A 401 proves the server is up,
// listening and routing - which is exactly what a caller waiting for a
// restarted service needs to know. Conflating the two failed a cutover
// that had in fact succeeded: the credentials supplied for the
// pre-migration instance were a config fallback-admin, which does not
// survive into v0.16 (its config is a store pointer, so the old
// [authentication.fallback-admin] block is simply gone), so every poll came
// back 401 and the phase reported the service as never having answered.
func (c *Client) WaitForResponse(ctx context.Context, timeout time.Duration) error {
url := strings.TrimRight(c.BaseURL, "/") + "/.well-known/jmap"
deadline := time.Now().Add(timeout)
var lastErr error
for {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return err
}
req.SetBasicAuth(c.Username, c.Password)
resp, err := c.httpClient().Do(req)
if err == nil {
resp.Body.Close()
return nil
}
lastErr = err
if !time.Now().Before(deadline) {
return fmt.Errorf("stalwartapi: %s did not respond within %s: %w", url, timeout, lastErr)
}
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(500 * time.Millisecond):
}
}
}