Two changes that arrived together: the cutover phase (ARCHITECTURE.md 4.5) is implemented, and the rollback phase is deleted. Recovery from a failed migration is now explicitly the operator's own snapshot or backup, and out of scope for this tool. internal/cutover implements 4.5 as seven checkpointed steps: verify the staged binary's version, install it, preserve and rewrite the service definition, reload, start, wait for a healthy JMAP session, recalculate quotas. The unit is rewritten in place rather than generated from a template. An operator's unit carries hardening options, limits and dependencies this tool has no business having an opinion about, and regenerating it would silently drop them. It repoints ExecStart (preserving systemd's -@:+! prefix characters and every argument after the executable), updates --config, and strips recovery-mode Environment lines - leaving STALWART_RECOVERY_MODE=1 set would recovery-boot the service on every restart, forever. It refuses on a unit with no ExecStart, and on an Environment line mixing a recovery variable with others: a line it only partly understands is one it must not edit. Quota recalculation is the one step allowed to fail without failing the phase. Its wire format is grounded in Stalwart's x:Task schema reference - Task/set creating one AccountMaintenance per account with maintenanceType recalculateQuota - but the upgrade guide only documents the WebUI path, so two details remain inferred and are called out in stalwartapi/task.go: whether the schema's "read-only" annotation on accountId/maintenanceType means "immutable after creation", and whether a finished task simply leaves the queue (TaskStatus documents Pending/Retry/Failed with no success state). Warning rather than failing is the honest response to that uncertainty, and stale counters are an accounting problem next to calling for a restore of a machine that is otherwise migrated and serving mail. Docker deployments are refused outright: cutting a container over means pulling an image and recreating it, not swapping a binary. On removing rollback. The implementation worked and was tested, and it was removed because restoring bytes correctly is not the hard part. It copied file contents and permissions and verified every restored file against a manifest - and did not preserve ownership. Run as root, as this tool requires, it would have produced a byte-perfect, checksum-verified, root-owned data directory that Stalwart, running as its own user, could not open, and it would have reported success. The PostgreSQL path was worse: pg_dump without --clean emits CREATE TABLE + COPY, which fails replaying into a database whose tables still exist, and the ON_ERROR_STOP=1 added so a half-applied restore couldn't be reported as success turned that into a hard failure. None of it had ever run against a real server. A filesystem snapshot has none of these failure modes, because it never lost the metadata to begin with. So cutover's gate is no longer rollback.CanRollBack but an explicit RecoveryPointConfirmed acknowledgement. That is an assertion, not a check - this tool cannot verify someone else's snapshot - and its only value is that nobody migrates a production mail server having never been asked the question. Two consequences are accepted deliberately: restoring any pre-migration recovery point discards mail delivered since, and a failed migration now stops and reports rather than undoing itself. What the tool still does to make a manual restore easier: the old binary is preserved and never deleted, the original service definition is preserved before the rewrite, the settings and principals dumps stay on disk, and every artifact path and checksum stays in the checkpoint where `status <run-id>` can print it. Also removed: the `confirm` command stub and RollbackWindowClosed, whose only purpose was closing a rollback window that no longer exists, and checkpoint.PhaseRollback. Old state.json files still load - JSON ignores the now-unknown field. Still open, and recorded in 8: cutover ignores systemd drop-ins, so an ExecStart or Environment override in stalwart.service.d/*.conf is invisible to the rewrite - including the recovery variable it exists to strip; nothing prevents concurrent runs on the same run-id; and nothing in this repo has ever run against a real Stalwart, real systemd, or a real store.
38 KiB
stalwart-migrator — Architecture
Status: design, no implementation yet. Scope: upgrade a Stalwart Mail Server in place from 0.15.5 to the current latest release (0.16.14 as of 2026-08-19) with no data loss, a working a recovery point the operator provides, and an automated post-migration validation pass.
1. Why this isn't a thin wrapper
Stalwart does not ship an automated upgrade tool today (planned for 1.0, targeted H1 2026, not yet released). 0.15.5 → 0.16.x is a major boundary, not a patch bump, and it is unusually dangerous to automate naively:
- The v0.15 → v0.16 config model changes completely: multiple TOML files plus
DB-resident settings collapse into one
config.jsonthat describes only the datastore connection, with everything else moved into JMAP-managed objects. - Account names change from bare usernames to full email addresses; DAV URLs
change (
/dav/cal/alice→/dav/cal/alice%40example.com). - On first v0.16 start, the server irreversibly deletes all directory records (users/groups/domains/tenants/OAuth clients), all settings, DMARC/ TLS/ARF reports, pending tasks, telemetry, spam training samples, and quota counters. Mail/calendar/contact data is untouched, but everything else is gone unless captured first.
- Migration requires a manual "recovery mode" boot of the new binary, then an
external tool (
stalwart-cli apply) replays a converted settings snapshot into it over HTTP while it's up in that special mode — a multi-process, multi-terminal, stateful procedure with no built-in resumability. - In a cluster, every node must be stopped before migration starts; one node left on v0.15 corrupts the shared store.
- Real-world failure mode already reported in the wild: post-migration WebUI
login breaks because the UI now requires HTTPS via
defaultHostname, not plain IP access — a config/DNS issue, not a data issue, but it reads as "the migration broke everything" to an operator. - Not all settings migrate automatically: SMTP listeners, routing, rate limits, spam rules, and auth backends are explicitly not carried over by Stalwart's own conversion script and must be recreated or replayed from a separately captured snapshot.
None of this is exotic — it's exactly what Stalwart's own
UPGRADING/v0_16.md
and resources/scripts/migrate_v016.py already do. This project's job is to
turn that fragile, manual, two-terminal runbook into a single supervised,
checkpointed, reversible operation — and to keep working as new releases land
on top of 0.16.x, most of which (0.16.1–0.16.14, per changelog) are pure
patch/feature releases with no schema migration, i.e. a binary swap +
smoke test, not a full migration.
2. Design goals / non-goals
Goals
- Zero data loss for mail, calendar, and contact content (the one thing Stalwart itself guarantees is untouched — everything else is on us).
- Nothing destructive happens until the operator has confirmed a recovery point exists. This tool does not implement the undo (see the non-goals and §4.8); it refuses to start without being told one is in place.
- Fully automated happy path; the operator answers a preflight confirmation once, then watches (or walks away and checks the report).
- Resumable: if the process dies mid-migration (crash, SSH drop, OOM), a re-run picks up from the last completed checkpoint instead of redoing or, worse, double-applying destructive steps.
- Works across the deployment shapes Stalwart actually supports: systemd + bare binary, Docker/Compose, and single-node vs. cluster — with embedded (RocksDB/SQLite) or external (PostgreSQL/MySQL/FoundationDB) stores.
- Extensible to future major boundaries (0.16 → 1.0 and beyond) without a rewrite: version-boundary logic is pluggable, not hardcoded into the core engine.
Non-goals
- Not a recovery tool. Restoring a failed migration is the operator's own snapshot or backup, by whatever method they already trust — ZFS/LVM/ btrfs snapshots, VM or volume snapshots, or a restorable backup. This tool does not take one, verify one, or restore from one. §4.8 explains why that turned out to be the right split.
- Not a general Stalwart config management tool (no drift detection, no day-2 ops beyond the migration window).
- Not a replacement for routine backups — it produces a migration-time
backup as a side effect, but ongoing backup policy is the operator's job
(Stalwart's own guidance: import/export is explicitly not a backup
substitute;
Vandelayper-account export is the documented backup tool). - Not a cross-major-version skip tool. If the source is older than 0.15.x, the tool requires stepping to 0.15.x first (this matches Stalwart's own stated constraint — see UPGRADING notes).
- No support for editing mail content during migration (no format conversion beyond what Stalwart's own store migration does).
3. High-level flow
┌─────────────┐ ┌───────────┐ ┌────────────┐ ┌───────────────┐ ┌────────────┐ ┌────────────┐
│ PREFLIGHT │──▶│ BACKUP │──▶│ STAGE NEW │──▶│ RECOVERY-MODE │──▶│ CUTOVER │──▶│ VALIDATE │
│ (checks, │ │ (defense │ │ BINARY + │ │ MIGRATE │ │ (swap, up, │ │ (functional│
│ dry-run) │ │ in depth) │ │ config │ │ (apply plan) │ │ smoke) │ │ + counts) │
└─────────────┘ └───────────┘ └────────────┘ └───────────────┘ └────────────┘ └────────────┘
│ │ │ │ │ │
└─────────────────┴────────────────┴─── on failure ──┴─────────────────┴──▶ STOP + REPORT
(operator restores
their own snapshot)
Each box is a phase; each phase is a sequence of idempotent, checkpointed steps. State is persisted to disk after every step (§5), so the whole pipeline can be killed and re-invoked safely.
For a pure patch bump within 0.16.x (no schema change per Stalwart's changelog through 0.16.14), the plan collapses to: PREFLIGHT → BACKUP → STAGE → CUTOVER → VALIDATE, skipping the recovery-mode phase entirely (see §4.6).
4. Phases
4.1 Preflight
Read-only. Aborts before touching anything if a hard blocker is found; warns and asks for confirmation on soft blockers.
- Detect current Stalwart version (
stalwart --version, or JMAPCore/echo/session endpoint if remote). - Refuse to run if current version is outside the tool's supported starting range (must be ≥0.15.0; older installs are told to upgrade to 0.15.x first, per Stalwart's own guidance).
- Detect topology: systemd unit vs. Docker container vs. Compose, single node vs. cluster member count (via config/cluster settings), reachable peer nodes.
- Cluster gate: refuse to proceed unless every node in the cluster is confirmed stopped (mirrors the documented hard requirement — one live v0.15 node during migration corrupts the store).
- Detect store backend(s): RocksDB, SQLite, FoundationDB, PostgreSQL, MySQL, plus configured blob store (local FS / S3-compatible) and FTS backend (native / Elasticsearch).
- Disk space check: require free space ≥ N× current data directory size (embedded stores need a full copy for the backup step; default threshold configurable, hard-fail below a safety floor).
- Resolve and download the target binary/image, verify checksum/signature against the published release.
- Fetch and pin the exact upstream
migrate_v016.py(or its 0.16-successor equivalent) revision, hash it, vendor the hash into the run's checkpoint record — we depend on it as an external, versioned dependency, not a static local copy that can silently drift from upstream. - Dry-run the settings dump against the live server (read-only JMAP calls) to confirm admin credentials and API reachability before anything destructive is scheduled.
- Snapshot pre-migration facts used later for validation: account count,
per-account mailbox message counts (IMAP
STATUS), domain list, DKIM key fingerprints, TLS cert fingerprints, listener port list. Stored alongside the checkpoint, not derived after the fact. - Emit a plain-language plan summary and require explicit confirmation
(
--yesto skip interactively, but never by default).
4.2 Backup — defense in depth
No single backup mechanism is trusted alone, because the risk profile is different at each layer:
- Filesystem/DB snapshot (infra-level, fast, whole-store):
- Embedded (RocksDB/SQLite): stop-the-world
cp -aof the data directory (or LVM/ZFS snapshot if available — preferred, since it doesn't require the copy to finish before the next step) to a sibling path (<datadir>.v0155-backup), never overwriting source. - External SQL (Postgres/MySQL): targeted dump of the critical table
set Stalwart's own guide calls out (
s d r h b g j f ufor Postgres; equivalent for MySQL), not a full-instance dump — matches the documented, tested restore path and stays fast on large installs. - FoundationDB:
fdbbackupagainst the configured cluster.
- Embedded (RocksDB/SQLite): stop-the-world
- Settings/principals export (the
migrate_v016.py dumpstep): captured during preflight and re-captured immediately before cutover, so the export used for the apply reflects the last-known-good state, not a stale preflight snapshot if time has passed. - Per-account content export (Vandelay/JMAP): for installations under
an operator-configurable account-count threshold, take a belt-and-suspenders
full
vandelay import(i.e. export-to-file) of every account into self-contained per-account SQLite archives. This is independent of storage backend and of the in-place migration path entirely — if everything else somehow goes wrong, mail content is recoverable via Stalwart's own documented import path into a clean instance. Skipped above the threshold by default (time cost), but available as--full-content-backupregardless of size. - Binary preservation: old binary is moved aside (
stalwart.v0155), never deleted, so putting the machine back by hand doesn't depend on re-downloading a specific old release under pressure.
Every backup artifact is checksummed and the checksum recorded in the checkpoint file. Before moving past this phase, the tool verifies the filesystem backup by opening it read-only with the old binary in a throwaway temp directory and confirming it reports the expected version and a sane account count — catching a corrupt or partial copy while the pre-migration instance is still up, rather than after it isn't.
4.3 Stage
- Install target binary alongside the old one (never overwrite in place).
- Run
migrate_v016.py convertagainst the fresh dump to produceconfig.json+export.json, applying path rewrites for Docker/volume layouts detected in preflight. - Additionally generate an apply-plan for the settings Stalwart's script
does not carry over — SMTP listeners, routing rules, rate limits, spam
rules, auth backend config — by diffing the old effective config against
the new schema and emitting a best-effort JMAP object set for
stalwart-cli apply. This is flagged clearly as best-effort and included in the final report for manual review; it's the one part of the documented procedure that's explicitly manual today, and silently getting it wrong (rather than flagging it) would be worse than not attempting it. - Stage new systemd unit / Compose file changes without activating them.
4.4 Recovery-mode migration
This is the phase most exposed to partial-failure — it drives an external
process (the new Stalwart binary) through an undocumented-duration startup,
then drives a second external process (stalwart-cli apply) against it over
HTTP. Both are supervised with explicit timeouts and health polling, not
fire-and-forget:
- Stop the old service.
- Start the new binary in the foreground with
STALWART_RECOVERY_MODE=1and a freshly generated one-timeSTALWART_RECOVERY_ADMINcredential (random, never the operator's real password, never logged). - Poll the recovery HTTP endpoint until healthy or a timeout elapses; on timeout, capture logs and stop rather than hanging indefinitely.
- Run
stalwart-cli apply --file export.json, then the generated best-effort settings plan from §4.3, capturing full output. - Verify the apply reported success for every object (the tool parses the apply-tool's structured output rather than trusting exit code alone — partial application with a zero exit code is exactly the kind of silent failure this tool exists to catch).
- Stop recovery mode cleanly (SIGTERM, not SIGKILL, to let it flush).
Checkpointed after each numbered step, so a crash between "apply succeeded" and "recovery mode stopped" resumes at step 6 instead of re-running apply against an already-migrated store.
4.5 Cutover
- Update the real systemd unit / Compose config to point at the new binary
and config, without the recovery env vars (leaving
STALWART_RECOVERY_MODE=1set is a documented footgun — it would recovery- boot on every restart). - Start the service normally.
- Wait for healthy JMAP session response.
- Trigger disk-quota (and tenant-quota, if multi-tenant) recalculation via the management API, and poll the task queue until it completes rather than firing and moving on.
Status: implemented (internal/cutover), but nothing calls it yet — see
§8. Notes on how it turned out:
- It refuses to run at all unless the operator has confirmed a recovery point exists (§4.8). That's an acknowledgement, not a check — this tool can't verify someone else's snapshot — but it makes the irreversibility of this phase impossible to walk into unasked.
- The unit is rewritten in place, not generated from a template: an
operator's unit carries hardening options, limits and dependencies this
tool has no business having an opinion about, and regenerating it would
silently drop them. It repoints
ExecStart(preserving systemd's-@:+!prefix characters and every argument after the executable), updates--configif asked, and strips recovery-modeEnvironment=lines. It refuses on a unit with noExecStart, and on anEnvironment=line that mixes a recovery variable with others — a line it only partly understands is one it must not edit. - The original unit is preserved and recorded as the
service-unitartifact before the rewrite, so an operator restoring by hand isn't reconstructing a unit file from memory. - Docker deployments are refused: cutting over a container means pulling a new image and recreating the container, not swapping a binary and rewriting a unit.
- Quota recalculation is the one step allowed to fail without failing the phase. Stale counters are an accounting problem; a failed cutover is one an operator has to respond to by restoring a machine that is otherwise migrated and serving mail correctly. Calling for that over a counter would be the worse outcome, so it warns and points at the WebUI's Tasks panel.
The quota call itself is grounded in Stalwart's x:Task schema reference
(docs/ref/object/task/), not guessed: x:Task/set creating one
AccountMaintenance variant per account with maintenanceType: "recalculateQuota", exactly as the WebUI's own "Recalculate disk quotas"
fans out. The upgrade guide only documents the WebUI path, so two details
remain unconfirmed against a live server and are called out in
internal/stalwartapi/task.go: whether the schema's "read-only" annotation
on accountId/maintenanceType means "immutable after creation" (it has
to, or the variant couldn't be created), and whether a finished task simply
leaves the queue (TaskStatus documents Pending/Retry/Failed with no
success state). That uncertainty is the reason this step warns rather than
fails.
4.6 Patch-bump fast path
For an already-0.16.x install moving to a newer 0.16.x patch (the common case after the initial major migration, and per the changelog the case for every release from 0.16.1 through 0.16.14 so far): preflight confirms no schema-migration flag is set for the target version, and the plan skips §4.4 entirely — binary swap, restart, same validation suite as §4.7. This is intentionally the same engine with a shorter plan, not a separate code path, so it doesn't rot independently.
4.7 Post-migration validation
Runs automatically after cutover; failure here stops the run, reports loudly, and exits non-zero, leaving the operator to decide what to restore (§4.8).
- Version check: reported server version matches the target exactly.
- Auth check: WebUI login succeeds over the configured hostname via HTTPS (not bare IP) — this directly targets the real-world post-0.16 login failure mode found in the field.
- Protocol reachability: JMAP session, IMAP, SMTP (submission + MTA), POP3, ManageSieve, CalDAV/CardDAV endpoints all accept a handshake on their configured ports.
- Directory integrity: account/domain/group counts match the preflight snapshot exactly (accounting for the bare-username → email-address rewrite, which the tool resolves by comparing normalized identities, not raw strings).
- Content integrity — the core no-data-loss check: per-account IMAP
STATUS (MESSAGES)compared against the preflight snapshot for every mailbox of every account (or a statistically sampled subset above a configurable account-count threshold, with the full sweep always available via--full-validation). Any mismatch is a hard failure. - DKIM/TLS check: key fingerprints and cert validity match or are intentionally rotated (new-key generation is an expected v0.16 behavior, not a bug — the check distinguishes "changed as documented" from "missing").
- DNS check: for domains under Stalwart's automatic DNS management (new in 0.16), diff expected vs. actual published records and flag drift rather than assume the automation ran correctly.
- Mail-flow smoke test: send one real message through SMTP submission to a dedicated canary mailbox and confirm it's retrievable via IMAP within a timeout — the one end-to-end check that nothing upstream can fake.
- Quota check: recalculation task (§4.5) completed and reported numbers are non-zero/sane where preflight showed non-zero usage.
Output is a single structured report (JSON + human summary): pass/fail per check, with enough detail to hand to the operator deciding whether to restore.
4.8 Recovery from a failed migration — out of scope
This tool does not undo a migration. Recovery is the operator's own snapshot or backup, taken by whatever method they already trust and know how to restore: ZFS/LVM/btrfs snapshots, a VM or volume snapshot, or a restorable backup. This tool does not take one, verify one, or restore from one. Cutover refuses to start until the operator confirms one exists (§4.5).
This replaced a working, tested rollback implementation, and the reasoning is worth recording because the deleted code looked good:
- Restoring bytes correctly is not the hard part; restoring everything else is. The implementation copied file contents and permissions and verified every restored file against a manifest — and did not preserve ownership. Run as root, it produced a byte-perfect, checksum-verified, root-owned data directory that Stalwart, running as its own user, could not open. It would have reported success. A filesystem snapshot has no such failure mode, because it never lost the metadata in the first place.
- The external-database path was worse.
pg_dumpwithout--cleanemitsCREATE TABLE+COPY; replaying that into a database whose tables still exist fails outright, andON_ERROR_STOP=1— added so a half-applied restore couldn't be reported as success — turned that into a hard failure. The two SQL paths were asymmetric and only one was plausibly correct. - It was never exercised against anything real. Every test drove fake
systemctl,psqlandstalwartbinaries. That's sound for logic and ordering, and it is not evidence about production. - Snapshots are already in the operator's runbook. They are atomic, metadata-preserving, cheap with copy-on-write, and cover the whole system — binary, unit file, config, data — rather than the subset one tool thought to capture.
What this tool keeps doing, so a manual restore is as easy as possible:
- The old binary is preserved next to the new one (§4.2), never deleted.
- The original service definition is preserved as a
service-unitartifact before cutover rewrites it, so the operator doesn't reconstruct a unit file from memory. - Every artifact path and checksum stays in the checkpoint, and
status <run-id>prints exactly which steps completed and which failed.
The mail-delivery gap is accepted. Restoring any pre-migration recovery point discards mail delivered since it was taken. This was equally true of the rollback implementation, is inherent to restoring a point in time, and is not something this tool can solve. Plan the migration window accordingly.
Two consequences worth being explicit about. First, the confirmation cutover requires is an assertion, not a check — an unverifiable promise is weaker than a guarantee, and the value is only that nobody migrates a production mail server having never been asked the question. Second, there is no longer an automatic response to a failed migration: a failure stops the run and reports, and a human decides what to restore. Both are deliberate trades for not shipping a recovery path that has never been tested against a real server.
4.9 Dry run
stalwart-migrate run --dry-run runs the real migration mechanics against a
disposable sandbox clone of the data, so an operator can get genuine
confidence before committing to a real cutover — not a simulation that
skips the fragile parts, the actual recovery-mode migration (§4.4) and a
post-migration boot check, just pointed somewhere disposable:
- Preflight (§4.1) runs for real, read-only, against the live instance.
- Backup (§4.2) runs for real too, with one exception:
SkipBinaryPreservationis set, so the production binary at the real install path is never moved aside. Taking a consistent filesystem snapshot of an embedded store still means the live service should be stopped first (the same requirement Stalwart's own export tooling has) — this tool doesn't automate that stop/start today (no systemd/Docker control exists yet), so a dry-run without a manual stop first is a best-effort snapshot of a live, in-use store, and the CLI says so. - Convert:
migrate_v016.py convertturns the settings/principals dump intoconfig.json+export.json, using the script's own documented--patch-paths <old>=<new>flag to point the generated config at the sandbox data directory instead of the real one. This is the officially documented mechanism for exactly this kind of path redirection — the tool deliberately does not try to rewriteconfig.json's contents itself, since depending on its exact schema (which has already changed once, 0.15 → 0.16) is a correctness risk this tool avoids wherever an official alternative exists. - The verified backup copy is cloned again into the sandbox directory (never reusing the same directory recovery mode is about to mutate as the one a manual restore would use).
- Recovery-mode migration (§4.4) runs for real against the sandbox:
the actual target binary, actual
STALWART_RECOVERY_MODE=1boot, actualstalwart-cli apply. - Boot check + content integrity: the migrated sandbox is started once
more as an ordinary boot (no recovery-mode env vars) and polled until its
HTTP listener answers, confirming the migrated store doesn't just accept
a settings apply but actually comes up cleanly afterward. If preflight
captured a pre-migration snapshot (§4.1, requires
--admin-url), the same boot is then used to capture a fresh post-migration snapshot and compare the two — this is the actual no-data-loss guarantee, not just "the mechanics ran": every account and mailbox from before must still be found afterward (matching by exact name, falling back to the part before@since v0.16's own migration rewrites bare usernames to full email addresses) with an identical message count. A mismatch or a missing account fails the check. This covers the message-count half of §4.7's full suite; DKIM/TLS fingerprint checks and a live mail-flow SMTP→IMAP smoke test are still open. - Every byte written by steps 2–6 (the fs-backup copy, settings/principals
dumps, downloaded
migrate_v016.py, sandbox clone, and generatedconfig.json/export.json) lives under one per-run directory (work-dir/<run-id>), which is removed on every exit path - success, a failed check partway through, or an early refusal - via a deferred cleanup, not just the happy path. The only thing left behind afterward is the checkpoint'sstate.jsonunder--state-dir: a small structured success/failure log (which check failed and why), not bulk data.--keep-artifactsopts out for inspecting a failure. Nothing at the real binary path, the real service, or the real data directory's contents is ever mutated by steps 2–6 in the first place.
A same-boundary patch bump (§4.6) has no recovery phase to simulate — dry run for that plan is just preflight + backup.
5. State machine / checkpointing
Every run gets a run-id and a checkpoint file
(/var/lib/stalwart-migrator/runs/<run-id>/state.json) written after each
step completes, containing: run-id, source/target version, current
phase/step, timestamps, artifact paths + checksums, and the preflight
snapshot facts used by validation. Steps are pure functions of
(checkpoint-state → new-state); re-invoking stalwart-migrate run with an
in-progress run-id resumes at the first incomplete step. Steps are written
to be safe to re-run if they were interrupted mid-execution (e.g. the
filesystem copy step checks for and resumes/redoes a partial copy rather
than trusting a checkpoint that says "started" as if it meant "done").
This is the same shape as a deployment pipeline's state file, deliberately — the risk profile (long-running, multi-process, must survive being killed) is the same problem.
6. CLI surface
stalwart-migrate preflight [--config PATH] ... # read-only, prints the report
stalwart-migrate run --dry-run [--target-binary PATH] ... # implemented — see §4.9
[--keep-artifacts]
(without --dry-run: refused today — see §8)
stalwart-migrate status [run-id] # implemented
stalwart-migrate report <run-id> [--json] # not yet implemented
run is the only command that mutates anything, and it always starts with
preflight. Nothing in this tool restores a failed migration (§4.8), so there
is no rollback command, and no confirm step to close a rollback window
that no longer exists.
The migration-time artifacts a run leaves behind — the preserved old binary,
the settings and principals dumps, the preserved service definition, and
(for the dry-run path) the filesystem copy — are never pruned automatically.
They're small next to the data directory, they're what a manual restore
reaches for first, and deleting them on a schedule to reclaim disk would be
the tool making a call that isn't its to make. Flags shown here are the
design intent; run stalwart-migrate <command> -h for the actual current
flag set.
7. Project layout (Go, matches this workspace's other CLI tools)
stalwart-migrator/
cmd/stalwart-migrate/ main.go, preflight.go, run.go, status.go — CLI entry + wiring
internal/plan/ version-boundary → ordered phase list (§4.6) [done]
internal/checkpoint/ run-id, state.json read/write, resume logic (§5) [done]
internal/preflight/ §4.1 checks [done]
internal/backup/ §4.2 — fs/db snapshot, settings dump+convert, Vandelay export [done]
internal/recovery/ §4.4 — recovery-mode process supervision + apply [done]
internal/cutover/ §4.5 — binary swap, unit rewrite, restart, quota rebuild [done, unwired]
internal/validate/ §4.7 — boot-check + content-integrity done; DKIM/TLS + mail-flow not yet [partial]
internal/service/ systemd/Docker start+stop, used by §4.5 [done]
internal/stalwartapi/ Ping, AccountSnapshot (per-mailbox counts via impersonation), quota tasks [done]
internal/config/ tool's own config (paths, thresholds, credentials handling) [not started]
docs/ this file + phase-specific notes as they get built out
There's no separate internal/stage package: the convert half of
migrate_v016.py lives in internal/backup next to dump (same script,
same invocation pattern), and the dry-run sandbox-cloning logic that stands
in for the rest of §4.3 currently lives directly in cmd/run.go rather than
its own package. Now that internal/cutover exists, that's the code a real
staging phase would be generalized out of.
There's no internal/rollback either, and that's a deliberate removal
rather than a gap — see §4.8.
internal/stalwartapi is deliberately the only thing that speaks JMAP/HTTP
to Stalwart — every other package depends on it, not on net/http directly,
so auth handling and retry/backoff live in one place. internal/service is
the same idea for the other external surface: it is the only thing that
shells out to systemctl or docker, so the commands that can take mail
delivery down sit in one auditable file rather than in each phase that
happens to need them. preflight.DeploymentKind is a type alias for
service.Kind, so detection and control can't drift apart.
8. Open questions for the next pass
- Credential handling: recovery-mode admin password and any stored JMAP credentials need a real secrets story (env var pass-through is fine for v1, but the checkpoint file must never contain them in plaintext).
- Cluster orchestration: §4.1's cluster gate assumes the operator stops other nodes manually; a v2 could SSH-coordinate that instead. Out of scope for v1.
migrate_v016.pydependency: pinning by hash is a start, but the script is Stalwart's, not ours — need a policy for what happens when it changes upstream (re-vendor + re-test before bumping the pin, never silently float tomain).- Best-effort settings apply-plan (§4.3): needs real-world testing
against a variety of existing SMTP/routing/spam configs before it's
trusted un-reviewed; v1 should probably always require operator sign-off
on that specific generated plan even with
--yesset for everything else. Not started — dry-run currently only replays whatmigrate_v016.pyitself converts. - Account/mailbox enumeration (
stalwartapi.Client.AccountSnapshot): implemented, including per-mailbox message counts. Account count and domains come fromx:Account/query+x:Account/getagainst Stalwart's management API (/api, capabilityurn:stalwart:jmap), confirmed againstcrates/jmap/src/principal/{get,query}.rsanddocs/ref/object/account.md. Per-mailbox counts needed a second research pass, because a superuser's own JMAP session does not implicitly grant cross-account access — confirmed by readingcrates/jmap/src/api/session.rs: the session'saccountsmap is built solely from the authenticated identity's own membership/sharing grants, unaffected by any admin flag. The real, documented mechanism is Stalwart'simpersonatepermission (docs/auth/authorization/administrator.md): an account holding it can log in as another account via the composite Basic-auth username<target>%<impersonator>, after which standard RFC 8621Mailbox/get(propertytotalEmails) works normally against that impersonated session's ownapiUrl(session-discovered per RFC 8620, not the/apimanagement endpoint — confirmed as a distinct endpoint indocs/ref/object/account.md).AccountSnapshotnow does this per account it finds; a single account's failure (most likely:impersonatenot granted) is recorded inSnapshot.MailboxErrorsrather than failing the whole snapshot, so one misconfigured account doesn't hide a working result for every other one. One resolved false alarm worth recording: an initial pass of this same research, reading Stalwart'smainbranch source directly, reportedx:Accountapparently replaced byx:Principal/x:Quota. Checking the published docs site directly (which has noprincipal.md/quota.mdpage, and still documentsx:Accountwith a working example) showed that was an unreleased/in-development refactor inmain, not the interface the current released version actually exposes — a reminder that "read the source" and "read what's actually shipped" can disagree, and it's worth checking both before changing already-working code on the strength of one. Preflight now populatesRunState.PreflightSnapshot.MailboxCountswhen--admin-urlis set, andvalidate.BootChecknow compares it against a fresh post-migration snapshot as part of the same boot (§4.9 step 6) — proven end-to-end with a live smoke test that deliberately made the "after" instance report fewer messages than the "before" snapshot and confirmed the dry run failed loudly with the exact before/after counts, rather than just trusting that. Still open: preflight/validate always attempt every account serially with no sampling/threshold, which could be slow on a large install —--full-validation's sampling idea from §4.7 hasn't been built yet for this; and DKIM/TLS fingerprint checks plus a live mail-flow SMTP→IMAP smoke test (the rest of §4.7's suite) aren't implemented. With this done, recovery, backup, dry-run, and account/ mailbox snapshotting all work end-to-end, and dry-run's comparison is now the closest thing to §4.7's actual no-data-loss guarantee this tool has — the remaining major gap is §4.3 staging and the production pipeline (below). - Cutover is built; nothing wires it into a production run yet.
internal/cutover(§4.5) andinternal/serviceare implemented and tested.runwithout--dry-runstill refuses, for one remaining reason: §4.3 stage doesn't exist, and neither does the production pipeline that would run preflight → backup → stage → recovery-mode → cutover → validate against real paths instead of a sandbox. What stage still needs: downloading and verifying the target binary into a staging path (preflight.ResolveReleaseandbackup.DownloadFilebetween them already have the pieces), running the convert step against real paths rather than the dry-run's patched sandbox ones, and the best-effort settings apply-plan, which is its own open question below. - Nothing has ever run against a real Stalwart. Every test in this
repo drives fake
systemctl,psqlandstalwartbinaries and httptest servers. That's sound for logic and ordering and is not evidence about production. One smoke test on a throwaway VM - real 0.15.5, real systemd unit, a few accounts with mail - would settle the quota wire format, systemd drop-in handling, and cutover's unit rewrite at once. It should happen before §4.3 is wired, not after. - Quota recalculation is grounded but unproven. The
x:Taskwire format comes from Stalwart's schema reference rather than a live server; §4.5 lists exactly which two details are inferred. A smoke test against a real 0.16 instance would settle both, and would let this step be promoted from "warns on failure" to a hard check. - Cutover doesn't handle Docker. It refuses container deployments outright, since cutting one over means pulling an image and recreating the container rather than swapping a binary and rewriting a unit.
- Cutover ignores systemd drop-ins. It rewrites only the main unit
file, so an
ExecStartorEnvironmentoverride in/etc/systemd/system/stalwart.service.d/*.confis invisible to it - including a recovery-mode variable set there, which is exactly the footgun the rewrite exists to prevent. Drop-ins are common enough that this needs handling before a production run, at minimum by detecting them and refusing. - Nothing prevents concurrent runs. Two invocations against the same run-id would both proceed; there's no lock file or equivalent.
- Dry-run's un-stopped backup snapshot (§4.9 step 2): a dry-run still
backs up a live, in-use store unless the operator stops it manually first.
internal/servicenow makes doing this properly possible - dry-run just hasn't been wired to offer it yet.
Sources
Grounded in Stalwart's own documentation and community reports as of 2026-08-19:
- UPGRADING/v0_16.md — exact migration procedure this tool automates
- UPGRADING/v0_15.md — prior breaking-change boundary
- Database Migration docs
- Backup docs (Vandelay)
- Upgrading guide
- v0.16 blog post
- Discussion #2892 — breaking changes overview
- Discussion #3004 — upgrading Q&A
- Discussion #3025 — real post-migration WebUI login failure
- CHANGELOG.md — confirms 0.16.1–0.16.14 carry no further schema migrations