Shorten the README; move the technical detail into docs/
The README keeps what the tool is, how to install it and the first commands, and points to the guide on docs.ihasmail.org. Everything else moves, whole, into docs/ and CONTRIBUTING.md, where it is organized for readers who want the detail. Where the old README disagreed with the code, the code wins.
This commit is contained in:
@@ -0,0 +1,86 @@
|
||||
# Docker deployments
|
||||
|
||||
How `run` migrates a Stalwart running in a Docker container: the flags it
|
||||
needs, what it carries across, what it refuses, and what it keeps for a manual
|
||||
restore. For the short version, see the [README](../README.md).
|
||||
|
||||
## The container path is unproven
|
||||
|
||||
**The container path has never completed a migration against a real Stalwart
|
||||
image.** What it inspects and what it assembles have been checked against one,
|
||||
which is how two problems were found and fixed (#11) — but a fake `docker`
|
||||
still proves only that the right commands are assembled, not that the image
|
||||
reads the config it is handed and comes up as the server it was.
|
||||
|
||||
`run` refuses a container deployment unless you pass
|
||||
`--container-path-unproven`, which is there so nobody reaches it without being
|
||||
told. Rehearse on a clone first ([rehearsal.md](rehearsal.md)); that advice
|
||||
goes double here.
|
||||
|
||||
## Flags
|
||||
|
||||
- `--target-image` names the image in full, e.g.
|
||||
`stalwartlabs/stalwart:v0.16.14`. It is never derived from the running
|
||||
container by swapping the tag — that is wrong for a digest-pinned image, a
|
||||
mirror or a fork, and being wrong means pulling the wrong software into a
|
||||
mail server.
|
||||
- `--container` names the container (default `stalwart`).
|
||||
- `--data-dir` must name the path **inside** the container, since that is
|
||||
where its data actually lives. `preflight` says so if it matches none of the
|
||||
container's mounts.
|
||||
|
||||
## What cutover carries across, and what it refuses
|
||||
|
||||
A container cannot be edited in place the way a unit file can, so cutting one
|
||||
over means rebuilding it. A container rebuilt without its capabilities, its
|
||||
custom network or its device mappings starts cleanly and is quietly not the
|
||||
server it was.
|
||||
|
||||
So cutover carries across what it understands and refuses outright when it
|
||||
finds anything else, naming what it found. It asks that question in
|
||||
`preflight`, while the server is still running, rather than only at cutover
|
||||
after it has stopped.
|
||||
|
||||
**It carries:** mounts, ports, environment, restart policy, labels, and
|
||||
anything the container overrides on its image — a `--user`, an
|
||||
`--entrypoint`, a command of your own. What the container merely *inherits*
|
||||
from its old image is left to the new one, whose own defaults are the ones
|
||||
that go with it.
|
||||
|
||||
**It refuses:**
|
||||
|
||||
- a container with settings it doesn't understand, such as extra
|
||||
capabilities, a custom network or device mappings;
|
||||
- a container whose data is not on a volume — an upgrade replaces the
|
||||
container, and the writable layer goes with it;
|
||||
- a container managed by Docker Compose, because recreating it out from under
|
||||
compose leaves the container and the compose file disagreeing about what is
|
||||
deployed, and the next `compose up` reverts the migration. Compose
|
||||
deployments are migrated by editing the image tag in the compose file and
|
||||
running `compose up -d`.
|
||||
|
||||
## Where the converted config goes
|
||||
|
||||
The converted v0.16 config is written into the host side of whichever mount
|
||||
covers `--data-dir`, named on the container side, and the recreated container
|
||||
is started with `--config` pointing at it.
|
||||
|
||||
It cannot go anywhere else: cutover recreates a container with the mounts it
|
||||
had and cannot invent a new one. The official image's own default is
|
||||
`--config /etc/stalwart/config.json`, which is a *different* volume, so a
|
||||
container left to that default would come up on whatever the old version had
|
||||
left there. If your container overrides its command, cutover refuses rather
|
||||
than merging the two — both are the container's argv and there is no honest
|
||||
way to guess.
|
||||
|
||||
## What it keeps
|
||||
|
||||
The old container is renamed rather than removed, the old image is never
|
||||
pruned, and the container's `docker inspect` is preserved as an artifact before
|
||||
anything is replaced. Together those are the manual restore path — see
|
||||
[recovery.md](recovery.md), which applies here exactly as it does to a binary
|
||||
install.
|
||||
|
||||
Measured on a full migration: the store converts in seconds, and the service
|
||||
was down for **6 seconds** end to end. Plan the window around verification,
|
||||
not data volume.
|
||||
@@ -0,0 +1,136 @@
|
||||
# Known Stalwart problems
|
||||
|
||||
Things about Stalwart's 0.15 → 0.16 upgrade that bite, and what to do about
|
||||
each: the two server fixes you must make first, the administrator account,
|
||||
the settings Stalwart's converter drops without saying so, the certificate on
|
||||
the mail ports, and the extra recovery boot some store migrations need. For
|
||||
the short version, see the [README](../README.md).
|
||||
|
||||
## Two things you must fix on the server first
|
||||
|
||||
Neither is something this tool can do for you, and both stop a migration dead.
|
||||
`preflight` refuses on both, while the mail server is still running — but they
|
||||
are worth knowing before you book a maintenance window, because fixing them is
|
||||
a change to your directory, not a flag.
|
||||
|
||||
### Remove or collapse multi-tenancy
|
||||
|
||||
v0.16 requires a tenant-scoped account to sit on a domain owned by that same
|
||||
tenant, for its primary domain and every alias. v0.15 imposed no such rule, so
|
||||
an install that is perfectly valid today can be unrepresentable in v0.16.
|
||||
|
||||
Run `stalwart-migrate tenants` to see who owns what.
|
||||
|
||||
- Where a domain has no tenant of its own and only one tenant's accounts use
|
||||
it, the conversion repairs it for you.
|
||||
- Where two tenants genuinely share a domain, nothing can. Resolve it in v0.15
|
||||
first: give each tenant its own domains, move the accounts into one tenant,
|
||||
or remove the tenants entirely.
|
||||
|
||||
### Migrate as a directory account, not the built-in admin
|
||||
|
||||
A `[authentication.fallback-admin]` from `config.toml` authenticates perfectly
|
||||
well right up to the moment the migration finishes, and then stops existing —
|
||||
v0.16 keeps its configuration in the store, so the block defining it is never
|
||||
read again. The migration itself still succeeds; what you lose is the ability
|
||||
to verify it, recalculate quotas, or administer the server afterwards. The
|
||||
next section has the detail.
|
||||
|
||||
## You need a named admin account before you migrate
|
||||
|
||||
**A config-file fallback admin will not survive the migration.** If the only
|
||||
administrator you have is an `[authentication.fallback-admin]` block in
|
||||
`config.toml` — which is what `stalwart --init` sets up — you will come out of
|
||||
the migration unable to administer the server.
|
||||
|
||||
Create a real account in the directory, with the admin role, and confirm you
|
||||
can log in as it *before* migrating. Four separate reasons, verified against a
|
||||
real 0.15.5 → 0.16.14 migration:
|
||||
|
||||
1. **v0.16 keeps its configuration in the store, not in a file.** After the
|
||||
migration the server is started with a config that is little more than a
|
||||
pointer at the data store, so the old `config.toml` — and the
|
||||
fallback-admin block inside it — is no longer read at all. That credential
|
||||
simply stops existing.
|
||||
2. **`migrate_v016.py` gives every migrated account the `User` role**,
|
||||
whatever it held before. An account that was an administrator in v0.15
|
||||
comes out authenticating normally and refused every management operation.
|
||||
`rehearse` generates the operation that restores it — but the account has
|
||||
to exist in the directory for there to be anything to restore.
|
||||
3. **The account's local part must be unambiguous.** v0.16 identifies an
|
||||
account by local part plus domain, so if `[email protected]` and
|
||||
`[email protected]` both exist, this tool refuses to restore either role
|
||||
rather than risk granting administrator rights to the wrong one. It says so
|
||||
rather than guessing; you then grant it by hand.
|
||||
4. **The role has to survive, not just the account.** `tenant-admin` has no
|
||||
v0.16 equivalent and is not restored, so an account whose rights came only
|
||||
from it authenticates afterwards and is still refused management
|
||||
operations. Preflight cannot check this — it cannot know which roles the
|
||||
converter will carry across — so confirm on a clone, or immediately
|
||||
afterwards, that the account can still administer the server.
|
||||
|
||||
`preflight` refuses to proceed if the account you authenticate with is not in
|
||||
the directory, so this is caught before anything is touched rather than after
|
||||
the migration completes.
|
||||
|
||||
**The practical check:** make sure you can authenticate to the admin API as a
|
||||
directory account — not as the fallback admin — that its local part is unique
|
||||
across your domains, and that it holds admin rights through a role other than
|
||||
`tenant-admin`.
|
||||
|
||||
## Stalwart's own converter silently drops ACME
|
||||
|
||||
`migrate_v016.py` consumes every `acme.*` setting and emits nothing for them.
|
||||
They are **not** reported as unmigrated either, so nothing warns you: the
|
||||
certificate carries over, the provider that renews it does not, and TLS keeps
|
||||
working until the certificate expires roughly ninety days later.
|
||||
|
||||
Check for `acme.*` in your dump before migrating, and recreate an
|
||||
`AcmeProvider` afterwards if there was one. `accountKey` is server-set in
|
||||
v0.16, so the existing ACME account cannot be carried over — the server
|
||||
registers a new one on first issuance.
|
||||
|
||||
This tool does not yet generate that object for you. It should: the
|
||||
supplemental plan already does the equivalent for listeners.
|
||||
[ARCHITECTURE.md](../ARCHITECTURE.md) §4.6a lists what else the converter drops.
|
||||
|
||||
## A certificate that serves HTTPS may not serve the mail ports
|
||||
|
||||
A `Certificate` object carried into v0.16 with the right SAN is picked up by
|
||||
the HTTP listener on its own. **IMAPS, SMTPS and POP3S are not**: they keep
|
||||
serving a self-signed certificate until `defaultCertificateId` is set on
|
||||
`SystemSettings` and the server is restarted.
|
||||
|
||||
This is the kind of thing that looks fine from a browser and surfaces as a
|
||||
mail client complaining days later, so check it as part of your
|
||||
post-migration verification: connect to 993 or 465 and confirm which
|
||||
certificate you are handed, not just to 443.
|
||||
|
||||
This tool does not set it for you. Like the `AcmeProvider` above, it should,
|
||||
and the supplemental plan is where it belongs. Reported by
|
||||
[@kaya-eu](https://github.com/Coffey-Labs/stalwart-migrator/issues/1).
|
||||
|
||||
## The store migration may need one more recovery boot
|
||||
|
||||
**This one is open, and it is the reason to rehearse on a clone.** The
|
||||
recovery cycle boots the target version once, replays your settings into it,
|
||||
and stops. Across three real migrations, two different failures showed up
|
||||
that one extra recovery-mode boot cured:
|
||||
|
||||
- the settings apply failing on its very first object with
|
||||
`primaryKeyViolation`, where re-running the identical apply against a fresh
|
||||
recovery boot went straight through; or
|
||||
- the next normal start panicking with *"Upgrading to version 0.16 is a
|
||||
multi-step process"*, where booting recovery mode once more, letting it come
|
||||
up and stopping it cleanly was enough.
|
||||
|
||||
Never both on the same run — whichever appeared, one more recovery boot
|
||||
before the real start got past it. That panic is Stalwart's own, and it
|
||||
suggests the store migration is not finished when the single boot exits.
|
||||
|
||||
The likely fix is a settle boot after the apply. **It is not in the tool**,
|
||||
because getting an extra recovery boot wrong is its own hazard — see [Do not
|
||||
boot recovery mode again
|
||||
afterwards](recovery.md#do-not-boot-recovery-mode-again-afterwards). If you hit
|
||||
either failure, the extra boot is a manual step. Reported by
|
||||
[@kaya-eu](https://github.com/Coffey-Labs/stalwart-migrator/issues/1).
|
||||
@@ -0,0 +1,103 @@
|
||||
# Recovery
|
||||
|
||||
What happens if a migration fails: why recovery is your snapshot and not this
|
||||
tool, what a restore costs, what the tool keeps to make a manual restore
|
||||
easier, and the one thing never to do on a migrated server. For the short
|
||||
version, see the [README](../README.md).
|
||||
|
||||
## Recovery is your job
|
||||
|
||||
**This tool does not undo a migration.** There is no `rollback` command.
|
||||
Recovery from a failed migration is your own snapshot or backup, taken by
|
||||
whatever method you already trust and know how to restore — a ZFS, LVM or
|
||||
btrfs snapshot, a VM or volume snapshot, or a restorable backup. Choosing that
|
||||
method, taking it, and verifying you can actually restore from it is out of
|
||||
scope for this tool: it does not take one, does not check that one exists, and
|
||||
cannot restore from one.
|
||||
|
||||
Cutover refuses to start until you confirm a recovery point exists, with
|
||||
`--recovery-point-confirmed`. That confirmation is an acknowledgement, not a
|
||||
check — nothing here can verify your snapshot. Its only purpose is that nobody
|
||||
migrates a production mail server having never been asked the question.
|
||||
|
||||
A failed run stops and reports; a person decides what to restore.
|
||||
|
||||
**Take the snapshot with the service stopped** if you want a clean one. A
|
||||
snapshot of a running Stalwart is crash-consistent rather than clean; RocksDB
|
||||
will usually recover from its WAL, but "usually" is doing real work in that
|
||||
sentence.
|
||||
|
||||
## Restoring from a snapshot loses mail delivered since
|
||||
|
||||
Reverting to any pre-migration recovery point discards mail delivered between
|
||||
taking it and restoring it. This is inherent to restoring a point in time and
|
||||
this tool cannot solve it — plan your migration window with that in mind, and
|
||||
consider holding inbound mail at a secondary MX for the duration if the gap
|
||||
matters to you.
|
||||
|
||||
## What the tool keeps to make a manual restore easier
|
||||
|
||||
- **The old binary is preserved**, never deleted, next to the new one as
|
||||
`<binary>.v<old-version>` — so putting things back doesn't depend on
|
||||
re-downloading a specific old release under pressure.
|
||||
- **The original service definition is preserved** as `<unit>.pre-<run-id>`
|
||||
before cutover rewrites it, so you aren't reconstructing a unit file from
|
||||
memory. For a container, the old container is renamed rather than removed
|
||||
and its `docker inspect` is kept; see [docker.md](docker.md).
|
||||
- **The settings and principals dumps, the apply plan and its supplement** are
|
||||
kept in `<state-dir>/<run-id>` — `/var/lib/stalwart-migrator/runs/<run-id>`
|
||||
unless you moved it. These four are the only files in a run that cannot be
|
||||
produced again afterwards: the dumps need a live pre-migration instance, and
|
||||
the plan is what was actually replayed into your store. They are kept
|
||||
whether or not the run succeeded and whether or not you passed
|
||||
`--keep-artifacts`.
|
||||
- **Every artifact path and checksum is in the checkpoint**, and
|
||||
`stalwart-migrate status <run-id>` prints exactly which steps completed and
|
||||
which failed — which is the first thing you want when deciding what to
|
||||
restore.
|
||||
|
||||
None of this is a substitute for the snapshot. It's what makes the twenty
|
||||
minutes after restoring one less unpleasant.
|
||||
|
||||
## Do not boot recovery mode again afterwards
|
||||
|
||||
The migration works by starting the new version once in recovery mode,
|
||||
replaying your settings into it, and stopping it. That is a one-time step in a
|
||||
migration, and it is not a general-purpose maintenance mode.
|
||||
|
||||
An operator who booted recovery mode again — the same way the migration does,
|
||||
`STALWART_RECOVERY_MODE=1` against the same data directory — for reasons
|
||||
unrelated to the migration, on a server that had migrated successfully days
|
||||
earlier, found that `Domain` and `Account` queries came back empty on the next
|
||||
normal start. This happened twice, on two different servers. It was not a
|
||||
stale read: creating a domain that had certainly existed a moment earlier
|
||||
succeeded, with no `primaryKeyViolation`, so the records were genuinely gone.
|
||||
Disk usage did not change.
|
||||
|
||||
What recovered it both times was re-applying that run's `export.json` and
|
||||
`supplement.json` against a fresh recovery boot, which is why those two files
|
||||
are kept for you. If you need to change something after a migration, use the
|
||||
admin API or `stalwart-cli` against the running server.
|
||||
|
||||
This is Stalwart's behavior rather than this tool's, and it is recorded here
|
||||
because this tool is where you learned the technique. Reported by
|
||||
[@kaya-eu](https://github.com/Coffey-Labs/stalwart-migrator/issues/1).
|
||||
|
||||
This is different from the one extra recovery boot some migrations need
|
||||
*during* the migration, before the first normal start; see [known Stalwart
|
||||
problems](known-stalwart-problems.md#the-store-migration-may-need-one-more-recovery-boot).
|
||||
|
||||
## Why it works this way
|
||||
|
||||
An earlier version of this tool implemented rollback itself: it restored the
|
||||
filesystem backup, verified every restored file against a manifest, replayed
|
||||
SQL dumps, reinstalled the old binary, and re-validated the result. It was
|
||||
tested and it looked good.
|
||||
|
||||
It was removed, because restoring bytes correctly is not the hard part. It
|
||||
copied contents and permissions but not *ownership*, so run as root it would
|
||||
have produced a byte-perfect, checksum-verified, root-owned data directory
|
||||
that Stalwart, running as its own user, could not open — and it would have
|
||||
reported success. A filesystem snapshot has no such failure mode, because it
|
||||
never lost the metadata to begin with. [ARCHITECTURE.md](../ARCHITECTURE.md)
|
||||
§4.8 records the full reasoning.
|
||||
@@ -0,0 +1,137 @@
|
||||
# Trying it safely, and rehearsing
|
||||
|
||||
How to find out what a migration will do before it does it: running
|
||||
`preflight` safely, what `rehearse` reports, and how to rehearse the whole
|
||||
migration on a clone of your server. For the short version, see the
|
||||
[README](../README.md).
|
||||
|
||||
## Running preflight
|
||||
|
||||
`preflight` is the sensible starting point. Its checks against the Stalwart
|
||||
installation are read-only:
|
||||
|
||||
```sh
|
||||
sudo go run ./cmd/stalwart-migrate preflight
|
||||
```
|
||||
|
||||
It still needs write access, because it records the run as a checkpoint before
|
||||
doing anything else. Without it, you see:
|
||||
|
||||
```
|
||||
create run: checkpoint: create run directory:
|
||||
mkdir /var/lib/stalwart-migrator: permission denied
|
||||
```
|
||||
|
||||
Checkpoints go to `/var/lib/stalwart-migrator/runs` by default
|
||||
(`checkpoint.DefaultBaseDir`). Either run as root, pre-create that directory
|
||||
writable, or point `--state-dir` somewhere writable. `run` and `rehearse` also
|
||||
take `--work-dir` for their scratch space, which is a different directory and
|
||||
does not move the checkpoint store.
|
||||
|
||||
Pass `--admin-url` and `--admin-user` for the reachability check, and the
|
||||
password with `--admin-password` or `STALWART_MIGRATE_ADMIN_PASSWORD`.
|
||||
|
||||
## What rehearse reports
|
||||
|
||||
```sh
|
||||
stalwart-migrate rehearse --admin-url https://mail.example.com \
|
||||
--admin-user admin --target 0.16.14
|
||||
```
|
||||
|
||||
It runs preflight, dumps your settings and principals, converts them with
|
||||
Stalwart's own `migrate_v016.py`, and reports **both halves** of the result:
|
||||
the apply plan of what will carry over, and the worklist of what will not.
|
||||
|
||||
It copies no data, clones nothing, starts no server, and never writes to the
|
||||
store, so it is safe to run against production repeatedly and without a
|
||||
maintenance window.
|
||||
|
||||
### The supplemental plan
|
||||
|
||||
It also generates a **supplemental plan** for the part it can rebuild
|
||||
automatically — currently your network listeners, which is the difference
|
||||
between a migrated server that answers and one that doesn't — and reports
|
||||
exactly how much of the worklist that covers (on a test instance: 24 of 3,505
|
||||
keys, and it says so rather than implying more).
|
||||
|
||||
`run` generates the same supplement and applies it after `export.json` for you.
|
||||
If you are replaying a rehearsal's plan by hand instead, review it, then apply
|
||||
it after `export.json`:
|
||||
|
||||
```sh
|
||||
stalwart-cli apply --file <state-dir>/<run-id>/supplement.json \
|
||||
--url https://mail.example.com
|
||||
```
|
||||
|
||||
It is applied after `export.json` rather than merged into it on purpose: the
|
||||
official conversion is the authority on everything it handles, and a generated
|
||||
plan that overlapped it could silently override a correct mapping with a
|
||||
guessed one.
|
||||
|
||||
### Reading the worklist
|
||||
|
||||
The worklist is long but mostly not work, and `rehearse` says which is which.
|
||||
Measured against a real production instance, `migrate_v016.py` carried
|
||||
**219 of 12,401 settings**. Of the 12,182 it left:
|
||||
|
||||
- **8,547** are runtime auto-ban state that repopulates itself
|
||||
- **3,337** are stock spam-filter and lookup data v0.16 ships its own copies
|
||||
of — restoring v0.15's would revert a year of upstream updates
|
||||
- **~224** were already carried another way, DKIM signatures included
|
||||
- **~293** genuinely need your eyes
|
||||
|
||||
`server.listener` is in that third group only because this tool regenerates it
|
||||
for you; without that a migrated instance answers on no ports at all.
|
||||
|
||||
### What it keeps
|
||||
|
||||
The conclusions are preserved under `<state-dir>/<run-id>/` —
|
||||
`/var/lib/stalwart-migrator/runs/<run-id>/` by default — as `export.json`,
|
||||
`unmigrated.txt` and `supplement.json`, even though the rest of the scratch
|
||||
directory is cleaned up (unless you pass `--keep-artifacts`).
|
||||
|
||||
### Why it replaced a dry run
|
||||
|
||||
`rehearse` replaced an earlier `run --dry-run` that cloned the data directory
|
||||
into a sandbox and migrated the copy. That proved the store opens, at the cost
|
||||
of copying it twice — while the half that found every real problem needed no
|
||||
copy at all. [ARCHITECTURE.md](../ARCHITECTURE.md) §4.9 has the reasoning.
|
||||
|
||||
## Rehearse on a clone first
|
||||
|
||||
`rehearse` is read-only and stops short of the half that matters: it converts
|
||||
your settings but never applies them, and applying is where a real migration
|
||||
fails. A clone closes that gap, and on a 2.4 GB store it costs about six
|
||||
seconds of production downtime to build.
|
||||
|
||||
1. **Copy the data directory with the service stopped.** RocksDB is
|
||||
single-writer, so a hot copy may be torn — and a rehearsal on a torn store
|
||||
fails for reasons production never would, or passes when it should not.
|
||||
Stop, `cp -a` the data dir to the same disk, start again, archive it
|
||||
afterwards; the service is down only for the local copy.
|
||||
2. **Give the guest no route off the host.** A clone of a live mail server
|
||||
will otherwise renew certificates for your real domains and deliver
|
||||
whatever is in the outbound queue. In libvirt that means a network with no
|
||||
`<forward>` element. Verify it from inside the guest rather than assuming.
|
||||
3. **Run the real thing**: `preflight`, then `run` with `--target-binary`,
|
||||
`--stalwart-cli` and `--migration-script` pointed at locally staged copies,
|
||||
since an isolated guest can download nothing.
|
||||
4. **Snapshot the guest while it is shut down**, so a failed attempt costs
|
||||
seconds to reset rather than a rebuild.
|
||||
|
||||
### What it caught
|
||||
|
||||
None of these could be expressed by the mock:
|
||||
|
||||
- a multi-tenant arrangement v0.16 cannot represent;
|
||||
- `--target-binary` never reaching preflight, so an air-gapped host failed on a
|
||||
release lookup;
|
||||
- an `admin` account that was a config fallback-admin and stopped working the
|
||||
moment the migration finished;
|
||||
- `tenant-admin` roles the converter does not restore.
|
||||
|
||||
### What isolation costs
|
||||
|
||||
Two things the isolation costs, so they are not mistaken for faults: the guest
|
||||
cannot fetch the v0.16 web interface, so `/account/` returns 404, and
|
||||
certificate renewal cannot be exercised at all.
|
||||
+106
@@ -0,0 +1,106 @@
|
||||
# Status and field reports
|
||||
|
||||
What works today, what `run` does, what has been tested against real servers,
|
||||
and the state of each package. For the short version, see the
|
||||
[README](../README.md).
|
||||
|
||||
## Commands
|
||||
|
||||
| Command | State |
|
||||
|---|---|
|
||||
| `stalwart-migrate preflight` | **Works** — read-only checks and a migration plan |
|
||||
| `stalwart-migrate rehearse` | **Works** — read-only; converts your settings and reports what won't carry over |
|
||||
| `stalwart-migrate run` | **Works** — performs the migration; `--recovery-point-confirmed --yes`. Container deployments additionally need `--container-path-unproven` (see [Docker deployments](docker.md)) |
|
||||
| `stalwart-migrate tenants` | **Works** — read-only; who owns which domain, and what would block a migration |
|
||||
| `stalwart-migrate status <id>` | **Works** |
|
||||
| `stalwart-migrate report <id>` | **Works** — prints what validation found for a run |
|
||||
|
||||
Run `stalwart-migrate <command> -h` for the current flags of each.
|
||||
|
||||
## What `run` does
|
||||
|
||||
`run` performs the migration in this order:
|
||||
|
||||
preflight → stage → dump → stop → convert → recovery-mode → cutover → validate
|
||||
|
||||
It needs two flags: `--yes` (intent) and `--recovery-point-confirmed` (a claim
|
||||
that you have a snapshot or backup you have verified you can restore — this
|
||||
tool cannot undo a migration and will not start without it). Start with
|
||||
`rehearse` first: it is read-only, needs no maintenance window, and tells you
|
||||
what `run` will and won't carry over. See [rehearsal.md](rehearsal.md).
|
||||
|
||||
An interrupted run resumes from its last completed step with
|
||||
`stalwart-migrate run --resume <run-id>` and the same flags; `status` lists run
|
||||
IDs and shows which steps completed.
|
||||
|
||||
### Validation after cutover
|
||||
|
||||
After cutover, `run` compares the migrated instance against the snapshot
|
||||
preflight took, and fails the command if an account that existed before is
|
||||
missing from it. A domain that no longer appears is reported as a warning
|
||||
rather than a failure: the two versions do not agree on what counts as a
|
||||
domain — principals on one side, `Domain` objects on the other — and failing a
|
||||
migration over that difference would abort runs that lost nothing.
|
||||
|
||||
The service is left running either way. By that point the store has been
|
||||
migrated in place, so stopping it would not undo anything; your recovery point
|
||||
is the way back (see [recovery.md](recovery.md)). `report <run-id>` prints the
|
||||
same finding again later. Where preflight had no admin URL to snapshot from,
|
||||
validation reports itself as skipped rather than passed.
|
||||
|
||||
## Field reports
|
||||
|
||||
### A production server, 2026-08-25
|
||||
|
||||
The tool took a live server — nine domains, six accounts, a 2.4 GB RocksDB
|
||||
store — from 0.15.5 to 0.16.19 with **8 seconds** of downtime, every phase
|
||||
green including post-cutover validation, mail flowing before and after. That
|
||||
run was preceded by a full dress rehearsal on a clone of the same server, which
|
||||
is the practice this project most recommends copying: see [Rehearse on a
|
||||
clone first](rehearsal.md#rehearse-on-a-clone-first).
|
||||
|
||||
### Three servers, reported in issue #1
|
||||
|
||||
[@kaya-eu](https://github.com/Coffey-Labs/stalwart-migrator/issues/1)
|
||||
reported three successful 0.15.5 → 0.16.19 migrations on three servers: a
|
||||
testing and a production instance, each 16 domains, 55 accounts and roughly
|
||||
**221 GB** of real mail, and an arm64 home server of 1.4 GB across 181 folders.
|
||||
|
||||
Read that with two qualifications:
|
||||
|
||||
- **They ran a commit predating the automated Docker cutover**, so they
|
||||
performed the cutover by hand. What those runs exercised is preflight, the
|
||||
dumps, the settings conversion and the recovery-mode store migration, not
|
||||
the container cutover.
|
||||
- **They hit things worth knowing about before you follow them:**
|
||||
- an arm64 binary that was fetched for the wrong architecture (fixed);
|
||||
- a store migration that needed one more recovery-mode boot than the tool
|
||||
performs (**not fixed** — see [The store migration may need one more
|
||||
recovery
|
||||
boot](known-stalwart-problems.md#the-store-migration-may-need-one-more-recovery-boot));
|
||||
- data loss from booting recovery mode again *after* a completed migration
|
||||
(see [Do not boot recovery mode
|
||||
again](recovery.md#do-not-boot-recovery-mode-again-afterwards)).
|
||||
|
||||
## Code
|
||||
|
||||
Roughly 14,600 lines of Go, standard library only, of which about 6,300 are
|
||||
tests. Every phase exists as a package, and `run` wires them into the
|
||||
migration described above.
|
||||
|
||||
Lines are implementation only; each package carries its tests alongside.
|
||||
|
||||
| Package | Lines | Tests |
|
||||
|---|---|---|
|
||||
| `internal/stalwartapi` | 1456 | yes |
|
||||
| `internal/backup` | 1333 | yes |
|
||||
| `internal/preflight` | 1035 | yes |
|
||||
| `internal/applyplan` | 913 | yes |
|
||||
| `internal/cutover` | 784 | yes |
|
||||
| `internal/recovery` | 431 | yes |
|
||||
| `internal/checkpoint` | 406 | yes |
|
||||
| `internal/validate` | 382 | yes |
|
||||
| `internal/stage` | 233 | yes |
|
||||
| `internal/service` | 201 | yes |
|
||||
| `internal/plan` | 130 | yes |
|
||||
| `internal/config` | stub | — |
|
||||
Reference in New Issue
Block a user