The container deployment's answer to downloading a release: pull the image
the operator named, then ask it what it is. That last part is the point of
the phase, exactly as it is for a binary - the tag, the registry and the
repository name are all assumptions about someone else's publishing
process, and the image's own answer is the only thing that settles what
arrived.
The image is never derived from the running container by swapping its tag.
That derivation is wrong for a digest-pinned image, wrong for a mirror and
wrong for a fork, and being wrong here means pulling the wrong software
into a mail server. It is named in full or the phase refuses.
What comes back is the image's ID rather than the tag it arrived under. A
tag can move between staging and cutover -- that is the whole reason latest
is a hazard -- and running the tag later would run something other than
what was verified here.
VersionFromOutput is exported from preflight so both paths parse a version
identically. Two copies of that regex could disagree about what they
staged, which is a difference nobody would look for.
One caveat recorded rather than hidden: asking an image its version means
running it with --version, which assumes its entrypoint is the server and
passes flags through. That has not been confirmed against a published
Stalwart image, there being none to hand. If the assumption is wrong this
fails loudly with the image's own output rather than staging something
unverified, and the fallback tries the binary by name before giving up.
SkipPull is for a host that loaded the image from a tarball, where a pull
cannot work and its failure would say nothing useful.
The phases have all existed for a while; nothing chained them. The order
here is the one arrived at by performing this migration by hand against a
clone of production before writing it down:
preflight -> stage -> dump -> preserve binary -> STOP ->
convert -> supplement -> recovery-mode migration -> cutover -> START
The dump runs before the stop because it reads settings over the admin API,
and a stopped server has no admin API. Everything from the stop to the end
of cutover is downtime.
internal/stage fills the last missing phase (4.3): resolve the release,
take the x86_64 linux-gnu server build and refuse to substitute another,
verify a pinned checksum if one was given, extract the binary - refusing
any archive entry that isn't a regular file, since a tarball is untrusted
input - and confirm the result reports the version its tag claimed.
Everything upstream of that last check is an assumption about someone
else's release process.
Two gates, separate on purpose. --yes is about intent. --recovery-point-
confirmed is a claim about the world: this tool cannot undo a migration
(4.8) and cannot check whether a snapshot exists, so a run that proceeded
without the operator asserting one would be proceeding on a hope.
Verified end to end against a real Stalwart 0.15.5 with email-style account
names, a named admin account, and seeded mail:
MIGRATION COMPLETE. Mail was down for 6s.
Every cutover step green, including recalculate-quotas ("rebuilt disk
quotas for 2 account(s)") - the first time the x:Task wire format inferred
from Stalwart's schema reference has actually been exercised. It works,
now that endpoint discovery and role restoration make it reachable. After
the migration the named admin still administers, alice logs in with
unchanged credentials to the same four messages, and new SMTP delivery is
accepted.
Both refusal gates were tested, as was the failure path: an apply that
fails leaves the run stopped with the store part-migrated, and the error
says to restore the recovery point rather than restart the old version
against it.