jcoffey-dev is traveling from Thursday 1 October through Sunday 4 October. Issues and pull requests are welcome, and will get an answer after that. Thanks for your patience.
export --dry-run counted what it would create and printed a table, but it
missed the failures it could see coming and always exited 0.
It now checks, before anything is written:
- a message larger than the target's maxSizeUpload;
- an object too large for one request under maxSizeRequest, as a contact
with its photo inlined can be;
- a Sieve script the target would reject, with SieveScript/validate, after
the vnd.stalwart names are renamed, so what is checked is what would be
uploaded. Sieve scripts are uploaded as blobs for that check, and nothing
else is; a target without the method gets one warning.
The counts are kept instead of being dropped, so a dry run in which
anything would fail exits 5, like the real run. The report is now a plan in
plain words -- per type, what would be created, updated, left unchanged,
deleted and would fail -- followed by each predicted failure and why.
When export created a default address book, it went on to send
AddressBook/set with onSuccessSetIsDefault -- in a dry run too, naming the
dry run's made-up id. A dry run writes nothing; the claim is now only
counted as planned.
Export wrote one message per request and said nothing while it did, so a
large mailbox took hours of silence.
Messages now go in batches of up to the target's maxObjectsInSet (at most
50), their blobs uploaded several at once, up to maxConcurrentUpload and no
more than --threads; each upload thread reads the archive through its own
read-only connection. Every message is still counted on its own: one the
target rejects fails alone, a request too large is split, and a method
error on the whole call is retried a message at a time.
A batch is sent once. If it ends without a clear answer -- a dropped
connection, a gateway timeout, a partial failure -- the target is read
again, the messages that arrived count as created, and only the rest are
imported again, so none is doubled.
A progress line (count, rate, time left) is printed every few seconds for
mail, contacts and events, and each type ends with a line of what was
created, updated, left unchanged and failed.
A POST that failed in transport -- including a timeout while waiting for the
answer -- is resent today, and so is one that got a 502 or 504. For a write
such as Email/import the server may already have applied the first copy,
so the resend can create a duplicate.
post_json_once and Request::send_once send it once: a transport failure or
a gateway error comes back to the caller, which can check the target before
trying again. 429 and 503, which mean the request was not processed, are
still retried.
Two things let a large mailbox fill memory. IMAP workers handed every
fetched message to the archive writer through an unbounded queue, so fast
workers could hold whole folders' worth of bodies while the single writer
caught up. And fetch batches were sized by count only: 20 items per EWS
GetItem, 8 connections at a time, is a few megabytes of ordinary mail and
several gigabytes of large attachments.
- The IMAP event queue is bounded at two events per worker, so a worker
waits for the writer instead of running ahead of it.
- Fetch batches are bounded by bytes as well as by count, with a shared
helper, `sync::batch::by_count_and_bytes`: 32 MiB by default, and a
single larger message goes alone.
- IMAP learns each new message's RFC822.SIZE on the control connection,
in 1000-UID metadata fetches, before the body fetch. A server that won't
say leaves the chunks sized by count. `--fetch-batch-mib` sets the cap.
- EWS asks FindItem for `item:Size` and packs GetItem batches by it,
fetched `--ews-connections` batches at a time. `--ews-getitem-batch-mib`
sets the cap. Items from an incremental SyncFolderItems run carry no
size and stay batched by count.
- docs/usage.md describes both options.
IMAP import wrote a whole folder in one transaction and stopped the folder
at the first message it could not import. A crash near the end of a large
INBOX kept nothing, and one bad INTERNALDATE lost the rest of the folder.
Worse, when a folder stopped early, fetches still in flight for it could be
filed into the next folder's mailbox.
- The transaction is committed after every fetch chunk. Each message is
written in its own savepoint, so what is committed is always whole, and a
rerun fetches only the UIDs still missing.
- A message that cannot be imported is rolled back on its own, logged with
its folder and UID, and counted as failed; the folder carries on, and the
message stays out of the UID map so the next run tries it again. Archive
and I/O errors still stop the run.
- Fetch jobs and events carry a folder generation. Moving to a new folder
cancels queued work for older ones, and any event from an older
generation is dropped, never filed. Shutdown drains in-flight events
before joining the workers, so it cannot hang on a blocked worker.
- INTERNALDATE month names are matched in any case.
- Maildir import gets the same per-message savepoint, and commits every 500
new messages instead of once per folder.
Export matched each item against the target and then skipped it, so a second
run -- the usual final pass of a cutover -- never carried anything that had
changed at the source since the first: read and flagged state, moves between
folders, edited contacts, events and Sieve scripts. It reported them as
skipped and exited 0, while the usage guide said matched items were updated.
Matched items are now updated, with one batched /set per type:
- Email: keywords are set to the archive's, added and removed, compared
case-insensitively. Memberships of folders this run migrated are added and
removed to match; folders that exist only on the target are left alone,
and a message is never left in no folder. Properties the server did not
report are not touched.
- Contacts and events: when both copies carry `updated`, the archive's is
written only if it is newer, compared as instants so an offset or a
fraction of a second is not taken for a change; otherwise each property the
archive writes is compared, and those that differ are sent whole.
- Sieve scripts: the target's copy is downloaded and compared byte for byte
with what export would write -- after renaming Stalwart's vendor names for
an inbuxa target -- and replaced when it differs, so a renamed script is
not re-uploaded on every run.
Updated items are counted as `updated`; unchanged ones stay `skipped`. The
usage guide now describes this.
The flag switched certificate checks off for every connection in the run.
That included the Microsoft sign-in endpoints, so a user passing it for a
self-signed source also sent refresh tokens, device codes and EWS client
secrets over unverified TLS. It also covered the export target, and any host
a server redirected to or named for its API, uploads or downloads.
It now applies only where the user pointed it: the host of --url, for the
source of an import or the target of an export. For an Exchange import with
no --url, it covers the mailbox's own domain, where on-premises Autodiscover
looks, and then only the EWS endpoint Autodiscover finds. The Microsoft and
Google sign-in and cloud hosts are always verified, with or without the flag.
Each HTTP client keeps a verifying agent and, only when the flag applies, a
second one that accepts invalid certificates, and picks per request by host.
The sign-in modules no longer take the flag at all. Autodiscover v2, which is
Microsoft's own service, is always verified.
An archive holds a whole mailbox, its contacts and calendars, and it was
created with SQLite's default mode, readable by every user on the machine.
On Unix a new archive is now created with mode 0600 before SQLite opens it;
SQLite gives the -wal and -shm files the database file's mode, so they
follow, which a test confirms. An existing archive that others can read is
left as it is, with a warning naming it and the chmod that fixes it.
Nothing changes on Windows.
No ureq agent set a timeout, and ureq sets none by default, so a connection
dropped silently mid-transfer (a NAT or load-balancer idle drop) hung a JMAP,
DAV, EWS or Graph run, or an export, with no error, and the retry logic never
got a chance to run. IMAP and ManageSieve already had read timeouts.
Every agent now takes its settings from a new net module: 30s to connect,
60s to send the request headers, 5 minutes for the server's first byte, and
30 minutes for a whole response body, which ureq counts as one budget for the
body rather than per read: enough for the 512 MiB limit at about 300 KB/s.
A JMAP upload's send budget grows with its size, from a 2 minute floor at an
assumed 64 KiB/s worst case.
A timeout is a transport error, and every client already retries those, so a
stalled transfer is now abandoned and retried. Tests pin that down for each
client. The TLS setup the seven agents repeated moves to one helper.
A source that files one message in several folders -- IMAP and Maildir
copies, Gmail labels -- leaves one archive row per folder, all pointing at the
same blob. Export imported each row on its own, so the message arrived on the
target once per folder. Worse, an interrupted export matched the rest by
Message-ID alone on the next run and skipped them, so their folders were never
added.
Export now folds rows with the same blob into one email, with the union of
their folders and keywords. On a match it adds, with one batched Email/set,
any migrated folder the target copy is missing; folders that exist only on the
target are left alone. Matching is still by Message-ID (or the fallback
digest), with size deciding between several messages that share one, so two
different messages are never folded onto one target email.
The archive is unchanged: the fold happens at export, so imports, resumes and
deletions work exactly as before.
inbuxa accepts vnd.inbuxa.while and vnd.inbuxa.expressions, and names its
environment items vnd.inbuxa.*, with no alias for Stalwart's vnd.stalwart.*
names. A script carried over unchanged failed to compile on the target, and
the account was left with no filtering behind a single warning.
When the target lists vnd.inbuxa extensions in its sieveExtensions, export
now renames Stalwart's names in the three places they are names: the strings
of a require list, the item name of an environment test, and
${env.vnd.stalwart.*} references inside strings. A small tokenizer skips
comments and reads quoted strings and text: blocks whole, so other text that
happens to contain the name is copied unchanged. Each rename is printed.
Against a target without the inbuxa names nothing changes.
The activation call is now checked. A method error, or an active script
that could not be created, is logged as an error naming the script and
counted as a failure, so export exits non-zero instead of leaving
filtering off with one warning line.
usage.md said items that match are updated. Export skips a match: a
change made at the source after the first export -- read state, a folder
move, an edited event, contact or Sieve script -- does not reach the target
on a second one. Say so until updating on re-export is built.
The .gitignore inherited from upstream ignores every dotfile except .github,
and every Markdown file except README and CHANGELOG. The Gitea workflows and
docs/usage.md, which the README links to, were therefore never committed.
Both are now excepted and added.
The README now says what the tool is in the suite's plain voice: the
archive in the middle, convergent reruns, --dry-run, the archive as a
backup, the sources it reads (Stalwart among the JMAP ones), the
INBUXA_MIGRATE_* credential variables, installing from the Gitea release
with SHA256SUMS, building and testing, and the license with the lineage
line. Upstream's logo, badges and install channels are gone.
The full command and flag reference moves to docs/usage.md, with the
renamed binary and variables. The CHANGELOG gains a section for this first
release, above upstream's history.
Upstream's cargo-dist pipeline published to npm and Homebrew under its own
names, built an MSI, and ran on every pull request. It goes, with its
dist-workspace.toml, the wix/ installer files and the README image.
CI now follows the rest of the suite. .gitea/workflows runs `cargo test
--locked` on Gitea's runners, or, when the org variable BUILD_ON is
'github', waits for the result GitHub reports on the commit. On GitHub,
.github/workflows/ci.yml runs the same tests, and for a v* tag on main
whose version matches Cargo.toml builds scripts/build-release.sh on native
amd64 and arm64 runners (Ubuntu 22.04, for the widest glibc range),
attaches inbuxa-migrate-linux-{amd64,arm64}.tar.gz and SHA256SUMS to the
Gitea Release, copies the Release to GitHub, and reports back. Releases are
built only on GitHub; with BUILD_ON unset a tag is tested but not released.
Published releases are announced on the forum.
Registry calls (x:Account, x:Domain and the rest of the x: types) were
always sent under urn:stalwart:jmap. inbuxa advertises the same registry
as urn:inbuxa:jmap:registry, so the capability now comes from the connected
server: Session::registry_urn() picks urn:inbuxa:jmap:registry when the
session advertises it, else urn:stalwart:jmap, so Stalwart servers still
work as a source. A request carrying an x: call against a server that
advertises neither fails with a MissingCapability error naming both,
treated like any other connection-level failure, instead of sending a
capability the server never offered.
The crate, binary, library path and every name the tool writes or reads
become inbuxa-migrate: the CLI name and about text, the HTTP user agent,
the credential variables (INBUXA_MIGRATE_PASSWORD, _TOKEN,
_EWS_CLIENT_SECRET and _GRAPH_TOKEN, with no fallback to the old names),
thread names, synthesized UIDs and calendar addresses
(urn:x-inbuxa-migrate:attendee:), and the Maildir file suffix used by the
export script. This is a new tool, so archives written by Vandelay are not
read back.
The version moves to the date scheme the release tags use (2026.9.30), and
the package metadata points at the inbuxa-migrate repository. The cargo-dist
and MSI metadata goes with the release tooling it served.
Modified files carry an added copyright line; the original notices stay.