Commit Graph
12 Commits
Author SHA1 Message Date
jcoffey-dev d3ce33b8e2 Merge pull request 'export: a dry run that predicts failures and exits as the real run would' (#11) from feat/predictive-dry-run into main
ci / test (push) Skipped
github/ci (branch) GitHub Actions
ci / github (push) Successful in 3m50s
ci / announce (push) Skipped
2026-09-30 19:49:53 +00:00
jcoffey-dev fde1f0f68b Merge pull request 'export: batch Email/import, upload in parallel, and show progress' (#10) from feat/export-progress-batching into main
ci / test (push) Skipped
ci / github (push) Canceled after 2m5s
ci / announce (push) Canceled after 0s
github/ci (branch) GitHub Actions
2026-09-30 19:47:46 +00:00
jcoffey-dev e47c074d43 export: a dry run that predicts failures and exits as the real run would
ci / test (pull_request) Skipped
github/ci (branch) GitHub Actions
ci / github (pull_request) Successful in 2m30s
ci / announce (pull_request) Skipped
export --dry-run counted what it would create and printed a table, but it
missed the failures it could see coming and always exited 0.

It now checks, before anything is written:
- a message larger than the target's maxSizeUpload;
- an object too large for one request under maxSizeRequest, as a contact
  with its photo inlined can be;
- a Sieve script the target would reject, with SieveScript/validate, after
  the vnd.stalwart names are renamed, so what is checked is what would be
  uploaded. Sieve scripts are uploaded as blobs for that check, and nothing
  else is; a target without the method gets one warning.

The counts are kept instead of being dropped, so a dry run in which
anything would fail exits 5, like the real run. The report is now a plan in
plain words -- per type, what would be created, updated, left unchanged,
deleted and would fail -- followed by each predicted failure and why.
2026-09-30 12:46:38 -07:00
jcoffey-dev 3f832e17b2 export: batch Email/import, upload in parallel, and show progress
ci / test (pull_request) Skipped
github/ci (branch) GitHub Actions
ci / github (pull_request) Successful in 3m11s
ci / announce (pull_request) Skipped
Export wrote one message per request and said nothing while it did, so a
large mailbox took hours of silence.

Messages now go in batches of up to the target's maxObjectsInSet (at most
50), their blobs uploaded several at once, up to maxConcurrentUpload and no
more than --threads; each upload thread reads the archive through its own
read-only connection. Every message is still counted on its own: one the
target rejects fails alone, a request too large is split, and a method
error on the whole call is retried a message at a time.

A batch is sent once. If it ends without a clear answer -- a dropped
connection, a gateway timeout, a partial failure -- the target is read
again, the messages that arrived count as created, and only the rest are
imported again, so none is doubled.

A progress line (count, rate, time left) is printed every few seconds for
mail, contacts and events, and each type ends with a line of what was
created, updated, left unchanged and failed.
2026-09-30 12:39:57 -07:00
jcoffey-dev 3cae9464f0 import: bound the memory an IMAP or EWS import holds
ci / test (pull_request) Skipped
github/ci (branch) GitHub Actions
ci / github (pull_request) Successful in 2m50s
ci / announce (pull_request) Skipped
Two things let a large mailbox fill memory. IMAP workers handed every
fetched message to the archive writer through an unbounded queue, so fast
workers could hold whole folders' worth of bodies while the single writer
caught up. And fetch batches were sized by count only: 20 items per EWS
GetItem, 8 connections at a time, is a few megabytes of ordinary mail and
several gigabytes of large attachments.

- The IMAP event queue is bounded at two events per worker, so a worker
  waits for the writer instead of running ahead of it.
- Fetch batches are bounded by bytes as well as by count, with a shared
  helper, `sync::batch::by_count_and_bytes`: 32 MiB by default, and a
  single larger message goes alone.
  - IMAP learns each new message's RFC822.SIZE on the control connection,
    in 1000-UID metadata fetches, before the body fetch. A server that won't
    say leaves the chunks sized by count. `--fetch-batch-mib` sets the cap.
  - EWS asks FindItem for `item:Size` and packs GetItem batches by it,
    fetched `--ews-connections` batches at a time. `--ews-getitem-batch-mib`
    sets the cap. Items from an incremental SyncFolderItems run carry no
    size and stay batched by count.
- docs/usage.md describes both options.
2026-09-30 12:32:29 -07:00
jcoffey-dev 687027c7cb export: bring matched items up to date on every run
ci / test (pull_request) Skipped
github/ci (branch) GitHub Actions
ci / github (pull_request) Successful in 2m50s
ci / announce (pull_request) Skipped
Export matched each item against the target and then skipped it, so a second
run -- the usual final pass of a cutover -- never carried anything that had
changed at the source since the first: read and flagged state, moves between
folders, edited contacts, events and Sieve scripts. It reported them as
skipped and exited 0, while the usage guide said matched items were updated.

Matched items are now updated, with one batched /set per type:

- Email: keywords are set to the archive's, added and removed, compared
  case-insensitively. Memberships of folders this run migrated are added and
  removed to match; folders that exist only on the target are left alone,
  and a message is never left in no folder. Properties the server did not
  report are not touched.
- Contacts and events: when both copies carry `updated`, the archive's is
  written only if it is newer, compared as instants so an offset or a
  fraction of a second is not taken for a change; otherwise each property the
  archive writes is compared, and those that differ are sent whole.
- Sieve scripts: the target's copy is downloaded and compared byte for byte
  with what export would write -- after renaming Stalwart's vendor names for
  an inbuxa target -- and replaced when it differs, so a renamed script is
  not re-uploaded on every run.

Updated items are counted as `updated`; unchanged ones stay `skipped`. The
usage guide now describes this.
2026-09-30 11:37:39 -07:00
jcoffey-dev 14a797cd65 Merge pull request 'Time out every connection, and scope --allow-invalid-certs' (#6) from fix/connection-safety into main
ci / test (push) Skipped
github/ci (branch) GitHub Actions
ci / github (push) Successful in 3m31s
ci / announce (push) Skipped
2026-09-30 18:34:50 +00:00
jcoffey-dev 01ac1a5b3e Merge pull request 'export: write a message once, in every folder it was in' (#4) from fix/one-message-many-folders into main
ci / test (push) Skipped
github/ci (branch) GitHub Actions
ci / github (push) Canceled after 1m36s
ci / announce (push) Canceled after 0s
2026-09-30 18:33:09 +00:00
jcoffey-dev 234b3203d7 Scope --allow-invalid-certs to the server the user named
ci / test (pull_request) Skipped
github/ci (branch) GitHub Actions
ci / github (pull_request) Successful in 2m33s
ci / announce (pull_request) Skipped
The flag switched certificate checks off for every connection in the run.
That included the Microsoft sign-in endpoints, so a user passing it for a
self-signed source also sent refresh tokens, device codes and EWS client
secrets over unverified TLS. It also covered the export target, and any host
a server redirected to or named for its API, uploads or downloads.

It now applies only where the user pointed it: the host of --url, for the
source of an import or the target of an export. For an Exchange import with
no --url, it covers the mailbox's own domain, where on-premises Autodiscover
looks, and then only the EWS endpoint Autodiscover finds. The Microsoft and
Google sign-in and cloud hosts are always verified, with or without the flag.

Each HTTP client keeps a verifying agent and, only when the flag applies, a
second one that accepts invalid certificates, and picks per request by host.
The sign-in modules no longer take the flag at all. Autodiscover v2, which is
Microsoft's own service, is always verified.
2026-09-30 11:31:44 -07:00
jcoffey-dev 5ae0625ee1 export: write a message once, in every folder it was in
ci / test (pull_request) Skipped
github/ci (branch) GitHub Actions
ci / github (pull_request) Successful in 2m11s
ci / announce (pull_request) Skipped
A source that files one message in several folders -- IMAP and Maildir
copies, Gmail labels -- leaves one archive row per folder, all pointing at the
same blob. Export imported each row on its own, so the message arrived on the
target once per folder. Worse, an interrupted export matched the rest by
Message-ID alone on the next run and skipped them, so their folders were never
added.

Export now folds rows with the same blob into one email, with the union of
their folders and keywords. On a match it adds, with one batched Email/set,
any migrated folder the target copy is missing; folders that exist only on the
target are left alone. Matching is still by Message-ID (or the fallback
digest), with size deciding between several messages that share one, so two
different messages are never folded onto one target email.

The archive is unchanged: the fold happens at export, so imports, resumes and
deletions work exactly as before.
2026-09-30 11:25:04 -07:00
jcoffey-dev 54fd18b410 docs: export matches and skips, it does not update
ci / test (pull_request) Skipped
github/ci (branch) GitHub Actions
ci / github (pull_request) Successful in 1m50s
ci / announce (pull_request) Skipped
usage.md said items that match are updated. Export skips a match: a
change made at the source after the first export -- read state, a folder
move, an edited event, contact or Sieve script -- does not reach the target
on a second one. Say so until updating on re-export is built.
2026-09-30 11:16:47 -07:00
jcoffey-dev 210c04997c Track the Gitea workflows and docs/usage.md
The .gitignore inherited from upstream ignores every dotfile except .github,
and every Markdown file except README and CHANGELOG. The Gitea workflows and
docs/usage.md, which the README links to, were therefore never committed.
Both are now excepted and added.
2026-09-30 10:23:30 -07:00