jcoffey-dev is traveling from Thursday 1 October through Sunday 4 October. Issues and pull requests are welcome, and will get an answer after that. Thanks for your patience.
Stacked on #8 (fix/imap-import-resilience): it bounds the event queue whose drain-on-shutdown #8 adds. Merge #8 first; this PR's diff then shrinks to its own commit.
Two things let a large mailbox fill memory. IMAP workers handed every
fetched message to the archive writer through an unbounded queue, so fast
workers could hold whole folders' worth of bodies while the single writer
caught up. And fetch batches were sized by count only: 20 items per EWS
GetItem, 8 connections at a time, is a few megabytes of ordinary mail and
several gigabytes of large attachments.
The IMAP event queue is bounded at two events per worker, so a worker
waits for the writer instead of running ahead of it.
Fetch batches are bounded by bytes as well as by count, with a shared
helper, sync::batch::by_count_and_bytes: 32 MiB by default, and a
single larger message goes alone.
IMAP learns each new message's RFC822.SIZE on the control connection,
in 1000-UID metadata fetches, before the body fetch. A server that won't
say leaves the chunks sized by count. --fetch-batch-mib sets the cap.
EWS asks FindItem for item:Size and packs GetItem batches by it,
fetched --ews-connections batches at a time. --ews-getitem-batch-mib
sets the cap. Items from an incremental SyncFolderItems run carry no
size and stay batched by count.
docs/usage.md describes both options.
Tests: 1392 pass (10 new: the batching helper, EWS byte-split GetItem and FindItem size parsing, IMAP byte-split body fetch, and a ten-message chunk through a two-event queue). fmt and clippy clean.
**Stacked on #8** (fix/imap-import-resilience): it bounds the event queue whose drain-on-shutdown #8 adds. Merge #8 first; this PR's diff then shrinks to its own commit.
Two things let a large mailbox fill memory. IMAP workers handed every
fetched message to the archive writer through an unbounded queue, so fast
workers could hold whole folders' worth of bodies while the single writer
caught up. And fetch batches were sized by count only: 20 items per EWS
GetItem, 8 connections at a time, is a few megabytes of ordinary mail and
several gigabytes of large attachments.
- The IMAP event queue is bounded at two events per worker, so a worker
waits for the writer instead of running ahead of it.
- Fetch batches are bounded by bytes as well as by count, with a shared
helper, `sync::batch::by_count_and_bytes`: 32 MiB by default, and a
single larger message goes alone.
- IMAP learns each new message's RFC822.SIZE on the control connection,
in 1000-UID metadata fetches, before the body fetch. A server that won't
say leaves the chunks sized by count. `--fetch-batch-mib` sets the cap.
- EWS asks FindItem for `item:Size` and packs GetItem batches by it,
fetched `--ews-connections` batches at a time. `--ews-getitem-batch-mib`
sets the cap. Items from an incremental SyncFolderItems run carry no
size and stay batched by count.
- docs/usage.md describes both options.
Tests: 1392 pass (10 new: the batching helper, EWS byte-split GetItem and FindItem size parsing, IMAP byte-split body fetch, and a ten-message chunk through a two-event queue). fmt and clippy clean.
IMAP import wrote a whole folder in one transaction and stopped the folder
at the first message it could not import. A crash near the end of a large
INBOX kept nothing, and one bad INTERNALDATE lost the rest of the folder.
Worse, when a folder stopped early, fetches still in flight for it could be
filed into the next folder's mailbox.
- The transaction is committed after every fetch chunk. Each message is
written in its own savepoint, so what is committed is always whole, and a
rerun fetches only the UIDs still missing.
- A message that cannot be imported is rolled back on its own, logged with
its folder and UID, and counted as failed; the folder carries on, and the
message stays out of the UID map so the next run tries it again. Archive
and I/O errors still stop the run.
- Fetch jobs and events carry a folder generation. Moving to a new folder
cancels queued work for older ones, and any event from an older
generation is dropped, never filed. Shutdown drains in-flight events
before joining the workers, so it cannot hang on a blocked worker.
- INTERNALDATE month names are matched in any case.
- Maildir import gets the same per-message savepoint, and commits every 500
new messages instead of once per folder.
Two things let a large mailbox fill memory. IMAP workers handed every
fetched message to the archive writer through an unbounded queue, so fast
workers could hold whole folders' worth of bodies while the single writer
caught up. And fetch batches were sized by count only: 20 items per EWS
GetItem, 8 connections at a time, is a few megabytes of ordinary mail and
several gigabytes of large attachments.
- The IMAP event queue is bounded at two events per worker, so a worker
waits for the writer instead of running ahead of it.
- Fetch batches are bounded by bytes as well as by count, with a shared
helper, `sync::batch::by_count_and_bytes`: 32 MiB by default, and a
single larger message goes alone.
- IMAP learns each new message's RFC822.SIZE on the control connection,
in 1000-UID metadata fetches, before the body fetch. A server that won't
say leaves the chunks sized by count. `--fetch-batch-mib` sets the cap.
- EWS asks FindItem for `item:Size` and packs GetItem batches by it,
fetched `--ews-connections` batches at a time. `--ews-getitem-batch-mib`
sets the cap. Items from an incremental SyncFolderItems run carry no
size and stay batched by count.
- docs/usage.md describes both options.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Stacked on #8 (fix/imap-import-resilience): it bounds the event queue whose drain-on-shutdown #8 adds. Merge #8 first; this PR's diff then shrinks to its own commit.
Two things let a large mailbox fill memory. IMAP workers handed every
fetched message to the archive writer through an unbounded queue, so fast
workers could hold whole folders' worth of bodies while the single writer
caught up. And fetch batches were sized by count only: 20 items per EWS
GetItem, 8 connections at a time, is a few megabytes of ordinary mail and
several gigabytes of large attachments.
waits for the writer instead of running ahead of it.
helper,
sync::batch::by_count_and_bytes: 32 MiB by default, and asingle larger message goes alone.
in 1000-UID metadata fetches, before the body fetch. A server that won't
say leaves the chunks sized by count.
--fetch-batch-mibsets the cap.item:Sizeand packs GetItem batches by it,fetched
--ews-connectionsbatches at a time.--ews-getitem-batch-mibsets the cap. Items from an incremental SyncFolderItems run carry no
size and stay batched by count.
Tests: 1392 pass (10 new: the batching helper, EWS byte-split GetItem and FindItem size parsing, IMAP byte-split body fetch, and a ten-message chunk through a two-event queue). fmt and clippy clean.
Two things let a large mailbox fill memory. IMAP workers handed every fetched message to the archive writer through an unbounded queue, so fast workers could hold whole folders' worth of bodies while the single writer caught up. And fetch batches were sized by count only: 20 items per EWS GetItem, 8 connections at a time, is a few megabytes of ordinary mail and several gigabytes of large attachments. - The IMAP event queue is bounded at two events per worker, so a worker waits for the writer instead of running ahead of it. - Fetch batches are bounded by bytes as well as by count, with a shared helper, `sync::batch::by_count_and_bytes`: 32 MiB by default, and a single larger message goes alone. - IMAP learns each new message's RFC822.SIZE on the control connection, in 1000-UID metadata fetches, before the body fetch. A server that won't say leaves the chunks sized by count. `--fetch-batch-mib` sets the cap. - EWS asks FindItem for `item:Size` and packs GetItem batches by it, fetched `--ews-connections` batches at a time. `--ews-getitem-batch-mib` sets the cap. Items from an incremental SyncFolderItems run carry no size and stay batched by count. - docs/usage.md describes both options.