jcoffey-dev is traveling from Thursday 1 October through Sunday 4 October. Issues and pull requests are welcome, and will get an answer after that. Thanks for your patience.
Phase 2 of the personal-data catalog spec.
resources/privacy/catalog.toml classifies every object in the schema
(316) and inbuxa's own JMAP objects (12): each property that can hold
personal data, with its categories, and for objects that hold any,
whose data it is, where it lives, its scope and what bounds its
retention (a named setting where there is one). Twenty sources that
are no object -- the log file, exporters, webhooks, spam lookups, the
Explain cache, relays and hooks, push, legacy-use records -- carry the
same facts plus the settings that turn them on, whether the data
leaves the host, and the code that writes it. Classifications of
objects that hold data about people are from the spec's source map;
the rest are typed from the schema alone (address, IP, secret).
tools/fork/privacy-check.py fails CI when an object or inbuxa object
has no entry, when a property the schema types as an address, IP or
secret is left to its object's default, when an entry names an
object, property, setting or code path that is gone, or when it uses
a word outside the catalog's vocabulary. --unlisted prints starting
entries. strip.py's report gains "Unclassified in the privacy
catalog": objects and fields new in an import and not classified,
informational like the Enterprise flags.
Tested: 13 unit tests (tools/fork/tests): the check passes on this
tree; fails on an unclassified object, an address hidden behind a
default, a secret in a set or object reference, stale properties,
objects, settings and code paths, an unlisted inbuxa object and a
word outside the vocabulary; --unlisted's entries; and the strip
report on a synthetic import. The check and the tests run in the
fork-checks job.
A webhook has a level (info by default) that nothing read: its events
were chosen by its list and policy alone. With the default policy,
exclude, and nothing listed, that meant every event type, including
smtp.raw-input (the raw SMTP bytes, DATA included) and the model's
reply to the spam classifier. The docs suggest a webhook to pass the
audit log to a SIEM; set up that way it would have received whole
messages. Found by the personal-data catalog investigation (finding 1).
Now an include list is sent as named, whatever each event's level:
naming an event is the choice. Otherwise a webhook gets only events at
or above its level, as a tracer does, and never a protocol's raw input
or output (IMAP, SMTP, POP3, ManageSieve, delivery, milter), which
carries whole messages and credentials; those go out only when named.
Tested: unit tests for the rule (level, raw I/O only when named, a
named event below the level, custom event levels, a webhook's own
errors); the telemetry system test, whose webhook names debug-level
connection events and still receives them.
The sidecar catalog; the Compliance Officer places and releases holds;
the Tenant Compliance Officer is built now; shortening audit retention
is recorded and surfaced, not gated on a second person; all seven
new-install defaults, in Phase 3; the webhook finding fixed now as a
bug; snapshots kept as long as the audit log.
Phase 1 of the GDPR auditor foundation: the investigation and the
design, committed before anything is built (SPEC.md §3 rule 3).
It maps every place the server stores or sends personal data found
in the code at de275ba, each with its categories, whose data it is,
the settings that control it, what bounds its retention, where it
lives, whether it leaves the host, its scope and the code that writes
it, and the default in a new install. It proposes a sidecar catalog
(resources/privacy/catalog.toml), since the schema and registry code
are upstream's generated output with no generator here; a CI check
modeled on name-check.py; a strip-report section; a read-only
inventory method with dated snapshots; a Compliance Officer role; and
the Compliance navigation with Overview and Data inventory.
Findings worth reading on their own: webhooks ignore levels and, at
their defaults, receive every event including raw SMTP input; log
files are never deleted; automatic bans never expire; some of the
fork's records outlive the account; spam training keeps whole
messages for 180 days; traces are on in a new install; the spam
filter sends IPs, domains, hashed addresses and body digests to
third-party services by default.
Proposed default changes (new installs only) and seven open questions
are for John to decide. No default is changed.
Per-protocol legacy switches (#79) and the hold export's exceptions
list (#75). The prepared Explain answers are relabeled for this
release; 706 carry over unchanged.
The legacy-protocols switch was all or nothing. An operator can now stop
POP3 and keep IMAP: each of IMAP, POP3 and ManageSieve has its own
switch, server-wide on inbuxa:ProtocolPolicy and per tenant on
inbuxa:TenantProtocolPolicy (properties imap, pop3, manageSieve).
legacyProtocols stays as the kill-all: setting it sets all three, and it
reads "disabled" exactly when all three are off. A policy stored before
this has only legacyProtocols and reads as all three at that value, so
existing servers and tenants carry over unchanged. In one /set, a
protocol named beside legacyProtocols overrides it.
SMTP submission keeps no switch of its own: sign-in over it is refused
only when all three are off, as the single switch did (LP-6), so
turning one protocol off never stops a mail app sending. For a tenant,
the server's switches and the tenant's count together.
Server-wide, a change closes the listeners of whatever is now off and
puts back the saved listeners of whatever is on again, both in one
change if asked; listeners of a protocol still off stay saved. Sign-in,
autoconfig, autodiscover, PACC (now prepared once per combination) and
the suggested DNS records all follow each protocol separately. A tenant
may turn a protocol on only while the server has it on (LP-9), and the
refusal names which. The JMAP session adds legacyAllowed, the protocols
still allowed for the account; legacyProtocols there keeps its meaning
for older webmail builds. Events name the switches ("pop3 disabled"),
and audit before/after reads every switch even from an older policy.
Tested: unit tests for the switches, the old-policy reading, the
server/tenant combination, the tenant refusal and listener refusal; and
tests/e2e/legacy_protocols.py against a running server, all 100 checks,
including new ones: POP3 alone off closes only its port and refuses
only its sign-in while IMAP and sending go on; only POP3 stops being
advertised; one change closes IMAP and reopens POP3; a tenant turns
POP3 off for itself, and can't turn IMAP on while the server has it off.
Cloudflare's DMARC report intake rejects every aggregate report we
send with "555 5.7.1 invalid_report_schema". Bisected against the live
endpoint: the only element it objects to is <disposition>pass</disposition>,
the value RFC 9990 added for mail that passed DMARC under an enforcing
policy. The RFC 9990 namespace, <np>, <discovery_method>, <testing> and
a missing <pct> are all accepted, and a report that differs only in
using "none" there goes through.
"none" (no action taken) is valid under both RFC 9990 and RFC 7489 and
says the same thing to the reader, so reports now go out with it. The
stored report keeps "pass"; only the serialized copy changes.
An item the hold covers whose stored record or content can't be read
goes in exceptions.csv with the path it would have had and the reason,
rather than being left out silently. The file is always in the ZIP, so a
header-only one shows nothing was missed, and manifest.sha256 carries
its hash beside the manifest's.
inbuxa:HoldExport/set takes a hold, optionally some of the accounts it
covers, and a reason; the collection runs in the background and get
says when it's ready. The ZIP has, per account, mail as .eml under its
folders, calendars as .ics, contacts as .vcf, files as stored, and the
archived items the hold keeps under archived/; a manifest.csv gives each
entry's account, kind, folder, date, whether it was archived, size and
SHA-256, and manifest.sha256 hashes the manifest. Accounts the hold
doesn't cover are left out, and items outside its date range are too:
live mail by arrival, events by start, and archived items the same way,
so an export doesn't carry deleted items that only another hold keeps.
The finished file is a blob of whoever started the export, so only they
download it, and it lasts as long as any upload (uploadTtl). Exports
are records under the hold (SUBSPACE_INBUXA H/e): never changed or
destroyed, each with its status, counts, size and checksum. Starting
one needs sysLegalHoldExport, an active hold and a reason, and is
recorded in the audit log like the audit log's own export.
The build is in memory and capped at 2 GB; bigger holds fail with a
message saying so, and are split by picking accounts.
Tested: unit tests for safe ZIP names and the manifest and its hash;
the legal_hold system test, on RocksDB, PostgreSQL and MySQL, exports a
hold end to end (live and archived mail, the manifest's hash, an asked-
for account the hold doesn't cover left out) and checks the refusals
(no reason, a user without the permission, a released hold) and the
audit record; and by hand from the console on a local server. Not
covered by a test: the archived-item date range with two holds of
different ranges over one account.
hickory 0.26.3 rejects two kinds of valid answers, and outbound
delivery then retries those hosts until the message expires:
- A zone delegated beneath an unsigned zone (l.google.com under
google.com). Proving the delegation insecure needs an SOA record in
the DS reply, and public resolvers often leave it out. Every Google
MX host behind a signed MX record was unreachable.
- A signed CNAME to a signed name that lacks the queried type. The
NSEC denial is checked against the original name, not the target's.
On a bogus verdict, follow a signed CNAME and repeat the lookup at its
target; otherwise look up the name's zone and its parents, nearest
first. A zone that validates as unsigned means nothing below it can be
signed, so the plain resolver answers and the result is insecure. A
zone that validates as signed first leaves the verdict standing.
Legal holds (#70): a hold on people, groups, domains, tenants or the
whole server keeps everything it covers from being destroyed, by anyone,
until it's released; deleted accounts keep their data. Audit records
name accounts by their full address and holds by their case name. The
daily clean-up of expired archived items works again.
Prepared Explain answers relabeled for this release; no setting changed
since 2026.9.27.2, so all 706 carry over.
An account's or mailing list's name is only its local part, so the log
said "Account ken.gosling" where two domains could each have one; it
now says [email protected]. A change to a legal hold was
recorded under its id; the hold's current state is now read first, so
the record carries its case name and each change reads before/after.
inbuxa:LegalHold/get answers accountsCovered, itemsHeld and sizeHeld
when asked: the accounts a hold reaches now (deleted ones it keeps
included) and the archived items it keeps, with their size. Worked out
in one pass over accounts and archive, only for requests that name them.
Held items stay out of the user's quota, as all archived copies do
(LH-9).
Destroying a held account removes the login, as offboarding needs, but
keeps its data as a deleted account with no expiry, whether or not
undelete keeps accounts; its addresses stay reserved and its holds name
it from then on. Destroy-now refuses it, and its DestroyAccount task
defers itself while it's held or its time hasn't come. Holds placed or
released later freeze or free kept accounts in the same settle pass,
with 30 days' grace after the last release (LH-8, LH-10).
Placing or widening a hold freezes what's already archived in its scope
and range, its old deadline noted; releasing one gives each item no other
hold covers that deadline back, or release plus 30 days if later. One
pass over the archive does both and changes nothing twice (LH-6, LH-10,
LH-11). A held archived item can't be destroyed; restoring still can,
and the hold is named only to callers who may see holds (LH-7). Audit
records about a held account survive the purge (AU-7).
Fixes the daily clean-up of expired archived items (UD-13), which never
found any: the registry's unfiltered query reads an all-ids index that
archived items aren't in. Items are now walked account by account, kept
deleted accounts included. Expired items were still removed whenever
their account's archive was read.
Every way of deleting mail (JMAP, IMAP EXPUNGE, POP3, mailbox removal,
Trash emptying) and Sieve scripts, events, contacts and files now asks
how the account's deletions are kept: a hold keeps them with no expiry
(archivedUntil 9999-12-31), even with undelete off; otherwise undelete's
period applies as before (LH-4).
A hold's date range decides by the item's own date (LH-3). Mail is noted
as held at deletion and settled when it's archived, once its received
date is known; outside the range it gets undelete's deadline or isn't
kept. Events go by their start, with a day's slack for time zones;
recurring events, contacts, files and scripts are held whole.
A groupware item's note now stays until its archive succeeds, and a
failure retries the task instead of being logged and lost (LH-5).
A hold reaches an account by name, through any of its addresses'
domains, its groups or its tenant, as they are now, so an account added
to a held domain later is held too. An account that leaves a held
domain, group or tenant stays held: the registry write hook adds it to
the hold by name on every account change, whoever makes it (LH-2).
Server::holds_on answers for the deletion paths, from the store each
time so a hold binds every node at once.
inbuxa:LegalHold get/set places a hold on accounts, groups, domains,
tenants or the whole server, with an optional date range. A hold's range
and scope can only widen, a released hold is read-only, and none is ever
deleted. Placing, changing and releasing each need a reason and are
audited (LH-1, LH-3, LH-10, AU-12).
Permissions 669-672 (see, place, widen or release, export held data)
go to server administrators only; the tenant ceiling always strips them,
as it does Impersonate (LH-13). Schema: Compliance > Legal Holds.
What a hold keeps comes next, through the undelete hooks.
Also moves the lock expiry helpers below the lock module's imports.
Delegates reach the whole locked account (#68): its calendars, contacts
and files as well as its mail, even a kind it holds none of yet, and
writing delegates may add at the top of its Files.
Prepared Explain answers relabeled for this release; no setting changed
since 2026.9.27.1, so all 706 carry over.
A shared account refuses top-level folders, so an organize or full
delegate couldn't add anything to a locked account with no folders. A
delegate who may write now can, as the owner could; the reconcile after
the create grants it the new folder. Read delegates still can't (AL-6,
AL-7).
A delegate's token listed the locked account only for kinds of data it
held grants on, so one with no files (or no calendar) was refused to the
delegate outright: "You do not have access to account". The token now
lists the locked account for mail, calendars, contacts and files alike,
so an empty kind reads as empty. What the delegate may see or change is
still each container's grant (AL-7).
The audit log (#64): every administrator change, admin sign-in and look
into someone else's data, recorded before it happens, chained per node
and checkable for tampering, exportable with a manifest, kept 2 years.
Locked accounts (#65, #66): an account that keeps receiving mail but
can't sign in and sends nothing on its own, handed to delegates at read,
organize or full, ending at a date when one is set.
Prepared Explain answers relabeled for this release; no setting changed
since 2026.9.27, so all 706 carry over.
A delegation with an end date dropped out of the delegate's token then,
but its folder grants stayed until the daily sweep, so the delegate kept
the account as an ordinary share for up to a day. Each node now sleeps
until the soonest end date, woken early by any lock write and at least
hourly, and re-applies that lock under a cluster-wide claim.
The sweep also had a second-run bug: a delegation past its date gave the
delegate back its earlier share, then dropped the note, so the next sweep
removed that share entirely. The note is now kept while the delegate is
still listed.
A locked account can't sign in (it fails as a wrong password does), its
sessions end on every node, refresh tokens stop working, and its Sieve
scripts forward and reply to nothing. Mail keeps arriving.
Delegates get real ACL grants on the account's mailboxes, calendars,
address books and files at read, organize or full, with the rights they
replaced restored on unlock. Folders made later are granted after the
create and in a daily sweep. Organize delegates can't destroy; send-as
needs organize or full. The JMAP session marks delegated accounts in
urn:inbuxa:jmap.
New inbuxa:AccountLock object with get/set, permissions 665-668, and a
Compliance > Locked Accounts entry in the schema. Lock, unlock and
delegate changes need a reason and are audited; delegate access and
writes are audited too (audit-hold-lock spec AL-1 to AL-12).
What administrators and the server itself do to the control plane is now
recorded, from inbuxa-drafts/specs/audit-hold-lock.md (AU-1 to AU-12):
settings, accounts, domains, roles and every other registry change, with
each field's before and after (secrets only as "changed"); the fork's own
settings objects; administrator sign-ins (and failed ones to administrator
accounts), master-user and recovery-admin sign-ins, once an hour per
account, method and address; access to another account's data through
impersonation or FetchAnyBlob, once an hour; exports and tamper checks;
and registry writes the server makes on its own, named by subsystem
(system:AcmeRenewal, system:auto-ban, system:directory-sync, ...), with a
spam rules update as one summary record.
No change without its record (AU-3): before a set method changes anything,
a pending record per requested create, update and destroy is written; if
that fails, the method is refused with serverFail. Its outcome follows as
a later entry. A change interrupted by a crash stays "unfinished".
Records live in the fork's subspace under L, as one SHA-256 hash chain per
node. The chain's head is stored, never cached, and every append asserts
it, so two writers can't take the same place. Nothing can edit or delete
a record; the daily purge removes the oldest past the retention (default
730 days, minimum 90) and records where the chain now starts, so
verification still passes. security.audit-recorded (647) copies each
record to webhooks, OpenTelemetry and the log; security.audit-write-failed
(648) reports a failed write.
New JMAP objects under urn:inbuxa:jmap: inbuxa:AuditEvent/get and /query
(filters: time, actor, action, target, account, tenant, outcome, address,
text), inbuxa:AuditSettings, inbuxa:AuditExport (CSV or JSON Lines built
on the server, each line with its chain hash, ending in a manifest; the
created object names the blob and its SHA-256) and
inbuxa:AuditVerification. New permissions sysAuditGet, sysAuditExport and
sysAuditSettingsUpdate: the Administrator role gets all three, the Tenant
Administrator role gets read and export, once, on existing installs too.
A tenant administrator sees records whose actor or target is in its
tenant, including a server administrator's changes there.
Sign-in method on the session: access tokens now remember how they signed
in (password, app password, API key, OAuth client, directory, master user,
recovery admin), including across the HTTP credential cache. New OAuth
access tokens carry their client id in the sealed claims; older ones show
as client "unknown" until they expire.
The schema gains the permissions, the two events and a Management >
Compliance > Audit Log link.
Stack: the request layer boxes every inner future where it's made. Without
that, a debug build overflowed the default 2 MB worker stack on a registry
set; measured with the same request, the branch and main now overflow at
the same stack size (between 1856 and 1920 KiB, debug), so the layer adds
nothing measurable.
Tests: unit tests in inbuxa-features and jmap; system::audit::audit_log_tests
(run with --ignored) passes on RocksDB, SQLite, PostgreSQL, PostgreSQL with a
read replica, MySQL, MySQL with a replica and FoundationDB. The system, JMAP
and SCIM suites pass. authorization.rs skipped fork permissions that guard
no registry object; the audit suite checks a plain user is refused instead.
inbuxa's own mark (#62): the kitten over a server with a bay for each
piece of the suite, on the built-in sign-in and RSVP pages, the web
logo and the email logo.
Prepared Explain answers relabeled for this release; no setting changed
since 2026.9.26.1, so all 706 carry over.
inbuxa's mark was ihasmail's cat-and-envelope reused unchanged. The new
one keeps the family's face, paws and colors, over a server with a bay
for each piece of the suite: the letter (webmail), a prompt (console),
status lights (server).
- The built-in sign-in and calendar RSVP pages, and the web logo, drew
the old cat as an embedded PNG. They now draw the mark as vector in
the same slot, keeping class="symbol"; each page is about 31 KB
lighter. The .min copies are updated the same way and the .min.gz
regenerated with gzip -9 -n, as minify_html.sh does.
- resources/branding: email-logo.png (the compact lockup at 380x80 on
white, as before) with its .b64 regenerated byte-for-byte in the old
76-column form, and favicon-64.png.
- img/brand: the logo bundle, now pure vector, with its README.
announce.yml runs coffey-labs/actions discourse-release on every published
release, posting it to this project's Announcements category on
community.coffeylabs.org. The release workflow also announces
from its own job, since a release made with the job token fires no
'on: release' workflow in Gitea.
Prepared Explain answers relabeled for this release; 11 settings whose
default is the time of creation drop out, since their answers could never
match.
ai-explain spec, amendment 1 (EX-22 to EX-28):
- answers are three or four sentences, max_tokens 160, cut at 700 chars;
- POST /api/explain streams the answer as server-sent events;
- each node remembers answers in memory (1,000, 24 h), keyed by the facts,
prompt version and model, shared by server-level administrators;
- resources/explain/settings.json.gz ships answers for settings at their
defaults, generated with prepare_setting_explanations (717 for 2026.9.27);
- the system prompt no longer carries the per-request marker, so a model
server can reuse it;
- inbuxa:Explanation gains source, answeredAt and preparedFor.
When a listener couldn't bind its address (a port below 1024 without
root, a port already in use, or the legacy-protocols switch putting a
listener back after privileges were dropped), the bind error was
reported but the socket was still passed to listen(). The kernel then
bound it itself, to a random port on every interface, and the server
logged the listener as started on the port it was configured with.
listen() now refuses a socket that isn't bound, so the listener is
reported with a listen error and skipped, and nothing opens anywhere
unexpected.
A singleton such as x:SpamSettings has no stored object until someone
saves it; /get shows its defaults instead. Explain looked only for the
stored object, so every setting still at its defaults answered "No
such x:SpamSettings." It now falls back to the defaults the same way.
A new method, inbuxa:Explanation/set, asks the node's local model for a
short plain-words reading of one thing an administrator is looking at:
a failed recipient in the queue, a Classify verdict, a log line or trace
event, or one setting with its saved value. The server builds the prompt
itself from stored data and the registry schema, never from text the
console sends, and grounds SMTP replies in RFC 3463 and RFC 5321.
What the model is never shown: secrets (including ones nested inside a
setting, like an AI model's HTTP auth), raw protocol events, and the
contents of any other event. A tag name that doesn't have a tag's shape is
refused before a model is asked.
Calls share the AI gate with spam classification, but mail always keeps
its slot, and Explain has its own hourly count per account and its own
on/off switch in inbuxa:AiLimits. The permission is sysAiExplain,
superuser only; tenant administrators can't use it. The session carries
an aiExplain flag so a console knows when to offer the button.
An install whose roles were stored before the permission existed gets it
added once, at start-up, to the roles that are administrators' alone,
not the User role their defaults share with every account. An operator
who removes it later isn't overruled.
Tests: unit tests in inbuxa-features and jmap, and ai_explain_tests
(run with --ignored) covering the acceptance tests and the upgrade.
Every node renews its lease once a minute instead of every 30 minutes,
so the lease works as a heartbeat. x:ClusterNode reports a node Stale
once it has gone three minutes without renewing (it used to take an
hour), and Inactive after a day, as before.
Taking over a lease still needs a full hour of silence. A node that is
slow rather than gone never loses its id to another host, so snowflake
ids stay unique.
The admin dashboard's Cluster Health card counts these statuses.
After #37, address fields on PostgreSQL are split into words as the
built-in index splits them, but language text (subject, body,
attachments) still goes straight to PostgreSQL's parser, which keeps a
URL, host, path or file name as tokens of its own:
"https://x.example/shipping-support/" becomes a url, a host and a
url_path, "invoice-2024.pdf" a file. So TEXT/BODY "shipping" missed
messages where the word appears only inside a link, while RocksDB and
the other built-in backends found them: 8 messages across a handful
of searches in the rehearsal.
On insert, language text is now indexed as it was, followed by the
word parts of each token that holds a URL separator (/ . @ : ? = & # _
% + ~ \), split with SpaceTokenizer as keyword_terms() splits addresses.
The parts go through the same text search configuration as the rest of
the text, so they are stemmed like the words around them. Plain words,
words that only carry punctuation ("end.", "(see") and hyphenated words
(the parser already splits those) add nothing, so text without links
is indexed exactly as before. Each part is added once per document.
On sample mail, the text vector of a short order notice with three
links grows from 546 to 716 bytes, a newsletter with 25 tracking links
from 5586 to 6430, and a plain letter not at all.
On search, a query word written as a URL, host, file or hyphenated word
also matches as its word parts, ORed with the query as written, so
"shipping-support" or "invoice-2024.pdf" match the new parts and
documents indexed before this change still match as they did.
Existing messages keep their old vectors until they are reindexed (the
reindexAccounts task); new and reindexed messages match at once.
store::search_tests gains test_url_word_search: five bodies, 19 body
searches for words found only in a URL path, query string, host or
file name, the tokens as written, plain words and non-matches, with the
same expected ids on every backend. It passes on RocksDB, SQLite,
MySQL and PostgreSQL; on main PostgreSQL fails at the first ("shipping"
finds [3], not [0, 3]). On PostgreSQL the suite then stops at the
account sort assertion (query.rs:689) exactly as it does on main.
A cluster rehearsal sent ten x:<Object>/set requests at once and got
ten full reloads on every node. #39's coalescing only joined writes
that queued behind a running reload, but the requests reached the
server about 33 ms apart and a reload takes tens of milliseconds, so
none overlapped one.
A full reload after a registry write now waits for writes to settle:
75 ms after the last one, and at most 250 ms after the first it
covers, so a steady stream still reloads at least four times a
second. 75 ms is a little over twice the gap the rehearsal saw between
requests. A single write pays it once: in the tests a settings write
takes about 140 ms instead of 60. The reload runs in a task of its
own, so a request that goes away doesn't cancel it for the others.
Each write takes the result of the first reload that started after it
was stored (the gate keeps the last 64 results), so applied true or
false still describes the reload that covered that write.
The 33 ms gap was a queue on the server, not password hashing: Basic
credentials are cached per Authorization header, so they are checked
once. Every authenticated HTTP request counted itself against the
account's rate limit by incrementing one counter per account in the
in-memory store, so parallel requests from one account queued on that
key: a row lock on PostgreSQL (a few round trips to the database
each) and conflict retries with a 50-300 ms backoff on RocksDB. An
account with the unlimitedRequests permission (administrators, by
default) passes the rate and concurrency limits anyway, so its
requests are no longer counted. Ten parallel Core/echo calls as the
admin now finish in 1-4 ms; before, they finished one after another
over 20 ms on a local PostgreSQL and 300-450 ms on RocksDB. Other
accounts still count every request.
system::auto_reload::settings_reload_tests: ten concurrent writes now
take one reload (the gate counts them; at most two allowed), all are
applied: true and in the running settings, and a single write takes
exactly one reload. RocksDB and PostgreSQL, 1 reload in 141-196 ms.
With the old behavior (no wait, requests counted) the same writes
took 5 reloads; without the wait but with the rate fix, 2.
cluster::broadcast (3 nodes, PostgreSQL + NATS) and system::reload
still pass.
A cluster rehearsal moved a Log tracer to another directory: the write
was reported x:settingsReload applied:true, but the tracer kept writing
to the old file until a restart. Telemetry::update only refreshed each
running tracer's events, level and lossiness; a tracer's own settings
(path, prefix, rotation, format, endpoint, headers, ...) stayed as built.
Each tracer now carries a hash of the registry object it was built
from, less the fields that change in place. The reload compares it with
the running tracer's: unchanged ones are updated in place as before,
changed ones are started over, new ones started and removed ones
stopped. Only tracers this server started are removed; upstream removed
every subscriber not in the settings, which also cut off live-tracing
streams on each reload.
Starting over is a swap in the collector, so no event is lost or
written twice: a subscriber registered under a running one's id
replaces it between two collection passes. The old one's batch is sent
first (what its full channel can't take moves to the new one), and
dropping it closes its channel, so its task writes what is queued and
ends. Per tracer kind:
- Log: a tracer started over on the same files (rotation or format
changed) waits for the old one to finish, so lines don't interleave.
- Webhook: the task held a sender of its own channel for retries, so
it never ended; retries now use a weak sender, and pending events are
posted when the channel closes.
- OpenTelemetry: pending logs and spans are exported when the channel
closes instead of dropped, and a span that was open across the swap
is exported by the new tracer with the events it saw.
- Console and journal: nothing kept between batches.
- Trace history: built from the tracing store, which takes a restart,
so it is never started over.
No kind needs a restart, so x:settingsReload doesn't gain one.
system::tracer_reload::tracer_reload_tests (new): a Log tracer created
over JMAP writes to its directory; its path is changed over JMAP while
2000 numbered events are emitted; after the reload, events land in the
new file and not the old one, each numbered event is in exactly one of
the two files, and a destroyed tracer writes nothing. On main the new
file never appears.
The publish workflow only built a tag whose commit is on main. That keeps
every image tied to reviewed code, but it means production can only get a
fix together with everything that has landed on main since its release.
A tag on a release/* branch is now accepted too. A hotfix branch starts at
an earlier release tag, takes fixes through pull requests into it (so the
code is still reviewed and CI-tested before it is tagged), bumps
brand_version! and is tagged there. The tag must still equal
v<brand_version!>, and the step prints which branch it was found on.
A tag runs the workflow file from its own commit, so a hotfix branch that
starts before this change needs this commit cherry-picked onto it before
its tag is pushed.
The report scheduler dropped DMARC and TLS events on a node whose role
lacks outboundMta (upstream never started it there, so they sat in a
channel nobody read). Mail received on a front node therefore never
reached an aggregate report, which is meant to cover all of a domain's
inbound mail, whichever node received it. In rehearsal, five messages
received on port 25 on a front node were missing from every report.
- The report scheduler records on every node. Recording is a store write
the nodes already share, so it needs nothing from the outbound MTA.
Building and sending a report (the DmarcReport and TlsReport tasks) stay
with outboundMta nodes, as the task manager already enforces.
- More nodes now append to one report at once. Appends already guard the
report's versioned primary key; a write that loses now retries up to ten
times after a short random pause, not three times at once.
- The node sending a report deletes it only if it is unchanged since it
was read, and reads it again otherwise, so a record another node appends
meanwhile goes out with the report instead of being deleted unsent.
Test: cluster::front_reports (PostgreSQL and MySQL). A front node's
results appear in the report the MTA node sends, alongside eight appended
at once from both nodes, and the front node never runs the report task.
It fails on main: the front node's results are never recorded.
Setting deliverAt on an internal DMARC or TLS report wrote the new task
queue row with the report's object type (0x21, 0x6e) instead of the task
type (7, 8), and left the task row at its old due. The task manager's scan
failed on that row with store.data-corruption ("Failed to iterate over task
queue"), and because the error ended the whole scan, every task due after
the row stopped running on every node.
- reschedule_ops writes the new queue row through schedule_task_with_id, so
it carries the task type and the task row gets the new due. It removes
the row the task is actually queued under (the task's due, which differs
from deliverAt once the task has been retried) and any row an earlier
reschedule left at deliverAt.
- x:DmarcInternalReport/set and x:TlsInternalReport/set lock the report's
task while they move it, as x:Task/set does, refuse while the report is
being sent, release the locks however the request ends, and wake the task
manager.
- The task manager logs a queue row it can't read (id, due, key, value) and
skips it instead of ending the scan. It then repairs the row from its task:
the row is rewritten with the task's type, and a row with no task behind
it is removed. A row holding a report's object type for a report task is
what the old reschedule wrote: the task is moved to that row's time, as
the reschedule intended, and its old queue row is removed. Stores that
already hold such a row recover on their own once it comes due.
- x:Task/query with a type filter skips an unreadable row instead of
failing.
Test: smtp::reporting::reschedule (RocksDB and PostgreSQL). It fails on
main: x:Task/get shows the old due, and with that check removed, neither
report nor a later task ever runs.
In cluster rehearsal 3, turning outboundMta off on node1's role was
reported applied (x:settingsReload applied: true), yet node1 kept
delivering mail, a report message included, until it was restarted.
The queue and report managers were started at boot only when the
node's role included outboundMta (crates/smtp/src/lib.rs), and the task
manager only when the role had some task type (spawn_task_manager).
After that nothing looked at the role again: a queue manager that was
running kept claiming and delivering, and one that wasn't never
started.
They now start on every node (outside recovery mode) and follow the
role live:
- Queue manager: before each scan it reads the role from the running
settings. Without outboundMta it claims nothing new; deliveries
already running finish and report back as usual, which releases
their locks. When the role comes back (a reload wakes the manager
with ReloadSettings, and it looks again every 30 s regardless) it
logs queue.started and scans the whole queue at once.
- Report scheduler: DMARC and TLS report events are handled only while
the role has outboundMta, as at boot; events arriving without it are
dropped, as they were on a node started without the role.
- Task manager: task_enabled already read the current role on every
scan. It now also runs on nodes whose role has no task type (the
scan returns at once until one is added), a job claimed before a
role change is handed back at once rather than run or held until
its lease lapses, and a settings reload wakes the manager so a role
that gained task types starts claiming them straight away.
Starting the queue manager on every node also drains the queue channel
on nodes without outboundMta. Upstream left that channel unread, so
each message queued there parked a refresh in it, and by the code,
queueing would block once 1024 had piled up (not reproduced here).
A role object edit reaches the nodes that name that role in
INBUXA_ROLE. Moving a node to another role still means changing its
environment, and so a restart. Listener changes in a role still need a
restart too (listeners bind at boot); this change is about tasks and
delivery.
cluster::live_roles::live_role_tests (new; PostgreSQL, two nodes over
one store):
1. A node started with outboundMta delivers and runs a TLS report
task; after its role loses outboundMta and the settings reload, a
new message isn't attempted and a new report task stays pending;
with the role back, both are taken up.
2. A node started with no task type at all gains outboundMta: a
waiting message is attempted and a report task runs.
On main the test fails at step 1 ("delivery attempted without
outboundMta"); with step 1 bypassed, step 2 fails (nothing picked the
message up in 20 s).
Cluster rehearsal 3: with PostgreSQL paused (docker pause, so its
kernel still answered TCP keepalives), requests on connections already
checked out hung until it came back, and /healthz/ready stayed 200
through the outage. #41 bounded getting a connection, not using one.
Client-side query limits (store::backend::query_timeout). Every
operation on a PostgreSQL or MySQL connection now runs under a time
limit. A server-side statement_timeout (or MySQL's MAX_EXECUTION_TIME,
which covers SELECTs only) can't do this: the server that would enforce
it is the one not answering. When an operation runs out, its connection
is closed instead of pooled, since a query may still be in flight on it
or a transaction open: deadpool's Object::take on PostgreSQL;
Conn::disconnect on MySQL, which marks the connection closed before it
sends anything, so the pool discards it even when the server never
answers.
- query, 2 minutes: reads, writes (the whole transaction with its
retries), blobs, SQL lookups, search queries and indexing. These take
milliseconds; two minutes leaves room for a large blob over a slow
link and still ends a hang.
- maintenance, 30 minutes: range deletes (account removal, purges),
unindexing, purge_store, and creating tables and indexes at startup,
which can legitimately run long in one statement. Their existing
chunked fallback for server-side statement timeouts is unchanged.
- iterate (exports, reindexing, maintenance scans) can run for hours,
so the query limit bounds each wait for the database (preparing, the
query starting, the next row) rather than the whole scan.
The limits are fixed, like the pool timeouts; the DataStore schema has
no field for them. Tests set them with Store::with_query_timeouts
(test_mode only).
Readiness. /healthz/ready answered 200 whenever a data store was
configured. It now reads one key from the data store with a 2 s limit
and reuses the answer for 2 s, so probes can't load the database;
while one probe runs, others get the last answer. The first failed
probe of an outage is logged. /healthz/live stays 200: restarting a
node doesn't bring its database back, and an orchestrator restarting on
failed liveness would restart every node at once. The container
HEALTHCHECK already uses /healthz/live.
Tests, store::pool_timeout (a proxy that stops forwarding while
keeping connections open plays the paused database):
- postgres_query_timeout, mysql_query_timeout (new): with four pooled
connections open, a read, a scan and a write each fail with "Query
timed out" 2.0 s after the pause (2 s test limit); once the proxy
forwards again the store answers. With the limits set to an hour
(upstream's behavior), the read was still waiting at the test's 20 s
limit.
- postgres_readiness (new, STORE=PostgreSql): a node's data store
goes through the proxy; /healthz/ready is 200, 503 about 4 s after
the pause while /healthz/live stays 200, and 200 again about 2 s
after it ends.
- postgres_pool_timeout, mysql_pool_timeout: pass as before.
store::store_tests (PostgreSql, MySql, including the MariaDB statement
timeout step) and store::task_locks (PostgreSql) pass;
store::search_tests (PostgreSql) fails at the same ordering assertion
(query.rs:684) as on main.
A three-node rehearsal on PostgreSQL saw searches take about 185 ms
with 80 to 260 pages in the full-text indexes' pending lists, 2 to 6 ms
right after gin_clean_pending_list() or VACUUM, then creep back up as
mail came in. The search tables' GIN indexes were created with the
default fastupdate=on: new entries wait in an unindexed pending list
that every search scans in full until VACUUM (or 4 MB of backlog)
merges it, and autovacuum only visits an insert-only table after
thousands of inserts.
The search GIN indexes are now created WITH (fastupdate = off), so an
insert pays its index update at once. The schema step runs at every
startup (create_search_tables, via SearchStore::create_indexes), so
indexes made before this change are switched there: when an index's
reloptions don't already turn fastupdate off, ALTER INDEX ... SET
(fastupdate = off) and one gin_clean_pending_list() merge its backlog.
The ALTER takes a SHARE UPDATE EXCLUSIVE lock, which blocks neither
reads nor writes; after the first startup the step is one catalog read
per index. A failure is logged and startup goes on (search still
works, only slower).
Per-table autovacuum settings for the search tables are left alone.
The pending list was the only reason the insert threshold mattered for
search; dead tuples and freezing are served by the defaults, and table
settings would override whatever tuning the DBA has done.
MySQL is unaffected: InnoDB FULLTEXT keeps new entries in an in-memory
cache that queries read directly, with no setting like fastupdate.
store::search_gin::postgres_gin_fastupdate (new, PostgreSQL) builds
the search schema in a schema of its own and checks pg_class.reloptions:
fastupdate=off on every GIN index of a fresh schema; then, with the
option reset to the default and 500 rows pending, one startup turns it
off everywhere and leaves no pending tuples (pgstatginindex); a second
startup changes nothing. On main it fails at the first check.
write_reload_target sent AllowedIp writes to the blocked-IP reload, but
that reload rebuilds only BlockedIps. Allowed IPs are parsed into the
core's security settings (Security::parse), which only a full reload
rebuilds, so an AllowedIp write reported x:settingsReload applied: true
while the change wasn't live until the next full reload.
AllowedIp now maps to the full reload, like the other settings objects;
BlockedIp keeps its targeted reload.
system::auto_reload::settings_reload_tests now creates an allowed IP
over JMAP and checks that is_ip_allowed sees it with no ReloadSettings,
and that destroying it takes it out again. On main it fails ("allowed
IP not in the running settings").
A 3-node rehearsal (PostgreSQL + NATS + Garage) found two ways a crash
leaves work stuck:
Pool hangs. The PostgreSQL pool (deadpool) was built with no timeouts,
so a request waited for a free connection, and for one to be opened or
recycled, for as long as it took: forever when the server stopped
answering. MySQL's pool (mysql_async) has no wait timeout at all.
- PostgreSQL: wait 30 s (or the store's timeout if longer), create the
store's timeout or 15 s (it bounds the whole handshake, where
tokio-postgres's connect_timeout covers only the TCP connect), recycle
10 s. The pool config is now always set, not only with
poolMaxConnections.
- MySQL: every connection is taken through MysqlStore::conn(), which
gives up after 30 s.
- Both: TCP keepalive after 60 s idle, so a server that vanished
without closing the connection is noticed in minutes rather than the
two-hour system default.
The DataStore schema has no pool timeout settings, so these are fixed
defaults; the store's own timeout bounds connecting on PostgreSQL.
Task locks. A task lock lasted an hour, so after a hard crash the dead
node's tasks waited up to an hour and five minutes. The lock is now a
five-minute lease: while this node runs a task, the task manager renews
its lock every third of the lifetime (InMemoryStore::renew_lock, a
compare-and-set on the store backends and SET XX EX on Redis, which
leaves a lock that already expired alone). A killed node's tasks run
elsewhere within about five minutes plus the claim recheck. A task this
node holds isn't handed to a worker again by the scan.
store::pool_timeout (new): a local listener that accepts connections
and never answers plays a hung server; a PostgreSQL store with a 2 s
timeout returns an error in about 4 s, and a MySQL store in 30 s.
Without the timeouts both wait for good. store::task_locks gains a
task held for 1.5 lock lifetimes: its lease is still held, and released
when the task ends.
A 3-node rehearsal found taskQueueProcessing didn't filter anything:
roles.task_manager only decided whether the task manager started, and
report, ACME, DKIM, DNS, calendar, thread-merge and restore tasks ran on
any node with a task manager (manager.rs returned true for them). A node
whose role left taskQueueProcessing off still ran them if it indexed or
did maintenance.
Every task type now answers to one ClusterTaskType (task_enabled):
- IndexDocument, UnindexDocument, IndexTrace: searchIndexing
- AccountMaintenance, TenantMaintenance, DestroyAccount:
accountMaintenance
- StoreMaintenance: storeMaintenance
- SpamFilterMaintenance: spamClassifierTraining
- DmarcReport, TlsReport: outboundMta. They build and send reports to
other domains (TLS reports can go straight to an HTTPS endpoint),
which is the outbound MTA's business.
- CalendarAlarmEmail, CalendarAlarmNotification, CalendarItipMessage,
MergeThreads, RestoreArchivedItem, AcmeRenewal, DkimManagement,
DnsManagement: taskQueueProcessing, the role for queue tasks with no
role of their own.
A node that may not run a task leaves it unclaimed (no lock), so a node
that may picks it up. The task manager also starts on a node whose only
task role is outboundMta, so reports still run there.
cluster::task_roles::task_role_tests (new, two task managers over one
PostgreSQL store): node A (taskQueueProcessing only) runs a DNS task and
leaves an unindex task and a TLS report pending; node B (searchIndexing
and outboundMta) comes up and runs those two; a DNS task scheduled next
stays pending on B and runs on A. On main node A runs the TLS report.
A 3-node rehearsal found that saving an MtaDeliverySchedule left it
unknown to the queue ("Queue strategy not found") until someone ran
x:Action ReloadSettings; only Directory and Authentication writes
reloaded (DIR-17). The admin UI has to remember a separate reload after
every save, and a script or API client that doesn't gets a server
running stale settings.
x:<Object>/set now reloads the running settings when it created,
updated or destroyed an object they are built from, and broadcasts the
same RegistryChange::Reload over the coordinator as ReloadSettings, so
every node applies it:
- Settings objects (MTA, spam filter, listeners, tracers, Sieve system
scripts, cluster roles, directories, ...: the object types the core,
telemetry, listener and directory builders read) get a full reload.
- Certificates, lookup stores and blocked/allowed IPs get their own
targeted reloads.
- Accounts, domains, roles and other data read as needed, stores (they
take a restart) and applications (their own reload action) get none.
Full reloads are coalesced: a write waits for a reload that started
after it was stored and joins one if it can, so a burst of writes, or
a request with many objects, costs one or two reloads, not one each.
The write itself is never undone. When the reload is refused (build
errors in objects that were working, the rule from the previous
commit), the set response says so in a new x:settingsReload field,
{"applied": false, "description": "Saved, but the running settings
were not reloaded. <object>: <error>"}; {"applied": true} otherwise.
The field is absent when the write needs no reload. The description
helper is shared with ReloadSettings' refusal.
Each reload sends the queue a ReloadSettings event, so the SMTP test
harness's read_event, try_read_event and assert_no_events now pass over
those; expect_reload_settings still waits for one.
system::auto_reload::settings_reload_tests (new): an MtaVirtualQueue
and an MtaDeliverySchedule created over JMAP are in the running
settings with no ReloadSettings, and gone once destroyed; eight
concurrent creates all land; a write whose reload fails is stored and
reported applied: false with the error; a domain write carries no
x:settingsReload. On main the new schedule is missing. The cluster
broadcast test (three nodes, PostgreSQL + NATS) now checks that every
node has a schedule created on node 0 without a reload.
A 3-node rehearsal found every settings reload refused, cluster-wide,
because one node couldn't resolve the Pyzor server:
- PyzorConfig::parse resolved the host while building the settings and
made a failed lookup a build error. It now keeps the host and port and
resolves when a message is checked (an IP address is used as is, a
name is reused for five minutes, the lookup counts against the Pyzor
timeout). A failure there is a Pyzor error for that message.
- A milter's hostname was resolved the same way, with a blocking
to_socket_addrs in async code. An IP address is kept; a name is now
resolved on each connection.
Other build-time I/O is already non-fatal: directories that can't
connect become unavailable with a warning (DIR-21), and the AI model
locality check only warns.
reload_registry swapped the core only when the whole build was free of
errors, while boot runs with whatever built. One failing object thus
refused every later reload, and the running settings went stale. Now a
reload is refused only for errors in objects that built when the
running settings were built (at boot or by the last applied reload):
applying it would lose those. Objects that already failed then are
missing from the running settings anyway, as at boot, so their errors
are logged and returned as known_errors but don't hold the reload back.
Refusing on new errors keeps a bad edit from taking a working object
out of service; the admin gets the error instead.
ReloadSettings now says "Settings were not reloaded." and names the
object and its error ("Tracer with id ...: Only one console tracer is
allowed"), with a count of any further errors. A refused reload after a
directory change logs its errors too.
system::reload::reload_tests (new): with Pyzor enabled on an
unresolvable host, ReloadSettings succeeds (on main it fails with
"Invalid address: failed to lookup address information"); an IP host
needs no lookup; a new build error refuses the reload, names the object
and leaves the running settings unchanged; the same error, once known
from the running settings' build, no longer blocks; once fixed, a new
error there blocks again. smtp::inbound::milter's session test now
names its milter "localhost", so the connect-time lookup is exercised.
A 3-node PostgreSQL rehearsal found IMAP SEARCH FROM "noreply" matched
0-2 messages where RocksDB matched 23 of 930. The message indexer hands
each address and display name of From/To/Cc/Bcc to the search store as
keyword text (Language::None). The built-in index splits keyword text
into lowercase runs of alphanumerics, so an address is found by its full
form, its local part, its domain or a display-name word. The SQL
backends didn't:
- PostgreSQL's text parser keeps "[email protected]" as one email
token (host names and URLs likewise), so neither "noreply" nor
"amazon.com" ever matched it. Keyword text is now split the same way
as the built-in index (SpaceTokenizer) before to_tsvector on insert
and before plainto_tsquery/phraseto_tsquery on search, still under
the 'simple' configuration, so the GIN index keeps serving the query.
The sort columns keep the raw text.
- MySQL's FULLTEXT parser already splits on punctuation, but InnoDB
never indexes its stopwords ("com", "de", "www", ...) or words under
innodb_ft_min_token_size (3), and a required +word it hasn't indexed
matches no row. So "amazon.com", "[email protected]" or "jane doe" found
nothing. Those words are now matched with a word-boundary REGEXP on
the rows the indexed words select. In language text (bodies,
subjects) they are dropped when other words remain, and only checked
when nothing else is left, so "the invoice" no longer finds nothing
either.
Existing PostgreSQL search indexes hold the old single-token vectors and
need a reindex (the reindexAccounts task) before address searches find
old messages. MySQL needs none: only the query changed.
store::search_tests gains test_address_search: five messages, 28
FROM/TO/CC/BCC searches by full address, local part, domain, domain
labels, display name and hyphenated local part, plus a TEXT-style OR,
with the same expected ids on every backend. It passes on RocksDB,
SQLite, PostgreSQL and MySQL; on main it fails on PostgreSQL (From
"noreply") and MySQL (From "[email protected]").
A node that started while NATS was down never got a coordinator. The
connect failed at boot, bootstrap recorded a build error and the node ran
with Coordinator::None until restarted. It had no broadcast subscriber
or publisher, so cross-node push and cache invalidation to it stayed
broken, and its healthcheck said nothing about it. Losing NATS after
startup was silent too.
- The NATS client now connects in the background
(retry_on_initial_connect): startup never waits on NATS or fails over
it, the node gets its coordinator, subscriber and publisher at once,
and the client keeps trying (async-nats's backoff, at most 4 s apart)
until NATS answers. Subscriptions made meanwhile start delivering when
it does. A configured maxReconnects still ends the attempts.
- Three new events report the connection: cluster.coordinator-connected
(info), cluster.coordinator-disconnected (warn: lost, closed, gave up,
or not connected within the connection timeout at startup) and
cluster.coordinator-error (warn: a failed attempt, reported once per
outage rather than every retry, and server errors, slow consumers and
lame duck mode). They are in the packaged schema, ids 644 to 646.
- GET /healthz/cluster reports the coordinator: 200
{"coordinator":"connected"}, 503 {"coordinator":"disconnected"}, or
200 with "none" (no coordinator) or "unknown" (a backend that doesn't
track its connection). /healthz/live and /healthz/ready are unchanged
on purpose: a node without its coordinator still serves mail, and
failing those would have orchestrators restart, or pull out of
service, every node at once whenever NATS is down.
Only NATS connects lazily; the other coordinator backends still fail at
boot as before.
cluster::coordinator::coordinator_reconnect_tests starts a node against a
NATS port with nothing behind it, checks it boots with a coordinator and
reports it disconnected, subscribes, then starts NATS on that port: the
node connects on its own and the subscription receives a message from a
second client. Stopping and restarting NATS shows disconnected, then
connected, and the same subscription keeps working.
A cluster rehearsal (PostgreSQL + NATS) left index tasks pending well
past the one-hour task lock after the node that claimed them was stopped
or killed. The exact cause there isn't confirmed; this closes every path
found in the task manager that stretches a takeover past the lock, or
keeps a task claimed without running it:
- A graceful stop never released the locks it held, so every task the
node had claimed stayed blocked for an hour. The server now tracks the
locks it holds (common::ipc::TaskLocks) and, once the shutdown signal
arrives, stops claiming and releases them before exiting.
- A node that failed to claim a task (another node held it) set its own
local hold for a full lock lifetime from that scan. If the holder
claimed it just after the scan began, or ran on a clock ahead, that
hold ran out a moment before the lock did and was set for another
hour: two hours in all. Such claims are now tried again every five
minutes (a twelfth of the lock lifetime), and the task manager wakes
up for them: before, a node without a coordinator could sleep up to
five minutes past the recheck, or until something else woke it.
- A worker that panicked took its task type down on that node for good,
while the scan kept claiming that type's tasks and failing to hand them
over, re-taking each lock as it expired and so starving every other
node of them. Each batch now runs on a task of its own; a panic is
logged, the batch's locks are released and the worker carries on. A
failed hand-over releases the lock too.
- A claimed task the worker couldn't read, or found gone, kept its lock
for the hour. It is released.
- An IndexDocument task for a file (not indexed) returned no result,
which shifted every later result in the batch onto the wrong task in
update_tasks. It returns Ignored. Nothing queues such a task today.
The lock lifetime stays one hour; it now lives per server so the tests
can shorten it.
store::task_locks::task_lock_tests plays a second node by writing its
locks straight into the in-memory store: tasks it claimed and abandoned
run here once its locks expire, including locks that outlive this node's
view of them, and a graceful stop hands this node's locks back at once
and claims nothing more. It passes on RocksDB, SQLite and PostgreSQL.
With the old recheck it fails.
Two release builds side by side on one machine each take twice as long,
and production only needs amd64. publish-amd64 now pushes :<version> as
soon as the amd64 build is done; publish-arm64 builds arm64 afterwards,
then replaces :<version> with the two-platform index and moves :latest.
Both jobs use one named BuildKit builder whose container outlives the
job, so the dependency layer (cargo chef cook) is reused until the
dependencies change. The release is created after amd64; the binaries
are attached once arm64 is in.
The trace index task wrote the event type (its name) and the queue id as
text, but the tracing search index types both as integers on every
backend: BIGINT on PostgreSQL and MySQL, long on Elasticsearch. On
PostgreSQL every batch holding a trace document failed with "cannot
convert between the Rust type String and the Postgres type int8", and
since a batch writes trace and email documents together, email indexing
stalled behind it.
The document is now built by trace_search_document(), which writes:
- the event type as the opening event's numeric id, the event
x:Trace/query's event filter already matches on;
- the queue id as an integer, the first one the trace names;
- every queue id into the keywords as well, since the column holds one
value and an SMTP session can queue several messages.
index_keyword() replaced the field on every call, so before this only the
last event type and queue id survived anyway.
x:Trace/query's queueId filter parses the id (a string, or now a number)
and matches the column or the keywords, so a session is found by any of
its queue ids on every backend. The monitoring spec says what is indexed.
Traces indexed before this on the built-in index keep their text values;
the reindexTelemetry maintenance task rebuilds them.
Tests: the search store suite builds trace documents with the index
task's code, indexes them and finds them by queue id, event type and
keyword (Sqlite, PostgreSQL, MySQL); the monitoring suite finds a real
trace by queueId through x:Trace/query.
The broadcast subscriber waited 1 << retry_count.max(6) seconds between
failed subscribe attempts. max(6) turns the cap into a floor: the first
retry waited 64 s instead of 1 s, and each later one doubled without a
bound (and would overflow the shift after enough failures).
The delay now comes from subscribe_retry_delay(), 1 s, 2 s, 4 s ... capped
at 64 s, and the retry counter saturates. A unit test pins the schedule
and the top of the range.
--export skipped three things, so a move from one database to another
(RocksDB to PostgreSQL, say) lost them without a word:
- archived items (subspace j), the records behind undelete;
- spam training samples (subspace w);
- the trained spam classifier and its trainer state, blobs stored under
fixed names that no blob link points at, so the walk over links never
reached them.
j and w now travel with the registry family, where their indexes and id
counters already were, so EXPORT_TYPES=registry keeps them consistent.
The two named blobs travel with the blob family. The file format is
unchanged and import reads any subspace it is given, so an export made
by an older binary still imports.
The full-text index (subspace z) stays out, on purpose. It belongs to one
search backend: PostgreSQL and MySQL index into their own tables and have
no z table at all, and external engines keep the index themselves. So
--import now returns the subspaces it wrote, and boot queues the
reindexAccounts and reindexTelemetry store maintenance tasks, the same
ones an administrator can queue by hand, to rebuild the index for
whichever search store the server runs with once it starts.
The round trip also turned up a loss in import itself: the SQL stores
add a negative amount with an UPDATE, which does nothing to a row that
isn't there yet, so every negative counter or quota vanished on import
into PostgreSQL, MySQL or SQLite. Import now creates the row first.
The in-memory subspaces (m, y) stay out: rate limits, locks, greylisting,
ACME challenge tokens and OAuth codes, all short-lived. Issued
certificates are registry objects and travel.
The store test now writes archived items, spam samples, directory
entries, the fork's own subspace and the named blobs, checks they come
back in place, then imports the same export into a fresh store of the
other local backend (RocksDB to SQLite, or SQLite to RocksDB), compares
it key for key and counter for counter, and checks the queued reindex.
It fails on the old export code ("Subspace j was not exported").
--help now says what an export holds.
#27 let the build context see vendor/, but the Dockerfile cooks the
dependencies before it copies the tree, from a recipe that carries only the
workspace's manifests. [patch.crates-io] points sieve-rs at vendor/, so the
cook failed the same way: failed to read /build/vendor/sieve-rs/Cargo.toml.
That's why 2026.9.24.2's publish failed. The builder stage now copies
vendor/ before cooking; a local build got past it into compiling the
dependencies.
context-check.py now also checks that each patched path is copied into the
cooking stage before the cook, and fails on the Dockerfile as it was.
Replaces 2026.9.24, whose tag predates the image build fix (#27) and never
published. Carries everything 2026.9.24 did -- upstream 0.16.23 and its
fixes, the scim release-profile fix -- and since then:
- identifiers renamed from the upstream name, with no aliases: the JMAP
registry capability is urn:inbuxa:jmap:registry, WebDAV tokens
urn:inbuxa:dav*, Sieve extensions vnd.inbuxa.*, the web interface client
inbuxa-webui; INBUXA_* settings only. Deploy with admin and webmail
releases that use the new names.
- the brand in lowercase where people see it.
- the spam filter rules bundled with the server; on first start they add
the AI classifier's LLM_* scores.
- a Local AI page link in Settings › Spam Filter, for the admin release
that draws it.
- two start-up migrations: the spam model moves to its renamed keys, and
the web interface's old OAuth client is retired.
Adds a link to CustomComponent/LocalAi in the packaged schema's Settings ›
Spam Filter, above LLM Classifier, and updates the schema hash so admins
fetch the new layout rather than a cached one.
INBUXA Admin draws the page (feature/local-ai-setup); this makes it
reachable. An admin from before that page would show "Unknown component"
here, so this lands after the admin release that carries it.
The server fetched upstream's latest published rules from GitHub at run
time: a version nobody here tested, code-like expressions from an account
we don't control, and the upstream name as a default in the admin form.
The published rules of spam-filter v3.0.2 are now embedded
(resources/spam-filter/, MIT, in THIRD-PARTY.md) and used whenever no other
source is configured. An empty setting and upstream's old default both mean
the bundled rules, so existing installs switch without a settings change;
the URL stays an operator override (https:// or file://). The schema default
is dropped and its description says what empty means, and the strip's
rename pass does the same to each import.
Rules load on first boot as before, and again whenever the bundled version
differs from the last one loaded, which only adds missing rules and tags.
That brings the AI classifier's LLM_* scores to installs that predate them:
production has none today.
upstream-watch now also opens an issue when spam-filter publishes a newer
release; resources/spam-filter/README.md says how to take it.
The antispam test now runs on the bundled rules, the path production
takes; SPAM_RULES_URL tests another set. Unit tests cover the URL handling
and that the bundled rules parse and score the AI tags as the AI spec says.
The rename pass vendored a patched sieve-rs and pointed Cargo.toml's
[patch.crates-io] at vendor/sieve-rs. .dockerignore ignores everything and
re-includes a short list that did not have vendor on it, so the image build
had no such directory and stopped at
failed to load source for dependency `sieve-rs`
failed to read /build/vendor/sieve-rs/Cargo.toml
CI could not have caught that: it builds from a checkout, where the
directory is simply there, and only the image build has a context to prune.
The first that was known about it was a tag that had already been pushed.
So: vendor is re-included, and tools/fork/context-check.py now asserts the
thing that was quietly assumed -- every path a [patch] section names exists
and survives .dockerignore. It runs beside the other fork checks and takes
no toolchain.
Also, the comments in .dockerignore started with // , which Docker does not
read as a comment: they were patterns that happened to match nothing. They
are # now.
v2026.9.24 was tagged on a commit CI had passed, and its release build could
not compile crates/scim at all:
error: queries overflow the depth limit!
= note: query depth increased by 130 when computing layout of
{async fn body of context::<impl ...>::writable_domain()}
The crate is ours, and the failure is profile-dependent: the release profile
computes those async fn layouts in one go and goes past rustc's default query
depth, while the dev profile never gets that far. CI builds dev, so CI was
green on a commit that could not be released. The tag produced no image and
no release, which is the one merciful part.
Two changes:
- #![recursion_limit = "256"] on the crate, which is what rustc itself
suggests, with a note saying why it only shows up in release. Proved by
building -p scim in release locally: it now finishes.
- CI builds the release profile too, on pushes to main. Pull requests stay
on dev, where the wait is worth less. A few minutes per merge is cheaper
than learning this from a tag, which throws away a multi-architecture
build and leaves a version half-cut.
The server sets a deferred recipient's next retry from its clock when the
attempt defers, in whole seconds. The test subtracted its own clock taken
when the loop next saw the message, after saving and reporting, so
whenever that lag crossed a second boundary the 2 s retry measured 1 s
and the test failed. Under load, after the other SMTP tests, that was
most runs.
It now measures from when the test started the attempt, which the
server's deferral can only follow, by under a second: each retry is its
interval or one more. Each position is still checked against its own
interval, so a wrong schedule still fails.
It listens on HTTP 19048, as the dkim2 DSN and report tests do. Those are
marked serial, this wasn't, so when the SMTP tests ran together (as
upstream's CI runs them) it could start beside them and requests reached
whichever server had the port: missing JMAP creates here, and 'You must
authenticate first' in dkim2_dsn_is_signed.
It failed everywhere but upstream's machines, for two reasons:
- The spam rules, which carry every score, came from a path on an
upstream developer's own disk. Without SPAM_RULES_URL none loaded, every
score was 0.00 and the combined case came out ham instead of spam at
13.70. The published rules of spam-filter v3.0.2 are now pinned beside
the test cases (Apache-2.0 or MIT, taken as MIT; in THIRD-PARTY.md).
SPAM_RULES_URL still overrides.
- The first combined case expects a Pyzor hit, and its digest (that of an
empty body) wasn't among the three the test mode answers, so it went to
a public Pyzor server: it failed offline and would drift with that
server's counts. Test mode now answers every digest from a fixed table,
with the empty body's added, and never reaches the network.
The test passes online and offline, alone and with the rest of the SMTP
tests. queue_retry, unrelated, still fails when it runs after the others
in one process, though it passes alone every time.
The name is inbuxa, lowercase, like the wordmark; INBUXA reads as an
acronym. The admin and webmail already changed. Here that's everything
the server shows people: the brand macro behind the protocol greetings,
the HTTP and SCIM realms, the startup banner and the calendar and contact
PRODID; the first-party OAuth client descriptions; the legacy-protocol
refusals; the default calendar and address book names and the SMTP
greeting default, in the code and the schema served to the admin
(checksum regenerated); startup and shutdown events; the User-Agent;
the sign-in and RSVP pages; the service units; the OpenAPI realm; the
crate descriptions and the README, where it's set in bold.
Identifiers that are uppercase for their own reasons stay: INBUXA_*
settings, SUBSPACE_INBUXA. So do code comments and the AGPL 5(a) notice
lines.
Tests follow: the IMAP ID name, the default collection names, the PRODID
in the iTIP fixtures and the CalDAV free-busy expectations, and the e2e
legacy-protocol refusals. The webdav, imap and jmap suites pass, so do
the unit tests of every crate touched, and 73 of 75 SMTP tests; of the
other two, antispam fails on main too, and queue_retry is a timing flake
that passes on its own.
Carries upstream 0.16.23 -- the DSN, POP3, Sieve, DMARC-report, ACME and
DNSSEC-resolver fixes in its own change log -- with the files it changed
marked under AGPL section 5(a), and one upstream test dropped that the fork's
routing makes meaningless.
It is also the first release whose tag attaches binaries: a host install can
now fetch inbuxa-linux-amd64.tar.gz or inbuxa-linux-arm64.tar.gz instead of
pulling the image and copying the file out of it.
Everything clients, users and operators meet now carries the fork's name,
with no aliases (SPEC.md §2.4, changed here from "protocol identifiers
stay"):
- JMAP: upstream's registry capability is urn:inbuxa:jmap:registry, beside
the fork's own urn:inbuxa:jmap.
- WebDAV lock and sync tokens are urn:inbuxa:dav*; clients resync once.
- Sieve: vnd.inbuxa.while and vnd.inbuxa.expressions. sieve-rs spells these
into its compiler, so it's vendored (vendor/sieve-rs, 0.7.3) and patched in;
a unit test fails if Cargo.lock ever moves past the vendored copy. The
trusted runtime now names itself too, rather than answering sieve-rs's
default.
- The web interface's OAuth client is inbuxa-webui. On every start the old
stalwart-webui client is removed and any application naming it is moved
over.
- The spam filter's blobs are INBUXA_SPAM_*; every start moves any left
under the old keys, so a trained model survives.
- SQL stores and log files default to inbuxa, in the code and in the
schema served to the admin (checksum regenerated).
- Settings are INBUXA_* only. A STALWART_* variable that's set where its
INBUXA_* one isn't stops the server at startup, naming it.
- The version-upgrade messages link docs.inbuxa.org's migration page, and
the OpenAPI description, smtp crate metadata and web-push test fixtures
lose the name.
Kept on purpose, allowlisted with reasons: the OAuth key-derivation
contexts (renaming them would end every session and invalidate every
sealed client id) and the hashed application prefix.
Also fixes a latent start-up failure: ensure_client updated an existing
first-party client with a revision of 0, which the registry's assertion
never matches, so adding a redirect URI or changing the webmail secret
failed start-up. And the principal session test now expects
legacyProtocols (C-1, added 2026-09-21), which it had missed.
Tested: the server builds without warnings; common's 106 unit tests,
including the vendoring check; a new integration test for the two
start-up migrations; and the webdav, jmap, imap and SMTP Sieve suites.
A release published an image and nothing else, so there was nothing for a
host install to download -- the only way to get the binary was to pull the
image and copy it out, which makes "install without Docker" depend on
Docker.
Each release now carries inbuxa-linux-amd64.tar.gz, inbuxa-linux-arm64.tar.gz
and SHA256SUMS, named as stalwart-migrator's are.
They are taken out of the image this pipeline just pushed rather than
compiled again. A second Rust build per architecture is the slowest thing
here, and it would leave two artifacts that are meant to be the same build
and only probably are. Extracting makes that identity a fact: the binary in
the tarball is the file the image runs. `docker create` starts nothing, so
copying a file out of an arm64 image on an amd64 runner needs no emulation.
One thing the extraction cannot carry: the image grants the binary
cap_net_bind_service, and a tar archive does not keep that xattr. The
release body says so, and says what to do instead -- setcap, or
AmbientCapabilities in the unit -- because a server that cannot bind 25 and
does not say why is a bad first hour.
Checked by hand against v2026.9.23 before this landed: both architectures
extract to the right ELF, and the amd64 binary runs on a bare Debian 13 with
every library resolved and reports its own version.
strip.py compiles the stripped tree, so a dual-licensed file that only
serves an Enterprise feature fails the import instead of the merge, as
v0.16.23's tests/src/directory/issuer.rs does. Upstream's tests of the
features the fork rebuilt are expected not to compile there and are listed
in build-check-known.txt; an error anywhere else fails the run. Checked
against both imports: v0.16.22 passes with its 16 expected errors, v0.16.23
fails on issuer.rs alone. Imports the strip leaves unused are reported.
It also renames the upstream name where clients, users or operators meet
it as an identifier, from tools/fork/renames.py: wire-protocol names, the
web interface's client id, store keys, configuration defaults and the
served schema. main is renamed with the same module, so a re-import
arrives purged and those lines don't conflict.
notice-check.py fails CI when an upstream file the fork changed, measured
against the upstream branch, lacks its AGPL 5(a) notice; --fix adds it.
It runs beside the name check in a renamed fork-checks job.
Also commits v0.16.23's strip report under docs/fork/strip-reports/, which
the import in #18 left out.
tests/src/directory/issuer.rs, new in v0.16.23, tests routing a bearer token
to a directory by its issuer. That routing is Enterprise-only upstream (the
body of get_directory_for_issuer), and the fork doesn't build it: a token
naming no address gets the server default (DIR-2). The test also calls a
helper from upstream's Enterprise-only OIDC test, so it can't compile here.
mta.rs imported types::id::Id for code inside an Enterprise snippet; the
stripped tree leaves it unused, upstream's as well as ours.
These upstream files were changed after the fork marked the files it had
modified, and never got the notice: six by the listener and schema-cache
work on 2026-09-20, two by the name check. Found by diffing against the
upstream snapshot branch, as before.
Five conflicts, resolved:
- crates/common/src/auth/authentication.rs: upstream's get_directory_for_token
and JwtClaims replace extract_jwt_domain; the per-domain directory code
(DIR-1, DIR-5 to DIR-7) is kept, and the token lookup routes through it.
The release's one new Enterprise snippet was the body of
get_directory_for_issuer, which stays returning None: a token naming no
address gets the server default, as DIR-2 specifies and as v0.16.22 did.
- crates/common/src/manager/application.rs: upstream's rewrite of the tests,
with the temp directory names renamed again, and the 5(a) notice the
name-purge change should have added.
- crates/common/src/network/mta.rs: both sides' imports.
- crates/main/Cargo.toml: the AGPL-only license kept, version 0.16.23.
- Cargo.lock: upstream's, with the fork's crates added by Cargo.
tools/fork/name-check.py reads every string literal in crates/ (comments
and test directories skipped) and fails on any that carries the upstream
name without an entry in name-allowlist.txt. An upstream merge can bring
such strings in without a conflict, so it runs on every push and PR.
The first run found three the earlier sweeps missed, fixed here: the SMTP
HELP reply pointed at upstream's website (now brand_url!), the event
collector thread was named after upstream, and the FreeBSD default data
path still said /var/db/stalwart/ where Linux already had /var/lib/inbuxa/.
Two operator-visible defaults are allowlisted as open, pending a decision:
the log file prefix and the SQL stores' default database and user.
Reads metadata only: upstream's releases list from GitHub's API and the
head of the upstream branch from Gitea's. Nothing of upstream's is
fetched, so its history can't land here. Daily at 06:17 UTC.
The first-party application descriptions and the telemetry service name and
instrumentation scope are shown to operators, and the unpacked-application
temp directory carried the name too.
Left alone deliberately: the OAuth key-derivation contexts (renaming them
would invalidate every sealed token and client id), the migration defaults
that read an upstream installation, links to upstream's upgrade guide, the
wire-protocol identifiers, and upstream's own license and templates.
brand_version_full! is user-visible -- --version, the startup banner, the
console, telemetry and the JMAP session's implementation field -- and the
name belongs only in copyright notices and the lineage line.
publish.yml replaces .github/workflows/publish.yml: on a v* tag it checks the
tag equals v<brand_version!> and is on main, builds the linux/amd64+arm64
image in one buildx run (the Dockerfile already cross-compiles, so only its
final stage goes through QEMU), pushes :<version> and :latest to the
registry, links the package, and creates the tag's release if it has none.
weekly-release.yml ports .github/workflows/release.yml: bump brand_version!
through the contents API, then create the release and so the tag, which
starts publish.yml. It only dry-runs until RELEASE_LIVE=1 and a
RELEASE_TOKEN secret exist.
GitHub took the organization's repos and GHCR offline on 2026-09-20. Repo,
release, raw-file and clone links now go to Gitea at git.coffeylabs.org,
container images to registry.coffeylabs.org, and GitLab-style /-/blob paths
to Gitea's /src/branch form. Go module paths are identifiers and stay as
they are; links to GitHub issues and pull requests are left as history.
The build now mounts the named volume inbuxa-server-cargo at /cache and keeps
CARGO_HOME and CARGO_TARGET_DIR there, so a push reuses the compiled
dependency tree (RocksDB included) instead of rebuilding it from scratch.
Both runners allow that one volume; each host keeps its own copy.
With the cache in place the job moves to runs-on: light, so it can run on
host2 as well. Cargo's parallelism now follows the job's CPU cap rather than
the host's core count, and the target dir is dropped past 25 GB.
Deleting a tenant now also removes its stored inbuxa:TenantProtocolPolicy,
in the same place the registry's other per-type clean-ups run. Without it
the row outlived the tenant, and a tenant that later came to have the same
id would have started with legacy protocols off.
The e2e deletes a tenant whose switch a server administrator had turned
off, and would check that a new tenant with the same id starts with them
on. On this build the registry hands out a fresh id instead ("d" after
"c"), so the reuse -- and with it the removal -- isn't observable over
JMAP; the test says so rather than passing silently. The risk it guards
was therefore smaller than feared, and the change is mostly about not
leaving an orphaned row behind. All 72 checks pass.
The Hardening link merged to main changed the packaged schema, which this
branch also changes. The file is gzipped, so the two can't be merged line
by line: this takes main's schema and adds security.legacy-protocols-changed
to it again, with the hash recomputed.