jcoffey-dev is traveling from Thursday 1 October through Sunday 4 October. Issues and pull requests are welcome, and will get an answer after that. Thanks for your patience.
The v2026.9.30 tag build's arm64 publish job was killed linking the
inbuxa binary (fat LTO, one codegen unit): cannot allocate memory on the
16 GB ubuntu-24.04-arm runner. index, ghcr, release, binaries and
announce were skipped. amd64 got through on the same size of runner.
Each publish job now adds a 16 GB swap file before the build; buildx's
container has no memory limit of its own, so the linker can use it.
Each node stores histograms as running totals since it started. A sample
didn't say which node wrote it (the node was only in the id's low bits),
so a reader couldn't diff totals per node, and the console diffed across
nodes: on the three-node production cluster the delivery attempt time
read 14.7 s over the last hour against 0.7 s from the nodes' own figures.
x:Metric/get now returns nodeId alongside timestamp, both from the id.
The telemetry suite checks every sample carries it.
The Gitea registry stays authoritative; GHCR becomes a copy of it, the way
the GitHub repository is a copy of the Gitea one. After the tag build has
pushed the release image to the registry, a new ghcr job copies it to
ghcr.io under the same version tag and :latest with `imagetools create` --
a copy, not a rebuild, so the digest on GHCR is the digest on the registry.
Anything still pulling the old ghcr.io name, including the TrueNAS app
submission, keeps receiving releases. The job uses the run's own token and is
left out of the status reported to Gitea, so a GHCR problem cannot fail a
release.
The mirror carries tags to GitHub but not releases, so the replica's
Releases page -- and anyone watching the repository there -- stopped at the
last release made on GitHub. After the tag build has published, a new
github-release job copies the tag's Gitea release to a GitHub release: the
same notes, with PR and issue numbers rewritten to Gitea links, the same
files, and a line pointing back to the Gitea release.
It uses the run's own token and is left out of the status reported to
Gitea, so it cannot fail a release. With no Gitea release for the tag it
does nothing.
GitHub scopes a run's Actions cache to its ref, so the cache a tag build
wrote could only ever be read by that same tag: the next release built cold
anyway. Each release also parked several GB of Rust layers in the
repository's 10 GB cache, enough to evict main's cargo cache and slow
everyday builds too. The image builds now run without a cache.
The github job only polls Gitea for GitHub's commit status, but it holds a
runner slot for as long as the GitHub build takes -- the better part of an
hour for a cold build. On the shared build runners a handful of those
could take every slot and stall real work, so it now runs on the `wait`
label: a runner of its own, with many slots, no docker socket and a small
CPU and memory cap.
The mirror can push one commit twice in quick succession. GitHub then
starts two runs and cancels the older, and that run's report job posted
"failure" for the commit. Gitea's github job, seeing the newest status,
failed the check while the surviving run was still building and later
passed.
A cancelled run now posts nothing and leaves the result to the run that
superseded it. A real failure still reports failure.
This repository is now push-mirrored to GitHub, where issues and pull
requests would never reach the maintainers. A note under the title says
where development happens, and sends issues to git.coffeylabs.org and
discussions to community.coffeylabs.org.
Gitea stays where the project lives and push-mirrors every branch and tag
to GitHub. With the Actions variable BUILD_ON set to 'github' on both
forges, the GitHub copy does the building and reports back to Gitea as a
commit status; unset, nothing changes and Gitea builds as before.
.github/workflows/ci.yml replaces the GitHub-era files. Branch pushes run
what Gitea's ci.yml checks (fork checks, dev build, test targets, the
release profile on main). v* tags run what publish.yml does, with the same
two guards: the image per architecture on native runners side by side,
the multi-arch index and :latest, the Gitea Release if the tag has none,
and the host-install binaries taken out of the image. A final job posts
"github/ci (branch)" or "github/ci (tag)" to the commit on Gitea.
On Gitea, the heavy jobs skip under BUILD_ON=github and a `github` job
waits for that status and passes or fails with it, so pull requests and
merges still look at a Gitea run. The weekly release, the upstream watch
and the announcement stay on Gitea.
Removed: cleanup.yml and publish.yml (GHCR), release.yml (a second weekly
schedule), and dependabot.yml, whose pull request branches every mirror
sync would delete.
Documents under C-5 why oAuthClientOverride counts only in bootstrap
and recovery mode, and adds tests/e2e/client_override.py: the
recovery administrator keeps the override in both modes; after setup,
an administrator gets no code for an unregistered client or a
redirect URI its client didn't register, and a device code approved
for an unregistered client can't be exchanged. The script fails
against a build without the change (3 of 8) and passes with it.
The recovery administrator signs in before any OAuth client is
registered, so it needs to skip the registration check. Outside
bootstrap and recovery mode, every account now signs in through a
registered client and one of its redirect URIs.
Anyone could host a copy of a front end on a server of their own,
collect a person's password there, and replay it as HTTP Basic against
JMAP or the API. Cross-origin rules don't stop that, since a server
isn't a browser, and neither does client registration, since Basic
never goes through OAuth (contract C-23).
JMAP (session, API, upload, download, event source, WebSocket), /api,
/auth/introspect, /auth/userinfo and authenticated /auth/register now
refuse an Authorization: Basic header before looking at the password,
with a 401 whose only challenge is Bearer. A wrong password gets the
same answer as the right one. CalDAV and CardDAV keep Basic, and their
401s still offer it. The sign-in page's /api/auth takes the password in
its body and is unaffected, as is the token endpoint's client
authentication.
Bootstrap and recovery mode accept Basic everywhere, as they keep
permissive CORS. INBUXA_HTTP_BASIC_AUTH=all puts it back everywhere;
dav is the default, and any other value logs a warning and keeps it.
Test builds accept Basic everywhere, since the integration suites sign
in with passwords, and legacy_protocols.py sets the variable.
Tested: unit tests for the paths, and tests/e2e/http_basic_auth.py
against the debug build, 26 checks, including both front ends' sign-in
path and a refused unregistered redirect.
Phase 4 of the journaling spec.
- inbuxa:JournalEntry/query and /get (sysJournalSearch): filter by time,
sender, recipient, either, direction, subject words, Message-ID and
journal, newest first; the whole report only when asked for.
- inbuxa:JournalExport/set (sysJournalExport): a reason is required; a
ZIP of the matching reports with manifest.csv, exceptions.csv and
manifest.sha256, up to 10,000 reports and 1 GB.
- inbuxa:JournalVerification/set (sysJournalGet): chains and reports
rechecked.
- Every search, listing, read, export and check is written to the audit
log before anything is returned, with existing actions only.
- Catalog entries for the three objects; spec as-built notes.
journal_tests: administrators can't search; a Compliance Officer searches,
lists, reads a report, exports (reason required) and checks the chain;
the officer can't change journals; each of those is in the audit log.
Phase 3 of the journaling spec.
- A journal's destination: builtIn (true for journals stored before) and
archiveAddress, at least one. Reports to an archive are queued from the
empty sender, one per address, flagged so they're never journaled.
- A pending record per report. When the queue lets go of one without
delivering it (refused, expired, deleted), it becomes its own entry in
the built-in journal under the sending journals' retention, the
journal's archiveFailures (count, last time, reason) goes up, and the
audit log records it; if that can't be written it stays queued.
- Journal it: a rule action naming a journal, on mail flow rules and
beside a DLP rule's block, warn or hold. A journal whose scope chooses
nobody takes only what rules send it.
- The report lists recipients a rule added or redirected to under
"Added by rule", by rule name.
- A rule's route is cleared between messages in one SMTP session, with the
new journal marks; a second message used to keep the first one's route.
tests/src/system/journal.rs: destination validation, a rule-only journal
fed by a rule that also adds a recipient, an unreachable archive's report
kept in the built-in journal with the failure counted, a report delivered
to an archive here and not journaled itself.
Phase 2 of the journaling spec.
- A copy of each message is taken in MessageWrapper::queue, after DLP and
transport rules, for every enabled journal that takes it (direction and
scope: everyone, or accounts, groups, domains, tenants). If the copy
can't be taken the message isn't queued (temporary failure).
- The journal report: the envelope one field a line (sender, To, Cc, Bcc
from the envelope, list members from their ORCPT, direction, held for
review), then the queued message byte for byte as message/rfc822.
- The built-in journal under J in the inbuxa subspace: one chain per node
whose links name each entry by SHA-256, so entries can expire out of
chain order; purge leaves a marker, and verify catches an entry changed
or removed early and a report that doesn't match.
- Retention per journal (30 to 3650 days); an entry keeps what it was
written with. The daily maintenance purges what's due, keeping entries
whose people a legal hold covers (deleted accounts a hold keeps too),
and records the counts in the audit log.
- inbuxa:Journal get/set, audited by the request layer. Permissions
680-683: administrators see and change journals; the Compliance Officer
sees, searches and exports. Whoever changes journals may grant search and
export without holding them, so officers can still be appointed.
- Catalog entries (inbuxa:Journal, source "journal"); spec as-built notes.
tests/src/system/journal.rs: validation, internal mail with a Bcc,
outgoing into two journals, incoming over LMTP, the report and its
original, tamper and early removal caught, hold-aware purge, retention
changes leave entries alone, disabled and removed journals take nothing.
senderGroup, senderTenant and recipientGroup conditions now read and write group and tenant ids as JMAP ids ("b", "c"…), like legal hold scopes and the rest of the API, so the console can use its object pickers; plain numbers are still read. Held as numbers for matching. Unit test for both forms and a bad id; mail_rules_tests round-trips a tenant condition over JMAP.
inbuxa:DlpSettings (singleton, urn:inbuxa:jmap): keepHeldDays, 1 to 90,
7 by default (settled answer 5 made it a setting). sysDlpPolicyGet reads
it, sysDlpPolicyUpdate changes it, server-level, audited by the request
layer. Each held message keeps the days it was given, and the sender's
notices say that number. Privacy catalog entry; spec §2.6 updated.
mail_rules_tests: 7 by default, 0 refused, 3 set and a message held
afterwards expires 3 days after it was held, the expiry notice says 3.
An EmailSubmission create's response carries inbuxa:held (dlp-and-mail-flow-rules spec, §2.6, §4): true when the message is held for review, false otherwise, so the webmail can say so at once. A sender can't read the review queue, and a held message's sendAt is its real send time, not the century-off release, so this is how the sender learns. mail_rules_tests checks both values.
Schema layout entries for the console pages of the DLP and mail flow rules spec (§3): Held Mail and Data Loss Prevention under Management > Compliance after Legal Holds, and Mail Flow Rules beside the server Sieve scripts (the console's nine-group Settings bar places it under Mail flow). An older console shows these as unknown pages, so this ships with the console that has them.
The hold action now holds (dlp-and-mail-flow-rules spec, §2.6), where
until now it blocked.
- At DATA a hold decision queues the message with its release a century
off (the queue's future-release mechanism, so the stored format is
unchanged and an older node just never sends it), transport rules
still applied, and replies 250 Held for review. A review record under
R/h + queue id keeps the sender, recipients, subject, size, rules and
detector counts. The sender is told when the rule asks.
- smtp/queue/held.rs: release (each recipient due now, its next notice
as far off as it was, its lifetime counted from the release), reject
(removed from the queue, the sender told, with the reviewer's note),
and expiry: the daily clean-up rejects what nobody reviewed in 7 days,
recorded as the server's doing.
- inbuxa:HeldMessage get/set: the review queue, sysDlpReviewGet to list
and read (preview, 64 KB of text, only when asked for and recorded as
blobAccess), sysDlpReviewUpdate to release or reject, a reason
required and audited by the request layer; no create or destroy;
server-level only.
- Guards: Emails > Queue refuses to change or delete held mail; the
sender can't unsend it.
- Privacy catalog entry for inbuxa:HeldMessage; spec §2.6 as built.
Tests: mail_rules_tests gains the whole flow (held and listed with
counts, sender notified and nothing delivered, queue and unsend
refused, preview recorded, reject needs a reason and tells the sender
the note, release delivers, expiry returns it, decisions audited with
reasons). smtp inbound, system_tests (after one BlobNotFound in
antispam, the known flake, then clean), features and common unit tests.
Phase 2g of the DLP and mail flow rules spec: transport rules now act,
on outgoing and incoming mail.
- features/mailflow/rewrite.rs: add or remove a header, prefix or set the
subject (an RFC 2047 word when not ASCII), add a disclaimer. A
disclaimer edits the message's main text and HTML bodies only, each
decoded, changed and written back as UTF-8 quoted-printable with its
other headers kept, top or bottom (after <body> or before </body> in
HTML); attachments and attached messages are left alone, and a
disclaimer already present isn't added again.
- smtp/inbound/mailflow.rs: the check runs for incoming mail too
(transport rules only; DLP stays outgoing). After DLP passes, each
matched transport rule's actions run in order: message edits,
add-recipient and redirect (envelope changes DATA applies), route (a
per-message queue ahead of the queue strategy), refuse (550 5.7.1
with the rule's text). The override tag is stripped with the same
subject writer, so a non-ASCII subject stays valid.
- Audit: refusals and changes to where mail goes are recorded (sender,
or system:mail-flow for incoming mail); wording and header changes
aren't, or a banner rule would record every message (spec §2.7).
Tests: rewrite unit tests (headers, encoded subjects, disclaimers on a
single part and on multipart/alternative with an attachment, once
only); mail_rules_tests gains the actions end to end: disclaimer,
header and subject prefix on a delivered message, a redirect, a
refusal, a banner on incoming LMTP mail that outgoing rules leave
alone, and which of those are audited.
Phase 2f of the DLP and mail flow rules spec: the rules now run on mail
an authenticated sender submits, after the DATA system script and
before headers and DKIM signing (§2.1).
- smtp/inbound/mailflow.rs: builds what the rules look at from the
message (subject, the text version of each body, one level of attached
messages, attachment text via the extractor, 10 MB of text at most)
and the envelope (sender's groups and tenant; each recipient local or
not, and its groups). Skipped entirely when no enabled rule applies to
outgoing mail. Rules that can't be loaded refuse with a 451: nothing
unchecked leaves.
- Block: 550 5.7.1 with the rule's notice. Warn: 550 5.7.1 with the
notice and how to override: "[override: reason]" at the start of the
subject, taken out before the message goes on (settled answer 1).
Until phase 3, a hold rule blocks rather than let mail through.
- JMAP: EmailSubmission takes inbuxa:dlpOverride {reason}; a refusal
comes back as inbuxa:dlpWarning or inbuxa:dlpBlocked with each rule's
name and notice (description too, for older clients).
- Audit: one record per DLP match, the sender as actor, action create,
target a message: the recipient domains, each rule with its detectors'
counts, the outcome, an override's reason. Never the matched text. No
new audit action: an older node that meets one fails its daily
clean-up, which would make rolling back unsafe (spec §2.7 updated).
Tests: mail_rules_tests gains the DLP flow over JMAP (no rules, warning
with rule and notice, local recipient not warned, override with a
reason, block that no reason passes, the subject tag stripped from the
delivered message, audit records with no card or key text). smtp
inbound tests pass; system_tests passed twice after one timeout in the
email delivery tests that didn't recur.
Phase 2e of the DLP and mail flow rules spec, the API half.
- inbuxa:MailRule/get and /set under urn:inbuxa:jmap. Rules convert
through serde, so what a client sends is the stored format. A create
or change is validated whole (Rule::validate) and refused with the
property at fault; id, createdBy, createdAt and updatedAt are the
server's. Every change goes through the request layer's audit record.
- Six permissions, ids 674-679 (enum and schema labels): mail flow rules
(sysMailRuleGet/Update), DLP rules (sysDlpPolicyGet/Update) and held
mail (sysDlpReviewGet/Update, for phase 3). Either kind's permission
gets through the gate; the handler shows and changes each rule only
with its own kind's. All server-level: a tenant is refused (settled
answer 3).
- Administrators get all six; the server-level Compliance Officer gets
DLP rules to see and held mail to review (settled answer 4), added
once to an existing server's officer role by the grant mechanism,
which gains an officer audience.
- Privacy catalog entry for inbuxa:MailRule.
tests/src/system/mail_rules.rs: create, list in order, validation,
server-set properties refused, update, kind-separated permissions for
an officer, destroy, audit records.
Phase 2e of the DLP and mail flow rules spec, in the features crate.
- rules.rs: a rule (§2.2) with its conditions (§2.3) and actions (§2.4),
as JSON under R/r in the fork's subspace. validate() enforces the
spec's shape: DLP rules check outgoing mail and have exactly one of
block, warn or hold; transport rules have neither those nor
detectors; lists, header names, header values (one line), addresses,
texts, word lists, patterns and detector ids are checked.
- engine.rs: rules compiled once (word lists to automata, patterns to
size-limited regexes) and run in priority order with exceptions and
stop processing. Each detector runs at most once per message and
only when a rule asks for it. The outcome lists what matched with
each detector's count, and decides DLP strictest first: block, hold,
warn; an override answers warnings only (§2.5).
- cache.rs: each node's compiled copy, refreshed after 30 seconds or at
once when this node changes a rule.
Nothing calls this yet: the JMAP object and the check at DATA follow.
55 unit tests in mailflow.
Phase 2b of the DLP and mail flow rules spec: every identifier in the
§2.3 catalog, each implemented from its issuer's published rules and
tested against published examples.
US (SSN, ITIN, EIN, ABA routing, driver's licenses, MBI, NPI, DEA), UK
(NI number, NHS number, UTR), Canada (SIN), Australia (TFN, Medicare),
the EU (Germany's tax ID and ID card, France's NIR, Spain's DNI/NIE,
Italy's codice fiscale, the Dutch BSN, Belgium's national number,
Poland's PESEL, Sweden's personnummer, Denmark's CPR, Finland's HETU,
Ireland's PPS, Portugal's NIF, Austria's SVNR), Norway, Switzerland,
India (Aadhaar, PAN), China, Japan, Singapore, South Korea, Brazil (CPF,
CNPJ), Mexico (CURP) and South Africa. 49 detectors in all, plus seven
templates named for what they find.
An identifier that is only digits and whose check about one random
number in ten passes counts alone only in its written form
(536-22-1234, 943 476 5919) and as bare digits only beside a word; ABA
routing numbers and NPIs always need one. Spec §2.3 records this.
A test runs every detector over an ordinary business email (order,
invoice and tracking numbers, dates, amounts, an address) and requires
nothing to fire but the contact detectors. 47 unit tests.
Phase 2a of the DLP and mail flow rules spec: pure functions in
crates/features/src/mailflow, nothing wired into the mail path yet.
- Detectors report distinct values found, each either checked by its
published check digit or counted only beside a corroborating word
within 50 characters. This PR adds the region-free ones: payment
cards (issuer prefixes, Luhn), IBAN (registry lengths, mod 97),
SWIFT/BIC, email addresses and phone numbers in bulk, dates of birth,
passport numbers, private keys and published service-token formats.
Regional identifiers follow, a region per PR.
- Word lists (Aho-Corasick, whole words, any case) and patterns (regex
with a compiled-size limit) count occurrences.
- Attachment text: text files with or without a UTF-16 mark, HTML,
DOCX/XLSX/PPTX, ODT/ODS/ODP and ZIP archives one level deep, read
with the zip and quick-xml crates the workspace already has.
Encrypted files, PDF, legacy binary Office files, nested archives
and anything past the limits come back as not inspectable, with why.
21 unit tests, against the networks' test card numbers and the IBAN
registry's own examples among others.
All six settled as recommended. Answer 6 ("and any other recognized and protected PII") becomes a catalog of identifiers with published formats and checks, grouped by region, each either checked by its check digit or counted only beside a corroborating word, plus templates named for what they find. Data with no number to find is covered by word lists and not claimed as detection. Office documents are read; PDF counts as can't be inspected.
Phase 1: one native rule engine at DATA, after the system Sieve script,
for both DLP policies and transport rules. DLP checks outgoing mail
with counted detectors (payment cards, IBAN, US SSN, word lists,
patterns) and blocks, warns with an audited override, or holds for
review. Held mail stays in the queue unscheduled, with its own review
record, so the queue's stored format is unchanged. Matches go to the
audit log without the matched text. Six questions for John at the end.
The automation suite's DNS test compared the published zone with one
copied from upstream v0.16.22. Its _ua-auto-config record carries a
SHA-256 of the account-configuration (PACC) document, and that document
names the provider as the brand, which the rebrand changed. The digest
the server publishes is right; the expected zone still held upstream's.
With the new digest the whole automation suite passes: ACME (including
the not-due reschedule check), DKIM, DNS and RFC 2136. It had been
failing at this point on main since the rebrand.
POST /api/directory/test takes a saved directory's id, an address and
optionally a password, and answers whether the directory opened, what a
recipient lookup of the address finds (account or group, with its
aliases, groups and name), and whether the password signs in. A wrong
password is told apart from a directory that can't be reached or is set
up wrong.
It calls the directory itself, below the sign-in path: a test never
creates or updates an account, never counts toward the sign-in ban and
doesn't depend on which domains use the directory. A password hash a
directory returns is never sent back. OIDC directories report their
discovered issuer; they take no passwords.
For server-level administrators with directory update permission. The
console's guided directory setup uses it to test a real person before
any domain is switched over.
The registry knows, for every expression field, the constants it may
evaluate to and the variables its conditions may read, and enforces both.
The schema served to INBUXA Admin described every one as a bare
x:Expression, so the console could offer nothing better than free text.
tools/fork/expr-schema.py reads those contexts from the generated registry
code and writes them onto each field's type as
expression: {constants, variables}. All 124 expression fields are covered.
CI runs it with --check so the schema can't drift from the registry.
The personal-data catalog and what it feeds: the data inventory and its
snapshots (#83, #89), the Compliance Officer roles (#88), the Compliance
Overview and Data Inventory menu entries (#92). Log file retention
(#87), privacy defaults for new installs (#85, #90), webhooks that send
only the events they name (#82), and upstream v0.16.24 (#84), whose
spam rules updates keep what an admin edited.
Explain: 12 settings asked about (upstream's new createdAt fields, the
certificate dates and the webhook events policy), with the release's
recommended model built locally; 705 answers carry over.
When a valid certificate already covered a domain's names (one stored
by hand before the domain was switched to automatic, for instance), the
renewal task ended with NotDue, which the task manager treats as a
permanent failure. Nothing rescheduled it, so the certificate expired
unrenewed. The renewal now returns a new AcmeRenewal task due when the
certificate falls due, the same way a successful renewal does, and logs
it as a backoff.
The ACME integration suite checks that renewing again right after
issuance hands back one AcmeRenewal for that domain, due at the
certificate's renewal point.
Seen in the console's first Overview. x:DmarcTroubleshoot and
x:SpamClassify are one-off actions whose results come back in the
response, not kept: object-life, not unbounded. Tasks go when done
(only a failed one's status may stay, still unconfirmed): object-life.
x:Log reads the log files, so it follows inbuxa:LogSettings.keepForDays.
x:TracerLog, x:WebHook and the OpenTelemetry tracers are configuration:
their credential fields stay classified, but they are no longer listed
in the inventory, where they counted as always sent off the server
even with none configured; the log-file, webhooks and otel-tracer
sources carry what they send.
Personal-data catalog spec, §8: the Compliance section in the console
reads Overview, Data Inventory, Legal Holds, Audit Log, Locked
Accounts. The two new entries are hand-built console pages
(CustomComponent/ComplianceOverview, CustomComponent/DataInventory),
shown to people with sysComplianceGet.
A console from before these pages shows the links and answers
"Unknown component", so the console that has them should be deployed
with the server release that carries this. Schema edited as the fork's
earlier Compliance entries were, hash updated.
Personal-data catalog spec, default D5 (settled 2026-09-28; built after
the v0.16.24 import's spam-rules loader landed). msbl.org's EBL is sent
a SHA-1 of every email address it's asked about. A new install's first
boot now leaves a note, and the rules update, once the bundled rules
are in, switches STWT_MSBL_EBL_EMAIL off and forgets the note, so it
happens once; the loader keeps that switch through later updates. An
existing server has no note and keeps every blocklist as it is.
Also fixes the data inventory's DNSBL endpoints: a zone is an
expression (`ip_reverse + '.zen.spamhaus.org'`, conditional branches,
`hash(email, 'sha1') + '.ebl.msbl.org'`), and the zone names are now
the quoted literals that start with a dot, from every branch, rather
than the expression's text.
Tested: unit test for the zone rule; the compliance system test (no
note, no change; the inventory lists ebl.msbl.org, not a hash; with the
note the blocklist goes off; the note works once); the system suite;
fork checks.
Personal-data catalog spec, §6 (Phase 3c).
inbuxa:DataInventory/get evaluates the catalog against the server's
live settings and says what this server holds: for each source and each
object that can hold personal data, its categories and whose data it
is, whether it is collected here at all, what bounds its retention (the
live value of the setting that does, or unbounded), whether it leaves
the host and to which endpoints, and a summary. Every host that
receives something is listed once as a candidate processor with what it
receives. Inside a tenant it answers with the tenant's slice and none
of the server's processors. Read-only, with sysComplianceGet.
inbuxa:InventorySnapshot/get is the history: a dated copy of the
evaluated inventory, recorded when it changes -- after a registry write
to an object the inventory reads, after inbuxa's log, audit or AI
settings change, and on the daily clean-up -- and kept as long as the
audit log's records. ids: null lists every snapshot, newest first; the
full inventory only when asked for.
The catalog is embedded and parsed at start (new dependency: toml,
MIT/Apache); the evaluation is a pure function of it and the live
facts, so each configuration is tested without a server. Loopback
endpoints stay on the host; any other configured endpoint leaves it.
Tested: unit tests for the evaluation (a new install's defaults, an
external blob store, a hosted AI endpoint, telemetry off, a tenant's
slice, hosts from URLs, loopback), snapshots, and the fact gathering's
store and duration rules; the compliance system test, extended (the
officer reads the inventory, a plain user is refused, a tenant's
officer sees its slice and no processors, a webhook to another host
becomes a processor and a snapshot names x:WebHook, a retention change
reads through); the system, audit, legal hold and account lock suites;
fork checks. The system suite failed once of three runs with an email
import's blob not found, in antispam.rs; the same happened once in
purge.rs on the previous branch. Nothing here touches uploads; noted
for a separate look.
The schema conflicted as a binary file: taken from main and the one
edit here re-applied (sysComplianceGet after sysLegalHoldExport). The
import kept the permission count at 673, so the new id stays 673.
Retested on the merged tree in its own target directory: the
compliance and system suites pass. One earlier system run failed in
purge.rs (an imported blob not found) and didn't recur.
Personal-data catalog spec, §7 (settled 2026-09-28).
sysComplianceGet (673) sees the data inventory and compliance
overview: superusers and, for their tenant's slice, tenant
administrators, by default and through the one-time grants on servers
that already have their roles stored.
A Compliance Officer role at server level holds it with reading and
exporting the audit log, placing, widening, releasing and exporting
legal holds, seeing account locks, and reading accounts, lists,
domains, tenants and roles. It changes no server setting, creates or
deletes no account, and can't shorten audit retention.
A tenant's accounts can hold only roles of their own tenant (MT-3), so
the tenant role is one "Compliance Officer" role per tenant, without
holds (LH-13): made once for every tenant a server has, and whenever a
tenant is created. While nobody holds it, it is removed with its tenant
so it doesn't block the delete, and put back if the delete is refused
for another reason. Both roles carry a user's own permissions too,
since roles given to a person replace the default user role, which a
tenant's accounts can't hold anyway.
Every server makes these once, new or existing -- the built-in roles
are only made on a server with none -- and records each under P c, so a
role an administrator deletes stays deleted.
Tested: unit tests (neither role changes a setting beyond a user's
own; holds for the server's officer only; per-place records); a new
compliance system test (one server-level role; an officer reads the
audit log, places and releases a hold, and is refused a setting, an
account and audit retention; a tenant gets its role, whose holder reads
the tenant's audit log and no holds; a tenant with an unused role is
deleted and the role goes with it); the system, audit, legal hold,
account lock and SCIM suites; fork checks. The directory suite needs
its LDAP container and wasn't run here.
Personal-data catalog spec, default D1 (settled 2026-09-28): log files
were never deleted. inbuxa:LogSettings.keepForDays says how many days
rotated log files are kept; unset (null) keeps every file, as before,
and a new install sets 30 days.
It is a fork-owned setting, stored under T + l as audit retention is,
not a field on x:TracerLog: that object is also stored inside
x:Bootstrap with a field after it, so a new field would change
x:Bootstrap's stored format. Server-level, with the tracers'
permissions (sysTracerGet, sysTracerUpdate); changes are in the audit
log, before and after.
Log files are local, so every node deletes its own: hourly, and at once
when the setting changes on that node. Only regular files named
<prefix>.<something> in each enabled log tracer's directory, last
changed more than the limit ago, are removed; the file being written is
never that old, and nothing else in the directory is touched. Minimum
one day. The catalog classifies inbuxa:LogSettings and points the log
file's retention at it.
Tested: unit tests for the file rule (only this log's old files; the
current file, other files and directories stay) and a purge on disk;
the system suite, which reads, sets, refuses zero, restores null and
checks the audit records; fork checks.
The schema, which both sides changed, merged as JSON with no conflicts.
The personal-data catalog (#83) gains upstream's new x:DnsServerPowerDns:
nothing personal but its API key, like the other DNS providers.
D1 becomes a fork-owned setting, as audit retention is, because a field
on x:TracerLog would change x:Bootstrap's stored format. D5 waits for
the v0.16.24 import's reworked spam-rules loader.
Personal-data catalog spec, defaults D2, D3, D4, D6 and D7 (settled
2026-09-28, new installs only):
- D2: automatic IP bans expire after 30 days instead of never; D3:
spam training samples, whole messages, are kept 90 days instead of
180; D4: Pyzor, which sends a digest of each message's text to a
public server, is off; D6: delivery history is kept 14 days instead
of 30. Written on the first boot of a new install only -- one with no
roles yet, the same test the built-in roles use -- by reading each
singleton, setting these fields and writing it back whole. A server
with roles keeps its settings, saved or default.
- D7: a webhook created from now on starts with the include policy and
no events, so it sends nothing until events are chosen (Rust default
and schema default, marked). The registry stores every field, so
existing webhooks keep their policy.
- Expired bans are also removed by the daily data clean-up. They
already stopped blocking and were deleted when settings next loaded;
a server that seldom reloads kept them.
D1 (log retention) and D5 (the hashed-address blocklist off) are held,
and the spec says why: x:TracerLog is stored inside x:Bootstrap with a
field after it, so adding one changes that object's stored format; and
the spam-rules loader D5 touches is being reworked by the v0.16.24
import. The spec also corrects finding 3: expired bans were deleted on
settings load; bans were permanent only because no period is set.
Tested: unit tests for the new-install values and that everything else
in each singleton stays; the system suite, whose security test now
purges an expired ban and checks its record is gone; the telemetry
test; common's unit tests; fork checks.
Phase 2 of the personal-data catalog spec.
resources/privacy/catalog.toml classifies every object in the schema
(316) and inbuxa's own JMAP objects (12): each property that can hold
personal data, with its categories, and for objects that hold any,
whose data it is, where it lives, its scope and what bounds its
retention (a named setting where there is one). Twenty sources that
are no object -- the log file, exporters, webhooks, spam lookups, the
Explain cache, relays and hooks, push, legacy-use records -- carry the
same facts plus the settings that turn them on, whether the data
leaves the host, and the code that writes it. Classifications of
objects that hold data about people are from the spec's source map;
the rest are typed from the schema alone (address, IP, secret).
tools/fork/privacy-check.py fails CI when an object or inbuxa object
has no entry, when a property the schema types as an address, IP or
secret is left to its object's default, when an entry names an
object, property, setting or code path that is gone, or when it uses
a word outside the catalog's vocabulary. --unlisted prints starting
entries. strip.py's report gains "Unclassified in the privacy
catalog": objects and fields new in an import and not classified,
informational like the Enterprise flags.
Tested: 13 unit tests (tools/fork/tests): the check passes on this
tree; fails on an unclassified object, an address hidden behind a
default, a secret in a set or object reference, stale properties,
objects, settings and code paths, an unlisted inbuxa object and a
word outside the vocabulary; --unlisted's entries; and the strip
report on a synthetic import. The check and the tests run in the
fork-checks job.
The per-protocol switches (#79) added legacyAllowed to the account's
urn:inbuxa:jmap capability, but this expectation wasn't updated, so
jmap_tests stopped here and the suites after it never ran.
A webhook has a level (info by default) that nothing read: its events
were chosen by its list and policy alone. With the default policy,
exclude, and nothing listed, that meant every event type, including
smtp.raw-input (the raw SMTP bytes, DATA included) and the model's
reply to the spam classifier. The docs suggest a webhook to pass the
audit log to a SIEM; set up that way it would have received whole
messages. Found by the personal-data catalog investigation (finding 1).
Now an include list is sent as named, whatever each event's level:
naming an event is the choice. Otherwise a webhook gets only events at
or above its level, as a tracer does, and never a protocol's raw input
or output (IMAP, SMTP, POP3, ManageSieve, delivery, milter), which
carries whole messages and credentials; those go out only when named.
Tested: unit tests for the rule (level, raw I/O only when named, a
named event below the level, custom event levels, a webhook's own
errors); the telemetry system test, whose webhook names debug-level
connection events and still receives them.
Eight conflicted files resolved, plus the lock file and the schema:
- crates/services/src/task_manager/spam_classifier.rs: upstream's rules
update now replaces existing rules, DNSBL servers, lookups and file
extensions, keeping only whether each is on. Taken, with one difference:
an object an admin edited is kept as it is. Every object an update writes
is fingerprinted (content without `enable`, SHA-256, stored under
SUBSPACE_INBUXA "Sf"), and only one that still matches is replaced.
Scores are never replaced, as upstream has it. The AU-1.10 summary record
now names what was added, replaced and kept, and the bundled rules are
marked applied only when the update fully succeeded, so a failure runs
again on the next start. The marker becomes "3.0.2+2", which runs the
update once on upgrade to fingerprint every rule still as bundled.
- crates/common/src/network/autoconfig/autodiscover.rs: upstream's rewrite
(implicit TLS first, labeled SSL), with the per-protocol switches (LP-7,
LP-14a) passed in as a filter.
- crates/store/src/backend/mysql/{search,write}.rs: upstream's chunked
deletes (no unbounded first DELETE, stop on a short chunk, halve the
chunk on the new chunk-too-large errors) inside the fork's query timeout.
- crates/smtp/src/lib.rs: the fork's queue spawn kept. It already fixed the
stall upstream fixes here (a node without outboundMta stops accepting
mail at about 1024 queued messages), and follows role changes live.
- crates/jmap/src/registry/mapping/bootstrap.rs: the log path stays
/var/log/inbuxa/; upstream's PowerDNS mapping taken.
- crates/main/Cargo.toml: the AGPL-only license kept, version 0.16.24.
- tests/src/jmap/principal/get.rs: the fork's capabilities kept.
- resources/schema/schema.json.gz: merged as JSON; upstream relabeled the
vendor Sieve extensions "(Stalwart)", kept as "(vnd.inbuxa)".
- Cargo.lock: upstream's, with the fork's crates added by Cargo.
Also:
- tests/src/smtp/inbound/spam_rules_kept.rs: an edited rule survives an
update, an unedited one is updated, rules from before fingerprints are
handled, and the audit summary says so. Upstream's own spam_rules test
passes unchanged.
- tests/src/smtp/reporting/reschedule.rs moves to port 19058; upstream's
new spam_rules test took 19057.
- tools/fork/renames.py renames the "(Stalwart)" labels and the default
log path, so neither conflicts again.
- tools/fork/notice-check.py compares against the newest snapshot in the
checked-out history instead of the upstream branch head, so moving the
branch no longer fails other open pull requests.
- tests/src/directory/issuer.rs (since v0.16.23) stays out, and is on the
build check's known list: it tests issuer-based directory routing, which
the fork doesn't have (DIR-2).
- Strip report: docs/fork/strip-reports/v0.16.24.{md,json}.
The sidecar catalog; the Compliance Officer places and releases holds;
the Tenant Compliance Officer is built now; shortening audit retention
is recorded and surfaced, not gated on a second person; all seven
new-install defaults, in Phase 3; the webhook finding fixed now as a
bug; snapshots kept as long as the audit log.
Phase 1 of the GDPR auditor foundation: the investigation and the
design, committed before anything is built (SPEC.md §3 rule 3).
It maps every place the server stores or sends personal data found
in the code at de275ba, each with its categories, whose data it is,
the settings that control it, what bounds its retention, where it
lives, whether it leaves the host, its scope and the code that writes
it, and the default in a new install. It proposes a sidecar catalog
(resources/privacy/catalog.toml), since the schema and registry code
are upstream's generated output with no generator here; a CI check
modeled on name-check.py; a strip-report section; a read-only
inventory method with dated snapshots; a Compliance Officer role; and
the Compliance navigation with Overview and Data inventory.
Findings worth reading on their own: webhooks ignore levels and, at
their defaults, receive every event including raw SMTP input; log
files are never deleted; automatic bans never expire; some of the
fork's records outlive the account; spam training keeps whole
messages for 180 days; traces are on in a new install; the spam
filter sends IPs, domains, hashed addresses and body digests to
third-party services by default.
Proposed default changes (new installs only) and seven open questions
are for John to decide. No default is changed.
Per-protocol legacy switches (#79) and the hold export's exceptions
list (#75). The prepared Explain answers are relabeled for this
release; 706 carry over unchanged.
The legacy-protocols switch was all or nothing. An operator can now stop
POP3 and keep IMAP: each of IMAP, POP3 and ManageSieve has its own
switch, server-wide on inbuxa:ProtocolPolicy and per tenant on
inbuxa:TenantProtocolPolicy (properties imap, pop3, manageSieve).
legacyProtocols stays as the kill-all: setting it sets all three, and it
reads "disabled" exactly when all three are off. A policy stored before
this has only legacyProtocols and reads as all three at that value, so
existing servers and tenants carry over unchanged. In one /set, a
protocol named beside legacyProtocols overrides it.
SMTP submission keeps no switch of its own: sign-in over it is refused
only when all three are off, as the single switch did (LP-6), so
turning one protocol off never stops a mail app sending. For a tenant,
the server's switches and the tenant's count together.
Server-wide, a change closes the listeners of whatever is now off and
puts back the saved listeners of whatever is on again, both in one
change if asked; listeners of a protocol still off stay saved. Sign-in,
autoconfig, autodiscover, PACC (now prepared once per combination) and
the suggested DNS records all follow each protocol separately. A tenant
may turn a protocol on only while the server has it on (LP-9), and the
refusal names which. The JMAP session adds legacyAllowed, the protocols
still allowed for the account; legacyProtocols there keeps its meaning
for older webmail builds. Events name the switches ("pop3 disabled"),
and audit before/after reads every switch even from an older policy.
Tested: unit tests for the switches, the old-policy reading, the
server/tenant combination, the tenant refusal and listener refusal; and
tests/e2e/legacy_protocols.py against a running server, all 100 checks,
including new ones: POP3 alone off closes only its port and refuses
only its sign-in while IMAP and sending go on; only POP3 stops being
advertised; one change closes IMAP and reopens POP3; a tenant turns
POP3 off for itself, and can't turn IMAP on while the server has it off.
Cloudflare's DMARC report intake rejects every aggregate report we
send with "555 5.7.1 invalid_report_schema". Bisected against the live
endpoint: the only element it objects to is <disposition>pass</disposition>,
the value RFC 9990 added for mail that passed DMARC under an enforcing
policy. The RFC 9990 namespace, <np>, <discovery_method>, <testing> and
a missing <pct> are all accepted, and a report that differs only in
using "none" there goes through.
"none" (no action taken) is valid under both RFC 9990 and RFC 7489 and
says the same thing to the reader, so reports now go out with it. The
stored report keeps "pass"; only the serialized copy changes.
An item the hold covers whose stored record or content can't be read
goes in exceptions.csv with the path it would have had and the reason,
rather than being left out silently. The file is always in the ZIP, so a
header-only one shows nothing was missed, and manifest.sha256 carries
its hash beside the manifest's.
inbuxa:HoldExport/set takes a hold, optionally some of the accounts it
covers, and a reason; the collection runs in the background and get
says when it's ready. The ZIP has, per account, mail as .eml under its
folders, calendars as .ics, contacts as .vcf, files as stored, and the
archived items the hold keeps under archived/; a manifest.csv gives each
entry's account, kind, folder, date, whether it was archived, size and
SHA-256, and manifest.sha256 hashes the manifest. Accounts the hold
doesn't cover are left out, and items outside its date range are too:
live mail by arrival, events by start, and archived items the same way,
so an export doesn't carry deleted items that only another hold keeps.
The finished file is a blob of whoever started the export, so only they
download it, and it lasts as long as any upload (uploadTtl). Exports
are records under the hold (SUBSPACE_INBUXA H/e): never changed or
destroyed, each with its status, counts, size and checksum. Starting
one needs sysLegalHoldExport, an active hold and a reason, and is
recorded in the audit log like the audit log's own export.
The build is in memory and capped at 2 GB; bigger holds fail with a
message saying so, and are split by picking accounts.
Tested: unit tests for safe ZIP names and the manifest and its hash;
the legal_hold system test, on RocksDB, PostgreSQL and MySQL, exports a
hold end to end (live and archived mail, the manifest's hash, an asked-
for account the hold doesn't cover left out) and checks the refusals
(no reason, a user without the permission, a released hold) and the
audit record; and by hand from the console on a local server. Not
covered by a test: the archived-item date range with two holds of
different ranges over one account.
hickory 0.26.3 rejects two kinds of valid answers, and outbound
delivery then retries those hosts until the message expires:
- A zone delegated beneath an unsigned zone (l.google.com under
google.com). Proving the delegation insecure needs an SOA record in
the DS reply, and public resolvers often leave it out. Every Google
MX host behind a signed MX record was unreachable.
- A signed CNAME to a signed name that lacks the queried type. The
NSEC denial is checked against the original name, not the target's.
On a bogus verdict, follow a signed CNAME and repeat the lookup at its
target; otherwise look up the name's zone and its parents, nearest
first. A zone that validates as unsigned means nothing below it can be
signed, so the plain resolver answers and the result is insecure. A
zone that validates as signed first leaves the verdict standing.
Legal holds (#70): a hold on people, groups, domains, tenants or the
whole server keeps everything it covers from being destroyed, by anyone,
until it's released; deleted accounts keep their data. Audit records
name accounts by their full address and holds by their case name. The
daily clean-up of expired archived items works again.
Prepared Explain answers relabeled for this release; no setting changed
since 2026.9.27.2, so all 706 carry over.
An account's or mailing list's name is only its local part, so the log
said "Account ken.gosling" where two domains could each have one; it
now says [email protected]. A change to a legal hold was
recorded under its id; the hold's current state is now read first, so
the record carries its case name and each change reads before/after.
inbuxa:LegalHold/get answers accountsCovered, itemsHeld and sizeHeld
when asked: the accounts a hold reaches now (deleted ones it keeps
included) and the archived items it keeps, with their size. Worked out
in one pass over accounts and archive, only for requests that name them.
Held items stay out of the user's quota, as all archived copies do
(LH-9).
Destroying a held account removes the login, as offboarding needs, but
keeps its data as a deleted account with no expiry, whether or not
undelete keeps accounts; its addresses stay reserved and its holds name
it from then on. Destroy-now refuses it, and its DestroyAccount task
defers itself while it's held or its time hasn't come. Holds placed or
released later freeze or free kept accounts in the same settle pass,
with 30 days' grace after the last release (LH-8, LH-10).
Placing or widening a hold freezes what's already archived in its scope
and range, its old deadline noted; releasing one gives each item no other
hold covers that deadline back, or release plus 30 days if later. One
pass over the archive does both and changes nothing twice (LH-6, LH-10,
LH-11). A held archived item can't be destroyed; restoring still can,
and the hold is named only to callers who may see holds (LH-7). Audit
records about a held account survive the purge (AU-7).
Fixes the daily clean-up of expired archived items (UD-13), which never
found any: the registry's unfiltered query reads an all-ids index that
archived items aren't in. Items are now walked account by account, kept
deleted accounts included. Expired items were still removed whenever
their account's archive was read.
Every way of deleting mail (JMAP, IMAP EXPUNGE, POP3, mailbox removal,
Trash emptying) and Sieve scripts, events, contacts and files now asks
how the account's deletions are kept: a hold keeps them with no expiry
(archivedUntil 9999-12-31), even with undelete off; otherwise undelete's
period applies as before (LH-4).
A hold's date range decides by the item's own date (LH-3). Mail is noted
as held at deletion and settled when it's archived, once its received
date is known; outside the range it gets undelete's deadline or isn't
kept. Events go by their start, with a day's slack for time zones;
recurring events, contacts, files and scripts are held whole.
A groupware item's note now stays until its archive succeeds, and a
failure retries the task instead of being logged and lost (LH-5).
A hold reaches an account by name, through any of its addresses'
domains, its groups or its tenant, as they are now, so an account added
to a held domain later is held too. An account that leaves a held
domain, group or tenant stays held: the registry write hook adds it to
the hold by name on every account change, whoever makes it (LH-2).
Server::holds_on answers for the deletion paths, from the store each
time so a hold binds every node at once.
inbuxa:LegalHold get/set places a hold on accounts, groups, domains,
tenants or the whole server, with an optional date range. A hold's range
and scope can only widen, a released hold is read-only, and none is ever
deleted. Placing, changing and releasing each need a reason and are
audited (LH-1, LH-3, LH-10, AU-12).
Permissions 669-672 (see, place, widen or release, export held data)
go to server administrators only; the tenant ceiling always strips them,
as it does Impersonate (LH-13). Schema: Compliance > Legal Holds.
What a hold keeps comes next, through the undelete hooks.
Also moves the lock expiry helpers below the lock module's imports.
Delegates reach the whole locked account (#68): its calendars, contacts
and files as well as its mail, even a kind it holds none of yet, and
writing delegates may add at the top of its Files.
Prepared Explain answers relabeled for this release; no setting changed
since 2026.9.27.1, so all 706 carry over.
A shared account refuses top-level folders, so an organize or full
delegate couldn't add anything to a locked account with no folders. A
delegate who may write now can, as the owner could; the reconcile after
the create grants it the new folder. Read delegates still can't (AL-6,
AL-7).
A delegate's token listed the locked account only for kinds of data it
held grants on, so one with no files (or no calendar) was refused to the
delegate outright: "You do not have access to account". The token now
lists the locked account for mail, calendars, contacts and files alike,
so an empty kind reads as empty. What the delegate may see or change is
still each container's grant (AL-7).
The audit log (#64): every administrator change, admin sign-in and look
into someone else's data, recorded before it happens, chained per node
and checkable for tampering, exportable with a manifest, kept 2 years.
Locked accounts (#65, #66): an account that keeps receiving mail but
can't sign in and sends nothing on its own, handed to delegates at read,
organize or full, ending at a date when one is set.
Prepared Explain answers relabeled for this release; no setting changed
since 2026.9.27, so all 706 carry over.
A delegation with an end date dropped out of the delegate's token then,
but its folder grants stayed until the daily sweep, so the delegate kept
the account as an ordinary share for up to a day. Each node now sleeps
until the soonest end date, woken early by any lock write and at least
hourly, and re-applies that lock under a cluster-wide claim.
The sweep also had a second-run bug: a delegation past its date gave the
delegate back its earlier share, then dropped the note, so the next sweep
removed that share entirely. The note is now kept while the delegate is
still listed.
A locked account can't sign in (it fails as a wrong password does), its
sessions end on every node, refresh tokens stop working, and its Sieve
scripts forward and reply to nothing. Mail keeps arriving.
Delegates get real ACL grants on the account's mailboxes, calendars,
address books and files at read, organize or full, with the rights they
replaced restored on unlock. Folders made later are granted after the
create and in a daily sweep. Organize delegates can't destroy; send-as
needs organize or full. The JMAP session marks delegated accounts in
urn:inbuxa:jmap.
New inbuxa:AccountLock object with get/set, permissions 665-668, and a
Compliance > Locked Accounts entry in the schema. Lock, unlock and
delegate changes need a reason and are audited; delegate access and
writes are audited too (audit-hold-lock spec AL-1 to AL-12).
What administrators and the server itself do to the control plane is now
recorded, from inbuxa-drafts/specs/audit-hold-lock.md (AU-1 to AU-12):
settings, accounts, domains, roles and every other registry change, with
each field's before and after (secrets only as "changed"); the fork's own
settings objects; administrator sign-ins (and failed ones to administrator
accounts), master-user and recovery-admin sign-ins, once an hour per
account, method and address; access to another account's data through
impersonation or FetchAnyBlob, once an hour; exports and tamper checks;
and registry writes the server makes on its own, named by subsystem
(system:AcmeRenewal, system:auto-ban, system:directory-sync, ...), with a
spam rules update as one summary record.
No change without its record (AU-3): before a set method changes anything,
a pending record per requested create, update and destroy is written; if
that fails, the method is refused with serverFail. Its outcome follows as
a later entry. A change interrupted by a crash stays "unfinished".
Records live in the fork's subspace under L, as one SHA-256 hash chain per
node. The chain's head is stored, never cached, and every append asserts
it, so two writers can't take the same place. Nothing can edit or delete
a record; the daily purge removes the oldest past the retention (default
730 days, minimum 90) and records where the chain now starts, so
verification still passes. security.audit-recorded (647) copies each
record to webhooks, OpenTelemetry and the log; security.audit-write-failed
(648) reports a failed write.
New JMAP objects under urn:inbuxa:jmap: inbuxa:AuditEvent/get and /query
(filters: time, actor, action, target, account, tenant, outcome, address,
text), inbuxa:AuditSettings, inbuxa:AuditExport (CSV or JSON Lines built
on the server, each line with its chain hash, ending in a manifest; the
created object names the blob and its SHA-256) and
inbuxa:AuditVerification. New permissions sysAuditGet, sysAuditExport and
sysAuditSettingsUpdate: the Administrator role gets all three, the Tenant
Administrator role gets read and export, once, on existing installs too.
A tenant administrator sees records whose actor or target is in its
tenant, including a server administrator's changes there.
Sign-in method on the session: access tokens now remember how they signed
in (password, app password, API key, OAuth client, directory, master user,
recovery admin), including across the HTTP credential cache. New OAuth
access tokens carry their client id in the sealed claims; older ones show
as client "unknown" until they expire.
The schema gains the permissions, the two events and a Management >
Compliance > Audit Log link.
Stack: the request layer boxes every inner future where it's made. Without
that, a debug build overflowed the default 2 MB worker stack on a registry
set; measured with the same request, the branch and main now overflow at
the same stack size (between 1856 and 1920 KiB, debug), so the layer adds
nothing measurable.
Tests: unit tests in inbuxa-features and jmap; system::audit::audit_log_tests
(run with --ignored) passes on RocksDB, SQLite, PostgreSQL, PostgreSQL with a
read replica, MySQL, MySQL with a replica and FoundationDB. The system, JMAP
and SCIM suites pass. authorization.rs skipped fork permissions that guard
no registry object; the audit suite checks a plain user is refused instead.
inbuxa's own mark (#62): the kitten over a server with a bay for each
piece of the suite, on the built-in sign-in and RSVP pages, the web
logo and the email logo.
Prepared Explain answers relabeled for this release; no setting changed
since 2026.9.26.1, so all 706 carry over.
inbuxa's mark was ihasmail's cat-and-envelope reused unchanged. The new
one keeps the family's face, paws and colors, over a server with a bay
for each piece of the suite: the letter (webmail), a prompt (console),
status lights (server).
- The built-in sign-in and calendar RSVP pages, and the web logo, drew
the old cat as an embedded PNG. They now draw the mark as vector in
the same slot, keeping class="symbol"; each page is about 31 KB
lighter. The .min copies are updated the same way and the .min.gz
regenerated with gzip -9 -n, as minify_html.sh does.
- resources/branding: email-logo.png (the compact lockup at 380x80 on
white, as before) with its .b64 regenerated byte-for-byte in the old
76-column form, and favicon-64.png.
- img/brand: the logo bundle, now pure vector, with its README.
announce.yml runs coffey-labs/actions discourse-release on every published
release, posting it to this project's Announcements category on
community.coffeylabs.org. The release workflow also announces
from its own job, since a release made with the job token fires no
'on: release' workflow in Gitea.
Prepared Explain answers relabeled for this release; 11 settings whose
default is the time of creation drop out, since their answers could never
match.
ai-explain spec, amendment 1 (EX-22 to EX-28):
- answers are three or four sentences, max_tokens 160, cut at 700 chars;
- POST /api/explain streams the answer as server-sent events;
- each node remembers answers in memory (1,000, 24 h), keyed by the facts,
prompt version and model, shared by server-level administrators;
- resources/explain/settings.json.gz ships answers for settings at their
defaults, generated with prepare_setting_explanations (717 for 2026.9.27);
- the system prompt no longer carries the per-request marker, so a model
server can reuse it;
- inbuxa:Explanation gains source, answeredAt and preparedFor.
When a listener couldn't bind its address (a port below 1024 without
root, a port already in use, or the legacy-protocols switch putting a
listener back after privileges were dropped), the bind error was
reported but the socket was still passed to listen(). The kernel then
bound it itself, to a random port on every interface, and the server
logged the listener as started on the port it was configured with.
listen() now refuses a socket that isn't bound, so the listener is
reported with a listen error and skipped, and nothing opens anywhere
unexpected.
A singleton such as x:SpamSettings has no stored object until someone
saves it; /get shows its defaults instead. Explain looked only for the
stored object, so every setting still at its defaults answered "No
such x:SpamSettings." It now falls back to the defaults the same way.
A new method, inbuxa:Explanation/set, asks the node's local model for a
short plain-words reading of one thing an administrator is looking at:
a failed recipient in the queue, a Classify verdict, a log line or trace
event, or one setting with its saved value. The server builds the prompt
itself from stored data and the registry schema, never from text the
console sends, and grounds SMTP replies in RFC 3463 and RFC 5321.
What the model is never shown: secrets (including ones nested inside a
setting, like an AI model's HTTP auth), raw protocol events, and the
contents of any other event. A tag name that doesn't have a tag's shape is
refused before a model is asked.
Calls share the AI gate with spam classification, but mail always keeps
its slot, and Explain has its own hourly count per account and its own
on/off switch in inbuxa:AiLimits. The permission is sysAiExplain,
superuser only; tenant administrators can't use it. The session carries
an aiExplain flag so a console knows when to offer the button.
An install whose roles were stored before the permission existed gets it
added once, at start-up, to the roles that are administrators' alone,
not the User role their defaults share with every account. An operator
who removes it later isn't overruled.
Tests: unit tests in inbuxa-features and jmap, and ai_explain_tests
(run with --ignored) covering the acceptance tests and the upgrade.
Every node renews its lease once a minute instead of every 30 minutes,
so the lease works as a heartbeat. x:ClusterNode reports a node Stale
once it has gone three minutes without renewing (it used to take an
hour), and Inactive after a day, as before.
Taking over a lease still needs a full hour of silence. A node that is
slow rather than gone never loses its id to another host, so snowflake
ids stay unique.
The admin dashboard's Cluster Health card counts these statuses.
After #37, address fields on PostgreSQL are split into words as the
built-in index splits them, but language text (subject, body,
attachments) still goes straight to PostgreSQL's parser, which keeps a
URL, host, path or file name as tokens of its own:
"https://x.example/shipping-support/" becomes a url, a host and a
url_path, "invoice-2024.pdf" a file. So TEXT/BODY "shipping" missed
messages where the word appears only inside a link, while RocksDB and
the other built-in backends found them: 8 messages across a handful
of searches in the rehearsal.
On insert, language text is now indexed as it was, followed by the
word parts of each token that holds a URL separator (/ . @ : ? = & # _
% + ~ \), split with SpaceTokenizer as keyword_terms() splits addresses.
The parts go through the same text search configuration as the rest of
the text, so they are stemmed like the words around them. Plain words,
words that only carry punctuation ("end.", "(see") and hyphenated words
(the parser already splits those) add nothing, so text without links
is indexed exactly as before. Each part is added once per document.
On sample mail, the text vector of a short order notice with three
links grows from 546 to 716 bytes, a newsletter with 25 tracking links
from 5586 to 6430, and a plain letter not at all.
On search, a query word written as a URL, host, file or hyphenated word
also matches as its word parts, ORed with the query as written, so
"shipping-support" or "invoice-2024.pdf" match the new parts and
documents indexed before this change still match as they did.
Existing messages keep their old vectors until they are reindexed (the
reindexAccounts task); new and reindexed messages match at once.
store::search_tests gains test_url_word_search: five bodies, 19 body
searches for words found only in a URL path, query string, host or
file name, the tokens as written, plain words and non-matches, with the
same expected ids on every backend. It passes on RocksDB, SQLite,
MySQL and PostgreSQL; on main PostgreSQL fails at the first ("shipping"
finds [3], not [0, 3]). On PostgreSQL the suite then stops at the
account sort assertion (query.rs:689) exactly as it does on main.
A cluster rehearsal sent ten x:<Object>/set requests at once and got
ten full reloads on every node. #39's coalescing only joined writes
that queued behind a running reload, but the requests reached the
server about 33 ms apart and a reload takes tens of milliseconds, so
none overlapped one.
A full reload after a registry write now waits for writes to settle:
75 ms after the last one, and at most 250 ms after the first it
covers, so a steady stream still reloads at least four times a
second. 75 ms is a little over twice the gap the rehearsal saw between
requests. A single write pays it once: in the tests a settings write
takes about 140 ms instead of 60. The reload runs in a task of its
own, so a request that goes away doesn't cancel it for the others.
Each write takes the result of the first reload that started after it
was stored (the gate keeps the last 64 results), so applied true or
false still describes the reload that covered that write.
The 33 ms gap was a queue on the server, not password hashing: Basic
credentials are cached per Authorization header, so they are checked
once. Every authenticated HTTP request counted itself against the
account's rate limit by incrementing one counter per account in the
in-memory store, so parallel requests from one account queued on that
key: a row lock on PostgreSQL (a few round trips to the database
each) and conflict retries with a 50-300 ms backoff on RocksDB. An
account with the unlimitedRequests permission (administrators, by
default) passes the rate and concurrency limits anyway, so its
requests are no longer counted. Ten parallel Core/echo calls as the
admin now finish in 1-4 ms; before, they finished one after another
over 20 ms on a local PostgreSQL and 300-450 ms on RocksDB. Other
accounts still count every request.
system::auto_reload::settings_reload_tests: ten concurrent writes now
take one reload (the gate counts them; at most two allowed), all are
applied: true and in the running settings, and a single write takes
exactly one reload. RocksDB and PostgreSQL, 1 reload in 141-196 ms.
With the old behavior (no wait, requests counted) the same writes
took 5 reloads; without the wait but with the rate fix, 2.
cluster::broadcast (3 nodes, PostgreSQL + NATS) and system::reload
still pass.
A cluster rehearsal moved a Log tracer to another directory: the write
was reported x:settingsReload applied:true, but the tracer kept writing
to the old file until a restart. Telemetry::update only refreshed each
running tracer's events, level and lossiness; a tracer's own settings
(path, prefix, rotation, format, endpoint, headers, ...) stayed as built.
Each tracer now carries a hash of the registry object it was built
from, less the fields that change in place. The reload compares it with
the running tracer's: unchanged ones are updated in place as before,
changed ones are started over, new ones started and removed ones
stopped. Only tracers this server started are removed; upstream removed
every subscriber not in the settings, which also cut off live-tracing
streams on each reload.
Starting over is a swap in the collector, so no event is lost or
written twice: a subscriber registered under a running one's id
replaces it between two collection passes. The old one's batch is sent
first (what its full channel can't take moves to the new one), and
dropping it closes its channel, so its task writes what is queued and
ends. Per tracer kind:
- Log: a tracer started over on the same files (rotation or format
changed) waits for the old one to finish, so lines don't interleave.
- Webhook: the task held a sender of its own channel for retries, so
it never ended; retries now use a weak sender, and pending events are
posted when the channel closes.
- OpenTelemetry: pending logs and spans are exported when the channel
closes instead of dropped, and a span that was open across the swap
is exported by the new tracer with the events it saw.
- Console and journal: nothing kept between batches.
- Trace history: built from the tracing store, which takes a restart,
so it is never started over.
No kind needs a restart, so x:settingsReload doesn't gain one.
system::tracer_reload::tracer_reload_tests (new): a Log tracer created
over JMAP writes to its directory; its path is changed over JMAP while
2000 numbered events are emitted; after the reload, events land in the
new file and not the old one, each numbered event is in exactly one of
the two files, and a destroyed tracer writes nothing. On main the new
file never appears.
The publish workflow only built a tag whose commit is on main. That keeps
every image tied to reviewed code, but it means production can only get a
fix together with everything that has landed on main since its release.
A tag on a release/* branch is now accepted too. A hotfix branch starts at
an earlier release tag, takes fixes through pull requests into it (so the
code is still reviewed and CI-tested before it is tagged), bumps
brand_version! and is tagged there. The tag must still equal
v<brand_version!>, and the step prints which branch it was found on.
A tag runs the workflow file from its own commit, so a hotfix branch that
starts before this change needs this commit cherry-picked onto it before
its tag is pushed.
The report scheduler dropped DMARC and TLS events on a node whose role
lacks outboundMta (upstream never started it there, so they sat in a
channel nobody read). Mail received on a front node therefore never
reached an aggregate report, which is meant to cover all of a domain's
inbound mail, whichever node received it. In rehearsal, five messages
received on port 25 on a front node were missing from every report.
- The report scheduler records on every node. Recording is a store write
the nodes already share, so it needs nothing from the outbound MTA.
Building and sending a report (the DmarcReport and TlsReport tasks) stay
with outboundMta nodes, as the task manager already enforces.
- More nodes now append to one report at once. Appends already guard the
report's versioned primary key; a write that loses now retries up to ten
times after a short random pause, not three times at once.
- The node sending a report deletes it only if it is unchanged since it
was read, and reads it again otherwise, so a record another node appends
meanwhile goes out with the report instead of being deleted unsent.
Test: cluster::front_reports (PostgreSQL and MySQL). A front node's
results appear in the report the MTA node sends, alongside eight appended
at once from both nodes, and the front node never runs the report task.
It fails on main: the front node's results are never recorded.
Setting deliverAt on an internal DMARC or TLS report wrote the new task
queue row with the report's object type (0x21, 0x6e) instead of the task
type (7, 8), and left the task row at its old due. The task manager's scan
failed on that row with store.data-corruption ("Failed to iterate over task
queue"), and because the error ended the whole scan, every task due after
the row stopped running on every node.
- reschedule_ops writes the new queue row through schedule_task_with_id, so
it carries the task type and the task row gets the new due. It removes
the row the task is actually queued under (the task's due, which differs
from deliverAt once the task has been retried) and any row an earlier
reschedule left at deliverAt.
- x:DmarcInternalReport/set and x:TlsInternalReport/set lock the report's
task while they move it, as x:Task/set does, refuse while the report is
being sent, release the locks however the request ends, and wake the task
manager.
- The task manager logs a queue row it can't read (id, due, key, value) and
skips it instead of ending the scan. It then repairs the row from its task:
the row is rewritten with the task's type, and a row with no task behind
it is removed. A row holding a report's object type for a report task is
what the old reschedule wrote: the task is moved to that row's time, as
the reschedule intended, and its old queue row is removed. Stores that
already hold such a row recover on their own once it comes due.
- x:Task/query with a type filter skips an unreadable row instead of
failing.
Test: smtp::reporting::reschedule (RocksDB and PostgreSQL). It fails on
main: x:Task/get shows the old due, and with that check removed, neither
report nor a later task ever runs.
In cluster rehearsal 3, turning outboundMta off on node1's role was
reported applied (x:settingsReload applied: true), yet node1 kept
delivering mail, a report message included, until it was restarted.
The queue and report managers were started at boot only when the
node's role included outboundMta (crates/smtp/src/lib.rs), and the task
manager only when the role had some task type (spawn_task_manager).
After that nothing looked at the role again: a queue manager that was
running kept claiming and delivering, and one that wasn't never
started.
They now start on every node (outside recovery mode) and follow the
role live:
- Queue manager: before each scan it reads the role from the running
settings. Without outboundMta it claims nothing new; deliveries
already running finish and report back as usual, which releases
their locks. When the role comes back (a reload wakes the manager
with ReloadSettings, and it looks again every 30 s regardless) it
logs queue.started and scans the whole queue at once.
- Report scheduler: DMARC and TLS report events are handled only while
the role has outboundMta, as at boot; events arriving without it are
dropped, as they were on a node started without the role.
- Task manager: task_enabled already read the current role on every
scan. It now also runs on nodes whose role has no task type (the
scan returns at once until one is added), a job claimed before a
role change is handed back at once rather than run or held until
its lease lapses, and a settings reload wakes the manager so a role
that gained task types starts claiming them straight away.
Starting the queue manager on every node also drains the queue channel
on nodes without outboundMta. Upstream left that channel unread, so
each message queued there parked a refresh in it, and by the code,
queueing would block once 1024 had piled up (not reproduced here).
A role object edit reaches the nodes that name that role in
INBUXA_ROLE. Moving a node to another role still means changing its
environment, and so a restart. Listener changes in a role still need a
restart too (listeners bind at boot); this change is about tasks and
delivery.
cluster::live_roles::live_role_tests (new; PostgreSQL, two nodes over
one store):
1. A node started with outboundMta delivers and runs a TLS report
task; after its role loses outboundMta and the settings reload, a
new message isn't attempted and a new report task stays pending;
with the role back, both are taken up.
2. A node started with no task type at all gains outboundMta: a
waiting message is attempted and a report task runs.
On main the test fails at step 1 ("delivery attempted without
outboundMta"); with step 1 bypassed, step 2 fails (nothing picked the
message up in 20 s).
Cluster rehearsal 3: with PostgreSQL paused (docker pause, so its
kernel still answered TCP keepalives), requests on connections already
checked out hung until it came back, and /healthz/ready stayed 200
through the outage. #41 bounded getting a connection, not using one.
Client-side query limits (store::backend::query_timeout). Every
operation on a PostgreSQL or MySQL connection now runs under a time
limit. A server-side statement_timeout (or MySQL's MAX_EXECUTION_TIME,
which covers SELECTs only) can't do this: the server that would enforce
it is the one not answering. When an operation runs out, its connection
is closed instead of pooled, since a query may still be in flight on it
or a transaction open: deadpool's Object::take on PostgreSQL;
Conn::disconnect on MySQL, which marks the connection closed before it
sends anything, so the pool discards it even when the server never
answers.
- query, 2 minutes: reads, writes (the whole transaction with its
retries), blobs, SQL lookups, search queries and indexing. These take
milliseconds; two minutes leaves room for a large blob over a slow
link and still ends a hang.
- maintenance, 30 minutes: range deletes (account removal, purges),
unindexing, purge_store, and creating tables and indexes at startup,
which can legitimately run long in one statement. Their existing
chunked fallback for server-side statement timeouts is unchanged.
- iterate (exports, reindexing, maintenance scans) can run for hours,
so the query limit bounds each wait for the database (preparing, the
query starting, the next row) rather than the whole scan.
The limits are fixed, like the pool timeouts; the DataStore schema has
no field for them. Tests set them with Store::with_query_timeouts
(test_mode only).
Readiness. /healthz/ready answered 200 whenever a data store was
configured. It now reads one key from the data store with a 2 s limit
and reuses the answer for 2 s, so probes can't load the database;
while one probe runs, others get the last answer. The first failed
probe of an outage is logged. /healthz/live stays 200: restarting a
node doesn't bring its database back, and an orchestrator restarting on
failed liveness would restart every node at once. The container
HEALTHCHECK already uses /healthz/live.
Tests, store::pool_timeout (a proxy that stops forwarding while
keeping connections open plays the paused database):
- postgres_query_timeout, mysql_query_timeout (new): with four pooled
connections open, a read, a scan and a write each fail with "Query
timed out" 2.0 s after the pause (2 s test limit); once the proxy
forwards again the store answers. With the limits set to an hour
(upstream's behavior), the read was still waiting at the test's 20 s
limit.
- postgres_readiness (new, STORE=PostgreSql): a node's data store
goes through the proxy; /healthz/ready is 200, 503 about 4 s after
the pause while /healthz/live stays 200, and 200 again about 2 s
after it ends.
- postgres_pool_timeout, mysql_pool_timeout: pass as before.
store::store_tests (PostgreSql, MySql, including the MariaDB statement
timeout step) and store::task_locks (PostgreSql) pass;
store::search_tests (PostgreSql) fails at the same ordering assertion
(query.rs:684) as on main.
A three-node rehearsal on PostgreSQL saw searches take about 185 ms
with 80 to 260 pages in the full-text indexes' pending lists, 2 to 6 ms
right after gin_clean_pending_list() or VACUUM, then creep back up as
mail came in. The search tables' GIN indexes were created with the
default fastupdate=on: new entries wait in an unindexed pending list
that every search scans in full until VACUUM (or 4 MB of backlog)
merges it, and autovacuum only visits an insert-only table after
thousands of inserts.
The search GIN indexes are now created WITH (fastupdate = off), so an
insert pays its index update at once. The schema step runs at every
startup (create_search_tables, via SearchStore::create_indexes), so
indexes made before this change are switched there: when an index's
reloptions don't already turn fastupdate off, ALTER INDEX ... SET
(fastupdate = off) and one gin_clean_pending_list() merge its backlog.
The ALTER takes a SHARE UPDATE EXCLUSIVE lock, which blocks neither
reads nor writes; after the first startup the step is one catalog read
per index. A failure is logged and startup goes on (search still
works, only slower).
Per-table autovacuum settings for the search tables are left alone.
The pending list was the only reason the insert threshold mattered for
search; dead tuples and freezing are served by the defaults, and table
settings would override whatever tuning the DBA has done.
MySQL is unaffected: InnoDB FULLTEXT keeps new entries in an in-memory
cache that queries read directly, with no setting like fastupdate.
store::search_gin::postgres_gin_fastupdate (new, PostgreSQL) builds
the search schema in a schema of its own and checks pg_class.reloptions:
fastupdate=off on every GIN index of a fresh schema; then, with the
option reset to the default and 500 rows pending, one startup turns it
off everywhere and leaves no pending tuples (pgstatginindex); a second
startup changes nothing. On main it fails at the first check.
write_reload_target sent AllowedIp writes to the blocked-IP reload, but
that reload rebuilds only BlockedIps. Allowed IPs are parsed into the
core's security settings (Security::parse), which only a full reload
rebuilds, so an AllowedIp write reported x:settingsReload applied: true
while the change wasn't live until the next full reload.
AllowedIp now maps to the full reload, like the other settings objects;
BlockedIp keeps its targeted reload.
system::auto_reload::settings_reload_tests now creates an allowed IP
over JMAP and checks that is_ip_allowed sees it with no ReloadSettings,
and that destroying it takes it out again. On main it fails ("allowed
IP not in the running settings").
A 3-node rehearsal (PostgreSQL + NATS + Garage) found two ways a crash
leaves work stuck:
Pool hangs. The PostgreSQL pool (deadpool) was built with no timeouts,
so a request waited for a free connection, and for one to be opened or
recycled, for as long as it took: forever when the server stopped
answering. MySQL's pool (mysql_async) has no wait timeout at all.
- PostgreSQL: wait 30 s (or the store's timeout if longer), create the
store's timeout or 15 s (it bounds the whole handshake, where
tokio-postgres's connect_timeout covers only the TCP connect), recycle
10 s. The pool config is now always set, not only with
poolMaxConnections.
- MySQL: every connection is taken through MysqlStore::conn(), which
gives up after 30 s.
- Both: TCP keepalive after 60 s idle, so a server that vanished
without closing the connection is noticed in minutes rather than the
two-hour system default.
The DataStore schema has no pool timeout settings, so these are fixed
defaults; the store's own timeout bounds connecting on PostgreSQL.
Task locks. A task lock lasted an hour, so after a hard crash the dead
node's tasks waited up to an hour and five minutes. The lock is now a
five-minute lease: while this node runs a task, the task manager renews
its lock every third of the lifetime (InMemoryStore::renew_lock, a
compare-and-set on the store backends and SET XX EX on Redis, which
leaves a lock that already expired alone). A killed node's tasks run
elsewhere within about five minutes plus the claim recheck. A task this
node holds isn't handed to a worker again by the scan.
store::pool_timeout (new): a local listener that accepts connections
and never answers plays a hung server; a PostgreSQL store with a 2 s
timeout returns an error in about 4 s, and a MySQL store in 30 s.
Without the timeouts both wait for good. store::task_locks gains a
task held for 1.5 lock lifetimes: its lease is still held, and released
when the task ends.
A 3-node rehearsal found taskQueueProcessing didn't filter anything:
roles.task_manager only decided whether the task manager started, and
report, ACME, DKIM, DNS, calendar, thread-merge and restore tasks ran on
any node with a task manager (manager.rs returned true for them). A node
whose role left taskQueueProcessing off still ran them if it indexed or
did maintenance.
Every task type now answers to one ClusterTaskType (task_enabled):
- IndexDocument, UnindexDocument, IndexTrace: searchIndexing
- AccountMaintenance, TenantMaintenance, DestroyAccount:
accountMaintenance
- StoreMaintenance: storeMaintenance
- SpamFilterMaintenance: spamClassifierTraining
- DmarcReport, TlsReport: outboundMta. They build and send reports to
other domains (TLS reports can go straight to an HTTPS endpoint),
which is the outbound MTA's business.
- CalendarAlarmEmail, CalendarAlarmNotification, CalendarItipMessage,
MergeThreads, RestoreArchivedItem, AcmeRenewal, DkimManagement,
DnsManagement: taskQueueProcessing, the role for queue tasks with no
role of their own.
A node that may not run a task leaves it unclaimed (no lock), so a node
that may picks it up. The task manager also starts on a node whose only
task role is outboundMta, so reports still run there.
cluster::task_roles::task_role_tests (new, two task managers over one
PostgreSQL store): node A (taskQueueProcessing only) runs a DNS task and
leaves an unindex task and a TLS report pending; node B (searchIndexing
and outboundMta) comes up and runs those two; a DNS task scheduled next
stays pending on B and runs on A. On main node A runs the TLS report.
A 3-node rehearsal found that saving an MtaDeliverySchedule left it
unknown to the queue ("Queue strategy not found") until someone ran
x:Action ReloadSettings; only Directory and Authentication writes
reloaded (DIR-17). The admin UI has to remember a separate reload after
every save, and a script or API client that doesn't gets a server
running stale settings.
x:<Object>/set now reloads the running settings when it created,
updated or destroyed an object they are built from, and broadcasts the
same RegistryChange::Reload over the coordinator as ReloadSettings, so
every node applies it:
- Settings objects (MTA, spam filter, listeners, tracers, Sieve system
scripts, cluster roles, directories, ...: the object types the core,
telemetry, listener and directory builders read) get a full reload.
- Certificates, lookup stores and blocked/allowed IPs get their own
targeted reloads.
- Accounts, domains, roles and other data read as needed, stores (they
take a restart) and applications (their own reload action) get none.
Full reloads are coalesced: a write waits for a reload that started
after it was stored and joins one if it can, so a burst of writes, or
a request with many objects, costs one or two reloads, not one each.
The write itself is never undone. When the reload is refused (build
errors in objects that were working, the rule from the previous
commit), the set response says so in a new x:settingsReload field,
{"applied": false, "description": "Saved, but the running settings
were not reloaded. <object>: <error>"}; {"applied": true} otherwise.
The field is absent when the write needs no reload. The description
helper is shared with ReloadSettings' refusal.
Each reload sends the queue a ReloadSettings event, so the SMTP test
harness's read_event, try_read_event and assert_no_events now pass over
those; expect_reload_settings still waits for one.
system::auto_reload::settings_reload_tests (new): an MtaVirtualQueue
and an MtaDeliverySchedule created over JMAP are in the running
settings with no ReloadSettings, and gone once destroyed; eight
concurrent creates all land; a write whose reload fails is stored and
reported applied: false with the error; a domain write carries no
x:settingsReload. On main the new schedule is missing. The cluster
broadcast test (three nodes, PostgreSQL + NATS) now checks that every
node has a schedule created on node 0 without a reload.
A 3-node rehearsal found every settings reload refused, cluster-wide,
because one node couldn't resolve the Pyzor server:
- PyzorConfig::parse resolved the host while building the settings and
made a failed lookup a build error. It now keeps the host and port and
resolves when a message is checked (an IP address is used as is, a
name is reused for five minutes, the lookup counts against the Pyzor
timeout). A failure there is a Pyzor error for that message.
- A milter's hostname was resolved the same way, with a blocking
to_socket_addrs in async code. An IP address is kept; a name is now
resolved on each connection.
Other build-time I/O is already non-fatal: directories that can't
connect become unavailable with a warning (DIR-21), and the AI model
locality check only warns.
reload_registry swapped the core only when the whole build was free of
errors, while boot runs with whatever built. One failing object thus
refused every later reload, and the running settings went stale. Now a
reload is refused only for errors in objects that built when the
running settings were built (at boot or by the last applied reload):
applying it would lose those. Objects that already failed then are
missing from the running settings anyway, as at boot, so their errors
are logged and returned as known_errors but don't hold the reload back.
Refusing on new errors keeps a bad edit from taking a working object
out of service; the admin gets the error instead.
ReloadSettings now says "Settings were not reloaded." and names the
object and its error ("Tracer with id ...: Only one console tracer is
allowed"), with a count of any further errors. A refused reload after a
directory change logs its errors too.
system::reload::reload_tests (new): with Pyzor enabled on an
unresolvable host, ReloadSettings succeeds (on main it fails with
"Invalid address: failed to lookup address information"); an IP host
needs no lookup; a new build error refuses the reload, names the object
and leaves the running settings unchanged; the same error, once known
from the running settings' build, no longer blocks; once fixed, a new
error there blocks again. smtp::inbound::milter's session test now
names its milter "localhost", so the connect-time lookup is exercised.
A 3-node PostgreSQL rehearsal found IMAP SEARCH FROM "noreply" matched
0-2 messages where RocksDB matched 23 of 930. The message indexer hands
each address and display name of From/To/Cc/Bcc to the search store as
keyword text (Language::None). The built-in index splits keyword text
into lowercase runs of alphanumerics, so an address is found by its full
form, its local part, its domain or a display-name word. The SQL
backends didn't:
- PostgreSQL's text parser keeps "[email protected]" as one email
token (host names and URLs likewise), so neither "noreply" nor
"amazon.com" ever matched it. Keyword text is now split the same way
as the built-in index (SpaceTokenizer) before to_tsvector on insert
and before plainto_tsquery/phraseto_tsquery on search, still under
the 'simple' configuration, so the GIN index keeps serving the query.
The sort columns keep the raw text.
- MySQL's FULLTEXT parser already splits on punctuation, but InnoDB
never indexes its stopwords ("com", "de", "www", ...) or words under
innodb_ft_min_token_size (3), and a required +word it hasn't indexed
matches no row. So "amazon.com", "[email protected]" or "jane doe" found
nothing. Those words are now matched with a word-boundary REGEXP on
the rows the indexed words select. In language text (bodies,
subjects) they are dropped when other words remain, and only checked
when nothing else is left, so "the invoice" no longer finds nothing
either.
Existing PostgreSQL search indexes hold the old single-token vectors and
need a reindex (the reindexAccounts task) before address searches find
old messages. MySQL needs none: only the query changed.
store::search_tests gains test_address_search: five messages, 28
FROM/TO/CC/BCC searches by full address, local part, domain, domain
labels, display name and hyphenated local part, plus a TEXT-style OR,
with the same expected ids on every backend. It passes on RocksDB,
SQLite, PostgreSQL and MySQL; on main it fails on PostgreSQL (From
"noreply") and MySQL (From "[email protected]").
A node that started while NATS was down never got a coordinator. The
connect failed at boot, bootstrap recorded a build error and the node ran
with Coordinator::None until restarted. It had no broadcast subscriber
or publisher, so cross-node push and cache invalidation to it stayed
broken, and its healthcheck said nothing about it. Losing NATS after
startup was silent too.
- The NATS client now connects in the background
(retry_on_initial_connect): startup never waits on NATS or fails over
it, the node gets its coordinator, subscriber and publisher at once,
and the client keeps trying (async-nats's backoff, at most 4 s apart)
until NATS answers. Subscriptions made meanwhile start delivering when
it does. A configured maxReconnects still ends the attempts.
- Three new events report the connection: cluster.coordinator-connected
(info), cluster.coordinator-disconnected (warn: lost, closed, gave up,
or not connected within the connection timeout at startup) and
cluster.coordinator-error (warn: a failed attempt, reported once per
outage rather than every retry, and server errors, slow consumers and
lame duck mode). They are in the packaged schema, ids 644 to 646.
- GET /healthz/cluster reports the coordinator: 200
{"coordinator":"connected"}, 503 {"coordinator":"disconnected"}, or
200 with "none" (no coordinator) or "unknown" (a backend that doesn't
track its connection). /healthz/live and /healthz/ready are unchanged
on purpose: a node without its coordinator still serves mail, and
failing those would have orchestrators restart, or pull out of
service, every node at once whenever NATS is down.
Only NATS connects lazily; the other coordinator backends still fail at
boot as before.
cluster::coordinator::coordinator_reconnect_tests starts a node against a
NATS port with nothing behind it, checks it boots with a coordinator and
reports it disconnected, subscribes, then starts NATS on that port: the
node connects on its own and the subscription receives a message from a
second client. Stopping and restarting NATS shows disconnected, then
connected, and the same subscription keeps working.
A cluster rehearsal (PostgreSQL + NATS) left index tasks pending well
past the one-hour task lock after the node that claimed them was stopped
or killed. The exact cause there isn't confirmed; this closes every path
found in the task manager that stretches a takeover past the lock, or
keeps a task claimed without running it:
- A graceful stop never released the locks it held, so every task the
node had claimed stayed blocked for an hour. The server now tracks the
locks it holds (common::ipc::TaskLocks) and, once the shutdown signal
arrives, stops claiming and releases them before exiting.
- A node that failed to claim a task (another node held it) set its own
local hold for a full lock lifetime from that scan. If the holder
claimed it just after the scan began, or ran on a clock ahead, that
hold ran out a moment before the lock did and was set for another
hour: two hours in all. Such claims are now tried again every five
minutes (a twelfth of the lock lifetime), and the task manager wakes
up for them: before, a node without a coordinator could sleep up to
five minutes past the recheck, or until something else woke it.
- A worker that panicked took its task type down on that node for good,
while the scan kept claiming that type's tasks and failing to hand them
over, re-taking each lock as it expired and so starving every other
node of them. Each batch now runs on a task of its own; a panic is
logged, the batch's locks are released and the worker carries on. A
failed hand-over releases the lock too.
- A claimed task the worker couldn't read, or found gone, kept its lock
for the hour. It is released.
- An IndexDocument task for a file (not indexed) returned no result,
which shifted every later result in the batch onto the wrong task in
update_tasks. It returns Ignored. Nothing queues such a task today.
The lock lifetime stays one hour; it now lives per server so the tests
can shorten it.
store::task_locks::task_lock_tests plays a second node by writing its
locks straight into the in-memory store: tasks it claimed and abandoned
run here once its locks expire, including locks that outlive this node's
view of them, and a graceful stop hands this node's locks back at once
and claims nothing more. It passes on RocksDB, SQLite and PostgreSQL.
With the old recheck it fails.
Two release builds side by side on one machine each take twice as long,
and production only needs amd64. publish-amd64 now pushes :<version> as
soon as the amd64 build is done; publish-arm64 builds arm64 afterwards,
then replaces :<version> with the two-platform index and moves :latest.
Both jobs use one named BuildKit builder whose container outlives the
job, so the dependency layer (cargo chef cook) is reused until the
dependencies change. The release is created after amd64; the binaries
are attached once arm64 is in.
The trace index task wrote the event type (its name) and the queue id as
text, but the tracing search index types both as integers on every
backend: BIGINT on PostgreSQL and MySQL, long on Elasticsearch. On
PostgreSQL every batch holding a trace document failed with "cannot
convert between the Rust type String and the Postgres type int8", and
since a batch writes trace and email documents together, email indexing
stalled behind it.
The document is now built by trace_search_document(), which writes:
- the event type as the opening event's numeric id, the event
x:Trace/query's event filter already matches on;
- the queue id as an integer, the first one the trace names;
- every queue id into the keywords as well, since the column holds one
value and an SMTP session can queue several messages.
index_keyword() replaced the field on every call, so before this only the
last event type and queue id survived anyway.
x:Trace/query's queueId filter parses the id (a string, or now a number)
and matches the column or the keywords, so a session is found by any of
its queue ids on every backend. The monitoring spec says what is indexed.
Traces indexed before this on the built-in index keep their text values;
the reindexTelemetry maintenance task rebuilds them.
Tests: the search store suite builds trace documents with the index
task's code, indexes them and finds them by queue id, event type and
keyword (Sqlite, PostgreSQL, MySQL); the monitoring suite finds a real
trace by queueId through x:Trace/query.
The broadcast subscriber waited 1 << retry_count.max(6) seconds between
failed subscribe attempts. max(6) turns the cap into a floor: the first
retry waited 64 s instead of 1 s, and each later one doubled without a
bound (and would overflow the shift after enough failures).
The delay now comes from subscribe_retry_delay(), 1 s, 2 s, 4 s ... capped
at 64 s, and the retry counter saturates. A unit test pins the schedule
and the top of the range.
--export skipped three things, so a move from one database to another
(RocksDB to PostgreSQL, say) lost them without a word:
- archived items (subspace j), the records behind undelete;
- spam training samples (subspace w);
- the trained spam classifier and its trainer state, blobs stored under
fixed names that no blob link points at, so the walk over links never
reached them.
j and w now travel with the registry family, where their indexes and id
counters already were, so EXPORT_TYPES=registry keeps them consistent.
The two named blobs travel with the blob family. The file format is
unchanged and import reads any subspace it is given, so an export made
by an older binary still imports.
The full-text index (subspace z) stays out, on purpose. It belongs to one
search backend: PostgreSQL and MySQL index into their own tables and have
no z table at all, and external engines keep the index themselves. So
--import now returns the subspaces it wrote, and boot queues the
reindexAccounts and reindexTelemetry store maintenance tasks, the same
ones an administrator can queue by hand, to rebuild the index for
whichever search store the server runs with once it starts.
The round trip also turned up a loss in import itself: the SQL stores
add a negative amount with an UPDATE, which does nothing to a row that
isn't there yet, so every negative counter or quota vanished on import
into PostgreSQL, MySQL or SQLite. Import now creates the row first.
The in-memory subspaces (m, y) stay out: rate limits, locks, greylisting,
ACME challenge tokens and OAuth codes, all short-lived. Issued
certificates are registry objects and travel.
The store test now writes archived items, spam samples, directory
entries, the fork's own subspace and the named blobs, checks they come
back in place, then imports the same export into a fresh store of the
other local backend (RocksDB to SQLite, or SQLite to RocksDB), compares
it key for key and counter for counter, and checks the queued reindex.
It fails on the old export code ("Subspace j was not exported").
--help now says what an export holds.
#27 let the build context see vendor/, but the Dockerfile cooks the
dependencies before it copies the tree, from a recipe that carries only the
workspace's manifests. [patch.crates-io] points sieve-rs at vendor/, so the
cook failed the same way: failed to read /build/vendor/sieve-rs/Cargo.toml.
That's why 2026.9.24.2's publish failed. The builder stage now copies
vendor/ before cooking; a local build got past it into compiling the
dependencies.
context-check.py now also checks that each patched path is copied into the
cooking stage before the cook, and fails on the Dockerfile as it was.
Replaces 2026.9.24, whose tag predates the image build fix (#27) and never
published. Carries everything 2026.9.24 did -- upstream 0.16.23 and its
fixes, the scim release-profile fix -- and since then:
- identifiers renamed from the upstream name, with no aliases: the JMAP
registry capability is urn:inbuxa:jmap:registry, WebDAV tokens
urn:inbuxa:dav*, Sieve extensions vnd.inbuxa.*, the web interface client
inbuxa-webui; INBUXA_* settings only. Deploy with admin and webmail
releases that use the new names.
- the brand in lowercase where people see it.
- the spam filter rules bundled with the server; on first start they add
the AI classifier's LLM_* scores.
- a Local AI page link in Settings › Spam Filter, for the admin release
that draws it.
- two start-up migrations: the spam model moves to its renamed keys, and
the web interface's old OAuth client is retired.
Adds a link to CustomComponent/LocalAi in the packaged schema's Settings ›
Spam Filter, above LLM Classifier, and updates the schema hash so admins
fetch the new layout rather than a cached one.
INBUXA Admin draws the page (feature/local-ai-setup); this makes it
reachable. An admin from before that page would show "Unknown component"
here, so this lands after the admin release that carries it.
The server fetched upstream's latest published rules from GitHub at run
time: a version nobody here tested, code-like expressions from an account
we don't control, and the upstream name as a default in the admin form.
The published rules of spam-filter v3.0.2 are now embedded
(resources/spam-filter/, MIT, in THIRD-PARTY.md) and used whenever no other
source is configured. An empty setting and upstream's old default both mean
the bundled rules, so existing installs switch without a settings change;
the URL stays an operator override (https:// or file://). The schema default
is dropped and its description says what empty means, and the strip's
rename pass does the same to each import.
Rules load on first boot as before, and again whenever the bundled version
differs from the last one loaded, which only adds missing rules and tags.
That brings the AI classifier's LLM_* scores to installs that predate them:
production has none today.
upstream-watch now also opens an issue when spam-filter publishes a newer
release; resources/spam-filter/README.md says how to take it.
The antispam test now runs on the bundled rules, the path production
takes; SPAM_RULES_URL tests another set. Unit tests cover the URL handling
and that the bundled rules parse and score the AI tags as the AI spec says.
The rename pass vendored a patched sieve-rs and pointed Cargo.toml's
[patch.crates-io] at vendor/sieve-rs. .dockerignore ignores everything and
re-includes a short list that did not have vendor on it, so the image build
had no such directory and stopped at
failed to load source for dependency `sieve-rs`
failed to read /build/vendor/sieve-rs/Cargo.toml
CI could not have caught that: it builds from a checkout, where the
directory is simply there, and only the image build has a context to prune.
The first that was known about it was a tag that had already been pushed.
So: vendor is re-included, and tools/fork/context-check.py now asserts the
thing that was quietly assumed -- every path a [patch] section names exists
and survives .dockerignore. It runs beside the other fork checks and takes
no toolchain.
Also, the comments in .dockerignore started with // , which Docker does not
read as a comment: they were patterns that happened to match nothing. They
are # now.
v2026.9.24 was tagged on a commit CI had passed, and its release build could
not compile crates/scim at all:
error: queries overflow the depth limit!
= note: query depth increased by 130 when computing layout of
{async fn body of context::<impl ...>::writable_domain()}
The crate is ours, and the failure is profile-dependent: the release profile
computes those async fn layouts in one go and goes past rustc's default query
depth, while the dev profile never gets that far. CI builds dev, so CI was
green on a commit that could not be released. The tag produced no image and
no release, which is the one merciful part.
Two changes:
- #![recursion_limit = "256"] on the crate, which is what rustc itself
suggests, with a note saying why it only shows up in release. Proved by
building -p scim in release locally: it now finishes.
- CI builds the release profile too, on pushes to main. Pull requests stay
on dev, where the wait is worth less. A few minutes per merge is cheaper
than learning this from a tag, which throws away a multi-architecture
build and leaves a version half-cut.
The server sets a deferred recipient's next retry from its clock when the
attempt defers, in whole seconds. The test subtracted its own clock taken
when the loop next saw the message, after saving and reporting, so
whenever that lag crossed a second boundary the 2 s retry measured 1 s
and the test failed. Under load, after the other SMTP tests, that was
most runs.
It now measures from when the test started the attempt, which the
server's deferral can only follow, by under a second: each retry is its
interval or one more. Each position is still checked against its own
interval, so a wrong schedule still fails.
It listens on HTTP 19048, as the dkim2 DSN and report tests do. Those are
marked serial, this wasn't, so when the SMTP tests ran together (as
upstream's CI runs them) it could start beside them and requests reached
whichever server had the port: missing JMAP creates here, and 'You must
authenticate first' in dkim2_dsn_is_signed.
It failed everywhere but upstream's machines, for two reasons:
- The spam rules, which carry every score, came from a path on an
upstream developer's own disk. Without SPAM_RULES_URL none loaded, every
score was 0.00 and the combined case came out ham instead of spam at
13.70. The published rules of spam-filter v3.0.2 are now pinned beside
the test cases (Apache-2.0 or MIT, taken as MIT; in THIRD-PARTY.md).
SPAM_RULES_URL still overrides.
- The first combined case expects a Pyzor hit, and its digest (that of an
empty body) wasn't among the three the test mode answers, so it went to
a public Pyzor server: it failed offline and would drift with that
server's counts. Test mode now answers every digest from a fixed table,
with the empty body's added, and never reaches the network.
The test passes online and offline, alone and with the rest of the SMTP
tests. queue_retry, unrelated, still fails when it runs after the others
in one process, though it passes alone every time.
The name is inbuxa, lowercase, like the wordmark; INBUXA reads as an
acronym. The admin and webmail already changed. Here that's everything
the server shows people: the brand macro behind the protocol greetings,
the HTTP and SCIM realms, the startup banner and the calendar and contact
PRODID; the first-party OAuth client descriptions; the legacy-protocol
refusals; the default calendar and address book names and the SMTP
greeting default, in the code and the schema served to the admin
(checksum regenerated); startup and shutdown events; the User-Agent;
the sign-in and RSVP pages; the service units; the OpenAPI realm; the
crate descriptions and the README, where it's set in bold.
Identifiers that are uppercase for their own reasons stay: INBUXA_*
settings, SUBSPACE_INBUXA. So do code comments and the AGPL 5(a) notice
lines.
Tests follow: the IMAP ID name, the default collection names, the PRODID
in the iTIP fixtures and the CalDAV free-busy expectations, and the e2e
legacy-protocol refusals. The webdav, imap and jmap suites pass, so do
the unit tests of every crate touched, and 73 of 75 SMTP tests; of the
other two, antispam fails on main too, and queue_retry is a timing flake
that passes on its own.
Carries upstream 0.16.23 -- the DSN, POP3, Sieve, DMARC-report, ACME and
DNSSEC-resolver fixes in its own change log -- with the files it changed
marked under AGPL section 5(a), and one upstream test dropped that the fork's
routing makes meaningless.
It is also the first release whose tag attaches binaries: a host install can
now fetch inbuxa-linux-amd64.tar.gz or inbuxa-linux-arm64.tar.gz instead of
pulling the image and copying the file out of it.
Everything clients, users and operators meet now carries the fork's name,
with no aliases (SPEC.md §2.4, changed here from "protocol identifiers
stay"):
- JMAP: upstream's registry capability is urn:inbuxa:jmap:registry, beside
the fork's own urn:inbuxa:jmap.
- WebDAV lock and sync tokens are urn:inbuxa:dav*; clients resync once.
- Sieve: vnd.inbuxa.while and vnd.inbuxa.expressions. sieve-rs spells these
into its compiler, so it's vendored (vendor/sieve-rs, 0.7.3) and patched in;
a unit test fails if Cargo.lock ever moves past the vendored copy. The
trusted runtime now names itself too, rather than answering sieve-rs's
default.
- The web interface's OAuth client is inbuxa-webui. On every start the old
stalwart-webui client is removed and any application naming it is moved
over.
- The spam filter's blobs are INBUXA_SPAM_*; every start moves any left
under the old keys, so a trained model survives.
- SQL stores and log files default to inbuxa, in the code and in the
schema served to the admin (checksum regenerated).
- Settings are INBUXA_* only. A STALWART_* variable that's set where its
INBUXA_* one isn't stops the server at startup, naming it.
- The version-upgrade messages link docs.inbuxa.org's migration page, and
the OpenAPI description, smtp crate metadata and web-push test fixtures
lose the name.
Kept on purpose, allowlisted with reasons: the OAuth key-derivation
contexts (renaming them would end every session and invalidate every
sealed client id) and the hashed application prefix.
Also fixes a latent start-up failure: ensure_client updated an existing
first-party client with a revision of 0, which the registry's assertion
never matches, so adding a redirect URI or changing the webmail secret
failed start-up. And the principal session test now expects
legacyProtocols (C-1, added 2026-09-21), which it had missed.
Tested: the server builds without warnings; common's 106 unit tests,
including the vendoring check; a new integration test for the two
start-up migrations; and the webdav, jmap, imap and SMTP Sieve suites.
A release published an image and nothing else, so there was nothing for a
host install to download -- the only way to get the binary was to pull the
image and copy it out, which makes "install without Docker" depend on
Docker.
Each release now carries inbuxa-linux-amd64.tar.gz, inbuxa-linux-arm64.tar.gz
and SHA256SUMS, named as stalwart-migrator's are.
They are taken out of the image this pipeline just pushed rather than
compiled again. A second Rust build per architecture is the slowest thing
here, and it would leave two artifacts that are meant to be the same build
and only probably are. Extracting makes that identity a fact: the binary in
the tarball is the file the image runs. `docker create` starts nothing, so
copying a file out of an arm64 image on an amd64 runner needs no emulation.
One thing the extraction cannot carry: the image grants the binary
cap_net_bind_service, and a tar archive does not keep that xattr. The
release body says so, and says what to do instead -- setcap, or
AmbientCapabilities in the unit -- because a server that cannot bind 25 and
does not say why is a bad first hour.
Checked by hand against v2026.9.23 before this landed: both architectures
extract to the right ELF, and the amd64 binary runs on a bare Debian 13 with
every library resolved and reports its own version.
strip.py compiles the stripped tree, so a dual-licensed file that only
serves an Enterprise feature fails the import instead of the merge, as
v0.16.23's tests/src/directory/issuer.rs does. Upstream's tests of the
features the fork rebuilt are expected not to compile there and are listed
in build-check-known.txt; an error anywhere else fails the run. Checked
against both imports: v0.16.22 passes with its 16 expected errors, v0.16.23
fails on issuer.rs alone. Imports the strip leaves unused are reported.
It also renames the upstream name where clients, users or operators meet
it as an identifier, from tools/fork/renames.py: wire-protocol names, the
web interface's client id, store keys, configuration defaults and the
served schema. main is renamed with the same module, so a re-import
arrives purged and those lines don't conflict.
notice-check.py fails CI when an upstream file the fork changed, measured
against the upstream branch, lacks its AGPL 5(a) notice; --fix adds it.
It runs beside the name check in a renamed fork-checks job.
Also commits v0.16.23's strip report under docs/fork/strip-reports/, which
the import in #18 left out.
tests/src/directory/issuer.rs, new in v0.16.23, tests routing a bearer token
to a directory by its issuer. That routing is Enterprise-only upstream (the
body of get_directory_for_issuer), and the fork doesn't build it: a token
naming no address gets the server default (DIR-2). The test also calls a
helper from upstream's Enterprise-only OIDC test, so it can't compile here.
mta.rs imported types::id::Id for code inside an Enterprise snippet; the
stripped tree leaves it unused, upstream's as well as ours.
These upstream files were changed after the fork marked the files it had
modified, and never got the notice: six by the listener and schema-cache
work on 2026-09-20, two by the name check. Found by diffing against the
upstream snapshot branch, as before.
Five conflicts, resolved:
- crates/common/src/auth/authentication.rs: upstream's get_directory_for_token
and JwtClaims replace extract_jwt_domain; the per-domain directory code
(DIR-1, DIR-5 to DIR-7) is kept, and the token lookup routes through it.
The release's one new Enterprise snippet was the body of
get_directory_for_issuer, which stays returning None: a token naming no
address gets the server default, as DIR-2 specifies and as v0.16.22 did.
- crates/common/src/manager/application.rs: upstream's rewrite of the tests,
with the temp directory names renamed again, and the 5(a) notice the
name-purge change should have added.
- crates/common/src/network/mta.rs: both sides' imports.
- crates/main/Cargo.toml: the AGPL-only license kept, version 0.16.23.
- Cargo.lock: upstream's, with the fork's crates added by Cargo.
tools/fork/name-check.py reads every string literal in crates/ (comments
and test directories skipped) and fails on any that carries the upstream
name without an entry in name-allowlist.txt. An upstream merge can bring
such strings in without a conflict, so it runs on every push and PR.
The first run found three the earlier sweeps missed, fixed here: the SMTP
HELP reply pointed at upstream's website (now brand_url!), the event
collector thread was named after upstream, and the FreeBSD default data
path still said /var/db/stalwart/ where Linux already had /var/lib/inbuxa/.
Two operator-visible defaults are allowlisted as open, pending a decision:
the log file prefix and the SQL stores' default database and user.
Reads metadata only: upstream's releases list from GitHub's API and the
head of the upstream branch from Gitea's. Nothing of upstream's is
fetched, so its history can't land here. Daily at 06:17 UTC.
The first-party application descriptions and the telemetry service name and
instrumentation scope are shown to operators, and the unpacked-application
temp directory carried the name too.
Left alone deliberately: the OAuth key-derivation contexts (renaming them
would invalidate every sealed token and client id), the migration defaults
that read an upstream installation, links to upstream's upgrade guide, the
wire-protocol identifiers, and upstream's own license and templates.
brand_version_full! is user-visible -- --version, the startup banner, the
console, telemetry and the JMAP session's implementation field -- and the
name belongs only in copyright notices and the lineage line.
publish.yml replaces .github/workflows/publish.yml: on a v* tag it checks the
tag equals v<brand_version!> and is on main, builds the linux/amd64+arm64
image in one buildx run (the Dockerfile already cross-compiles, so only its
final stage goes through QEMU), pushes :<version> and :latest to the
registry, links the package, and creates the tag's release if it has none.
weekly-release.yml ports .github/workflows/release.yml: bump brand_version!
through the contents API, then create the release and so the tag, which
starts publish.yml. It only dry-runs until RELEASE_LIVE=1 and a
RELEASE_TOKEN secret exist.
GitHub took the organization's repos and GHCR offline on 2026-09-20. Repo,
release, raw-file and clone links now go to Gitea at git.coffeylabs.org,
container images to registry.coffeylabs.org, and GitLab-style /-/blob paths
to Gitea's /src/branch form. Go module paths are identifiers and stay as
they are; links to GitHub issues and pull requests are left as history.
The build now mounts the named volume inbuxa-server-cargo at /cache and keeps
CARGO_HOME and CARGO_TARGET_DIR there, so a push reuses the compiled
dependency tree (RocksDB included) instead of rebuilding it from scratch.
Both runners allow that one volume; each host keeps its own copy.
With the cache in place the job moves to runs-on: light, so it can run on
host2 as well. Cargo's parallelism now follows the job's CPU cap rather than
the host's core count, and the target dir is dropped past 25 GB.
Deleting a tenant now also removes its stored inbuxa:TenantProtocolPolicy,
in the same place the registry's other per-type clean-ups run. Without it
the row outlived the tenant, and a tenant that later came to have the same
id would have started with legacy protocols off.
The e2e deletes a tenant whose switch a server administrator had turned
off, and would check that a new tenant with the same id starts with them
on. On this build the registry hands out a fresh id instead ("d" after
"c"), so the reuse -- and with it the removal -- isn't observable over
JMAP; the test says so rather than passing silently. The risk it guards
was therefore smaller than feared, and the change is mostly about not
leaving an orphaned row behind. All 72 checks pass.
The Hardening link merged to main changed the packaged schema, which this
branch also changes. The file is gzipped, so the two can't be merged line
by line: this takes main's schema and adds security.legacy-protocols-changed
to it again, with the hash recomputed.
The impact panel's data. Every successful sign-in over IMAP, POP3,
ManageSieve or SMTP AUTH records, per account and per protocol, one
timestamp -- nothing else: no address, no IP, no client. It is written at
most once an hour per account and protocol, so a mail app polling every
minute costs a read per sign-in and a write an hour. A record that can't be
written is logged and the sign-in goes ahead.
Both switches serve it as a read-only property, recentLegacyUse, as
wouldClose serves the confirmation: a list of {accountId, name, protocol,
lastUsedAt} for sign-ins in the last 30 days, most recent first.
inbuxa:ProtocolPolicy lists every account; inbuxa:TenantProtocolPolicy
lists only its tenant's own (MT-1). Accounts since deleted are left out. It
is computed only when the property is asked for.
The recording sits where the tenant check already runs once the account is
known, which becomes admit_legacy_session: refuse if the account's tenant
has legacy protocols off, otherwise record. A refused sign-in is never
recorded.
The spec leaves the interface to the implementation; a property on each
switch keeps the panel's data behind the same permission as the switch
itself, with no new object.
Unit tests hold the 30-day window to acceptance test 11 (three days ago
listed, forty not), the hourly throttle and the keys. The e2e proves on a
running server that the admin's IMAP and submission sign-ins are listed
with their time, that a second sign-in within the hour isn't written again,
and that a tenant's list holds its own user and nobody outside the tenant.
All 70 checks pass.
The urn:inbuxa:jmap capability on the signed-in principal's own account
gains legacyProtocols: "enabled" or "disabled", the stricter of the
server's switch and the account's tenant's (legacy-protocols spec,
Interfaces). It is what the webmail needs to tell someone why their phone's
mail app won't connect (LP-19), and it closes acceptance test 13.
contract.md's C-1 gains the line. It is an optional field added, which
C-3 says doesn't bump the contract version.
tests/e2e/legacy_protocols.py reads it back from the session on a running
server: enabled for the tenant's user while both switches are on, disabled
once its tenant turns legacy protocols off while an account outside the
tenant still reads enabled, disabled for everyone while the server switch
is off, and enabled again at the end. All 67 checks pass.
The tenant switch. A tenant's administrator turns legacy mail protocols
off for its own tenant, and from then on sign-in over IMAP, POP3,
ManageSieve and SMTP AUTH is refused for every address on the tenant's
domains, while every other domain on the server carries on. No port
closes, since other tenants share them (LP-13): it is one stored fact per
tenant, read at sign-in and when client configuration is answered.
inbuxa:TenantProtocolPolicy/get and /set, one per tenant, id the tenant's:
- Inside a tenant, a principal reaches only its own tenant's switch
(MT-1): /get with no ids answers with it, another tenant's is notFound
and can't be changed. At server level /get with no ids lists every
tenant's.
- Turning it off is always allowed. Turning it back on is refused with
forbidden, naming inbuxa:ProtocolPolicy, while the server has legacy
protocols off (LP-9).
- A change raises security.legacy-protocols-changed with policy = tenant,
the tenant's id, the new value and who made it (LP-14).
- It takes sysDomainGet and sysDomainUpdate, not the two new permissions
the spec names. The switch governs sign-in on the tenant's domains, so
whoever manages those domains may turn it -- and the default Tenant
Administrator role already holds both, where new permissions would reach
no role already stored on a server (MT-12's note), leaving today's
tenant administrators without the switch until someone edited their
role by hand. The same trade inbuxa:AiLimits and inbuxa:ProtocolPolicy
made. /query is not built yet; /get with no ids covers listing.
Sign-in (LP-10 to LP-12). Before the credentials are looked at, the name
given is resolved to its domain and the domain to its tenant, so a real
account and a made-up address on the domain get the same refusal, with a
right password or a wrong one, counted as no failed sign-in (LP-11). The
words are the spec's: "Your organization allows only INBUXA webmail and
JMAP apps...", in each protocol's form. A bearer token needn't name an
account, so after authentication the account's own tenant is checked too;
a token that named nobody can't slip past.
The refusal carries policy = tenant and the domain, not the tenant's id:
IMAP answers a command's tag from the Id key, so an error holding one was
sent under the wrong tag and the mail app hung waiting for its reply. The
first live run found that; a unit test now holds the refusal to it.
Client configuration (LP-14a). Autoconfig, autodiscover, PACC and the
suggested DNS records now ask whether legacy services are off for the
domain being answered for -- the server's switch, or the domain's
tenant's -- so a tenant's domains stop offering IMAP, POP3 and
submission while others still do.
tests/e2e/legacy_protocols.py builds a tenant with its own domain, a user
and a tenant administrator, and a second tenant, and proves on a running
server: the admin sees and changes only its own tenant's switch (test 10);
turning it off is an event (test 14); the tenant's user is refused over
IMAP with the right password and a wrong one, a made-up address on the
domain the same (tests 6, 7); POP3 and submission refuse in their own
forms and JMAP still works (test 8); an account on another domain signs in
normally (test 6); autoconfig drops IMAP for the tenant's domain only; with
the server off, the tenant can't turn it back on (test 9); and once back
on, the user signs in again. All 62 checks pass.
Turning legacy mail protocols off or back on raises
security.legacy-protocols-changed (id 643, info level, also in the packaged
schema), with the scope (policy = server), the new value, who made the
change (accountId), whether listeners closed or reopened (details), which
ones (listenerId), and -- only when a listener could not be put back --
which and why (reason).
It is raised in Server::set_protocol_policy rather than by the JMAP
method, so whatever turns the switch is reported. A /set that changes
nothing -- the switch already where it was asked to be, nothing to close
or reopen -- is not a change and raises nothing.
The event is never an error, but jmap's exhaustive map from security
events to HTTP errors has to name it; it joins the other two that can't
occur there. rustfmt now also wraps LP-6's two over-long lines in
enums_impl.rs, which it flagged along with this change's.
tests/e2e/legacy_protocols.py now gives the server a stdout tracer and
reads events from the container's log: turning the switch off is exactly
one event naming the scope, value, author and listeners closed; setting it
off again raises none; turning it on is one event naming the listeners
reopened. It also proves LP-6's side: seven refused submission sign-ins
are seven auth.legacy-protocol-refused events, and there is no auth.failed
or auth.too-many-attempts among them. All checks pass.
While the switch is off, the answers that tell a mail app where to connect
stop offering what the switch closed, so a new phone or desktop app is not
sent to a port that is shut or a sign-in that will be refused:
- Thunderbird-style autoconfig (/mail/config-v1.1.xml and its other
paths) and Outlook autodiscover leave out IMAP, POP3 and SMTP
submission.
- PACC (/.well-known/user-agent-configuration.json) offers JMAP, CalDAV,
CardDAV and WebDAV, and no IMAP, POP3, SMTP or ManageSieve. The document
is rendered once per configuration load, so the JMAP-only version is
rendered beside it and chosen per request; the _ua-auto-config digest in
the suggested zone follows, since it hashes the same document.
- The suggested zone publishes _imap, _imaps, _pop3, _pop3s, _submission
and _submissions with target "." -- "not offered", RFC 6186 section 3.4 --
the spec's decision, rather than dropping them: a client that looks is
told, and an automatically managed zone replaces the old records instead
of leaving them behind.
- It also drops the TLSA records for ports 993 and 995. A TLS pin for a
port the switch has closed advertises a service that is not there.
Submission's 465 keeps its record: the SMTP lock keeps that port open.
The switch is read per answer, as sign-in reads it, so every node agrees
the moment it turns. Inbound mail, MX records and the JMAP, CalDAV and
CardDAV answers are untouched.
tests/e2e/legacy_protocols.py checks all four on a running server: with the
switch on they offer IMAP, POP3 and SMTP (the control); while it is off
they offer none of them and every legacy SRV name has target "."; and once
it is back on, autoconfig and the zone read as they did before. All checks
pass.
LP-6 added auth.legacy-protocol-refused to the packaged schema, which this
branch also changes. The file is gzipped, so the two can't be merged line
by line: this takes main's schema and adds the Settings › Security ›
Hardening link to it again, with the hash recomputed.
While legacy mail protocols are off, x:NetworkListener/set refuses to
create a listener the switch would close, and refuses an update that would
turn an existing one into such a listener -- otherwise changing a
listener's protocol would walk straight past the check. The refusal is
invalidProperties on protocol (or on bind, for a submission listener once
SMTP is unlocked, since its port is what makes it one), and its description
names inbuxa:ProtocolPolicy and says to turn legacy protocols back on first.
The rule is the switch's own, listeners::closes, so what can't be added is
exactly what the switch would close: locked protocols (SMTP, LMTP, HTTP)
and the inbound port are never refused. Putting saved listeners back
(LP-5) writes through the registry, not /set, so it is unaffected.
The e2e changes with it. LP-6's check that an IMAP listener "created by
mistake" refuses sign-in can't be set up any more -- LP-4 is what stops
that listener existing -- so that step now proves test 4 instead: creating
an IMAP listener is refused, naming the policy; an SMTP listener can still
be created; and updating it to IMAP is refused. LP-6 stays proven live over
submission, and its IMAP wording by unit tests. All checks pass.
The second lock. While legacy mail protocols are off, a sign-in over IMAP,
POP3, ManageSieve or SMTP AUTH is refused for every account, so a listener
that exists by mistake -- or submission, which the SMTP lock keeps open --
still lets nobody in.
The check sits at the top of each protocol's sign-in, before the
credentials are looked at. So the answer is the same for a right password,
a wrong one and an account that doesn't exist; it isn't auth.failed, so it
counts nothing against the account and never feeds the auto-ban; and the
session stays open, since the mail app is being told, not thrown off.
Mail apps read the spec's words (LP-12, at server scope):
IMAP NO [ALERT] This server allows only INBUXA webmail and JMAP
apps. This mail app can't sign in.
POP3 -ERR [AUTH] ...the same...
ManageSieve NO "This server allows only INBUXA webmail and JMAP apps."
SMTP 535 5.7.0 This server allows only INBUXA webmail and JMAP
apps. This mail app can't send.
SMTP AUTH is refused on every SMTP listener, port 25 included: only mail
apps authenticate, so inbound delivery is untouched. LMTP is left alone.
The policy is read from the store on each sign-in rather than cached, so
every node of a cluster answers the same the moment the switch turns.
Each refusal raises a new event, auth.legacy-protocol-refused (id 642, info
level, also in the packaged schema), with the protocol as source, the
policy's scope and the domain -- never the account. The session adds the
listener and remote IP.
tests/e2e/legacy_protocols.py now also proves, on a running server: a
normal IMAP and submission sign-in works with the switch on, before and
after; while off, submission refuses the right password and six wrong ones
with the same words and without hanging up; and an IMAP listener created by
mistake while off refuses the right password, a wrong one and an account
that doesn't exist. All 33 checks pass. SMTP sign-ins in the script wait
out a second first: every connection arrives from Docker's gateway, and the
stock inbound throttle takes five a second from one IP.
Adds a link to CustomComponent/LegacyProtocols in the packaged schema's
Settings › Security, between Settings and Blocked IPs, and updates the
schema hash so admins fetch the new layout rather than a cached one.
INBUXA Admin draws the screen; this is what makes it reachable. An admin
from before that screen would show "Unknown component" here, so this
lands after the admin release that carries it.
Ports .github/workflows/ci.yml after the GitHub account was suspended and
Actions stopped being reachable. Same checks, same order, with the image
pinned by digest in place of the workflow's SHA-pinned actions.
cleanup.yml is not ported: it pruned GHCR through an action, and GitLab
keeps that as a container registry cleanup policy on the project rather
than as a pipeline. publish.yml and release.yml are larger and follow
separately.
The Actions workflows stay in the tree as the reference.
.gitignore blanket-ignores dotfiles, so .gitlab-ci.yml is negated there the
same way .github already is.
main is protected as of today -- no force-push, no deletion, and a pull
request with a green build to merge -- and GITHUB_TOKEN is not a bypass
actor. `git push origin HEAD:main` in the cut job would have been refused
from Monday, on a scheduled run nobody watches.
GitHub would not take the obvious fix. Adding the Actions integration as a
bypass actor is rejected ("must be part of the ruleset source or owner
organization") because the organization has no app installations. The
other two routes -- an organization-level ruleset, a deploy key with write
access -- both amount to handing the release a credential that outranks
the rule, which is a worse thing to own than a slower Monday.
So the bump lands the way every other change does. It commits to
release/v<version>, opens a pull request, waits for the build the ruleset
requires, merges, and tags what came out. The waiting is not merely the
rule being satisfied: a release cut from a tree that does not compile is
the failure this whole arrangement exists to prevent, and until now
nothing checked.
Three details that would each have produced a wrong tag. The sha comes
from GitHub's merge commit, not the tip that was pushed, because a rebase
merge rewrites it. The pull request is tracked by number, not by branch,
because the branch is deleted on merge and a deleted branch no longer
resolves to its pull request. And a failed or slow build leaves the pull
request open and cuts nothing, rather than tagging whatever main happened
to hold.
Quiet weeks are unaffected: the tag still names the bump commit, so
`previous..HEAD` is still zero when nothing else has landed.
The cost is a Monday run that now takes as long as a full build -- about
25 minutes at the moment, most of it saving the cache.
main has a ruleset as of today: no force-push, no deletion, and a pull
request with a green build to merge. CONTRIBUTING said nothing about any
of it, and a contributor's first clue would have been a rejected push.
No approving review is required. A review gate nobody can pass is not a
gate, and this is a project with one maintainer; the build is the part
that has to hold.
The section also says why the rule exists rather than only what it is.
The weekly release cuts from main on a Monday and ships whatever is there,
so main is expected to be releasable continuously -- which makes "not
finished" a thing that belongs behind a default-off switch or off main
altogether, not a state main passes through on a Thursday.
Administrators can bypass. That is written down as being for correcting
the tree, not for skipping the path, because an undocumented bypass
becomes the normal route.
tests/e2e/legacy_protocols.py boots the debug binary in a container, turns
the switch off and on, and checks the ports themselves. Everything below it
was unit-tested and none of it could have told us this worked.
What it establishes: IMAP and POP3 stop answering while inbound SMTP,
submission and JMAP keep going (LP-1, LP-2, LP-3); the listeners are saved
whole (LP-1); a restart does not reopen them, which is the point of taking
the objects away rather than only the sockets; both come back on their own
without a restart (LP-5); savedListeners empties; and asking to close
submission is overruled to false and reported, with 465 still answering
(LP-21, acceptance test 18). wouldClose named imaps, pop3s and sieve, and
those were exactly the three that closed (LP-16).
One caveat about the method, because it nearly produced a false pass in
reverse. A published Docker port accepts connections whether or not
anything is listening in the container, so connecting proves nothing. The
first run of this script reported IMAP still open after the switch, and
that was the script being wrong, not the server. Each port now has to
speak: a TLS handshake on 993, 995 and 465, a greeting on 25.
It lives under tests/ because target/ is ignored and this is worth keeping.
It derives its own root, needs Docker and a debug build, and clears its
state directory first -- a half-bootstrapped one from an earlier run is no
longer in bootstrap mode and the recovery admin stops working.
Found while setting up the live check, which is the only place it could
have shown: every unit test passes without it.
spawn_restored_listeners re-parsed the listeners and spawned them, but
never bound their sockets. Binding is not part of parsing -- it happens in
bind_and_drop_priv, once, at startup -- so listen() would have failed on an
unbound socket and the port would have stayed shut while the policy
recorded it as reopened. LP-5 would have been a promise the server did not
keep, and the operator's only clue a log line.
bind() is now split out of bind_and_drop_priv and called on its own here.
It cannot be the whole of bind_and_drop_priv, because that also drops
privileges, which must happen once at startup and never again.
That split has a consequence worth stating: a listener on a port below 1024
cannot be bound again once privileges are gone. Ports 143 and 110 are the
realistic cases. Rather than leave such a listener parsed, spawned and
silently dead, the bind errors are read back and those listeners are
reported as needing a restart -- which is the "cannot be recreated" case
LP-5 already anticipated, and it stays saved for another try.
Re-parsing is also narrowed to the listeners being restored, so putting one
back cannot bind a port another listener already holds.
The fork had CI and nothing after it. v2026.9.20 was tagged and released
by hand, and there has never been an image: running INBUXA meant
building the tree yourself, or using install.sh to do it for you.
This adds the three workflows ihasmail already runs -- weekly release,
publish, prune.
Monday 10:07 UTC, and nothing on a quiet week. Last of the three, so
INBUXA Admin and the webmail release ahead of the server they talk to,
and staggered so a bad Monday names one repository rather than three.
The version is the difference from ihasmail. ihasmail derives its
version from the commit it builds, so its release only reads. INBUXA's
lives in the brand_version! macro, deliberately apart from Cargo.toml so
upstream's bumps merge without conflicts -- so the release writes it:
the bump is committed to main and the tag names that commit. The tree a
tag points at therefore reports the version the tag claims, which a tag
placed beside an unbumped macro cannot promise.
Both the bump and the read are scoped to the macro body and fail if they
do not match exactly once. branding.rs holds other string literals, and
a bump that silently edited one of those, or an image tagged from one,
would be worse than a run that stops.
The existing Dockerfile needs nothing: it cross-compiles from
BUILDPLATFORM and takes no arguments beyond TARGETPLATFORM, so each
architecture builds on its own native runner as ihasmail's does, without
docker-bake.hcl. `docker build --check` is clean.
Two things to expect from the first run. GHCR creates a package private
the first time even in a public repository, and no workflow can change
that, so the first image will refuse an anonymous pull until its
visibility is set by hand. And a full Rust build of this tree is long;
the per-platform GitHub Actions cache is what keeps the second one from
being just as long, and it is worth watching that it stays inside the
cache limit.
The switch is now reachable. /get and /set on a server-level singleton,
wired through jmap-proto the way inbuxa:AiLimits is: object, method names,
request and response variants, reference resolution and evaluation.
/set does not write the policy. It hands what was asked to
Server::set_protocol_policy, which applies the locks, moves the listener
objects and opens or closes their sockets, and reports what happened. So
the method cannot drift from what the switch actually does.
Two properties exist for the screen rather than the server. lockedProtocols
serves LP-21's locked set, so the selector renders SMTP and JMAP locked
from what the server says instead of a list the front end carries -- and
unlocking later needs no admin release. wouldClose answers LP-16: exactly
which listeners turning the switch on would close, by name and port, before
anything happens. It is computed against a hypothetical disabled policy, so
it reads the same whichever way the switch is set, and the registry is only
asked when the property was requested.
savedListeners, changedAt, changedBy and both of those are the server's to
say; a client that sets one gets invalidProperties naming it. closeSubmission
is different: locked, not immutable, so it is overruled rather than refused
and the response hands back what was really stored (false). JMAP already has
the place for that, the value beside an updated id.
Permissions reuse SysNetworkListenerGet and SysNetworkListenerUpdate rather
than adding to a schema-generated enum -- the same choice AiLimits made with
the classifier's. It also reads right: this takes listeners away and puts
them back, so whoever may edit a listener may turn the switch.
changedBy stores the account id, not the name, which survives a rename.
Still no screen, no sign-in refusal (LP-6) and no event (LP-8).
The join: the policy decides, features owns the listener objects,
ListenerControl owns the running sockets, and only Server has both.
Server::set_protocol_policy is what a click performs. It applies the locks
to what was asked before storing anything (LP-21), so what is recorded is
what the server allows. Closing removes each listener object and then stops
its socket; opening puts the object back and then spawns it. The order is
the point in both directions -- a socket stopped while its object remains
returns on the next restart, and a socket spawned before its object exists
has nothing to come back to.
saved_listeners is carried over from the stored policy rather than taken
from the request. A client never sets it, and a /set that omitted it would
otherwise lose the listeners still waiting to come back.
Putting a listener back has to bind a fresh socket, so it re-parses from
the registry -- the objects are already back by then -- rather than trying
to revive the saved one. Only main knows which session manager a protocol
wants, so it leaves a spawner behind at startup and spawn_listener is now
shared between that and the initial spawn. Without a spawner a restored
listener is reported as pending a restart rather than promised, which is
what the test servers will see.
A listener that cannot be put back does not stop the others and stays
saved for another try (LP-5).
Still nothing an operator can reach: no JMAP method calls this yet, and no
sign-in is refused. What it does do is close and reopen a port on a
running server, which is the part that did not exist this morning.
John, 2026-09-20: "SMTP and JMAP should be shown with the selector locked,
we want to prevent those two protocols from being shutdown for now."
The selector lists every mail protocol the server speaks, so the operator
sees the whole surface at once; SMTP and JMAP sit in it named and visibly
not switchable. JMAP was never closeable -- closing it locks everyone out
of their mail and the operator out of INBUXA Admin, with no way back but
the host -- and is now visibly so. SMTP is locked whole. LP-3 already
spared inbound on 25; this extends that to submission on 465 and 587,
which LP-1 would otherwise have closed by default.
So closeSubmission has no effect while the lock stands, and is forced to
false. A client that asks for true is not refused: the value is recorded,
overruled, and the overrule reported, because the field is specified and
the lock is meant to be temporary. is_locked() is consulted before
anything else in closes(), so no phrasing of a request reaches past it.
This costs the feature nothing. Submission's ports stay open and sign-in
over them is still refused once LP-6 lands, which is the case acceptance
test 2 already described: a mail app reaching 465 is told it cannot sign
in rather than finding nothing listening. The operator also keeps a port
they may well be forwarding, which is the LP-20 problem in miniature.
The locked set is a server constant the front ends read, not a list they
carry, so unlocking later is a server change and no admin release. The
LP-3 tests stay as they are, to keep it covered if the lock is lifted.
Recorded as LP-21, with acceptance tests 17 and 18.
LP-1 and LP-5, the registry half. close() removes every listener object
the policy closes, saving each one whole first; reopen() puts them back.
The switch removes the listener objects, not just their sockets. A stopped
socket returns on the next restart, which would reopen every port the
operator had just closed, and the operator would have no way to know. A
removed object stays removed, and a server that boots with the switch on
never spawns those listeners at all -- so there is no boot-time special
case to write or to forget.
LP-3 is decided here, on the object rather than the running socket, and
sees every address a listener binds: a submission listener that also binds
25 is inbound and stays. lmtp and http are never candidates.
A listener that cannot be put back does not stop the others; it comes back
with its reason and stays saved for another try (LP-5). A delete the
registry declines is reported as not removed, so the policy never claims a
port is closed while it is still accepting.
Stopping the running socket is still a separate step in common, which owns
the listener registry. Nothing calls any of this yet.
The server-wide legacy-protocols policy: the switch, whether submission
closes with it, the listeners taken away to honour it, and who last
changed it. Stored like inbuxa:AiLimits, as JSON in the fork's subspace,
so an unset field reads as its default and an old record still loads.
closes() is where LP-3 lives. imap, pop3 and manageSieve are named
outright; smtp is not, because an SMTP listener is inbound or submission
depending on its port and nothing else can tell them apart. A listener
bound to 25 is inbound whatever it is called, including one that also
binds 465, so it stays. http and lmtp are never candidates at all.
savedListeners keeps each listener's registry object whole rather than a
few fields of it. LP-5 promises the listeners come back exactly as they
were, and a listener carries proxy networks, TLS timeouts and socket
options that no one should have to re-derive -- a field this code has
never heard of has to survive the round trip too, and a test holds that.
The module is under security/ rather than beside the rebuilt features,
because this one is not a rebuild: upstream has nothing like it.
Still only a fact. Nothing reads this policy yet, so no port closes and
no sign-in is refused; the acting code needs the listener registry and
the config store, which live above this crate.
Left over from this morning's dependabot merges: the minor-and-patch group
freed sequoia-openpgp to use base64 0.22.1, but the lockfile still pinned
0.21.7 for it. Any cargo invocation rewrites the line, so it was showing up
as spurious drift in unrelated diffs.
No manifest changed and nothing is upgraded here; this only writes down
what cargo already resolves.
The legacy-protocols switch has to close the IMAP, POP3 and ManageSieve
ports and leave everything else accepting. The server could not do that.
Two findings from the source, both now recorded in the spec. A settings
reload never closes a port: cache/reload.rs parses the listeners only to
collect configuration errors and drops the result, and sockets are bound
once at startup through init.servers.spawn in main.rs. And there is only
one shutdown signal -- Listeners::spawn makes a single watch channel and
hands every listener a clone -- so the one thing the server could do was
stop all of them at once, port 25 included. That answers the spec's open
question 1, and the answer was neither of the two it offered.
So each listener gets its own channel. ListenerControl holds the sending
ends keyed by listener id; firing one breaks that accept loop, which drops
its TcpListener and closes the socket. The accept loop itself is unchanged
-- it already did the right thing, it just had no way to be told about one
listener. stop_matching takes a predicate and a keep list, because the
inbound listener shares its protocol with submission and telling them
apart is the caller's job (LP-3), not this registry's.
spawn_with_control is a second method rather than a change to spawn. The
registry owns the senders, so a dropped registry would stop every listener
at once; the four test callers pass no registry and keep the old shared
channel exactly as it was.
Whole-server shutdown now fires the per-listener channels too, since the
returned sender no longer reaches them.
No policy, no JMAP and no screen yet: this is only the mechanism, with
seven tests over stopping one, stopping many, sparing port 25 and sparing
submission. It closes no port on its own, and it does not touch the host's
firewall or any port-forward -- that is LP-20, and stays the operator's.
decancer 4.0 changes CuredString's Deref target from String to str. That
is all it takes to break two call sites in the classifier: .as_str() used
to resolve to String::as_str through one deref, and now resolves to the
inherent str::as_str, which is still unstable (rust-lang #130366). Stable
rustc rejects it, so the whole crate fails to compile -- the two E0658s
that are currently red on the decancer bump in PR #4.
Neither call site wanted an inherent method, only a &str. Deref coercion
gives that under either target, so dropping the .as_str() fixes 4.0 and
keeps 3.3.3 building; cargo check passes against both. The result is
identical either way, so no behaviour changes here.
Committed against 3.3.3, which is still what the lockfile pins. The bump
itself stays PR #4's to carry, and rebases onto this.
Translation::String going from Cow<'static, str> to CuredString, the other
breaking change in the 4.0 notes, touches nothing: the type appears
nowhere in the tree.
brand_version! goes to 2026.9.20 (SPEC.md 2.6: YYYY.M.D), the version this
release is tagged at. Verified from the built binary rather than the
source: --version prints "2026.9.20 (Stalwart 0.16.22)" and --help leads
with "INBUXA Server 2026.9.20 (Stalwart 0.16.22)".
install.sh still said "See https://inbuxa.org once it's up", which stopped
being true when the site went up this morning, and offered nothing but a
cargo line. It now names both ways to build, says what a server with no
configuration does, and points at the releases page and the docs.
The installer it stands in for is still unbuilt (SPEC.md 6.1), and the
script says so plainly: an installer trusted with a mail host is not a
thing to improvise, so it declines rather than half-doing one.
README: inbuxa.org is up, so stop saying it is not, and link the docs.
John, 2026-09-20: committed, remove it. Done the same day as the cutover and
before the certificate-renewal gate this page proposed, which is the
operator's call to make and is recorded as such.
Archived first and the archive verified off-host by checksum, then the tree,
the unit and its drop-in removed. The stalwart user stays: redis-server runs
as it, which step 0's pgrep had already shown and which is exactly the kind
of thing that makes "remove the service user" a bad reflex.
Two things the doing taught, both for migration.md. The unit does not live
inside the tree it manages, so an archive of /opt/stalwart alone is not a
restorable rollback -- stalwart.service and its drop-in had to be saved
separately, and a tool that archives before retiring has to take them too.
And retiring changes what a rollback means: up to that moment it was a
service swap against a store still on disk, minutes and no restore; after
it, an untar, a chown, a unit to reinstate and a webmail image that is no
longer on the host. Still possible, slower, and no longer what the Rollback
section describes.
INBUXA is on the fork. The window was 78 seconds, 2.2 G of store copied in
1.4, mail queued at senders and nothing lost, both front ends up within the
hour, and the rollback never needed.
The page stops being a plan and becomes the record of one, which is what
migration.md is built from. Written up by kind rather than in order, because
nobody reading it later wants the chronology.
What step 0 was worth, most of all. Reading `systemctl cat stalwart` before
touching anything found a network namespace nothing in this document knew
about, and that was two failures rather than one: a collision with nginx on
443, loud and quickly understood, and egress from the wrong address, which
would have cost the provider's port 25 exemption and failed outbound mail
at every receiver with no local symptom and no logs to find it in.
What the copied store brought with it, three times in three guises: a
tracer still writing to the old tree, Stalwart's own web interface being
served from the registry's Application entries, and — from the other
direction — a front end configured by copying variable names the fork had
renamed. A front end reporting healthy is not a front end talking to the
right server; the health check passed while it pointed at example.com.
What this document had wrong: `systemctl mask` cannot mask a unit that
lives in /etc/systemd/system; there is no "let mail flow" gate, because the
fork takes port 25 as it starts and the free-rollback window closes there;
and IMAP's INBOX is not JMAP's account, so the counts differ before and
after alike.
And the bug it found, which only exists when §5.3 is followed: INBUXA Admin
hosted off the mail server cannot fetch its schema, because that response
was publicly cacheable and immutable for a year while its CORS headers vary
by origin. Fixed in 7c4add8.
The Open section loses the two the run settled and gains the two it
created, and names the ACME date: ~28 October, because R12 renews at the
halfway point and nothing brings that forward.
INBUXA Admin, hosted off the mail server as SPEC.md §5.3 requires, signs in
and then cannot load: "Failed to load the admin panel configuration. Failed
to fetch." Every other endpoint works from the same origin with the same
token; only /api/schema fails, and it is the one thing a schema-driven
interface cannot do without.
It is Chrome's cache, not CORS. Measured from the page itself: a normal
fetch fails, while cache: "reload", cache: "no-store" and a cache-busted URL
all return 200. The server never sees the failing request, which is why the
logs had nothing to show and why it looked like a CORS fault for so long.
Two things made that possible, and both are fixed here.
The schema response was `public, max-age=31536000, immutable`. It is served
behind authenticate_headers and its CORS headers vary by Origin, so it is
neither public nor safe to freeze for a year on a hash-named URL that never
changes. It is now `private`, matching what DownloadResponse already does
for the same reason. The other caller of with_immutable_cache serves the
applications' static bundles, which really are public, and keeps it.
And `Vary: Origin` was only emitted when an origin list existed. Before the
front ends are configured that list is empty, so a response cached in that
window carries neither CORS headers nor Vary, and a cache will later replay
it to an origin that should have been allowed. Vary now goes on every
response, so entries key on the origin whatever the configuration was when
they were stored.
Verified against a bootstrapped server in restrictive CORS mode, from a
browser on a separate origin: /api/account, /api/schema and the hashed
target all return 200, with `private, max-age=31536000, immutable` and
`Vary: Origin`.
Nobody hit this before because the admin has always been served from the
mail host at /admin, where it is same-origin and no CORS applies. The first
deployment that follows §5.3 meets it immediately.
cutover-run.md step 4 says to write down what has to be true afterwards
while the old server can still be asked, and step 10 checks against it.
Done by hand it gets skipped, and skipping it turns "each mailbox holds
what was recorded" into "each mailbox holds something", which is a
different check and will not catch a partial copy.
It reuses record-compat.py's client rather than growing a second one, so
the guard that makes it safe to point at a live server — call() refuses
any method that is not a /get or a /query — covers this too. Verified
that it bites: x:Account/set is refused before anything is sent.
From the administrator alone it records every account with its address,
aliases, tenant and usedDiskQuota, which is the number that moves if mail
goes missing, plus the domains and tenants. Exact per-mailbox counts need
the mailbox's own credentials, since an administrator has reach over an
account but not always into it, so --as takes one and repeats. For a
handful of mailboxes that is worth it: it makes step 10 an equality
rather than an estimate.
Aliases are resolved to the domain's name rather than its id, because an
id is not what anyone checks against at 2am.
SPEC 2.2a says INBUXA writes its own when the repository is first published,
and it is. Until now the public repository carried Stalwart's: a security
policy telling people to report vulnerabilities to Stalwart Labs, and a
contributing guide whose policy is that pull requests from anyone not on
upstream's vouched list are closed automatically. Neither is this project's,
and both were being offered to anyone who looked.
So: a security policy that says where to send a report, and what happens if
it turns out to be upstream's bug rather than ours; a contributing guide that
says what a fork of someone else's code needs from a contributor, including
the clean-room question, since the record has to stay true; the Contributor
Covenant; and a sponsor link. Upstream's two security documents move to
.github-upstream/ beside its workflows -- kept, not used, not presented as
ours.
CI builds the server and compiles every test target, and deliberately runs
no suite. The unit tests only build with the integration crate in the graph,
and the integration suites want a STORE, fixed ports and a container apiece,
so running them here would mean a tick that skipped everything or a cross
that means "the runner has no Redis". The workflow says as much, so nobody
has to rediscover it.
Also ignores /artifact: two hand-built binaries, ~190 MB, one `git add -A`
away from a public repository.
The AGPL asks a modified version to carry prominent notices saying it was
modified, and giving a date. Publishing the source is the conveyance that
asks for it, so it wants doing before the repository is public rather than
at the release.
Every upstream file the fork changed now says so in its header, beneath the
notice it came with: 164 files, found by diffing against the upstream
snapshot branch rather than by guessing, so the list is what actually
differs. Files the fork wrote itself already carry their own copyright and
need nothing. Upstream's notices are untouched, which its licence requires
and which was already true.
The README says the same thing in prose, since the obligation is on the
work as a whole and not only its Rust files.
Builds unchanged: the server and the test binary both compile.
cutover.md is the reasoning and is too long to read at 2am. This is the
same sequence as commands, for this install: eight mailboxes, two people
and a printer.
That scale settles three things the general plan leaves open. The store is
small enough that the two-pass rsync buys nothing, so it is one cp inside
the window and a simpler sequence when it matters. "Every account still
works" is two sign-ins. And a reboot inside the window is affordable, which
is the only honest proof that the fork comes up on boot and the old unit
does not — is-enabled says what is configured, a reboot says what happens.
The rollback leads with chattr -i, because step 7's guard stops the
Enterprise build exactly as it stops the fork, and finding that out during
a rollback costs the worst ten minutes of the night.
And it names the printer as its own check. It is the one user that cannot
report a fault: a hardcoded credential and an old TLS stack, of the kind a
stricter default quietly refuses. The people will phone; the printer will
just stop, and nobody will notice for a fortnight.
Taken from the cutover page, where the reasoning is worked out: with no
filesystem snapshot to take, a single copy inside the window makes every
byte downtime. A first pass while the server still serves moves the bulk
and is deliberately inconsistent; a delta pass after the process has exited
makes it consistent and moves little, because a RocksDB store is mostly
immutable SST files.
That is the difference between a window proportional to the store and one
proportional to the delta, which is the number this tool exists to
advertise. Also carried over: hand the copy to the user the fork runs as,
and make the original unwritable before the fork starts, since the old
server's store lock was the only thing holding that line until it stopped.
AGPL section 13 starts at the cutover, not at the announcement: the fork is
a modified AGPL program and its users reach it over a network. Settled
today that the source is released after the cutover, with the links live
then, and in the meantime the server's users are the operator's household,
so the people owed an offer and the people holding the repository are the
same people.
Anyone migrating their own server inherits that obligation on their first
day and has no such overlap, so the migration tool says so at the end of a
successful run instead of leaving it to be discovered.
Two decisions §2 implied but never settled.
§2.4 said no "Stalwart" in UI text; §2.6 requires the startup banner, the
JMAP implementation string and OpenTelemetry's service.version to name the
base. Taken literally, the first would strip exactly what the second exists
to keep. Version and build metadata are now exempt, with the distinction
written down: §2.4's first bullet governs identity, the new one governs
provenance.
A second bullet covers material outside the product. The name appears with
its trademark attribution; the fork relationship is stated once in the
provenance or license section; the migration path names the server it
migrates from, because an operator searching for it has to find it. The base
version stays out of taglines, page titles and social previews, where it
reads as a source identifier rather than a fact. No comparison in either
direction: what INBUXA offers is stated on its own terms.
§2.6 said the base drops out of the version string "when it happens", which
left someone judging the moment. The trigger is now the first release that
isn't a rebase on an upstream tag. Because the reason for publishing the
base is one-way store conversion, and that outlives the string, the
amendment routes it to the upgrade documentation rather than letting it go.
§8 no longer asks whether the fork follows upstream's version numbers. §2.6
answered that on 2026-09-18: it has its own.
Monitoring, SCIM, scale-out storage and per-domain directories each say
"Built 2026-09-19" in their own implementation-status sections, and the
code is where those sections say it is. The §4 table still listed them as
specs awaiting a build, which made the whole feature set look half-finished
to anyone reading the table alone.
Each row now names where the feature landed, as rows 1 to 5 already did.
Monitoring and per-domain directories sit outside crates/features, so their
rows name the paths rather than the crate.
§2.2b already recorded that all nine were rebuilt by 2026-09-19 and every
pending-rebuild gate came off. Only the table lagged.
The section warned in general and so warned about nothing. An operator
reading "it can fail in ways it cannot undo" learns less than one reading
that the window is usually longer than guessed, that opening the source
store with the new server ends the rollback permanently, that a rollback
after mail has flowed does not bring that mail with it, and that a
certificate which stops renewing says nothing for ninety days. Each of
those has been measured or seen; each has something the operator can do
about it.
Also sharpens the part that matters most and is easiest to get wrong: the
old install is a service safety net, not a data one. One copy, same
machine, one moment. A backup is a copy elsewhere that has been restored
from, and anyone who cannot say when they last restored one does not yet
know whether they have one.
And replaces the flat line about nobody else being responsible with what
it was trying to say: the operator carries the outcome, because this is
software running against a server it has never seen, holding data somebody
else depends on.
Asked for by John, 2026-09-19. A tool that stops somebody's mail server
should say so while there is still time to stop it, rather than leaving the
licence to have said it in a file nobody opens. AGPL-3.0 §15 and §16
already disclaim warranty and liability and this narrows neither; it is the
same thing at the moment it is useful.
Specific rather than blanket, because a blanket one protects less and helps
nobody: what the tool does to the server, what a rollback does not return,
and that backups and recovery are the operator's. Keeping the source
install is not a backup — it is one copy, on one machine, of one moment,
and the same disk failure takes both.
Paired with what the tool does to earn the trust it is asking for, because
that is the half that reduces the friction: a dry run the real run refuses
to start without, never writing to the source, verification before mail
flows with automatic rollback, the old install kept, and every phase timed.
--yes skips the prompt, not the dry run.
John, 2026-09-19, on both counts. The old install is kept, shut down, not
removed: its unit installed and disabled, its store read-only, started
again if a rollback is ever wanted. When it stops being worth the disk the
tool asks — keep or delete — rather than deciding, because it does not
remove the thing its own rollback depends on.
And the new stack depends on nothing in it. That is the shape's purpose:
/opt/stalwart is a reference, everything needed is copied to new paths, and
when it goes nothing notices. Nothing in the fork works against that —
inbuxa.service substitutes its own prefix, no path names the old tree, and
certificates and ACME keys are in the registry inside the store — so a
dependency, if one appears, was made by hand during the move.
Which is worth proving rather than asserting, and reversibly: nothing open
under the old tree, then rename it and leave it a day under real traffic.
Deleting proves the same thing and cannot be undone.
Two corrections this forces. The rollback has to make the original store
writable again first: the guard of step 3 blocks the Enterprise build
exactly as it blocks the fork, and finding that out during a rollback is
the worst time. And the copy has to be chowned — rsync -a preserves
ownership, so it arrives owned by the old service user while the unit runs
as User=inbuxa.
Also drops the stale "untested" wording about carrying the data back. It
was tested; it is impossible.
The mail host is ext4 (John, 2026-09-19). There is no filesystem snapshot
to take, so the sequence as written puts the whole store inside the
downtime: stop, copy everything, start.
An rsync before the stop and a second one after it moves the bulk while
mail is still flowing and leaves only the delta in the window. The first
pass is knowingly inconsistent and exists only as a warm-up; the second,
once the process has actually exited, is what makes the copy consistent.
A RocksDB store suits this, being mostly immutable SST files: what changes
between the passes is the WAL, the MANIFEST and any compaction output.
Step 3 now says so, and says to time both during the rehearsal, because
the second pass is the window and nobody knows yet how long it is.
John, 2026-09-19: the fork takes a copy of the config rather than pointing
at the old one, and /opt/stalwart goes away once the migration is
confirmed.
That turns step 4's store path from a free choice into a constraint. The
step said an existing install keeps whatever its configuration names, which
is true of the server and no longer true of this migration: nothing the
fork runs on may sit under a directory that is going to be deleted. Worth
checking before starting rather than after removing.
Removing it strands nothing else. ACME account keys and issued certificates
are written to the registry, inside the store, so they came across with the
copy; /opt/stalwart holds the old binary, its config and its data and
nothing the fork reads.
What it does end is the rollback, so the new section says when. The
rollback stops being one within hours anyway — after mail has flowed,
going back means losing what arrived since — so the real question is how
long to keep a cold copy of the pre-cutover state. The gate is the first
certificate renewal, which is the one check in "The first week" whose
failure would send anyone back; forcing a renewal closes it in a day
rather than ninety. Archive the store off-host before removing the
directory.
Step 2 stops stalwart.service and step 4 points the fork at a store path.
Between those two moments nothing protects the original, and the rehearsal
had already shown that one open by the fork costs the rollback for good.
What protects it until then turns out to be the running server itself:
RocksDB refuses a second opener with "While lock file: LOCK: Resource
temporarily unavailable". So the intuition that a service shutdown prevents
the mistake is backwards — the shutdown is what enables it.
Measured, in probe_guard.py, in the three states that matter: held by the
running server, the fork is refused and the rollback is intact; stopped but
read-only, the fork is refused while rotating its own log and the Enterprise
build still starts on it afterwards; stopped and writable, the fork opens,
adds its column family, and upstream never starts again.
So step 3 now makes the original read-only as soon as the copy is taken,
which turns a discipline problem into a one-line one, and the answered
section carries the table.
The runbook said nothing in it had been rehearsed. Now the sequence has
been, on data made up for the purpose: upstream 0.16.22 in a container as
the install running today, the fork beside it, both unprivileged with
CAP_NET_BIND_SERVICE. It rehearses the sequence, not the data, which is
what the compat tests are for. 27 of 27 checks passed and the rollback
took 1.5 seconds.
The question §"Open" asked about carrying a store back is answered, and
the answer is no. Upstream refuses to start on a store the fork has
opened: "Column families not opened: _". The fork adds one RocksDB column
family for masked email (SUBSPACE_INBUXA = b'_') and opens with
create_missing_column_families, so it creates it on first open; upstream
has no descriptor for it and RocksDB will not open a database holding one
it was not told about.
That makes step 3's "a copy, not a move" load-bearing in a way the step
did not say. One open by the fork is enough: pointing it at the original
even once, to check something, leaves the Enterprise install unable to
start, and there is no rollback after that. It fails loudly and before
reading anything, which is the good version of this failure, but it is
not recoverable.
Two things the rehearsal found that would have wasted time on the day:
memberTenantId does not come down from the domain and is refused on
create, so a tenant "admin" set up the obvious way is a server
administrator and the check passes while proving nothing; and IMAP's
INBOX is not JMAP's account, because mail from an unauthenticated sender
is filed as spam, so the two counts differ before and after alike.
What the rehearsal does not cover is in its README and in §"Open":
systemd and `systemctl disable stalwart` above all, ACME renewal, load,
the front ends, and INBUXA's own data.
Asked for by John, 2026-09-19. Public ihasmail is Stalwart-facing and knows
nothing about INBUXA, so running an unmodified one against the migrated
server checks something the fork's own suites cannot: that a client written
for upstream still works.
Each difference it finds is one of two things, and the point is to say
which: a regression against upstream's contract, which the fork's tests
would not catch because they test the fork; or a feature that now expects
INBUXA's own front ends, which belongs in the contract and the release
notes rather than in a user's surprise.
After mail is flowing, not as a gate. It informs the contract; it doesn't
block a cutover.
INBUXA's cutover is the first run of something other operators will want:
an existing Stalwart server becoming an INBUXA one with nothing re-entered
and nothing re-issued. Accounts, passwords, app passwords, OAuth sessions,
aliases, tenants, DNS records and provider settings, certificates and ACME
state, Sieve scripts, the queue and the mail all live in the store, so a
migration that copies the store carries them.
Offered beside the fresh-install workflow, which has different questions to
ask, so §6.1 now names both.
It rolls back, which is where it parts company with stalwart-migrator:
that one upgrades in place and says outright it cannot undo a migration.
This one never writes to what it migrates from, so going back is stopping
one service and starting another. Rollback is automatic when verification
fails, available on demand while the old install stands, honest about the
mail that stays behind, and never points the old server at the store the
fork has written.
Every phase is timed, and the number to advertise is the downtime, phases
2 to 7, not the total that preflight and the copy dominate. The report
writes both as JSON so a release note quotes something measured.
John's plan, 2026-09-19: stop the Enterprise server and the ihasmail
container, install the fork at its own path, copy the data across, bring up
INBUXA Admin and the new webmail, and every account carries on.
That is a better shape than the in-place swap this draft assumed, because
the rollback becomes a service swap rather than a restore: the old install
and its data are untouched, so going back is stopping one unit and starting
another. What it costs is whatever the fork accepted in between, since the
two stores diverge the moment the fork starts.
It buys one failure the in-place swap couldn't produce: both servers on one
set of ports, each with its own store, if a reboot brings the old unit back.
So the unit is disabled, not just stopped.
Also written down: the front ends' OAuth clients travel inside the store, so
a front end that keeps its client id and redirect URIs keeps working and one
deployed fresh needs them set up, which fails looking like an account
problem when it isn't.
/etc/ufw/user.rules is world-readable, so this needed no privilege after
all. The default input policy is DROP and 8899 is not among the allowed
ports, so pebble's connection to the suite's listener is dropped. That
matches the live run, where twelve probes from a container completed no
handshake while the same openssl reached pebble's own TLS port.
The nc readings that pointed the other way — instant refusals on closed
ports, where a DROP should hang — are still unexplained, and are left on
the page as unexplained rather than quietly dropped, since they are what
sent an earlier pass through this page in the wrong direction.
Nobody has added the allow rule and re-run the suite, so the fix is
written down as a prediction. The cutover draft says the same: the test
failing here is not evidence that renewal works on the host.
Steps 1 to 3 of SPEC §7 are met, so the remaining one is running the fork
as the mail server. The draft covers the sequence on the host, what to
check before letting mail flow, and what to watch in the first week.
Two things it refuses to gloss: rolling back stops being a snapshot restore
the moment the fork accepts a message, because nobody has tested whether
the Enterprise build reads a store the fork has written; and certificate
renewal is the failure that arrives 90 days late and quietly, on the one
path the suites couldn't settle.
Nothing in it has been rehearsed. The rehearsal on a copy is step 2 of
"Before the day", and it is what turns "the data opens" into "the server
runs on it".
Run from a copy of the stopped server's RocksDB store. Eight green lines,
which are worth reading carefully: six carry weight, and masked_email and
undelete carry none, because there are no masked addresses and retention is
off, so they iterate an empty list. tenant_compat checked the tenant, its
quotas and its members, but not what a tenant administrator can see, which
needs a --tenant-admin recording.
ai_compat is the one that might have looked vacuous and isn't: the twelve
LLM_ tags are there with the scores that were observed.
Three things had to be fixed first, each failing all eight identically and
none about the data: listener names, privileged ports, pending tasks. The
next import's copy will bring the same three, so they are written down.
SPEC §7: cutover steps 1 to 3 are met. Step 4 remains.
The run against INBUXA's store hung printing "Waiting for pending task
AcmeRenewal(...)": the copy carries that server's task queue, and a renewal
due in 2026-11 will not come due while a test watches it.
Under NO_INSERT the wait now skips tasks that aren't due and ones that have
permanently failed, which leaves the tasks the test itself caused — a
restore in undelete_compat comes due at once — and gives up after a minute
with the offending task printed. A test that was really waiting on its own
work now fails on its assertion, which says more than a spinner.
Ordinary runs are untouched: system_tests, which waits on tasks throughout,
still passes in 135s.
The first run against INBUXA's store failed all eight tests identically,
before reading a single record: the copy carries that server's listeners on
25, 443, 465, 587, 110, 143, 993 and 995, and nothing in a test run is
root, so each one failed with "Permission denied (os error 13)".
The builder now remembers the listeners it adds, and under NO_INSERT drops
build errors for any it didn't. Every other error still stands, including a
bind failing on one of its own, so this can't hide the case where the
harness's own port is taken.
The copy isn't edited for this: its listeners are simply not what a compat
run needs, and it reaches the server over the compat- ones instead. Checked
that a NO_INSERT run still boots and that scim_tests, which takes the
ordinary path, still passes.
Rehearsed the run against a real RocksDB store, and it died at startup
before checking anything: the harness inserts listeners of its own, the
registry keys them by name, and a real server already has a "jmap" and an
"imap". The message was "Primary key conflict on property name with
existing object NetworkListener", which says nothing about what to do.
Under NO_INSERT the harness now calls its listeners compat-jmap and so on,
and the same run gets through to the test's own checks.
run-compat.sh copies the store for each test and removes the copy after,
because several of these write to what they open: monitoring_compat purges
the history it reads and undelete_compat restores what it finds. The source
stays untouched, which matters when it is the only copy of a production
store anyone took that day.
INBUXA runs RocksDB, so a copy is a directory copy. The SQL backends would
need more than this: the harness builds its own container and connects to
fixed local credentials, so it cannot open a dump in place.
record-compat.py ran against the live server as a server-level
administrator: 8 accounts, which matches the dashboard, so it reached all
of them. One tenant with its quotas and members; no masked addresses, no
archived items.
The empty files are right rather than short. INBUXA has no masked
addresses and retention is off, so masked_email_compat and undelete_compat
iterate an empty list: they pass without comparing anything, which is worth
saying plainly, because a green run from either would otherwise read as
evidence of compatibility. That puts them where scim_compat and
per_domain_directory_compat already sit.
So one recording carries weight, expected.json, and it is made. SPEC §7's
cutover steps 2 and 3 come down to the tenant, the domains and the
accounts until either feature is switched on.
"You are not an owner of account X" is not a permission the server is
withholding; it is how far that identity can see. Answering it with "needs
sysMaskedEmailGet" sends you off to grant something that changes nothing.
A refusal that mentions ownership now says so, and says which account the
run wants: the administrator with the run of the server, with tenant
administrators passed as --tenant-admin.
A recording runs against a server that may not be up again soon, so an
administrator missing one permission shouldn't throw away the whole pass.
Each of the three is recorded on its own now: what the server allows is
written, what it refuses is named at the end with the permission it wants,
and the exit is still non-zero so an incomplete recording can't pass for a
finished one.
A refusal used to print the raw JMAP error. It now reads, for example,
"[email protected] may not x:Tenant/get: You are not authorized to perform
this action (needs sysTenantGet)".
Checked both ways: the refusal path against a stubbed client, where the
other two sections still record; the whole thing against a live test
server, which recorded 3 tenants and 53 masked addresses and exited 0.
A wrong tenant-administrator password failed only once the script had
already enumerated every account, on a server that may not be up for long.
All the identities are now checked against the session endpoint first, and
a refusal names each one that failed, with the two things that usually
explain it: basic authentication wants the account's name rather than its
email address, and an account with two-factor or OAuth-only sign-in needs
an app password. It also gives the curl line to test one on its own.
The unreachable message no longer suggests --insecure for a refused
connection; that hint is now only for a certificate it couldn't verify.
Three of the eight compat tests check INBUXA's data against a recording of
how the Enterprise server read it, and that recording can only be made
while that server is still up. SPEC §7 gives it 45 days from the notice, so
the capture shouldn't wait on the cutover being scheduled.
record-compat.py writes all three files: the tenants with their quotas and
members and what each tenant administrator sees, every masked address and
its state, and every archived item whole, since undelete_compat compares
every property it recorded. It only reads, and refuses to send a method
that isn't /get or /query, because it is the one tool here that runs
against the live server. Queries follow their pages, so a server that caps
one doesn't leave a short recording behind.
Exercised against the fork's own test server, which answers the same JMAP:
3 tenants with members, 8 masked addresses and 3 archived items, each in
the shape its test reads.
Probing further contradicted the previous two commits. During a live run,
when the suite is certainly listening on 8899, twelve handshakes from a
container completed nothing, with or without the ACME ALPN, while the same
openssl in the same container talks to pebble's TLS port and prints its
certificate. A listener that is up but unreachable from a container is what
a ufw DROP looks like. An instant refusal on a closed port, which nc saw
from two images, is not. Both were observed minutes apart.
So the ufw suspicion is neither confirmed nor dismissed, and the page now
says that rather than picking the reading that suits the last probe.
Settling it needs `sudo ufw status verbose` and a listener bound by hand,
neither of which this session could do.
What stands on pebble's own log, and does not depend on any of this: it
runs validations and marks the authorizations invalid, so "never validates,
stays pending" was wrong.
The previous commit called the bridge open on the strength of one nc run.
A later openssl s_client against the same closed port hung for its whole
timeout instead of reporting the refusal nc had just seen. Repeating the nc
test from a second image, with a control port and two closed ports, agreed
with the first: immediate refusal, which is not what a DROP looks like. The
openssl behavior is still unexplained, so the claim now carries what was
measured and the anomaly beside it.
The ALPN probe is recorded as proving nothing, for the same reason: it
hangs against a port with nothing behind it, so its silence during a
renewal says nothing about acme-tls/1.
Pebble's log is untouched by any of this: it validates and marks the
authorizations invalid, which is what makes the old explanation wrong.
The page blamed ufw for blocking the docker bridge, and said pebble never
validates, so the authorizations stay pending. All three are wrong.
From a container on the ACME network the host answers on both gateway
addresses: port 22 connects, and 8899 refuses at once with nothing
listening, where a DROP would hang. No firewall rule was read or changed to
establish that. Pebble's own log shows 20 validation attempts in the
regression run, five for each of the four tls.org names, each ending in
"INVALID by completed challenge". The challenges are answered and refused.
So the order goes invalid, no certificate is issued, and the test unwraps a
None. What's left to explain is the TLS-ALPN handshake. The responder is
intact and listen.rs picks it per connection from has_acme_tls_challenge,
which is computed when the network config is parsed, while the test adds
its provider after boot. That's written down as a hypothesis, not a
finding: it hasn't been tested.
None of the eight had ever executed, so all eight ran against an empty store
with synthetic inputs. The plumbing works: the documented JSON shapes parse,
and NO_INSERT stops each one before the harness touches the store, which a
sentinel file in each store directory confirmed — it survived every run,
including the one launched without NO_INSERT.
Two things the runbook got wrong, both of which would have cost a day on the
day the copy exists:
- TMPDIR is the copy's parent, not the copy. The harness opens
$TMPDIR/<test name>, so a TMPDIR pointing at the copy gets an empty store
created beside it and the test calls INBUXA's data missing.
- masked_email_compat and undelete_compat need INBUXA_COMPAT_MASKS and
INBUXA_COMPAT_ARCHIVED, which only the tests' doc comments mentioned.
Every run ended on a 401 raised as "Missing list in response", which reads
as INBUXA's data being wrong when the login is what's wrong. Each test now
authenticates once first and names the variable that failed.
Every one of the 14 was `#[cfg(not(feature = "enterprise"))]` on the arm the
fork always compiles: the Enterprise arms went with the import, and nothing
turns the feature on. Removing the attribute leaves the same code, now
unconditional, in 11 files.
Two of them looked like behavior worth checking before touching: the
`validate_tenant_quota` stub that always passes, and the refusal to cancel a
pending DestroyAccount task. The stub is vestigial — the rebuilt
multi-tenancy enforces quotas in `crates/features/src/tenancy/quota.rs` for
those objects and more — and the refusal is undelete's open question, which
this change leaves exactly as it was.
The binary builds with no new warnings, and `system_tests` and `jmap_tests`,
which cover the touched registry, task-manager and auth paths, both pass.
The feature definitions stay in the manifests, inert: taking them out would
widen every sync's diff for nothing.
The feature held the shared tests of features the fork hadn't rebuilt yet.
All nine are built and the last gate came off with per-domain directories,
so the feature was defined and documented but gated nothing.
The crate builds and holds the same tests without it: 113 by default, 118
with postgres, mysql and redis. SPEC §2.2b keeps its account of the first
import and now says when the gates came off, so a spec that still describes
a suite as gated reads as the record of its own date, which is what
monitoring.md's gate table already calls itself.
Ran it single-threaded on RocksDb: 87 passed, 3 failed, 23 ignored, all 113
tests a default build holds. lmtp_delivery passed this time, which is what
the note about queue timing under a sequential run predicted, so the four
documented failures are three. The ACME suite logged no 400 at all, so the
renewal fix holds; it still ends on the certificate that never came,
because pebble can't validate across the bridge with ufw up.
The line it replaces claimed 86 passed, 4 failed, 27 ignored from the same
command. That totals 117, and no feature set of this tree produces 117 —
113 by default, 118 with all three backends — with no test added or removed
since. Said so rather than presenting the two as a trend.
A few of upstream's dual-licensed files carry code from other projects
under MIT or BSD terms. The fork redistributes it, so their licenses
require the notices to travel with it. THIRD-PARTY.md reproduces them.
strip.py now reads the stripped tree's comments for another copyright
holder, another license, or a note that code came from somewhere else, and
names any file THIRD-PARTY.md doesn't cover. It reports, never fails: the
notice goes in with the merge that brings the release in.
On v0.16.22 it finds 14 files, all of them covered. The rest of the report
is byte-for-byte what the committed one says, so the scan disturbs nothing
it already did.
Re-ran all nine on containers removed beforehand, one suite at a time.
All nine pass. mysql_replica_position_tests failed the first time on test
18's lag assertion, with the pair half a minute old and still catching up,
and passed on a second run against the same containers; it fails before
the point where it changes the replica's settings, so it leaves nothing to
restore. Noted both, with each suite's time.
The table's STORE column said "default" for five suites, which reads as
"leave it unset". The harness has no default: it panics with "Missing or
invalid store type" before the suite starts. They run on RocksDb, as the
regression does, so the column now names it.
The renewal loop re-posted the challenge every time it polled and found the
authorization still pending. RFC 8555 section 7.5.1 has the client post a
challenge once to say it is ready and then poll; a server that has already
moved the challenge to "processing" refuses a second post, and pebble
answers 400 malformed, which failed the whole renewal.
Verified against pebble: the 400s are gone and the client polls. The suite
still can't finish on this machine, because pebble never reaches the test
server to validate the challenge; that path is the environment, and the
runbook now says so.
`cargo test -p tests` runs none of the fork's own feature suites except
what's inside `system_tests`: SCIM, per-domain directories, the sharded
stores and the four replica suites are all ignored, each needing a
container, a STORE the harness only builds on request, or both. The
runbook lists what each one needs, and the container-reuse trap.
`allowScimProvisioning` is cached with the domain as DOMAIN_FLAG_SCIM, but
it wasn't among the fields whose change drops the cached entry, so flipping
the flag changed nothing until something else evicted the domain. SCIM-60
says the change takes effect without a restart.
Found by the two acceptance checks that were written but never called from
the driver, so neither had ever run: `authority` (SCIM-58 to SCIM-60,
through `synchronize_account` itself) and `rate_limits` (SCIM-14). Both are
wired in now, and `scim_tests` passes with them.
Two long-standing flakes in system_tests, both timing:
- security: the ban period was set to one second, but the test makes a
hundred more requests before it checks that a valid password from the
banned address is refused, so under load the ban had already expired.
Five seconds, and the expiry check sleeps six.
- task: a task scheduled one second out was queried for straight away,
and under load the query landed after the task manager had run and
removed it. Three seconds.
Seven consecutive system_tests runs, neither recurred.
docs/spec/compat-tests.md lists each compat test, what it needs, what it
checks and what a failure means, and says plainly that they run against a
copy only, since the monitoring one purges the history it reads. The
scale-out and per-domain statuses record the acceptance tests that now
run.
A token the OIDC directory rejects is an authentication failure, so it
counts toward the ban; a network, provider or configuration fault stays
an error and doesn't. Before, a rejected token was an error too, so bad
tokens never led to a ban.
The Keycloak container now imports a second realm, so test 10 checks
/api/discover and the PACC record answer with each domain's own provider.
Test 18 checks that eight sign-ins during an outage don't ban the client,
while bad tokens do.
A FETCH with CHANGEDSINCE presents its mod-sequence to the read scope, so
a replica must have that change before it answers, as a JMAP sinceState
already did. replica_cluster_tests covers test 11: with the replica's
replay paused, a write on one node is read back through a second store
with its own marks, sharing through Redis.
Composite stores nest store futures deeply enough to pass rustc's default
query depth once both postgres and redis are compiled in, so the server
crates raise their recursion limit.
Two source-and-replica pairs run in containers: one replicating with
GTIDs, one by binary log position. mysql_replica_tests covers test 17
(tests 9, 10 and 12 with GTIDs) and mysql_replica_position_tests covers
test 18 (lag from Seconds_Behind_Source, and a replica whose account
lacks REPLICATION CLIENT getting no reads) and test 19 (a parallel
replica without replica_preserve_commit_order left out at startup).
The lag reader now takes Seconds_Behind_Source whatever numeric type the
server returns, accepts the older column name, and says in the log why it
gave up measuring.
Where the routing lives and what a read scope covers, which acceptance
tests the container suite runs, what isn't exercised (two nodes, the
primary stopped, and the MySQL paths, which are built but unrun), what
was settled from the code, and the known limits.
A data store with readReplicas becomes a replicated store. Writes,
operator-written SQL and everything outside a read scope go to the
primary. JMAP reads before a request's first write, IMAP LIST, STATUS,
SEARCH, SORT and FETCH, POP3 RETR and TOP, DAV GET, PROPFIND and REPORT,
and blob downloads run in a read scope. Only account data (properties,
indexes, change logs, counters, ACLs, blobs, the search index) is read
from a replica; the registry, in-memory values, the task queue and the
rest stay on the primary.
In a scope, the first read picks a replica round-robin among those up
and under the lag limit, and only if it has every change this node has
written or heard of for the scope's accounts: marks come from write
results, the cluster's state-change broadcasts, a sinceState the client
presents, and, with more than one node, Redis. A write inside the scope
sends the rest of it to the primary. A miss on a replica is looked up on
the primary, and a replica error retries the read there and marks the
replica down.
Each node samples lag every second (WAL positions on PostgreSQL; GTID
sets or Seconds_Behind_Source on MySQL), stops reading from a replica
over 5 s and starts again under 2.5 s, and probes a down replica every
10 s. At startup a replica is left out if it's the primary, isn't
read-only, applies out of commit order, or doesn't show a marker written
to the primary within six tries.
replica_tests (postgres, STORE=PostgreSqlReplicated) runs a primary and a
streaming hot standby in containers: tests 9, 10, 12, 13, 14 and 15 pass.
per_domain_directory_compat, ignored, checks a copy of INBUXA's data has
no directory, no server default and no domain with its own directory, as
observed. The status names where the rules live, which suite covers each
acceptance test and how far, what was settled from the code, and the
known limits. SCIM's status notes its test 5 now passes.
directory_tests now runs a new oidc module in place of the removed one,
with Keycloak as example.org's own directory: first sign-in creates the
account with its name and group, an existing account is reused, a token
named as another user is refused, password sign-in is refused, forged
JWTs (HS256, unknown kid, another issuer, expired) are refused, an OIDC
address is a recipient only once an administrator creates it, and sync
can't pass a tenant's account limit. scim_oidc_tests, deferred until
this feature, passes.
A write to x:Directory or Authentication reloads the directories at once,
here and across the cluster. A directory that fails to open is logged as
a warning against its id and becomes unavailable; before, it was a build
error, and any build error stopped every later reload from applying.
per_domain_directory_tests covers acceptance tests 1, 3, 4, 6, 7, 9, 11,
15 and 19 over SQL directories on SQLite files: each domain against its
own directory and the default, no fallback to internal passwords, a
directory answering for another directory's domain, aliases and groups
dropped, recipients through the directory and a 4xx while it's down, app
passwords while it's down, password changes refused and then allowed
after a move to the internal directory, linked directories, tenant
foreign keys, and changes taking effect without a reload.
The two lookups every caller uses now honor Domain.directoryId, then the
server default, then the internal directory, so sign-in, bearer routing,
recipient lookup, discovery, the PACC record and the refusal of password
changes on external accounts all follow the domain. A directoryId, or a
server default, naming a directory that doesn't exist is unavailable,
never the internal directory.
A directory speaks only for the domains it serves: an account it returns
on another directory's domain is refused, for sign-in and recipients
alike, and aliases and group claims on such domains are dropped with a
warning. A bearer token must belong to the user the client names, or the
name must be one of its aliases with alias sign-in allowed. Accounts and
groups that sync creates pass the tenant checks, limits included.
Where the code lives, which suite covers each acceptance test, what isn't
exercised, what was settled from the code, and that read-replica routing
(ST-5 to ST-15) waits for per-domain directories.
A Sharded blob store places each blob on xxh3(key) mod N, over the whole
key; reads fall back to the other members, so blobs placed under an
earlier member list stay readable, and deletes find them wherever they
are. The member list is recorded in the data store (secrets left out):
added or reordered members are a warning, a missing one refuses to open.
Blobs are compressed and marked before they reach a member. A Sharded
in-memory or lookup store sends each key to its home Redis member, and
prefix deletes and purges to all; a node whose member list differs from
the recorded one logs an error and runs on. Members are checked for
duplicates and must all open.
Until read-replica routing is built, each configured replica is reported
at startup instead of being silently ignored, and nothing connects to it.
The new scaleout_blob_tests covers tests 2 to 7, and the existing blob
suite passes against three FileSystem members (BLOB_STORE=Sharded).
Where each part lives, which suite covers each acceptance test, test 5
deferred to per-domain directories and test 31 unrun until a copy of
INBUXA's data, what was settled from the code, and the known limits.
A principal's key without unlimitedRequests gets 429 with Retry-After
over the limit, while its unlimited key still works. scim_compat, ignored,
checks a copy of INBUXA's data reads back as observed: no domain open to
SCIM and no externalId.
scim2-client builds its models from /Schemas, so meta is described there
as the mapping tables give it, without lastModified. The tester container
now shares the host's network, so a host firewall that drops the Docker
bridge doesn't block the test server. With SCIM_CONFORMANCE=1, the
scim2-client lifecycle (12 steps), the eight replayed Okta, Keycloak and
Entra payloads, and scim2-tester (errors only for its generated
non-address userName) all pass.
The push router records which account each subscription belongs to, and
a new Revoke event drops every subscription the account holds, so its
IMAP IDLE, JMAP event streams and WebSockets on this node end. Other
users' subscriptions to its shared mailboxes stay. Test 28 now checks an
open IDLE is ended, and that the account's password over HTTP and its
own API key, both already cached, are refused on the next request.
On such a domain, sign-in sync creates no account (the person gets an
ordinary authentication failure), changes nothing on an existing one,
and creates no group from a groups claim. Without the flag sync works as
before, and turning it off hands accounts back to sync with no restart.
Checked through synchronize_account itself; acceptance test 5 does the
same over OIDC once per-domain directories exist.
Every SCIM operation becomes the x:Account get, query or set JMAP makes,
as the service principal, so permissions, tenant scope and limits,
address uniqueness and account destruction are enforced in one place.
Discovery is anonymous; everything else takes an API key as a bearer
token and nothing else. Domains open to SCIM carry a flag in the domain
cache. Filters take eq and and, answered from the account indexes, with
unindexed attributes checked on at most 200 candidates. Cursors are
stateless, HMAC-sealed under the server key. PATCH applies to the
resource in memory and saves it as a PUT, so it is all or nothing.
Groups get an address from their display name on the principal's
domain; membership is written on each user.
Every write emits one of five new scim.* events (ids 637 to 641), also
added to the packaged schema. The helpers the surviving SCIM suites
import are rebuilt from the spec; scim_tests runs the new acceptance
suite and the surviving tenant isolation suite, and both pass.
The filter parser reads all of RFC 7644's grammar, so a server supporting
only eq and and can name the construct it refuses. PATCH paths take a
schema URN prefix, sub-attributes and value filters.
monitoring_compat, ignored, checks the observed settings against a copy
of INBUXA's data, reads the old history without an error, and purges it.
The status names where each part lives, which suite covers each
acceptance test, the tests not exercised, and the known limits.
Inside system_tests the quota suite leaves uploads with a 1-second
lifetime, so test 6's Sieve script upload could expire before its set
(BlobNotFound), intermittently. The suite now resets upload quota, count
and lifetime to their defaults first.
system::monitoring::monitoring_tests covers the defaults, metric
sampling, the Prometheus gauges, trace history over SMTP and LMTP
(probe sessions skipped, no raw I/O), trace destroy, indexTelemetry off,
live metrics and live tracing with their tokens and stream limit,
alerts by event and by email, and tenant admins refused. Also removes a
leftover enterprise-only attribute on the Prometheus counters loop,
which would have dropped counters from an enterprise-feature build.
Every minute the node that calculates metrics evaluates each enabled
x:Alert, read from the registry, with metric('name') or the underscore
name. An alert fires when its condition turns true: it emits
telemetry.alert-event with the rendered message, and queues one email per
recipient through the outbound queue, placeholders filled and
Auto-Submitted set. An alert naming an unknown metric is refused on save.
The shared alerts suite runs; every telemetry suite is un-gated.
GET /api/token/tracing and /api/token/metrics issue a 60-second token for
holders of liveTracing or liveMetrics outside a tenant. /api/live/tracing
streams event: trace frames of x:TraceEvents, never raw I/O, filtered by
text or by key, with a ping while idle; /api/live/metrics streams event:
metrics frames of current totals every interval. At most eight streams run
per node, each for 30 minutes. Upstream's documented paths are aliases.
A lossy collector subscriber keeps each inbound SMTP session that reached
MAIL FROM and each delivery attempt, info level and above and never raw
I/O, at most 1,000 events with strings cut at 4 KiB, and writes it when
the span closes as an x:Trace under the telemetry key class, scheduling
its indexing. The index task builds a document of event types, queue ids
and keywords when indexTelemetry is on. x:Trace/get derives timestamp,
from, to and size; /query filters by opening event, text, queueId and
time; /set destroys only. The data purge honours holdTracesFor. The shared
tracing and webhook suites run.
Every node writes a sample per metric on metricsCollectionInterval:
counters as the increase since its last sample, gauges always, histograms
as totals when changed. Samples are x:Metric in the registry's encoding
under the telemetry key class, ids time-ordered. x:Metric/get and /query
read them with metric and timestamp filters and full paging, hide what's
past holdMetricsFor, and the data purge deletes it. The is_enterprise
split is gone, so every gauge and histogram is collected and exported, and
queue.count is set from the queue itself. The shared metrics suite runs.
The calibration harness (tests/src/system/ai_calibration.rs, ignored) sends
what the classifier sends to a real local model and scores its answers with
the classifier's parser.
The classifier sends only the subject and text, between unforgeable markers
after the operator's prompt, to an OpenAI-compatible endpoint the operator
configured; nothing is preset. Its answer maps to an LLM_ tag whose score is
clamped (+5.0, -1.0 by default) and can never discard or reject on its own;
X-Spam-LLM is sanitized, encoded and folded, and a planted one is removed.
Failures, timeouts past the ceiling, a full slot or a paused model leave
mail flowing untagged. llm_prompt answers trusted scripts, and accounts
holding interactAi within an hourly limit. Redirects aren't followed and no
content or secret is logged. The limits live in inbuxa:AiLimits.
Acceptance tests 1 and 3 to 21; test 2 as the re-enabled shared llm case,
whose setup no longer waits on a rules file from a developer's own path;
test 22 written as the ignored ai_compat.
Expiry counts whole seconds, so a 1s access token, code or device code could
lapse before a debug build's next request: 4 of 8 baseline runs failed at
six different points. Lifetimes are now 5s (tokens, codes), 15s (refresh)
and 10s (renewal), with the expiry checks' sleeps scaled to match. OIDC then
passed in 12 of 12 runs.
Logos resolve domain, then tenant, then server-wide, then the built-in, with
subdomains finding their domain. GET /logo serves a data-URL image, redirects
to a URL logo without fetching it, sandboxes SVG, and answers 404 when no
custom logo applies. Emails embed the first PNG, JPEG or GIF logo. Logo and
template writes are checked; stored templates are read at send time, always
escaped, and fall back to the built-in with a build warning when they don't
parse. The RSVP page is served byte for byte with a CSP and no-referrer. The
sign-in and RSVP pages load the logo through an image element. MT-22's
session logo follows the chain to the server-wide logo.
Acceptance tests 1 to 17; test 18 written as the ignored branding_compat.
With archiveDeletedAccountsFor set, a destroyed account's record is kept in
the fork subspace with its id, its DestroyAccount task is due at the end of
the period, and its shares are suspended both ways. Its addresses can't be
taken by new accounts, aliases, lists or masks. inbuxa:DeletedAccount/get
lists kept accounts to server and tenant administrators; /set restores one
with a new password (same id, task cancelled, shares reinstated) or destroys
it now. The destroy task also clears undelete's own records.
Acceptance test 14; test 16 written as the ignored undelete_compat.
Files, events and contacts are noted when deleted for good and archived by
the unindex task when retention is on; Sieve scripts are archived at
deletion. A restore goes back to its folder, calendar or address book if it
still exists, takes a free " (restored)" name, comes back inactive for
scripts, and is refused over quota with the item left archived.
Acceptance tests 6, 8 and 12.
Every way of deleting mail for good (JMAP, IMAP expunge, POP3, Trash
emptying, mailbox removal) notes the message's mailboxes and keywords while
archiving is on, fixing its deadline then; when its data is finally removed
it becomes an x:ArchivedItem record, written as upstream writes them, with
its copy held until the deadline. Retention is read at deletion time, so a
change applies at once. Restore puts a message back in the mailboxes it was
in (Trash only if that was all), with its keywords, and removes the record;
over quota it stays archived. x:ArchivedItem/get returns status and
accountId; query filters on type, archivedAt and text; set requests a
restore once or destroys; /changes is a fork addition. Expired items go in
the data purge. The shared account-access rule moves to jmap::inbuxa::access.
system_tests now calls undelete::test, and the archiving gate is gone.
Found by running system_tests, which masked email no longer stops:
- rcpt_resolve rewrites a live mask to its owner's address, so
Delivered-To names the account; delivery recognizes the mask from the
original recipient when it belongs to that account.
- x:MaskedEmail/set create responses carry the server-set email.
- x:MaskedEmail/query returns every mask to a server-level impersonate
holder, and filters on accountId.
- The refusal for an unlinked emailDomain uses upstream's wording.
- The shared delivery test checks the fork's address format (ME-13).
- The masked email test's tenant domain uses manual DKIM, so its cleanup
leaves nothing behind.
tests/src/system/masked_email.rs runs from system_tests and alone as
masked_email_tests. Fastmail's MaskedEmail/set updates from the stored
object, so the registry's revision check holds.
Advertised as https://www.fastmail.com/dev/maskedemail in the session and on
every account that may hold masks. Masks created through it start pending
unless the create sets a state; pending can't be set again once left; state
and the other mutable fields map onto the same records the x: API uses.
The fork's per-account change log answers /changes, collapsing a mask
created and destroyed in the window. /changes on a registry type needs that
type's get permission.
A live mask accepts mail at RCPT TO and delivers to its owner, with an
X-Masked-Email header naming it. A disabled mask files straight to Trash,
past the owner's Sieve script. Deleted and expired masks are refused like
unknown addresses, without being cached as unknown. Mail moves lastMessageAt
and turns a pending mask enabled. Sub-addresses on a mask work.
x:MaskedEmail is no longer refused as unbuilt. Creates generate the address
on an allowed domain and check the prefix, maxMaskedAddresses and the create
rate; updates keep server-set fields; enabled reads and writes map to the
shared state; query filters on enabled, forDomain and text; a tenant
administrator reaches its tenant's accounts' masks.
Subspace _ holds the fork's own data, with its own SQL table and RocksDB
column family, and is part of backup. The masked_email module keeps each
mask's state, last mail and pending deadline beside upstream's record, an
index from address to mask with tombstones, and a per-account change log;
and generates addresses in the fork's format.
tests/src/system/tenant.rs is new, and system_tests calls it again without
the pending-rebuild gate. tenant_tests runs it alone; tenant_compat is test 15,
ignored until a copy of INBUXA's data is provided.
JMAP shareWith refuses a grantee outside the owner's tenant, or missing, with
invalidForeignKey naming the account. WebDAV ACL answers AllowedPrincipal and
IMAP SETACL answers as for an unknown account.
The session's own account carries urn:inbuxa:jmap with logo: its domain's
logo, else its tenant's, as stored. The session lists urn:inbuxa:jmap as a
server capability too.
Inside a tenant, server-level object types are forbidden and x:Tenant reads
return only the caller's own tenant, which it can't change. Registry writes
refuse links across tenant boundaries in both directions, give a new
principal its domain's tenant, move a domain's principals and DKIM keys with
it into a tenant, refuse moves out of a tenant while its people remain, and
refuse creates past a tenant's count limits with overQuota and
limit.tenant-quota.
New crate crates/features (inbuxa-features), AGPL-3.0-only, holding the
tenancy rules. Hooks in common: a tenant's roles and permission lists cap
its people's permissions; a change to a tenant, or to a role a tenant holds,
drops its members' cached permissions; delivery and every other write check
the tenant's maxDiskQuota; usedDiskQuota reads the tenant usage counter.
The default Tenant Administrator role gains sysTenantGet and sysTenantQuery.
Test 3 expects invalidForeignKey, as MT-3 says. A tenant admin reads its own
tenant through sysTenantGet and sysTenantQuery, added to the default Tenant
Administrator role for new installs only. The server adds nothing for MT-7's
dashboard list. The MT-19a submission warning is deferred. MT-22's logo is
the urn:inbuxa:jmap account capability's logo field (contract C-1).
Discovery and a contract version in the session; front ends configured once
(x:FrontEnds); OAuth with required registration, first-party clients,
server-hosted sign-in and consent for everything else; per-grant revocation;
cross-origin limited to the front ends; an admin lane by scope; push
unchanged. Records what upstream does today, including that it accepts any
client and redirect URI by default, and the phishing that allows.
inbuxa --version, the banner, startup events, OpenTelemetry and the JMAP
implementation string read "2026.9.18 (Stalwart 0.16.22)". The base comes
from Cargo, which keeps following upstream so version bumps merge cleanly.
Received headers and IMAP ID carry the INBUXA version.
Administration is no longer moving wholesale into ihasmail. INBUXA Admin, a
fork of the schema-driven webui (no Enterprise-only code, ordinary fork),
covers the whole server, setup and recovery from its own deployment.
ihasmail keeps its account and tenant administration. Contract additions:
the inbuxa-admin OAuth client, cross-origin access limited to the two
front ends (the server currently allows every origin), and relative OAuth
endpoints on servers with no public URL.
- The recovery administrator (INBUXA_RECOVERY_ADMIN, or STALWART_RECOVERY_ADMIN)
is honored only in bootstrap and recovery mode. On a configured server it's
ignored with a startup warning. Before, it was a standing full-admin login
for as long as the variable stayed set.
- The Enterprise upsell error is replaced by "This feature isn't available in
INBUXA yet" for the features still to be rebuilt.
- Workspace warnings: 25 to 0. cargo fix removed the unused imports. The
seven places where Enterprise code used to plug in keep their parameters,
each with an inbuxa: comment naming the rebuild that uses it again. The
antispam test's mock-server imports are back behind pending-rebuild.
- crates/main: package and [[bin]] renamed to inbuxa; homepage inbuxa.org;
license AGPL-3.0-only (upstream is dual; the fork takes the AGPL).
- types::branding::env_var reads INBUXA_<name>, falling back to
STALWART_<name> with a warning, for all nine server settings.
STALWART_APP_ and STALWART_SPAM_* storage keys are unchanged.
- New-install default paths /var/lib/inbuxa and /var/log/inbuxa.
- Dockerfiles, systemd unit, launchd plist and AppArmor profile renamed.
- Upstream's .github moved to .github-upstream so none of it runs.
- install.sh stubbed: upstream's would install Stalwart.
- Two missed brand strings: the SMTP Received header and the utils user agent.
The early return for a missing redirect_uri came before the logo loader,
so the logo stayed hidden behind data-loading. An upstream bug, surfaced by
the rebrand check.
Upstream's Dockerfiles and CI pass --features "... enterprise" on the cargo
command line, which the manifest edits don't reach. 11 occurrences at
v0.16.22. Dockerfile.build and Dockerfile.fdb are now scanned too, whatever
their extension.
One branding module (types::brand!) supplies the name to every protocol
greeting, the HTTP realm, the startup banner, Received headers, the
user agent, the IMAP ID response and calendar PRODIDs. The sign-in and
calendar pages take ihasmail's palette and font and the INBUXA logo, and
calendar emails embed the INBUXA lockup. Default calendar and address book
names follow. Protocol identifiers (urn:stalwart:jmap, vnd.stalwart Sieve
extensions) and upstream copyright notices are unchanged.
Thank you for your interest in contributing to Stalwart. We appreciate the support and enthusiasm of the open-source community. To keep the project maintainable and the review process sustainable, contributions are subject to the policies described below. Please read them in full before opening a pull request.
Patches, bug reports and questions are welcome.
## Vouched Contributors Only
## Before a pull request
Due to the high volume of low-quality, AI-generated submissions, pull requests are limited to a list of vouched contributors. Pull requests opened by anyone who is not on this list are closed automatically.
**Open an issue first for anything substantial.** A feature or a refactor is
worth agreeing on before it is written, because this is a fork that tracks
upstream: a change that moves code around costs a conflict on every import,
and it should be worth that.
To be added as a vouched contributor, post a message at [support.stalw.art](https://support.stalw.art) explaining the code changes you would like to submit, and include a link to the proposed change (a branch, diff, or draft). Once a maintainer has reviewed your request and vouched for you, you will be able to open pull requests directly.
Small fixes — a bug, a typo, a test — need no ceremony. Send them.
This policy lets us focus limited review capacity on contributions from people who have taken the time to understand the codebase and discuss their changes first.
## How a change lands
## What Contributions Are Accepted
`main` is protected. It cannot be force-pushed or deleted, and a change
reaches it through a pull request whose `build` check has passed. No approving
review is required — this is a small project and a gate nobody can pass is not
a gate — but the build is not optional.
At this stage of the project we accept a narrow set of contributions:
So the shape of a change is: a branch, a pull request, a green CI run, a merge.
Branches are deleted on merge. Repository administrators can bypass the rule,
which exists so the maintainer can correct the tree quickly, not so that the
ordinary path can be skipped; use it for an emergency, not for convenience.
- **Bug fixes.** Corrections to existing, incorrect behavior are welcome. Please include steps to reproduce the bug and describe the fix.
- **Translations.** Additions and corrections to existing translations are welcome.
Releases are cut weekly from `main` by `.github/workflows/release.yml`, on
Monday morning UTC, and nothing is released on a quiet week. That is the reason
the rule matters: whatever is on `main` when the run starts is what ships, so
`main` is expected to be releasable at all times rather than at the end of a
piece of work. A change that is not finished should be behind something that
defaults to off, or it should not be on `main` yet.
New features are generally **not** accepted, unless they involve only a few lines of code. Larger features fall outside the scope of what we can review and integrate while the architecture is still evolving.
## What this repository is
If you would like to see a new feature, please request it at [support.stalw.art](https://support.stalw.art) under the **Feature Ideas** category rather than opening a pull request. This lets the community discuss and prioritize ideas before any code is written.
INBUXA is a fork of Stalwart, taken under the AGPL-3.0-only half of its dual
licence, with nine features rebuilt independently. Two things follow:
## No AI-Generated Code
- **The clean room is real.** The rebuilt features in `crates/features` were
written from specifications in `docs/spec/features/`, by people who had not
read Stalwart's Enterprise source. If you have read it, say so in the pull
request and it will be reviewed with that in mind, or declined for the parts
it touches. Nothing about this is personal: the project's defence of
independent creation is a record, and the record has to be true.
- **Upstream files stay recognisable.** Changes to files that came from
upstream are kept small and marked with an `inbuxa:` comment saying which
requirement they serve, so the next import merges cleanly and a reader can
tell fork from base. New work belongs in the fork's own crates where it can.
AI-generated code is not accepted in this project.
## Licence and provenance
Even the most advanced models write inefficient Rust code. Beyond raw performance, AI creates technical debt by generating large amounts of code that not even the authors who submitted it can fully understand or maintain. Reviewing and untangling such contributions costs the maintainers far more time than it saves.
Contributions are under AGPL-3.0-only. Keep upstream's copyright headers where
they are; if you change a file that came from upstream, leave its "Modified by
Coffey Labs" line in place. New files carry:
Using AI as a fancy autocomplete is perfectly fine. What matters is that every line generated by a model is read, understood, and reviewed by a human before it is submitted. You are responsible for every line in your pull request, regardless of how it was produced. If you cannot explain why a change is written the way it is, it is not ready to be submitted.
```
/*
* SPDX-FileCopyrightText: 2026 Coffey Labs
*
* SPDX-License-Identifier: AGPL-3.0-only
*/
```
## Pull Request Process
If you bring in code from another project, it stays under its own licence and
its notice goes in `THIRD-PARTY.md`. `tools/fork/strip.py` reports any file
that is missing from there on every import.
Once you are a vouched contributor:
## Running the tests
1. Keep each pull request small and focused on a single logical change.
2. Match the style and conventions of the surrounding code.
3. Make sure the project builds and the test suite passes before opening the pull request.
4. In the pull request description, explain what the change does and why, and link to the [support.stalw.art](https://support.stalw.art) discussion where the change was vouched.
`cargo test -p tests` runs what needs nothing but a store on disk. The rest
need containers, a particular backend, or a copy of real data, and are
`#[ignore]`d:
## Code of Conduct
-`docs/spec/container-tests.md` — the suites that need containers, with the
`STORE` each one wants and what a plain regression leaves failing.
-`docs/spec/compat-tests.md` — the compatibility set, which needs a copy of a
real server's data.
We as members, contributors, and leaders pledge to make participation in our community a harassment-free experience for everyone, regardless of age, body size, visible or invisible disability, ethnicity, sex characteristics, gender identity and expression, level of experience, education, socio-economic status, nationality, personal appearance, race, religion, or sexual identity and orientation. We pledge to act and interact in ways that contribute to an open, welcoming, diverse, inclusive, and healthy community.
Run one suite at a time. They bind fixed ports, and the timing checks flake if
two run at once.
You can read the full Code of Conduct [here](https://github.com/stalwartlabs/.github/blob/main/CODE_OF_CONDUCT.md).
## Commit messages
## Licensing
This project is licensed under the Affero General Public License (AGPL) version 3.0. By contributing to this project, you agree that your contributions will be licensed under the AGPL-3.0 license.
## Fiduciary Contributor License Agreement
Before making any contributions, all contributors are required to sign the Fiduciary Contributor License Agreement (FLA). The FLA is a legal agreement that assigns the copyright of contributions to a designated fiduciary, who manages these rights on behalf of the project. This arrangement ensures that the software remains free and open, even as contributors come and go.
Key points of the FLA:
- Ensures the software remains free and open source
- Protects the project from potential copyright issues
- Includes a reversion clause: if the fiduciary violates Free Software principles, rights revert to the original contributors
For more details about the FLA, please refer to the [FLA FAQ](https://fsfe.org/activities/fla/fla.en.html).
Say what changed and why, in prose, wrapped at 72 characters or so. The why is
the part that is hard to recover later. No tool trailers.
FROM--platform=$BUILDPLATFORM docker.io/lukemathwalker/cargo-chef:latest-rust-slim-trixie AS chef
FROM--platform=$BUILDPLATFORM docker.io/lukemathwalker/cargo-chef:latest-rust-slim-trixie@sha256:38dfdbf4fda95c516f873f33032e490baa988b75f7d83c7d12f788f770785b36 AS chef
WORKDIR/build
FROM--platform=$BUILDPLATFORM chef AS planner
@@ -19,9 +19,13 @@ RUN export DEBIAN_FRONTEND=noninteractive && \
g++-x86-64-linux-gnu binutils-x86-64-linux-gnu
RUN rustup target add "$(cat /target.txt)"
COPY --from=planner /recipe.json /recipe.json
# inbuxa: [patch.crates-io] points sieve-rs at vendor/, and the recipe only
# carries the workspace's own manifests, so cooking the dependencies needs the
# vendored crate itself (the context allows it since #27; this puts it here).
> Development happens on [git.coffeylabs.org/inbuxa/inbuxa-server](https://git.coffeylabs.org/inbuxa/inbuxa-server); the copy on GitHub is a read-only mirror.
> Report issues at **[git.coffeylabs.org/inbuxa/inbuxa-server/issues](https://git.coffeylabs.org/inbuxa/inbuxa-server/issues)**, and join discussions at **[community.coffeylabs.org](https://community.coffeylabs.org)**.
## Features
**inbuxa** is a mail and collaboration server: JMAP, IMAP, POP3, SMTP,
CalDAV, CardDAV and WebDAV, in one Rust binary, with ihasmail as its web front
end. It is a fork of [Stalwart](https://github.com/stalwartlabs/stalwart).
**Stalwart** is an open-source mail & collaboration server with JMAP, IMAP4, POP3, SMTP, CalDAV, CardDAV and WebDAV support and a wide range of modern features. It is written in Rust and designed to be secure, fast, robust and scalable.
Stalwart ships some features only in a paid Enterprise Edition: multi-tenancy,
masked email, undelete and others. **inbuxa** ships everything to everybody under
the AGPL-3.0, rebuilding those features independently and without using any
of Stalwart's Enterprise code.
Key features:
## What's different from Stalwart
- **Email** server with complete protocol support:
- JMAP:
* [JMAP for Mail](https://datatracker.ietf.org/doc/html/rfc8621) server.
* [JMAP for Sieve Scripts](https://www.ietf.org/archive/id/draft-ietf-jmap-sieve-22.html).
* [WebSocket](https://datatracker.ietf.org/doc/html/rfc8887), [Blob Management](https://www.rfc-editor.org/rfc/rfc9404.html) and [Quotas](https://www.rfc-editor.org/rfc/rfc9425.html) extensions.
- IMAP:
* [IMAP4rev2](https://datatracker.ietf.org/doc/html/rfc9051) and [IMAP4rev1](https://datatracker.ietf.org/doc/html/rfc3501) server.
- [STLS](https://datatracker.ietf.org/doc/html/rfc2595) and [SASL](https://datatracker.ietf.org/doc/html/rfc5034) support as well as other [extensions](https://datatracker.ietf.org/doc/html/rfc2449).
- SMTP:
* SMTP server with built-in [DMARC](https://datatracker.ietf.org/doc/html/rfc7489), [DKIMv2](https://datatracker.ietf.org/doc/draft-ietf-dkim-dkim2-spec/), [DKIMv1](https://datatracker.ietf.org/doc/html/rfc6376), [SPF](https://datatracker.ietf.org/doc/html/rfc7208) and [ARC](https://datatracker.ietf.org/doc/html/rfc8617) support for message authentication.
* Strong transport security through [DANE](https://datatracker.ietf.org/doc/html/rfc6698), [MTA-STS](https://datatracker.ietf.org/doc/html/rfc8461) and [SMTP TLS](https://datatracker.ietf.org/doc/html/rfc8460) reporting.
* Automated DKIM key rotation and management.
* Inbound throttling and filtering with granular configuration rules, sieve scripting, MTA hooks and milter integration.
* Distributed virtual queues with delayed delivery, priority delivery, quotas, routing rules and throttling support.
* Envelope rewriting and message modification.
- **Collaboration** server:
- Calendaring and scheduling:
- [CalDAV](https://datatracker.ietf.org/doc/html/rfc4791) and [CalDAV Scheduling](https://datatracker.ietf.org/doc/html/rfc6638) support.
- [JMAP for Calendars](https://datatracker.ietf.org/doc/html/draft-ietf-jmap-calendars-24) support.
- Comprehensive set of filtering **rules** on par with popular solutions.
- LLM-driven spam filtering and message analysis.
- Statistical **spam classifier** with collaborative filtering, automatic training capabilities and address book integration.
- DNS Blocklists (**DNSBLs**) checking of IP addresses, domains, and hashes.
- Collaborative digest-based spam filtering with **Pyzor**.
- **Phishing** protection against homographic URL attacks, sender spoofing and other techniques.
- Trusted **reply** tracking to recognize and prioritize genuine e-mail replies.
- Sender **reputation** monitoring by IP address, ASN, domain and email address.
- **Greylisting** to temporarily defer unknown senders.
- **Spam traps** to set up decoy email addresses that catch and analyze spam.
- **Flexible**:
- Pluggable storage backends with **RocksDB**, **FoundationDB**, **PostgreSQL**, **mySQL**, **SQLite**, **S3-Compatible**, **Azure** and **Redis** support.
- Full-text search available in 17 languages using the built-in search engine or via **Meilisearch**, **ElasticSearch**, **OpenSearch**, **PostgreSQL** or **mySQL** backends.
- Sieve scripting language with support for all [registered extensions](https://www.iana.org/assignments/sieve-extensions/sieve-extensions.xhtml).
- Email aliases, mailing lists, subaddressing and catch-all addresses support.
- Automated DNS management.
- Automatic account configuration and discovery with [PACC](https://datatracker.ietf.org/doc/draft-ietf-mailmaint-pacc/), [autoconfig](https://datatracker.ietf.org/doc/draft-ietf-mailmaint-autoconfig/) and [autodiscover](https://learn.microsoft.com/en-us/exchange/architecture/client-access/autodiscover?view=exchserver-2019).
- Multi-tenancy support with domain and tenant isolation.
- Disk quotas per user and tenant.
- **Secure and robust**:
- Encryption at rest with **S/MIME** or **OpenPGP**.
- Automatic TLS certificate provisioning with [ACME](https://datatracker.ietf.org/doc/html/rfc8555) using `TLS-ALPN-01`, `DNS-01`, `DNS-PERSIST-01` or `HTTP-01` challenges.
- Automated blocking of IP addresses that attack, abuse or scan the server for exploits.
- Rate limiting.
- Security audited (read the [report](https://stalw.art/blog/security-audit)).
- Memory safe (thanks to Rust).
- **Scalable and fault-tolerant**:
- Designed to handle growth seamlessly, from small setups to large-scale deployments of thousands of nodes.
- Built with **fault tolerance** and **high availability** in mind, recovers from hardware or software failures with minimal operational impact.
- Peer-to-peer cluster coordination or with **Kafka**, **Redpanda**, **NATS** or **Redis**.
- **Kubernetes**, **Apache Mesos** and **Docker Swarm** support for automated scaling and container orchestration.
- Read replicas, sharded blob storage and in-memory data stores for high performance and low latency.
- **Authentication and Authorization**:
- **OpenID Connect** authentication.
- OAuth 2.0 authorization with [authorization code](https://www.rfc-editor.org/rfc/rfc8628) and [device authorization](https://www.rfc-editor.org/rfc/rfc8628) flows.
- **LDAP**, **OIDC**, **SQL** or built-in authentication backend support.
- System for Cross-domain Identity Management ([SCIM](https://www.rfc-editor.org/info/rfc7643/)) v2 for automated provisioning.
- Two-factor authentication with Time-based One-Time Passwords (`2FA-TOTP`)
- Application passwords (App Passwords).
- Roles and permissions.
- Access Control Lists (ACLs).
- **Observability**:
- Logging and tracing with **OpenTelemetry**, journald, log files and console support.
- Metrics with **OpenTelemetry** and **Prometheus** integration.
- Webhooks for event-driven automation.
- Alerts with email and webhook notifications.
- Live tracing and metrics.
- **Web-based administration**:
- Dashboard with real-time statistics and monitoring.
- Account, domain, group and mailing list management.
- SMTP queue management for messages and outbound DMARC and TLS reports.
- Report visualization interface for received DMARC, TLS-RPT and Failure (ARF) reports.
- Configuration of every aspect of the mail server.
- Log viewer with search and filtering capabilities.
- Self-service portal for password reset and encryption-at-rest key management.
- **Every feature, one edition.** No license key, no edition checks, no
upsell. See `docs/spec/SPEC.md` §4 for the features being rebuilt, and
`docs/spec/features/` for each one's specification.
- **Webmail and administration by ihasmail,** as a separate service that can
run beside the server or elsewhere. Stalwart's own web interface is removed,
so there's no web front end on the mail host.
- **Clean-room rebuilds.** Enterprise-only code is stripped from every
upstream release before it's imported. The rebuilt features are written
from specifications that use only public sources (`docs/spec/SPEC.md` §3).
## Screenshots
## How the fork is kept
<img src="./img/demo.gif">
Upstream releases arrive as stripped snapshots, never with upstream's git
history, which contains Enterprise code. `tools/fork/strip.py` builds each
snapshot on top of upstream's own `ossify.py`, then verifies it independently.
The report for every import is in `docs/fork/strip-reports/`. See
`docs/spec/SPEC.md` §2.
## Presentation
## Building
**Want a deeper dive?** Need to explain to your boss why Stalwart is the perfect fit? Whether you're evaluating options, making a case to your team, or simply curious about how it all works under the hood, these slides walk you through the key features, architecture, and benefits of Stalwart. Browse the [slides](https://stalw.art/slides) to see what makes it stand out.
```bash
cargo build --release -p inbuxa # the binary is target/release/inbuxa
docker build -t inbuxa . # or the container image
```
## Get Started
Settings are read from `INBUXA_*` environment variables. An existing Stalwart
install's `STALWART_*` variables aren't read: the server stops at startup and
names each one to rename.
New installs keep their data in `/var/lib/inbuxa` and logs in
`/var/log/inbuxa`. Existing installs keep the paths their configuration
already names, so none of their data moves.
Install Stalwart on your server by following the instructions for your platform:
Coffey Labs in 2026**. Upstream's copyright notices are kept on every file
they cover, and every upstream file this fork changed says so in its header,
under the notice it came with. Stalwart's files are dual-licensed
AGPL-3.0-only or Stalwart's Enterprise License, and **inbuxa** takes them under
the AGPL-3.0 only. A few of those files also carry code from other projects
under MIT or BSD licenses, which stays under those licenses;
[THIRD-PARTY.md](./THIRD-PARTY.md) lists it with its notices. "Stalwart" is
Stalwart Labs' name. **inbuxa** isn't affiliated with or endorsed by Stalwart
Labs.
## Support
If you are having problems running Stalwart, found a bug, or just have a question, please head to the [Stalwart Support Portal](https://support.stalw.art) at [support.stalw.art](https://support.stalw.art).
Additionally, you may purchase an [Enterprise License](https://stalw.art/enterprise) to obtain priority support from Stalwart Labs LLC, including response-time commitments and a private Priority Support area on the portal.
## Contributing
We welcome contributions, but to keep the project maintainable there are a few things to know before opening a pull request. Because of the high volume of low-quality, AI-generated submissions, pull requests are limited to a list of vouched contributors; to be added, post at [support.stalw.art](https://support.stalw.art) describing the change you would like to submit, together with a link to the proposed change. At this stage only bug fixes and translations are accepted, and new features are not, unless they involve just a few lines of code.
For the full guidelines, please read [CONTRIBUTING.md](CONTRIBUTING.md).
## Roadmap
Stalwart has reached an exciting point in its journey, it’s now **feature complete**. All the core functionality and open standard email and collaboration protocols that we set out to support are in place. In other words, Stalwart already does everything you’d expect from a modern, standards-compliant mail and collaboration platform.
The next major milestone is all about refinement: finalizing the database schema and focusing on performance optimizations to ensure everything runs as efficiently and reliably as possible. Once that’s done, we’ll be ready to roll out version **1.0**.
Of course, development doesn’t stop there. The community has contributed hundreds of great ideas for improvements and new features, everything from subtle usability tweaks to entirely new integrations. You can see the full list of proposals over on our [GitHub issues](https://github.com/stalwartlabs/stalwart/issues?q=is%3Aissue+is%3Aopen+sort%3Areactions-%2B1-desc+label%3Aenhancement). If there’s something you’d like to see prioritized, just give it a thumbs up as we plan to implement enhancements based on the community’s votes.
## Sponsorship
Your support is crucial in helping us continue to improve the project, add new features, and maintain the highest level of quality. By [becoming a sponsor](https://opencollective.com/stalwart), you help fund the development and future of Stalwart. As a thank-you, sponsors who contribute $5 per month or more will automatically receive a [Enterprise edition](https://stalw.art/enterprise/) license. And, sponsors who contribute $30 per month or more, also have access to [Premium Support](https://stalw.art/support) from Stalwart Labs.
## Funding
Part of the development of this project was funded through:
- [NGI0 Entrust Fund](https://nlnet.nl/entrust), a fund established by [NLnet](https://nlnet.nl/) with financial support from the European Commission's [Next Generation Internet](https://ngi.eu/) programme, under the aegis of DG Communications Networks, Content and Technology under grant agreement No 101069594.
- [NGI Zero Core](https://nlnet.nl/NGI0/), a fund established by [NLnet](https://nlnet.nl/) with financial support from the European Commission's programme, under the aegis of DG Communications Networks, Content and Technology under grant agreement No 101092990.
If you find the project useful you can help by [becoming a sponsor](https://opencollective.com/stalwart). Thank you!
## License
This project is dual-licensed under the **GNU Affero General Public License v3.0** (AGPL-3.0; as published by the Free Software Foundation) and the **Stalwart Enterprise License v2 (SELv2)**:
- The [GNU Affero General Public License v3.0](./LICENSES/AGPL-3.0-only.txt) is a free software license that ensures your freedom to use, modify, and distribute the software, with the condition that any modified versions of the software must also be distributed under the same license.
- The [Stalwart Enterprise License v2 (SELv2)](./LICENSES/LicenseRef-SEL.txt) is a proprietary license designed for commercial use. It offers additional features and greater flexibility for businesses that do not wish to comply with the AGPL-3.0 license requirements.
Each file in this project contains a license notice at the top, indicating the applicable license(s). The license notice follows the [REUSE guidelines](https://reuse.software/) to ensure clarity and consistency. The full text of each license is available in the [LICENSES](./LICENSES/) directory.
## Copyright
Copyright (C) 2020, Stalwart Labs LLC
The **inbuxa** mark reuses ihasmail's cat-and-envelope artwork.
We provide security updates for the following versions of Stalwart:
INBUXA is developed on `main`, and security fixes are applied there and in
the latest release. Older tags are not backported.
| Version | Supported | End of Support |
| ------- | ------------------ | -------------- |
| 0.16.x | :white_check_mark: | TBD |
| 0.15.x | :white_check_mark: | 2026-12-01 |
| < 0.14 | :x: | Ended |
| Version | Supported |
| --- | --- |
| `main` and the latest release | :white_check_mark: |
| Older releases | :x: |
**Note**: We typically support the current major version and one previous major version. Users are strongly encouraged to upgrade to the latest version for the best security posture.
## Reporting a vulnerability
## Reporting a Vulnerability
**Please don't open a public issue for a security problem.** An issue is
visible to everyone, including whoever would use it, before there is a fix.
We take the security of Stalwart very seriously. If you believe you've found a security vulnerability, we encourage you to inform us responsibly through coordinated disclosure.
Report it privately by email to:
### How to Report
**johnellisATlinuxDOTcom**
**Do not report security vulnerabilities through public GitHub issues, discussions, or social media.**
Include as much as you can of:
Instead, please use one of these secure channels:
- what the vulnerability is, and what it lets someone do;
- how to reproduce it, or a proof of concept;
- the version or commit affected;
- anything about the deployment that matters — backend, front ends, whether
| `crates/nlp/src/tokenizers/types.rs` | test cases from [linkify](https://github.com/robinst/linkify) (MIT or Apache-2.0) | Copyright (c) 2017 Robin Stocker |
| `resources/spam-filter/spam-filter-rules.json.gz` | the published rules of [spam-filter](https://github.com/stalwartlabs/spam-filter) v3.0.2, unmodified, built into the server as its default spam rules (MIT or Apache-2.0) | Copyright (C) 2024, Stalwart Labs LLC |
Each notice above applies with this permission notice:
> Permission is hereby granted, free of charge, to any person obtaining a copy
> of this software and associated documentation files (the "Software"), to deal
> in the Software without restriction, including without limitation the rights
> copies of the Software, and to permit persons to whom the Software is
> furnished to do so, subject to the following conditions:
>
> The above copyright notice and this permission notice shall be included in
> all copies or substantial portions of the Software.
>
> THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
> IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
> FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
> AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
> LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
> OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
> SOFTWARE.
## Under the BSD 3-Clause license
| Where | From | Notice |
|---|---|---|
| `crates/jmap-proto/src/types/date.rs`, `crates/registry/src/types/datetime.rs` | [upb](https://github.com/protocolbuffers/upb/blob/22182e6e/upb/json_decode.c), the date parsing marked in each file | Copyright (c) 2009-2011, Google Inc. All rights reserved. |
```text
Copyright (c) 2009-2011, Google Inc.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
* Redistributions of source code must retain the above copyright
notice, this list of conditions and the following disclaimer.
* Redistributions in binary form must reproduce the above copyright
notice, this list of conditions and the following disclaimer in the
documentation and/or other materials provided with the distribution.
* Neither the name of Google Inc. nor the names of any other
contributors may be used to endorse or promote products
derived from this software without specific prior written permission.
THIS SOFTWARE IS PROVIDED BY GOOGLE INC. ``AS IS'' AND ANY EXPRESS OR IMPLIED
WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF
MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE DISCLAIMED. IN NO
EVENT SHALL GOOGLE INC. BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL,
SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO,
PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR
BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER
IN CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE)
ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE
POSSIBILITY OF SUCH DAMAGE.
```
## Credited algorithms
These files implement published algorithms and credit their source. No code
is copied, so there's no notice to carry. They're listed so the strip report
pubmodstore;// inbuxa: monitoring history (MON-10 to MON-17)
useregistry::{
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.