jcoffey-dev is traveling from Thursday 1 October through Sunday 4 October. Issues and pull requests are welcome, and will get an answer after that. Thanks for your patience.
Phase 1: one native rule engine at DATA, after the system Sieve script,
for both DLP policies and transport rules. DLP checks outgoing mail
with counted detectors (payment cards, IBAN, US SSN, word lists,
patterns) and blocks, warns with an audited override, or holds for
review. Held mail stays in the queue unscheduled, with its own review
record, so the queue's stored format is unchanged. Matches go to the
audit log without the matched text. Six questions for John at the end.
Personal-data catalog spec, default D5 (settled 2026-09-28; built after
the v0.16.24 import's spam-rules loader landed). msbl.org's EBL is sent
a SHA-1 of every email address it's asked about. A new install's first
boot now leaves a note, and the rules update, once the bundled rules
are in, switches STWT_MSBL_EBL_EMAIL off and forgets the note, so it
happens once; the loader keeps that switch through later updates. An
existing server has no note and keeps every blocklist as it is.
Also fixes the data inventory's DNSBL endpoints: a zone is an
expression (`ip_reverse + '.zen.spamhaus.org'`, conditional branches,
`hash(email, 'sha1') + '.ebl.msbl.org'`), and the zone names are now
the quoted literals that start with a dot, from every branch, rather
than the expression's text.
Tested: unit test for the zone rule; the compliance system test (no
note, no change; the inventory lists ebl.msbl.org, not a hash; with the
note the blocklist goes off; the note works once); the system suite;
fork checks.
Personal-data catalog spec, §6 (Phase 3c).
inbuxa:DataInventory/get evaluates the catalog against the server's
live settings and says what this server holds: for each source and each
object that can hold personal data, its categories and whose data it
is, whether it is collected here at all, what bounds its retention (the
live value of the setting that does, or unbounded), whether it leaves
the host and to which endpoints, and a summary. Every host that
receives something is listed once as a candidate processor with what it
receives. Inside a tenant it answers with the tenant's slice and none
of the server's processors. Read-only, with sysComplianceGet.
inbuxa:InventorySnapshot/get is the history: a dated copy of the
evaluated inventory, recorded when it changes -- after a registry write
to an object the inventory reads, after inbuxa's log, audit or AI
settings change, and on the daily clean-up -- and kept as long as the
audit log's records. ids: null lists every snapshot, newest first; the
full inventory only when asked for.
The catalog is embedded and parsed at start (new dependency: toml,
MIT/Apache); the evaluation is a pure function of it and the live
facts, so each configuration is tested without a server. Loopback
endpoints stay on the host; any other configured endpoint leaves it.
Tested: unit tests for the evaluation (a new install's defaults, an
external blob store, a hosted AI endpoint, telemetry off, a tenant's
slice, hosts from URLs, loopback), snapshots, and the fact gathering's
store and duration rules; the compliance system test, extended (the
officer reads the inventory, a plain user is refused, a tenant's
officer sees its slice and no processors, a webhook to another host
becomes a processor and a snapshot names x:WebHook, a retention change
reads through); the system, audit, legal hold and account lock suites;
fork checks. The system suite failed once of three runs with an email
import's blob not found, in antispam.rs; the same happened once in
purge.rs on the previous branch. Nothing here touches uploads; noted
for a separate look.
Personal-data catalog spec, §7 (settled 2026-09-28).
sysComplianceGet (673) sees the data inventory and compliance
overview: superusers and, for their tenant's slice, tenant
administrators, by default and through the one-time grants on servers
that already have their roles stored.
A Compliance Officer role at server level holds it with reading and
exporting the audit log, placing, widening, releasing and exporting
legal holds, seeing account locks, and reading accounts, lists,
domains, tenants and roles. It changes no server setting, creates or
deletes no account, and can't shorten audit retention.
A tenant's accounts can hold only roles of their own tenant (MT-3), so
the tenant role is one "Compliance Officer" role per tenant, without
holds (LH-13): made once for every tenant a server has, and whenever a
tenant is created. While nobody holds it, it is removed with its tenant
so it doesn't block the delete, and put back if the delete is refused
for another reason. Both roles carry a user's own permissions too,
since roles given to a person replace the default user role, which a
tenant's accounts can't hold anyway.
Every server makes these once, new or existing -- the built-in roles
are only made on a server with none -- and records each under P c, so a
role an administrator deletes stays deleted.
Tested: unit tests (neither role changes a setting beyond a user's
own; holds for the server's officer only; per-place records); a new
compliance system test (one server-level role; an officer reads the
audit log, places and releases a hold, and is refused a setting, an
account and audit retention; a tenant gets its role, whose holder reads
the tenant's audit log and no holds; a tenant with an unused role is
deleted and the role goes with it); the system, audit, legal hold,
account lock and SCIM suites; fork checks. The directory suite needs
its LDAP container and wasn't run here.
D1 becomes a fork-owned setting, as audit retention is, because a field
on x:TracerLog would change x:Bootstrap's stored format. D5 waits for
the v0.16.24 import's reworked spam-rules loader.
Personal-data catalog spec, defaults D2, D3, D4, D6 and D7 (settled
2026-09-28, new installs only):
- D2: automatic IP bans expire after 30 days instead of never; D3:
spam training samples, whole messages, are kept 90 days instead of
180; D4: Pyzor, which sends a digest of each message's text to a
public server, is off; D6: delivery history is kept 14 days instead
of 30. Written on the first boot of a new install only -- one with no
roles yet, the same test the built-in roles use -- by reading each
singleton, setting these fields and writing it back whole. A server
with roles keeps its settings, saved or default.
- D7: a webhook created from now on starts with the include policy and
no events, so it sends nothing until events are chosen (Rust default
and schema default, marked). The registry stores every field, so
existing webhooks keep their policy.
- Expired bans are also removed by the daily data clean-up. They
already stopped blocking and were deleted when settings next loaded;
a server that seldom reloads kept them.
D1 (log retention) and D5 (the hashed-address blocklist off) are held,
and the spec says why: x:TracerLog is stored inside x:Bootstrap with a
field after it, so adding one changes that object's stored format; and
the spam-rules loader D5 touches is being reworked by the v0.16.24
import. The spec also corrects finding 3: expired bans were deleted on
settings load; bans were permanent only because no period is set.
Tested: unit tests for the new-install values and that everything else
in each singleton stays; the system suite, whose security test now
purges an expired ban and checks its record is gone; the telemetry
test; common's unit tests; fork checks.
The sidecar catalog; the Compliance Officer places and releases holds;
the Tenant Compliance Officer is built now; shortening audit retention
is recorded and surfaced, not gated on a second person; all seven
new-install defaults, in Phase 3; the webhook finding fixed now as a
bug; snapshots kept as long as the audit log.
Phase 1 of the GDPR auditor foundation: the investigation and the
design, committed before anything is built (SPEC.md §3 rule 3).
It maps every place the server stores or sends personal data found
in the code at de275ba, each with its categories, whose data it is,
the settings that control it, what bounds its retention, where it
lives, whether it leaves the host, its scope and the code that writes
it, and the default in a new install. It proposes a sidecar catalog
(resources/privacy/catalog.toml), since the schema and registry code
are upstream's generated output with no generator here; a CI check
modeled on name-check.py; a strip-report section; a read-only
inventory method with dated snapshots; a Compliance Officer role; and
the Compliance navigation with Overview and Data inventory.
Findings worth reading on their own: webhooks ignore levels and, at
their defaults, receive every event including raw SMTP input; log
files are never deleted; automatic bans never expire; some of the
fork's records outlive the account; spam training keeps whole
messages for 180 days; traces are on in a new install; the spam
filter sends IPs, domains, hashed addresses and body digests to
third-party services by default.
Proposed default changes (new installs only) and seven open questions
are for John to decide. No default is changed.
The trace index task wrote the event type (its name) and the queue id as
text, but the tracing search index types both as integers on every
backend: BIGINT on PostgreSQL and MySQL, long on Elasticsearch. On
PostgreSQL every batch holding a trace document failed with "cannot
convert between the Rust type String and the Postgres type int8", and
since a batch writes trace and email documents together, email indexing
stalled behind it.
The document is now built by trace_search_document(), which writes:
- the event type as the opening event's numeric id, the event
x:Trace/query's event filter already matches on;
- the queue id as an integer, the first one the trace names;
- every queue id into the keywords as well, since the column holds one
value and an SMTP session can queue several messages.
index_keyword() replaced the field on every call, so before this only the
last event type and queue id survived anyway.
x:Trace/query's queueId filter parses the id (a string, or now a number)
and matches the column or the keywords, so a session is found by any of
its queue ids on every backend. The monitoring spec says what is indexed.
Traces indexed before this on the built-in index keep their text values;
the reindexTelemetry maintenance task rebuilds them.
Tests: the search store suite builds trace documents with the index
task's code, indexes them and finds them by queue id, event type and
keyword (Sqlite, PostgreSQL, MySQL); the monitoring suite finds a real
trace by queueId through x:Trace/query.
The server fetched upstream's latest published rules from GitHub at run
time: a version nobody here tested, code-like expressions from an account
we don't control, and the upstream name as a default in the admin form.
The published rules of spam-filter v3.0.2 are now embedded
(resources/spam-filter/, MIT, in THIRD-PARTY.md) and used whenever no other
source is configured. An empty setting and upstream's old default both mean
the bundled rules, so existing installs switch without a settings change;
the URL stays an operator override (https:// or file://). The schema default
is dropped and its description says what empty means, and the strip's
rename pass does the same to each import.
Rules load on first boot as before, and again whenever the bundled version
differs from the last one loaded, which only adds missing rules and tags.
That brings the AI classifier's LLM_* scores to installs that predate them:
production has none today.
upstream-watch now also opens an issue when spam-filter publishes a newer
release; resources/spam-filter/README.md says how to take it.
The antispam test now runs on the bundled rules, the path production
takes; SPAM_RULES_URL tests another set. Unit tests cover the URL handling
and that the bundled rules parse and score the AI tags as the AI spec says.
Everything clients, users and operators meet now carries the fork's name,
with no aliases (SPEC.md §2.4, changed here from "protocol identifiers
stay"):
- JMAP: upstream's registry capability is urn:inbuxa:jmap:registry, beside
the fork's own urn:inbuxa:jmap.
- WebDAV lock and sync tokens are urn:inbuxa:dav*; clients resync once.
- Sieve: vnd.inbuxa.while and vnd.inbuxa.expressions. sieve-rs spells these
into its compiler, so it's vendored (vendor/sieve-rs, 0.7.3) and patched in;
a unit test fails if Cargo.lock ever moves past the vendored copy. The
trusted runtime now names itself too, rather than answering sieve-rs's
default.
- The web interface's OAuth client is inbuxa-webui. On every start the old
stalwart-webui client is removed and any application naming it is moved
over.
- The spam filter's blobs are INBUXA_SPAM_*; every start moves any left
under the old keys, so a trained model survives.
- SQL stores and log files default to inbuxa, in the code and in the
schema served to the admin (checksum regenerated).
- Settings are INBUXA_* only. A STALWART_* variable that's set where its
INBUXA_* one isn't stops the server at startup, naming it.
- The version-upgrade messages link docs.inbuxa.org's migration page, and
the OpenAPI description, smtp crate metadata and web-push test fixtures
lose the name.
Kept on purpose, allowlisted with reasons: the OAuth key-derivation
contexts (renaming them would end every session and invalidate every
sealed client id) and the hashed application prefix.
Also fixes a latent start-up failure: ensure_client updated an existing
first-party client with a revision of 0, which the registry's assertion
never matches, so adding a redirect URI or changing the webmail secret
failed start-up. And the principal session test now expects
legacyProtocols (C-1, added 2026-09-21), which it had missed.
Tested: the server builds without warnings; common's 106 unit tests,
including the vendoring check; a new integration test for the two
start-up migrations; and the webdav, jmap, imap and SMTP Sieve suites.
`allowScimProvisioning` is cached with the domain as DOMAIN_FLAG_SCIM, but
it wasn't among the fields whose change drops the cached entry, so flipping
the flag changed nothing until something else evicted the domain. SCIM-60
says the change takes effect without a restart.
Found by the two acceptance checks that were written but never called from
the driver, so neither had ever run: `authority` (SCIM-58 to SCIM-60,
through `synchronize_account` itself) and `rate_limits` (SCIM-14). Both are
wired in now, and `scim_tests` passes with them.
docs/spec/compat-tests.md lists each compat test, what it needs, what it
checks and what a failure means, and says plainly that they run against a
copy only, since the monitoring one purges the history it reads. The
scale-out and per-domain statuses record the acceptance tests that now
run.
Where the routing lives and what a read scope covers, which acceptance
tests the container suite runs, what isn't exercised (two nodes, the
primary stopped, and the MySQL paths, which are built but unrun), what
was settled from the code, and the known limits.
per_domain_directory_compat, ignored, checks a copy of INBUXA's data has
no directory, no server default and no domain with its own directory, as
observed. The status names where the rules live, which suite covers each
acceptance test and how far, what was settled from the code, and the
known limits. SCIM's status notes its test 5 now passes.
Where the code lives, which suite covers each acceptance test, what isn't
exercised, what was settled from the code, and that read-replica routing
(ST-5 to ST-15) waits for per-domain directories.
Where each part lives, which suite covers each acceptance test, test 5
deferred to per-domain directories and test 31 unrun until a copy of
INBUXA's data, what was settled from the code, and the known limits.
monitoring_compat, ignored, checks the observed settings against a copy
of INBUXA's data, reads the old history without an error, and purges it.
The status names where each part lives, which suite covers each
acceptance test, the tests not exercised, and the known limits.
The classifier sends only the subject and text, between unforgeable markers
after the operator's prompt, to an OpenAI-compatible endpoint the operator
configured; nothing is preset. Its answer maps to an LLM_ tag whose score is
clamped (+5.0, -1.0 by default) and can never discard or reject on its own;
X-Spam-LLM is sanitized, encoded and folded, and a planted one is removed.
Failures, timeouts past the ceiling, a full slot or a paused model leave
mail flowing untagged. llm_prompt answers trusted scripts, and accounts
holding interactAi within an hourly limit. Redirects aren't followed and no
content or secret is logged. The limits live in inbuxa:AiLimits.
Acceptance tests 1 and 3 to 21; test 2 as the re-enabled shared llm case,
whose setup no longer waits on a rules file from a developer's own path;
test 22 written as the ignored ai_compat.
Logos resolve domain, then tenant, then server-wide, then the built-in, with
subdomains finding their domain. GET /logo serves a data-URL image, redirects
to a URL logo without fetching it, sandboxes SVG, and answers 404 when no
custom logo applies. Emails embed the first PNG, JPEG or GIF logo. Logo and
template writes are checked; stored templates are read at send time, always
escaped, and fall back to the built-in with a build warning when they don't
parse. The RSVP page is served byte for byte with a CSP and no-referrer. The
sign-in and RSVP pages load the logo through an image element. MT-22's
session logo follows the chain to the server-wide logo.
Acceptance tests 1 to 17; test 18 written as the ignored branding_compat.
With archiveDeletedAccountsFor set, a destroyed account's record is kept in
the fork subspace with its id, its DestroyAccount task is due at the end of
the period, and its shares are suspended both ways. Its addresses can't be
taken by new accounts, aliases, lists or masks. inbuxa:DeletedAccount/get
lists kept accounts to server and tenant administrators; /set restores one
with a new password (same id, task cancelled, shares reinstated) or destroys
it now. The destroy task also clears undelete's own records.
Acceptance test 14; test 16 written as the ignored undelete_compat.
Found by running system_tests, which masked email no longer stops:
- rcpt_resolve rewrites a live mask to its owner's address, so
Delivered-To names the account; delivery recognizes the mask from the
original recipient when it belongs to that account.
- x:MaskedEmail/set create responses carry the server-set email.
- x:MaskedEmail/query returns every mask to a server-level impersonate
holder, and filters on accountId.
- The refusal for an unlinked emailDomain uses upstream's wording.
- The shared delivery test checks the fork's address format (ME-13).
- The masked email test's tenant domain uses manual DKIM, so its cleanup
leaves nothing behind.
Test 3 expects invalidForeignKey, as MT-3 says. A tenant admin reads its own
tenant through sysTenantGet and sysTenantQuery, added to the default Tenant
Administrator role for new installs only. The server adds nothing for MT-7's
dashboard list. The MT-19a submission warning is deferred. MT-22's logo is
the urn:inbuxa:jmap account capability's logo field (contract C-1).