afffa0fc960aab1a6b920ed4097100ac53568c56
123
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b20b09f81a |
New installs start with the hashed-address blocklist off, and DNSBL zones read right
Personal-data catalog spec, default D5 (settled 2026-09-28; built after the v0.16.24 import's spam-rules loader landed). msbl.org's EBL is sent a SHA-1 of every email address it's asked about. A new install's first boot now leaves a note, and the rules update, once the bundled rules are in, switches STWT_MSBL_EBL_EMAIL off and forgets the note, so it happens once; the loader keeps that switch through later updates. An existing server has no note and keeps every blocklist as it is. Also fixes the data inventory's DNSBL endpoints: a zone is an expression (`ip_reverse + '.zen.spamhaus.org'`, conditional branches, `hash(email, 'sha1') + '.ebl.msbl.org'`), and the zone names are now the quoted literals that start with a dot, from every branch, rather than the expression's text. Tested: unit test for the zone rule; the compliance system test (no note, no change; the inventory lists ebl.msbl.org, not a hash; with the note the blocklist goes off; the note works once); the system suite; fork checks. |
||
|
|
a8fb10458b |
Evaluate the personal-data catalog: the data inventory and its history
Personal-data catalog spec, §6 (Phase 3c). inbuxa:DataInventory/get evaluates the catalog against the server's live settings and says what this server holds: for each source and each object that can hold personal data, its categories and whose data it is, whether it is collected here at all, what bounds its retention (the live value of the setting that does, or unbounded), whether it leaves the host and to which endpoints, and a summary. Every host that receives something is listed once as a candidate processor with what it receives. Inside a tenant it answers with the tenant's slice and none of the server's processors. Read-only, with sysComplianceGet. inbuxa:InventorySnapshot/get is the history: a dated copy of the evaluated inventory, recorded when it changes -- after a registry write to an object the inventory reads, after inbuxa's log, audit or AI settings change, and on the daily clean-up -- and kept as long as the audit log's records. ids: null lists every snapshot, newest first; the full inventory only when asked for. The catalog is embedded and parsed at start (new dependency: toml, MIT/Apache); the evaluation is a pure function of it and the live facts, so each configuration is tested without a server. Loopback endpoints stay on the host; any other configured endpoint leaves it. Tested: unit tests for the evaluation (a new install's defaults, an external blob store, a hosted AI endpoint, telemetry off, a tenant's slice, hosts from URLs, loopback), snapshots, and the fact gathering's store and duration rules; the compliance system test, extended (the officer reads the inventory, a plain user is refused, a tenant's officer sees its slice and no processors, a webhook to another host becomes a processor and a snapshot names x:WebHook, a retention change reads through); the system, audit, legal hold and account lock suites; fork checks. The system suite failed once of three runs with an email import's blob not found, in antispam.rs; the same happened once in purge.rs on the previous branch. Nothing here touches uploads; noted for a separate look. |
||
|
|
7285b3e38a |
Merge main (upstream v0.16.24) into feature/compliance-roles
The schema conflicted as a binary file: taken from main and the one edit here re-applied (sysComplianceGet after sysLegalHoldExport). The import kept the permission count at 673, so the new id stays 673. Retested on the merged tree in its own target directory: the compliance and system suites pass. One earlier system run failed in purge.rs (an imported blob not found) and didn't recur. |
||
|
|
9893452ca2 | Merge pull request 'Merge upstream v0.16.24' (#84) from merge/upstream-v0.16.24 into main | ||
|
|
63adb4e2b8 |
Add the compliance permission and the Compliance Officer roles
Personal-data catalog spec, §7 (settled 2026-09-28). sysComplianceGet (673) sees the data inventory and compliance overview: superusers and, for their tenant's slice, tenant administrators, by default and through the one-time grants on servers that already have their roles stored. A Compliance Officer role at server level holds it with reading and exporting the audit log, placing, widening, releasing and exporting legal holds, seeing account locks, and reading accounts, lists, domains, tenants and roles. It changes no server setting, creates or deletes no account, and can't shorten audit retention. A tenant's accounts can hold only roles of their own tenant (MT-3), so the tenant role is one "Compliance Officer" role per tenant, without holds (LH-13): made once for every tenant a server has, and whenever a tenant is created. While nobody holds it, it is removed with its tenant so it doesn't block the delete, and put back if the delete is refused for another reason. Both roles carry a user's own permissions too, since roles given to a person replace the default user role, which a tenant's accounts can't hold anyway. Every server makes these once, new or existing -- the built-in roles are only made on a server with none -- and records each under P c, so a role an administrator deletes stays deleted. Tested: unit tests (neither role changes a setting beyond a user's own; holds for the server's officer only; per-place records); a new compliance system test (one server-level role; an officer reads the audit log, places and releases a hold, and is refused a setting, an account and audit retention; a tenant gets its role, whose holder reads the tenant's audit log and no holds; a tenant with an unused role is deleted and the role goes with it); the system, audit, legal hold, account lock and SCIM suites; fork checks. The directory suite needs its LDAP container and wasn't run here. |
||
|
|
1d5a49409f |
Keep rotated log files for a set number of days
Personal-data catalog spec, default D1 (settled 2026-09-28): log files were never deleted. inbuxa:LogSettings.keepForDays says how many days rotated log files are kept; unset (null) keeps every file, as before, and a new install sets 30 days. It is a fork-owned setting, stored under T + l as audit retention is, not a field on x:TracerLog: that object is also stored inside x:Bootstrap with a field after it, so a new field would change x:Bootstrap's stored format. Server-level, with the tracers' permissions (sysTracerGet, sysTracerUpdate); changes are in the audit log, before and after. Log files are local, so every node deletes its own: hourly, and at once when the setting changes on that node. Only regular files named <prefix>.<something> in each enabled log tracer's directory, last changed more than the limit ago, are removed; the file being written is never that old, and nothing else in the directory is touched. Minimum one day. The catalog classifies inbuxa:LogSettings and points the log file's retention at it. Tested: unit tests for the file rule (only this log's old files; the current file, other files and directories stay) and a purge on disk; the system suite, which reads, sets, refuses zero, restores null and checks the audit records; fork checks. |
||
|
|
80d6c09c59 |
Merge main into merge/upstream-v0.16.24
The schema, which both sides changed, merged as JSON with no conflicts. The personal-data catalog (#83) gains upstream's new x:DnsServerPowerDns: nothing personal but its API key, like the other DNS providers. |
||
|
|
a0ffdb8071 |
New-install privacy defaults, and expired bans purged daily
Personal-data catalog spec, defaults D2, D3, D4, D6 and D7 (settled 2026-09-28, new installs only): - D2: automatic IP bans expire after 30 days instead of never; D3: spam training samples, whole messages, are kept 90 days instead of 180; D4: Pyzor, which sends a digest of each message's text to a public server, is off; D6: delivery history is kept 14 days instead of 30. Written on the first boot of a new install only -- one with no roles yet, the same test the built-in roles use -- by reading each singleton, setting these fields and writing it back whole. A server with roles keeps its settings, saved or default. - D7: a webhook created from now on starts with the include policy and no events, so it sends nothing until events are chosen (Rust default and schema default, marked). The registry stores every field, so existing webhooks keep their policy. - Expired bans are also removed by the daily data clean-up. They already stopped blocking and were deleted when settings next loaded; a server that seldom reloads kept them. D1 (log retention) and D5 (the hashed-address blocklist off) are held, and the spec says why: x:TracerLog is stored inside x:Bootstrap with a field after it, so adding one changes that object's stored format; and the spam-rules loader D5 touches is being reworked by the v0.16.24 import. The spec also corrects finding 3: expired bans were deleted on settings load; bans were permanent only because no period is set. Tested: unit tests for the new-install values and that everything else in each singleton stays; the system suite, whose security test now purges an expired ban and checks its record is gone; the telemetry test; common's unit tests; fork checks. |
||
|
|
eac3db34e9 |
Principal get test: expect legacyAllowed
The per-protocol switches (#79) added legacyAllowed to the account's urn:inbuxa:jmap capability, but this expectation wasn't updated, so jmap_tests stopped here and the suites after it never ran. |
||
|
|
b2453d066b |
Merge upstream v0.16.24
Eight conflicted files resolved, plus the lock file and the schema:
- crates/services/src/task_manager/spam_classifier.rs: upstream's rules
update now replaces existing rules, DNSBL servers, lookups and file
extensions, keeping only whether each is on. Taken, with one difference:
an object an admin edited is kept as it is. Every object an update writes
is fingerprinted (content without `enable`, SHA-256, stored under
SUBSPACE_INBUXA "Sf"), and only one that still matches is replaced.
Scores are never replaced, as upstream has it. The AU-1.10 summary record
now names what was added, replaced and kept, and the bundled rules are
marked applied only when the update fully succeeded, so a failure runs
again on the next start. The marker becomes "3.0.2+2", which runs the
update once on upgrade to fingerprint every rule still as bundled.
- crates/common/src/network/autoconfig/autodiscover.rs: upstream's rewrite
(implicit TLS first, labeled SSL), with the per-protocol switches (LP-7,
LP-14a) passed in as a filter.
- crates/store/src/backend/mysql/{search,write}.rs: upstream's chunked
deletes (no unbounded first DELETE, stop on a short chunk, halve the
chunk on the new chunk-too-large errors) inside the fork's query timeout.
- crates/smtp/src/lib.rs: the fork's queue spawn kept. It already fixed the
stall upstream fixes here (a node without outboundMta stops accepting
mail at about 1024 queued messages), and follows role changes live.
- crates/jmap/src/registry/mapping/bootstrap.rs: the log path stays
/var/log/inbuxa/; upstream's PowerDNS mapping taken.
- crates/main/Cargo.toml: the AGPL-only license kept, version 0.16.24.
- tests/src/jmap/principal/get.rs: the fork's capabilities kept.
- resources/schema/schema.json.gz: merged as JSON; upstream relabeled the
vendor Sieve extensions "(Stalwart)", kept as "(vnd.inbuxa)".
- Cargo.lock: upstream's, with the fork's crates added by Cargo.
Also:
- tests/src/smtp/inbound/spam_rules_kept.rs: an edited rule survives an
update, an unedited one is updated, rules from before fingerprints are
handled, and the audit summary says so. Upstream's own spam_rules test
passes unchanged.
- tests/src/smtp/reporting/reschedule.rs moves to port 19058; upstream's
new spam_rules test took 19057.
- tools/fork/renames.py renames the "(Stalwart)" labels and the default
log path, so neither conflicts again.
- tools/fork/notice-check.py compares against the newest snapshot in the
checked-out history instead of the upstream branch head, so moving the
branch no longer fails other open pull requests.
- tests/src/directory/issuer.rs (since v0.16.23) stays out, and is on the
build check's known list: it tests issuer-based directory routing, which
the fork doesn't have (DIR-2).
- Strip report: docs/fork/strip-reports/v0.16.24.{md,json}.
|
||
|
|
f59b084ce5 |
Import upstream v0.16.24, stripped
Upstream commit: af37a234981722493b74623a983581691d2b70b6 Enterprise-only files removed or emptied: 63 Enterprise-only snippets removed: 118 in 50 files Dangling module declarations removed: 5 Edits turning enterprise off: 25 Third-party code: 14 files, 0 not in THIRD-PARTY.md Renamed identifiers: 62 in 18 files Verification: clean The same Enterprise footprint as v0.16.23. The build check fails only on tests/src/directory/issuer.rs, unchanged since v0.16.23: it calls a helper from upstream's Enterprise-only OIDC test, and tests issuer-based directory routing, an Enterprise feature. main has never carried it. |
||
|
|
f5888d79b0 | Merge pull request 'Give IMAP, POP3 and ManageSieve a switch each' (#79) from feature/per-protocol-switches into main | ||
|
|
8e9cedbe97 |
Give IMAP, POP3 and ManageSieve a switch each
The legacy-protocols switch was all or nothing. An operator can now stop
POP3 and keep IMAP: each of IMAP, POP3 and ManageSieve has its own
switch, server-wide on inbuxa:ProtocolPolicy and per tenant on
inbuxa:TenantProtocolPolicy (properties imap, pop3, manageSieve).
legacyProtocols stays as the kill-all: setting it sets all three, and it
reads "disabled" exactly when all three are off. A policy stored before
this has only legacyProtocols and reads as all three at that value, so
existing servers and tenants carry over unchanged. In one /set, a
protocol named beside legacyProtocols overrides it.
SMTP submission keeps no switch of its own: sign-in over it is refused
only when all three are off, as the single switch did (LP-6), so
turning one protocol off never stops a mail app sending. For a tenant,
the server's switches and the tenant's count together.
Server-wide, a change closes the listeners of whatever is now off and
puts back the saved listeners of whatever is on again, both in one
change if asked; listeners of a protocol still off stay saved. Sign-in,
autoconfig, autodiscover, PACC (now prepared once per combination) and
the suggested DNS records all follow each protocol separately. A tenant
may turn a protocol on only while the server has it on (LP-9), and the
refusal names which. The JMAP session adds legacyAllowed, the protocols
still allowed for the account; legacyProtocols there keeps its meaning
for older webmail builds. Events name the switches ("pop3 disabled"),
and audit before/after reads every switch even from an older policy.
Tested: unit tests for the switches, the old-policy reading, the
server/tenant combination, the tenant refusal and listener refusal; and
tests/e2e/legacy_protocols.py against a running server, all 100 checks,
including new ones: POP3 alone off closes only its port and refuses
only its sign-in while IMAP and sending go on; only POP3 stops being
advertised; one change closes IMAP and reopens POP3; a tenant turns
POP3 off for itself, and can't turn IMAP on while the server has it off.
|
||
|
|
5ba54e8fb7 |
Merge pull request 'List what a hold export can't read instead of skipping it (LH-12)' (#75) from fix/hold-export-exceptions into main
Reviewed-on: #75 |
||
|
|
e1e8a9aeb0 |
Send "none" instead of "pass" as the DMARC report disposition
Cloudflare's DMARC report intake rejects every aggregate report we send with "555 5.7.1 invalid_report_schema". Bisected against the live endpoint: the only element it objects to is <disposition>pass</disposition>, the value RFC 9990 added for mail that passed DMARC under an enforcing policy. The RFC 9990 namespace, <np>, <discovery_method>, <testing> and a missing <pct> are all accepted, and a report that differs only in using "none" there goes through. "none" (no action taken) is valid under both RFC 9990 and RFC 7489 and says the same thing to the reader, so reports now go out with it. The stored report keeps "pass"; only the serialized copy changes. |
||
|
|
5c506b9d2b |
List what a hold export can't read instead of skipping it (LH-12)
An item the hold covers whose stored record or content can't be read goes in exceptions.csv with the path it would have had and the reason, rather than being left out silently. The file is always in the ZIP, so a header-only one shows nothing was missed, and manifest.sha256 carries its hash beside the manifest's. |
||
|
|
68dd749291 |
Export what a legal hold keeps as a ZIP (LH-12)
inbuxa:HoldExport/set takes a hold, optionally some of the accounts it covers, and a reason; the collection runs in the background and get says when it's ready. The ZIP has, per account, mail as .eml under its folders, calendars as .ics, contacts as .vcf, files as stored, and the archived items the hold keeps under archived/; a manifest.csv gives each entry's account, kind, folder, date, whether it was archived, size and SHA-256, and manifest.sha256 hashes the manifest. Accounts the hold doesn't cover are left out, and items outside its date range are too: live mail by arrival, events by start, and archived items the same way, so an export doesn't carry deleted items that only another hold keeps. The finished file is a blob of whoever started the export, so only they download it, and it lasts as long as any upload (uploadTtl). Exports are records under the hold (SUBSPACE_INBUXA H/e): never changed or destroyed, each with its status, counts, size and checksum. Starting one needs sysLegalHoldExport, an active hold and a reason, and is recorded in the audit log like the audit log's own export. The build is in memory and capped at 2 GB; bigger holds fail with a message saying so, and are split by picking accounts. Tested: unit tests for safe ZIP names and the manifest and its hash; the legal_hold system test, on RocksDB, PostgreSQL and MySQL, exports a hold end to end (live and archived mail, the manifest's hash, an asked- for account the hold doesn't cover left out) and checks the refusals (no reason, a user without the permission, a released hold) and the audit record; and by hand from the console on a local server. Not covered by a test: the archived-item date range with two holds of different ranges over one account. |
||
|
|
c1b5bf956c |
Audit records name accounts in full, and holds by name
An account's or mailing list's name is only its local part, so the log said "Account ken.gosling" where two domains could each have one; it now says [email protected]. A change to a legal hold was recorded under its id; the hold's current state is now read first, so the record carries its case name and each change reads before/after. |
||
|
|
3217aae4e8 |
LegalHold/get takes coveringAccount
Only the active holds covering one account, live or deleted and kept, through any route: for the console's Held badge (LH-14). |
||
|
|
538ae107d7 |
Legal holds, step 6: what each hold keeps
inbuxa:LegalHold/get answers accountsCovered, itemsHeld and sizeHeld when asked: the accounts a hold reaches now (deleted ones it keeps included) and the archived items it keeps, with their size. Worked out in one pass over accounts and archive, only for requests that name them. Held items stay out of the user's quota, as all archived copies do (LH-9). |
||
|
|
39707cd2e8 |
Legal holds, step 5: held accounts can't be destroyed
Destroying a held account removes the login, as offboarding needs, but keeps its data as a deleted account with no expiry, whether or not undelete keeps accounts; its addresses stay reserved and its holds name it from then on. Destroy-now refuses it, and its DestroyAccount task defers itself while it's held or its time hasn't come. Holds placed or released later freeze or free kept accounts in the same settle pass, with 30 days' grace after the last release (LH-8, LH-10). |
||
|
|
8d3e99bc00 |
Legal holds, step 4: freezing, release, and the audit log
Placing or widening a hold freezes what's already archived in its scope and range, its old deadline noted; releasing one gives each item no other hold covers that deadline back, or release plus 30 days if later. One pass over the archive does both and changes nothing twice (LH-6, LH-10, LH-11). A held archived item can't be destroyed; restoring still can, and the hold is named only to callers who may see holds (LH-7). Audit records about a held account survive the purge (AU-7). Fixes the daily clean-up of expired archived items (UD-13), which never found any: the registry's unfiltered query reads an all-ids index that archived items aren't in. Items are now walked account by account, kept deleted accounts included. Expired items were still removed whenever their account's archive was read. |
||
|
|
7b97efbb7f |
Legal holds, step 3: deleted items in a held account are kept
Every way of deleting mail (JMAP, IMAP EXPUNGE, POP3, mailbox removal, Trash emptying) and Sieve scripts, events, contacts and files now asks how the account's deletions are kept: a hold keeps them with no expiry (archivedUntil 9999-12-31), even with undelete off; otherwise undelete's period applies as before (LH-4). A hold's date range decides by the item's own date (LH-3). Mail is noted as held at deletion and settled when it's archived, once its received date is known; outside the range it gets undelete's deadline or isn't kept. Events go by their start, with a day's slack for time zones; recurring events, contacts, files and scripts are held whole. A groupware item's note now stays until its archive succeeds, and a failure retries the task instead of being logged and lost (LH-5). |
||
|
|
318783f444 |
Legal holds, step 2: who a hold covers
A hold reaches an account by name, through any of its addresses' domains, its groups or its tenant, as they are now, so an account added to a held domain later is held too. An account that leaves a held domain, group or tenant stays held: the registry write hook adds it to the hold by name on every account change, whoever makes it (LH-2). Server::holds_on answers for the deletion paths, from the store each time so a hold binds every node at once. |
||
|
|
5d2e35b2dc |
Legal holds, step 1: the hold itself
inbuxa:LegalHold get/set places a hold on accounts, groups, domains, tenants or the whole server, with an optional date range. A hold's range and scope can only widen, a released hold is read-only, and none is ever deleted. Placing, changing and releasing each need a reason and are audited (LH-1, LH-3, LH-10, AU-12). Permissions 669-672 (see, place, widen or release, export held data) go to server administrators only; the tenant ceiling always strips them, as it does Impersonate (LH-13). Schema: Compliance > Legal Holds. What a hold keeps comes next, through the undelete hooks. Also moves the lock expiry helpers below the lock module's imports. |
||
|
|
9f6761c9dd |
Writing delegates may add at the top of a locked account's Files
A shared account refuses top-level folders, so an organize or full delegate couldn't add anything to a locked account with no folders. A delegate who may write now can, as the owner could; the reconcile after the create grants it the new folder. Read delegates still can't (AL-6, AL-7). |
||
|
|
d4d127fa7d |
Delegates reach the whole locked account
A delegate's token listed the locked account only for kinds of data it held grants on, so one with no files (or no calendar) was refused to the delegate outright: "You do not have access to account". The token now lists the locked account for mail, calendars, contacts and files alike, so an empty kind reads as empty. What the delegate may see or change is still each container's grant (AL-7). |
||
|
|
a36236efff |
End a locked account's delegation at its date
A delegation with an end date dropped out of the delegate's token then, but its folder grants stayed until the daily sweep, so the delegate kept the account as an ordinary share for up to a day. Each node now sleeps until the soonest end date, woken early by any lock write and at least hourly, and re-applies that lock under a cluster-wide claim. The sweep also had a second-run bug: a delegation past its date gave the delegate back its earlier share, then dropped the note, so the next sweep removed that share entirely. The note is now kept while the delegate is still listed. |
||
|
|
447229f871 |
Lock accounts: keep receiving mail, no sign-in, hand to delegates
A locked account can't sign in (it fails as a wrong password does), its sessions end on every node, refresh tokens stop working, and its Sieve scripts forward and reply to nothing. Mail keeps arriving. Delegates get real ACL grants on the account's mailboxes, calendars, address books and files at read, organize or full, with the rights they replaced restored on unlock. Folders made later are granted after the create and in a daily sweep. Organize delegates can't destroy; send-as needs organize or full. The JMAP session marks delegated accounts in urn:inbuxa:jmap. New inbuxa:AccountLock object with get/set, permissions 665-668, and a Compliance > Locked Accounts entry in the schema. Lock, unlock and delegate changes need a reason and are audited; delegate access and writes are audited too (audit-hold-lock spec AL-1 to AL-12). |
||
|
|
86d7ebd982 |
Audit log: a permanent, tamper-evident record of admin actions
What administrators and the server itself do to the control plane is now recorded, from inbuxa-drafts/specs/audit-hold-lock.md (AU-1 to AU-12): settings, accounts, domains, roles and every other registry change, with each field's before and after (secrets only as "changed"); the fork's own settings objects; administrator sign-ins (and failed ones to administrator accounts), master-user and recovery-admin sign-ins, once an hour per account, method and address; access to another account's data through impersonation or FetchAnyBlob, once an hour; exports and tamper checks; and registry writes the server makes on its own, named by subsystem (system:AcmeRenewal, system:auto-ban, system:directory-sync, ...), with a spam rules update as one summary record. No change without its record (AU-3): before a set method changes anything, a pending record per requested create, update and destroy is written; if that fails, the method is refused with serverFail. Its outcome follows as a later entry. A change interrupted by a crash stays "unfinished". Records live in the fork's subspace under L, as one SHA-256 hash chain per node. The chain's head is stored, never cached, and every append asserts it, so two writers can't take the same place. Nothing can edit or delete a record; the daily purge removes the oldest past the retention (default 730 days, minimum 90) and records where the chain now starts, so verification still passes. security.audit-recorded (647) copies each record to webhooks, OpenTelemetry and the log; security.audit-write-failed (648) reports a failed write. New JMAP objects under urn:inbuxa:jmap: inbuxa:AuditEvent/get and /query (filters: time, actor, action, target, account, tenant, outcome, address, text), inbuxa:AuditSettings, inbuxa:AuditExport (CSV or JSON Lines built on the server, each line with its chain hash, ending in a manifest; the created object names the blob and its SHA-256) and inbuxa:AuditVerification. New permissions sysAuditGet, sysAuditExport and sysAuditSettingsUpdate: the Administrator role gets all three, the Tenant Administrator role gets read and export, once, on existing installs too. A tenant administrator sees records whose actor or target is in its tenant, including a server administrator's changes there. Sign-in method on the session: access tokens now remember how they signed in (password, app password, API key, OAuth client, directory, master user, recovery admin), including across the HTTP credential cache. New OAuth access tokens carry their client id in the sealed claims; older ones show as client "unknown" until they expire. The schema gains the permissions, the two events and a Management > Compliance > Audit Log link. Stack: the request layer boxes every inner future where it's made. Without that, a debug build overflowed the default 2 MB worker stack on a registry set; measured with the same request, the branch and main now overflow at the same stack size (between 1856 and 1920 KiB, debug), so the layer adds nothing measurable. Tests: unit tests in inbuxa-features and jmap; system::audit::audit_log_tests (run with --ignored) passes on RocksDB, SQLite, PostgreSQL, PostgreSQL with a read replica, MySQL, MySQL with a replica and FoundationDB. The system, JMAP and SCIM suites pass. authorization.rs skipped fork permissions that guard no registry object; the audit suite checks a plain user is refused instead. |
||
|
|
ad648d8d12 | Calibration test: pass the new stream argument to request::body | ||
|
|
866d7d3ed5 |
Explain a setting that was never saved, from its defaults
A singleton such as x:SpamSettings has no stored object until someone saves it; /get shows its defaults instead. Explain looked only for the stored object, so every setting still at its defaults answered "No such x:SpamSettings." It now falls back to the defaults the same way. |
||
|
|
d9a6db025b |
Explain this: the local model reads delivery failures, verdicts, logs and settings
A new method, inbuxa:Explanation/set, asks the node's local model for a short plain-words reading of one thing an administrator is looking at: a failed recipient in the queue, a Classify verdict, a log line or trace event, or one setting with its saved value. The server builds the prompt itself from stored data and the registry schema, never from text the console sends, and grounds SMTP replies in RFC 3463 and RFC 5321. What the model is never shown: secrets (including ones nested inside a setting, like an AI model's HTTP auth), raw protocol events, and the contents of any other event. A tag name that doesn't have a tag's shape is refused before a model is asked. Calls share the AI gate with spam classification, but mail always keeps its slot, and Explain has its own hourly count per account and its own on/off switch in inbuxa:AiLimits. The permission is sysAiExplain, superuser only; tenant administrators can't use it. The session carries an aiExplain flag so a console knows when to offer the button. An install whose roles were stored before the permission existed gets it added once, at start-up, to the roles that are administrators' alone, not the User role their defaults share with every account. An operator who removes it later isn't overruled. Tests: unit tests in inbuxa-features and jmap, and ai_explain_tests (run with --ignored) covering the acceptance tests and the upgrade. |
||
|
|
5927dda7e2 |
PostgreSQL search: find words inside URLs and file names in body text
After #37, address fields on PostgreSQL are split into words as the built-in index splits them, but language text (subject, body, attachments) still goes straight to PostgreSQL's parser, which keeps a URL, host, path or file name as tokens of its own: "https://x.example/shipping-support/" becomes a url, a host and a url_path, "invoice-2024.pdf" a file. So TEXT/BODY "shipping" missed messages where the word appears only inside a link, while RocksDB and the other built-in backends found them: 8 messages across a handful of searches in the rehearsal. On insert, language text is now indexed as it was, followed by the word parts of each token that holds a URL separator (/ . @ : ? = & # _ % + ~ \), split with SpaceTokenizer as keyword_terms() splits addresses. The parts go through the same text search configuration as the rest of the text, so they are stemmed like the words around them. Plain words, words that only carry punctuation ("end.", "(see") and hyphenated words (the parser already splits those) add nothing, so text without links is indexed exactly as before. Each part is added once per document. On sample mail, the text vector of a short order notice with three links grows from 546 to 716 bytes, a newsletter with 25 tracking links from 5586 to 6430, and a plain letter not at all. On search, a query word written as a URL, host, file or hyphenated word also matches as its word parts, ORed with the query as written, so "shipping-support" or "invoice-2024.pdf" match the new parts and documents indexed before this change still match as they did. Existing messages keep their old vectors until they are reindexed (the reindexAccounts task); new and reindexed messages match at once. store::search_tests gains test_url_word_search: five bodies, 19 body searches for words found only in a URL path, query string, host or file name, the tokens as written, plain words and non-matches, with the same expected ids on every backend. It passes on RocksDB, SQLite, MySQL and PostgreSQL; on main PostgreSQL fails at the first ("shipping" finds [3], not [0, 3]). On PostgreSQL the suite then stops at the account sort assertion (query.rs:689) exactly as it does on main. |
||
|
|
71ce11c57d |
Settings writes: wait for a burst to settle before reloading
A cluster rehearsal sent ten x:<Object>/set requests at once and got ten full reloads on every node. #39's coalescing only joined writes that queued behind a running reload, but the requests reached the server about 33 ms apart and a reload takes tens of milliseconds, so none overlapped one. A full reload after a registry write now waits for writes to settle: 75 ms after the last one, and at most 250 ms after the first it covers, so a steady stream still reloads at least four times a second. 75 ms is a little over twice the gap the rehearsal saw between requests. A single write pays it once: in the tests a settings write takes about 140 ms instead of 60. The reload runs in a task of its own, so a request that goes away doesn't cancel it for the others. Each write takes the result of the first reload that started after it was stored (the gate keeps the last 64 results), so applied true or false still describes the reload that covered that write. The 33 ms gap was a queue on the server, not password hashing: Basic credentials are cached per Authorization header, so they are checked once. Every authenticated HTTP request counted itself against the account's rate limit by incrementing one counter per account in the in-memory store, so parallel requests from one account queued on that key: a row lock on PostgreSQL (a few round trips to the database each) and conflict retries with a 50-300 ms backoff on RocksDB. An account with the unlimitedRequests permission (administrators, by default) passes the rate and concurrency limits anyway, so its requests are no longer counted. Ten parallel Core/echo calls as the admin now finish in 1-4 ms; before, they finished one after another over 20 ms on a local PostgreSQL and 300-450 ms on RocksDB. Other accounts still count every request. system::auto_reload::settings_reload_tests: ten concurrent writes now take one reload (the gate counts them; at most two allowed), all are applied: true and in the running settings, and a single write takes exactly one reload. RocksDB and PostgreSQL, 1 reload in 141-196 ms. With the old behavior (no wait, requests counted) the same writes took 5 reloads; without the wait but with the rate fix, 2. cluster::broadcast (3 nodes, PostgreSQL + NATS) and system::reload still pass. |
||
|
|
a891667149 |
Tracers whose settings change start over on reload
A cluster rehearsal moved a Log tracer to another directory: the write was reported x:settingsReload applied:true, but the tracer kept writing to the old file until a restart. Telemetry::update only refreshed each running tracer's events, level and lossiness; a tracer's own settings (path, prefix, rotation, format, endpoint, headers, ...) stayed as built. Each tracer now carries a hash of the registry object it was built from, less the fields that change in place. The reload compares it with the running tracer's: unchanged ones are updated in place as before, changed ones are started over, new ones started and removed ones stopped. Only tracers this server started are removed; upstream removed every subscriber not in the settings, which also cut off live-tracing streams on each reload. Starting over is a swap in the collector, so no event is lost or written twice: a subscriber registered under a running one's id replaces it between two collection passes. The old one's batch is sent first (what its full channel can't take moves to the new one), and dropping it closes its channel, so its task writes what is queued and ends. Per tracer kind: - Log: a tracer started over on the same files (rotation or format changed) waits for the old one to finish, so lines don't interleave. - Webhook: the task held a sender of its own channel for retries, so it never ended; retries now use a weak sender, and pending events are posted when the channel closes. - OpenTelemetry: pending logs and spans are exported when the channel closes instead of dropped, and a span that was open across the swap is exported by the new tracer with the events it saw. - Console and journal: nothing kept between batches. - Trace history: built from the tracing store, which takes a restart, so it is never started over. No kind needs a restart, so x:settingsReload doesn't gain one. system::tracer_reload::tracer_reload_tests (new): a Log tracer created over JMAP writes to its directory; its path is changed over JMAP while 2000 numbered events are emitted; after the reload, events land in the new file and not the old one, each numbered event is in exactly one of the two files, and a destroyed tracer writes nothing. On main the new file never appears. |
||
|
|
59e631eded | Merge pull request 'Every node records DMARC and TLS results for the aggregate reports' (#47) from fix/front-node-dmarc into main | ||
|
|
5dde9793eb |
Every node records DMARC and TLS results for the aggregate reports
The report scheduler dropped DMARC and TLS events on a node whose role lacks outboundMta (upstream never started it there, so they sat in a channel nobody read). Mail received on a front node therefore never reached an aggregate report, which is meant to cover all of a domain's inbound mail, whichever node received it. In rehearsal, five messages received on port 25 on a front node were missing from every report. - The report scheduler records on every node. Recording is a store write the nodes already share, so it needs nothing from the outbound MTA. Building and sending a report (the DmarcReport and TlsReport tasks) stay with outboundMta nodes, as the task manager already enforces. - More nodes now append to one report at once. Appends already guard the report's versioned primary key; a write that loses now retries up to ten times after a short random pause, not three times at once. - The node sending a report deletes it only if it is unchanged since it was read, and reads it again otherwise, so a record another node appends meanwhile goes out with the report instead of being deleted unsent. Test: cluster::front_reports (PostgreSQL and MySQL). A front node's results appear in the report the MTA node sends, alongside eight appended at once from both nodes, and the front node never runs the report task. It fails on main: the front node's results are never recorded. |
||
|
|
1a7859a8cc |
Report reschedules keep the task queue readable
Setting deliverAt on an internal DMARC or TLS report wrote the new task
queue row with the report's object type (0x21, 0x6e) instead of the task
type (7, 8), and left the task row at its old due. The task manager's scan
failed on that row with store.data-corruption ("Failed to iterate over task
queue"), and because the error ended the whole scan, every task due after
the row stopped running on every node.
- reschedule_ops writes the new queue row through schedule_task_with_id, so
it carries the task type and the task row gets the new due. It removes
the row the task is actually queued under (the task's due, which differs
from deliverAt once the task has been retried) and any row an earlier
reschedule left at deliverAt.
- x:DmarcInternalReport/set and x:TlsInternalReport/set lock the report's
task while they move it, as x:Task/set does, refuse while the report is
being sent, release the locks however the request ends, and wake the task
manager.
- The task manager logs a queue row it can't read (id, due, key, value) and
skips it instead of ending the scan. It then repairs the row from its task:
the row is rewritten with the task's type, and a row with no task behind
it is removed. A row holding a report's object type for a report task is
what the old reschedule wrote: the task is moved to that row's time, as
the reschedule intended, and its old queue row is removed. Stores that
already hold such a row recover on their own once it comes due.
- x:Task/query with a type filter skips an unreadable row instead of
failing.
Test: smtp::reporting::reschedule (RocksDB and PostgreSQL). It fails on
main: x:Task/get shows the old due, and with that check removed, neither
report nor a later task ever runs.
|
||
|
|
e00978c0b4 |
Cluster role changes apply to delivery and tasks without a restart
In cluster rehearsal 3, turning outboundMta off on node1's role was
reported applied (x:settingsReload applied: true), yet node1 kept
delivering mail, a report message included, until it was restarted.
The queue and report managers were started at boot only when the
node's role included outboundMta (crates/smtp/src/lib.rs), and the task
manager only when the role had some task type (spawn_task_manager).
After that nothing looked at the role again: a queue manager that was
running kept claiming and delivering, and one that wasn't never
started.
They now start on every node (outside recovery mode) and follow the
role live:
- Queue manager: before each scan it reads the role from the running
settings. Without outboundMta it claims nothing new; deliveries
already running finish and report back as usual, which releases
their locks. When the role comes back (a reload wakes the manager
with ReloadSettings, and it looks again every 30 s regardless) it
logs queue.started and scans the whole queue at once.
- Report scheduler: DMARC and TLS report events are handled only while
the role has outboundMta, as at boot; events arriving without it are
dropped, as they were on a node started without the role.
- Task manager: task_enabled already read the current role on every
scan. It now also runs on nodes whose role has no task type (the
scan returns at once until one is added), a job claimed before a
role change is handed back at once rather than run or held until
its lease lapses, and a settings reload wakes the manager so a role
that gained task types starts claiming them straight away.
Starting the queue manager on every node also drains the queue channel
on nodes without outboundMta. Upstream left that channel unread, so
each message queued there parked a refresh in it, and by the code,
queueing would block once 1024 had piled up (not reproduced here).
A role object edit reaches the nodes that name that role in
INBUXA_ROLE. Moving a node to another role still means changing its
environment, and so a restart. Listener changes in a role still need a
restart too (listeners bind at boot); this change is about tasks and
delivery.
cluster::live_roles::live_role_tests (new; PostgreSQL, two nodes over
one store):
1. A node started with outboundMta delivers and runs a TLS report
task; after its role loses outboundMta and the settings reload, a
new message isn't attempted and a new report task stays pending;
with the role back, both are taken up.
2. A node started with no task type at all gains outboundMta: a
waiting message is attempted and a report task runs.
On main the test fails at step 1 ("delivery attempted without
outboundMta"); with step 1 bypassed, step 2 fails (nothing picked the
message up in 20 s).
|
||
|
|
ad58c35f39 | Merge pull request 'SQL queries time out; readiness follows the data store' (#45) from fix/query-timeouts into main | ||
|
|
08f29926d4 |
SQL queries time out; readiness follows the data store
Cluster rehearsal 3: with PostgreSQL paused (docker pause, so its kernel still answered TCP keepalives), requests on connections already checked out hung until it came back, and /healthz/ready stayed 200 through the outage. #41 bounded getting a connection, not using one. Client-side query limits (store::backend::query_timeout). Every operation on a PostgreSQL or MySQL connection now runs under a time limit. A server-side statement_timeout (or MySQL's MAX_EXECUTION_TIME, which covers SELECTs only) can't do this: the server that would enforce it is the one not answering. When an operation runs out, its connection is closed instead of pooled, since a query may still be in flight on it or a transaction open: deadpool's Object::take on PostgreSQL; Conn::disconnect on MySQL, which marks the connection closed before it sends anything, so the pool discards it even when the server never answers. - query, 2 minutes: reads, writes (the whole transaction with its retries), blobs, SQL lookups, search queries and indexing. These take milliseconds; two minutes leaves room for a large blob over a slow link and still ends a hang. - maintenance, 30 minutes: range deletes (account removal, purges), unindexing, purge_store, and creating tables and indexes at startup, which can legitimately run long in one statement. Their existing chunked fallback for server-side statement timeouts is unchanged. - iterate (exports, reindexing, maintenance scans) can run for hours, so the query limit bounds each wait for the database (preparing, the query starting, the next row) rather than the whole scan. The limits are fixed, like the pool timeouts; the DataStore schema has no field for them. Tests set them with Store::with_query_timeouts (test_mode only). Readiness. /healthz/ready answered 200 whenever a data store was configured. It now reads one key from the data store with a 2 s limit and reuses the answer for 2 s, so probes can't load the database; while one probe runs, others get the last answer. The first failed probe of an outage is logged. /healthz/live stays 200: restarting a node doesn't bring its database back, and an orchestrator restarting on failed liveness would restart every node at once. The container HEALTHCHECK already uses /healthz/live. Tests, store::pool_timeout (a proxy that stops forwarding while keeping connections open plays the paused database): - postgres_query_timeout, mysql_query_timeout (new): with four pooled connections open, a read, a scan and a write each fail with "Query timed out" 2.0 s after the pause (2 s test limit); once the proxy forwards again the store answers. With the limits set to an hour (upstream's behavior), the read was still waiting at the test's 20 s limit. - postgres_readiness (new, STORE=PostgreSql): a node's data store goes through the proxy; /healthz/ready is 200, 503 about 4 s after the pause while /healthz/live stays 200, and 200 again about 2 s after it ends. - postgres_pool_timeout, mysql_pool_timeout: pass as before. store::store_tests (PostgreSql, MySql, including the MariaDB statement timeout step) and store::task_locks (PostgreSql) pass; store::search_tests (PostgreSql) fails at the same ordering assertion (query.rs:684) as on main. |
||
|
|
fde43774b4 |
PostgreSQL search GIN indexes without a pending list
A three-node rehearsal on PostgreSQL saw searches take about 185 ms with 80 to 260 pages in the full-text indexes' pending lists, 2 to 6 ms right after gin_clean_pending_list() or VACUUM, then creep back up as mail came in. The search tables' GIN indexes were created with the default fastupdate=on: new entries wait in an unindexed pending list that every search scans in full until VACUUM (or 4 MB of backlog) merges it, and autovacuum only visits an insert-only table after thousands of inserts. The search GIN indexes are now created WITH (fastupdate = off), so an insert pays its index update at once. The schema step runs at every startup (create_search_tables, via SearchStore::create_indexes), so indexes made before this change are switched there: when an index's reloptions don't already turn fastupdate off, ALTER INDEX ... SET (fastupdate = off) and one gin_clean_pending_list() merge its backlog. The ALTER takes a SHARE UPDATE EXCLUSIVE lock, which blocks neither reads nor writes; after the first startup the step is one catalog read per index. A failure is logged and startup goes on (search still works, only slower). Per-table autovacuum settings for the search tables are left alone. The pending list was the only reason the insert threshold mattered for search; dead tuples and freezing are served by the defaults, and table settings would override whatever tuning the DBA has done. MySQL is unaffected: InnoDB FULLTEXT keeps new entries in an in-memory cache that queries read directly, with no setting like fastupdate. store::search_gin::postgres_gin_fastupdate (new, PostgreSQL) builds the search schema in a schema of its own and checks pg_class.reloptions: fastupdate=off on every GIN index of a fresh schema; then, with the option reset to the default and 500 rows pending, one startup turns it off everywhere and leaves no pending tuples (pgstatginindex); a second startup changes nothing. On main it fails at the first check. |
||
|
|
fcef4b1c3f |
Allowed IPs take the full settings reload after a write
write_reload_target sent AllowedIp writes to the blocked-IP reload, but
that reload rebuilds only BlockedIps. Allowed IPs are parsed into the
core's security settings (Security::parse), which only a full reload
rebuilds, so an AllowedIp write reported x:settingsReload applied: true
while the change wasn't live until the next full reload.
AllowedIp now maps to the full reload, like the other settings objects;
BlockedIp keeps its targeted reload.
system::auto_reload::settings_reload_tests now creates an allowed IP
over JMAP and checks that is_ip_allowed sees it with no ReloadSettings,
and that destroying it takes it out again. On main it fails ("allowed
IP not in the running settings").
|
||
|
|
6e50ba25a9 |
SQL pools time out; task locks are a renewed five-minute lease
A 3-node rehearsal (PostgreSQL + NATS + Garage) found two ways a crash leaves work stuck: Pool hangs. The PostgreSQL pool (deadpool) was built with no timeouts, so a request waited for a free connection, and for one to be opened or recycled, for as long as it took: forever when the server stopped answering. MySQL's pool (mysql_async) has no wait timeout at all. - PostgreSQL: wait 30 s (or the store's timeout if longer), create the store's timeout or 15 s (it bounds the whole handshake, where tokio-postgres's connect_timeout covers only the TCP connect), recycle 10 s. The pool config is now always set, not only with poolMaxConnections. - MySQL: every connection is taken through MysqlStore::conn(), which gives up after 30 s. - Both: TCP keepalive after 60 s idle, so a server that vanished without closing the connection is noticed in minutes rather than the two-hour system default. The DataStore schema has no pool timeout settings, so these are fixed defaults; the store's own timeout bounds connecting on PostgreSQL. Task locks. A task lock lasted an hour, so after a hard crash the dead node's tasks waited up to an hour and five minutes. The lock is now a five-minute lease: while this node runs a task, the task manager renews its lock every third of the lifetime (InMemoryStore::renew_lock, a compare-and-set on the store backends and SET XX EX on Redis, which leaves a lock that already expired alone). A killed node's tasks run elsewhere within about five minutes plus the claim recheck. A task this node holds isn't handed to a worker again by the scan. store::pool_timeout (new): a local listener that accepts connections and never answers plays a hung server; a PostgreSQL store with a 2 s timeout returns an error in about 4 s, and a MySQL store in 30 s. Without the timeouts both wait for good. store::task_locks gains a task held for 1.5 lock lifetimes: its lease is still held, and released when the task ends. |
||
|
|
1543ea5a9e |
Task manager: every task type follows the node's cluster role
A 3-node rehearsal found taskQueueProcessing didn't filter anything: roles.task_manager only decided whether the task manager started, and report, ACME, DKIM, DNS, calendar, thread-merge and restore tasks ran on any node with a task manager (manager.rs returned true for them). A node whose role left taskQueueProcessing off still ran them if it indexed or did maintenance. Every task type now answers to one ClusterTaskType (task_enabled): - IndexDocument, UnindexDocument, IndexTrace: searchIndexing - AccountMaintenance, TenantMaintenance, DestroyAccount: accountMaintenance - StoreMaintenance: storeMaintenance - SpamFilterMaintenance: spamClassifierTraining - DmarcReport, TlsReport: outboundMta. They build and send reports to other domains (TLS reports can go straight to an HTTPS endpoint), which is the outbound MTA's business. - CalendarAlarmEmail, CalendarAlarmNotification, CalendarItipMessage, MergeThreads, RestoreArchivedItem, AcmeRenewal, DkimManagement, DnsManagement: taskQueueProcessing, the role for queue tasks with no role of their own. A node that may not run a task leaves it unclaimed (no lock), so a node that may picks it up. The task manager also starts on a node whose only task role is outboundMta, so reports still run there. cluster::task_roles::task_role_tests (new, two task managers over one PostgreSQL store): node A (taskQueueProcessing only) runs a DNS task and leaves an unindex task and a TLS report pending; node B (searchIndexing and outboundMta) comes up and runs those two; a DNS task scheduled next stays pending on B and runs on A. On main node A runs the TLS report. |
||
|
|
2c684be5c9 |
Registry writes apply to the running settings without ReloadSettings
A 3-node rehearsal found that saving an MtaDeliverySchedule left it
unknown to the queue ("Queue strategy not found") until someone ran
x:Action ReloadSettings; only Directory and Authentication writes
reloaded (DIR-17). The admin UI has to remember a separate reload after
every save, and a script or API client that doesn't gets a server
running stale settings.
x:<Object>/set now reloads the running settings when it created,
updated or destroyed an object they are built from, and broadcasts the
same RegistryChange::Reload over the coordinator as ReloadSettings, so
every node applies it:
- Settings objects (MTA, spam filter, listeners, tracers, Sieve system
scripts, cluster roles, directories, ...: the object types the core,
telemetry, listener and directory builders read) get a full reload.
- Certificates, lookup stores and blocked/allowed IPs get their own
targeted reloads.
- Accounts, domains, roles and other data read as needed, stores (they
take a restart) and applications (their own reload action) get none.
Full reloads are coalesced: a write waits for a reload that started
after it was stored and joins one if it can, so a burst of writes, or
a request with many objects, costs one or two reloads, not one each.
The write itself is never undone. When the reload is refused (build
errors in objects that were working, the rule from the previous
commit), the set response says so in a new x:settingsReload field,
{"applied": false, "description": "Saved, but the running settings
were not reloaded. <object>: <error>"}; {"applied": true} otherwise.
The field is absent when the write needs no reload. The description
helper is shared with ReloadSettings' refusal.
Each reload sends the queue a ReloadSettings event, so the SMTP test
harness's read_event, try_read_event and assert_no_events now pass over
those; expect_reload_settings still waits for one.
system::auto_reload::settings_reload_tests (new): an MtaVirtualQueue
and an MtaDeliverySchedule created over JMAP are in the running
settings with no ReloadSettings, and gone once destroyed; eight
concurrent creates all land; a write whose reload fails is stored and
reported applied: false with the error; a domain write carries no
x:settingsReload. On main the new schedule is missing. The cluster
broadcast test (three nodes, PostgreSQL + NATS) now checks that every
node has a schedule created on node 0 without a reload.
|
||
|
|
999ae12cc7 |
Settings reload: no DNS at build time, don't refuse over old failures
A 3-node rehearsal found every settings reload refused, cluster-wide,
because one node couldn't resolve the Pyzor server:
- PyzorConfig::parse resolved the host while building the settings and
made a failed lookup a build error. It now keeps the host and port and
resolves when a message is checked (an IP address is used as is, a
name is reused for five minutes, the lookup counts against the Pyzor
timeout). A failure there is a Pyzor error for that message.
- A milter's hostname was resolved the same way, with a blocking
to_socket_addrs in async code. An IP address is kept; a name is now
resolved on each connection.
Other build-time I/O is already non-fatal: directories that can't
connect become unavailable with a warning (DIR-21), and the AI model
locality check only warns.
reload_registry swapped the core only when the whole build was free of
errors, while boot runs with whatever built. One failing object thus
refused every later reload, and the running settings went stale. Now a
reload is refused only for errors in objects that built when the
running settings were built (at boot or by the last applied reload):
applying it would lose those. Objects that already failed then are
missing from the running settings anyway, as at boot, so their errors
are logged and returned as known_errors but don't hold the reload back.
Refusing on new errors keeps a bad edit from taking a working object
out of service; the admin gets the error instead.
ReloadSettings now says "Settings were not reloaded." and names the
object and its error ("Tracer with id ...: Only one console tracer is
allowed"), with a count of any further errors. A refused reload after a
directory change logs its errors too.
system::reload::reload_tests (new): with Pyzor enabled on an
unresolvable host, ReloadSettings succeeds (on main it fails with
"Invalid address: failed to lookup address information"); an IP host
needs no lookup; a new build error refuses the reload, names the object
and leaves the running settings unchanged; the same error, once known
from the running settings' build, no longer blocks; once fixed, a new
error there blocks again. smtp::inbound::milter's session test now
names its milter "localhost", so the connect-time lookup is exercised.
|
||
|
|
639a415a4f |
Search: find addresses by local part, domain or name on PostgreSQL and MySQL
A 3-node PostgreSQL rehearsal found IMAP SEARCH FROM "noreply" matched 0-2 messages where RocksDB matched 23 of 930. The message indexer hands each address and display name of From/To/Cc/Bcc to the search store as keyword text (Language::None). The built-in index splits keyword text into lowercase runs of alphanumerics, so an address is found by its full form, its local part, its domain or a display-name word. The SQL backends didn't: - PostgreSQL's text parser keeps "[email protected]" as one email token (host names and URLs likewise), so neither "noreply" nor "amazon.com" ever matched it. Keyword text is now split the same way as the built-in index (SpaceTokenizer) before to_tsvector on insert and before plainto_tsquery/phraseto_tsquery on search, still under the 'simple' configuration, so the GIN index keeps serving the query. The sort columns keep the raw text. - MySQL's FULLTEXT parser already splits on punctuation, but InnoDB never indexes its stopwords ("com", "de", "www", ...) or words under innodb_ft_min_token_size (3), and a required +word it hasn't indexed matches no row. So "amazon.com", "[email protected]" or "jane doe" found nothing. Those words are now matched with a word-boundary REGEXP on the rows the indexed words select. In language text (bodies, subjects) they are dropped when other words remain, and only checked when nothing else is left, so "the invoice" no longer finds nothing either. Existing PostgreSQL search indexes hold the old single-token vectors and need a reindex (the reindexAccounts task) before address searches find old messages. MySQL needs none: only the query changed. store::search_tests gains test_address_search: five messages, 28 FROM/TO/CC/BCC searches by full address, local part, domain, domain labels, display name and hyphenated local part, plus a TEXT-style OR, with the same expected ids on every backend. It passes on RocksDB, SQLite, PostgreSQL and MySQL; on main it fails on PostgreSQL (From "noreply") and MySQL (From "[email protected]"). |
||
|
|
95f0445d83 |
Coordinator: join the cluster when NATS comes up, report the connection
A node that started while NATS was down never got a coordinator. The
connect failed at boot, bootstrap recorded a build error and the node ran
with Coordinator::None until restarted. It had no broadcast subscriber
or publisher, so cross-node push and cache invalidation to it stayed
broken, and its healthcheck said nothing about it. Losing NATS after
startup was silent too.
- The NATS client now connects in the background
(retry_on_initial_connect): startup never waits on NATS or fails over
it, the node gets its coordinator, subscriber and publisher at once,
and the client keeps trying (async-nats's backoff, at most 4 s apart)
until NATS answers. Subscriptions made meanwhile start delivering when
it does. A configured maxReconnects still ends the attempts.
- Three new events report the connection: cluster.coordinator-connected
(info), cluster.coordinator-disconnected (warn: lost, closed, gave up,
or not connected within the connection timeout at startup) and
cluster.coordinator-error (warn: a failed attempt, reported once per
outage rather than every retry, and server errors, slow consumers and
lame duck mode). They are in the packaged schema, ids 644 to 646.
- GET /healthz/cluster reports the coordinator: 200
{"coordinator":"connected"}, 503 {"coordinator":"disconnected"}, or
200 with "none" (no coordinator) or "unknown" (a backend that doesn't
track its connection). /healthz/live and /healthz/ready are unchanged
on purpose: a node without its coordinator still serves mail, and
failing those would have orchestrators restart, or pull out of
service, every node at once whenever NATS is down.
Only NATS connects lazily; the other coordinator backends still fail at
boot as before.
cluster::coordinator::coordinator_reconnect_tests starts a node against a
NATS port with nothing behind it, checks it boots with a coordinator and
reports it disconnected, subscribes, then starts NATS on that port: the
node connects on its own and the subscription receives a message from a
second client. Stopping and restarting NATS shows disconnected, then
connected, and the same subscription keeps working.
|