9653219c53fc4f0c253a4b5158dc2b8e44e6e96e
16
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e35fc3e6d6 | Ports: each node checks the others' ports from outside | ||
|
|
a8fb10458b |
Evaluate the personal-data catalog: the data inventory and its history
Personal-data catalog spec, §6 (Phase 3c). inbuxa:DataInventory/get evaluates the catalog against the server's live settings and says what this server holds: for each source and each object that can hold personal data, its categories and whose data it is, whether it is collected here at all, what bounds its retention (the live value of the setting that does, or unbounded), whether it leaves the host and to which endpoints, and a summary. Every host that receives something is listed once as a candidate processor with what it receives. Inside a tenant it answers with the tenant's slice and none of the server's processors. Read-only, with sysComplianceGet. inbuxa:InventorySnapshot/get is the history: a dated copy of the evaluated inventory, recorded when it changes -- after a registry write to an object the inventory reads, after inbuxa's log, audit or AI settings change, and on the daily clean-up -- and kept as long as the audit log's records. ids: null lists every snapshot, newest first; the full inventory only when asked for. The catalog is embedded and parsed at start (new dependency: toml, MIT/Apache); the evaluation is a pure function of it and the live facts, so each configuration is tested without a server. Loopback endpoints stay on the host; any other configured endpoint leaves it. Tested: unit tests for the evaluation (a new install's defaults, an external blob store, a hosted AI endpoint, telemetry off, a tenant's slice, hosts from URLs, loopback), snapshots, and the fact gathering's store and duration rules; the compliance system test, extended (the officer reads the inventory, a plain user is refused, a tenant's officer sees its slice and no processors, a webhook to another host becomes a processor and a snapshot names x:WebHook, a retention change reads through); the system, audit, legal hold and account lock suites; fork checks. The system suite failed once of three runs with an email import's blob not found, in antispam.rs; the same happened once in purge.rs on the previous branch. Nothing here touches uploads; noted for a separate look. |
||
|
|
318783f444 |
Legal holds, step 2: who a hold covers
A hold reaches an account by name, through any of its addresses' domains, its groups or its tenant, as they are now, so an account added to a held domain later is held too. An account that leaves a held domain, group or tenant stays held: the registry write hook adds it to the hold by name on every account change, whoever makes it (LH-2). Server::holds_on answers for the deletion paths, from the store each time so a hold binds every node at once. |
||
|
|
86d7ebd982 |
Audit log: a permanent, tamper-evident record of admin actions
What administrators and the server itself do to the control plane is now recorded, from inbuxa-drafts/specs/audit-hold-lock.md (AU-1 to AU-12): settings, accounts, domains, roles and every other registry change, with each field's before and after (secrets only as "changed"); the fork's own settings objects; administrator sign-ins (and failed ones to administrator accounts), master-user and recovery-admin sign-ins, once an hour per account, method and address; access to another account's data through impersonation or FetchAnyBlob, once an hour; exports and tamper checks; and registry writes the server makes on its own, named by subsystem (system:AcmeRenewal, system:auto-ban, system:directory-sync, ...), with a spam rules update as one summary record. No change without its record (AU-3): before a set method changes anything, a pending record per requested create, update and destroy is written; if that fails, the method is refused with serverFail. Its outcome follows as a later entry. A change interrupted by a crash stays "unfinished". Records live in the fork's subspace under L, as one SHA-256 hash chain per node. The chain's head is stored, never cached, and every append asserts it, so two writers can't take the same place. Nothing can edit or delete a record; the daily purge removes the oldest past the retention (default 730 days, minimum 90) and records where the chain now starts, so verification still passes. security.audit-recorded (647) copies each record to webhooks, OpenTelemetry and the log; security.audit-write-failed (648) reports a failed write. New JMAP objects under urn:inbuxa:jmap: inbuxa:AuditEvent/get and /query (filters: time, actor, action, target, account, tenant, outcome, address, text), inbuxa:AuditSettings, inbuxa:AuditExport (CSV or JSON Lines built on the server, each line with its chain hash, ending in a manifest; the created object names the blob and its SHA-256) and inbuxa:AuditVerification. New permissions sysAuditGet, sysAuditExport and sysAuditSettingsUpdate: the Administrator role gets all three, the Tenant Administrator role gets read and export, once, on existing installs too. A tenant administrator sees records whose actor or target is in its tenant, including a server administrator's changes there. Sign-in method on the session: access tokens now remember how they signed in (password, app password, API key, OAuth client, directory, master user, recovery admin), including across the HTTP credential cache. New OAuth access tokens carry their client id in the sealed claims; older ones show as client "unknown" until they expire. The schema gains the permissions, the two events and a Management > Compliance > Audit Log link. Stack: the request layer boxes every inner future where it's made. Without that, a debug build overflowed the default 2 MB worker stack on a registry set; measured with the same request, the branch and main now overflow at the same stack size (between 1856 and 1920 KiB, debug), so the layer adds nothing measurable. Tests: unit tests in inbuxa-features and jmap; system::audit::audit_log_tests (run with --ignored) passes on RocksDB, SQLite, PostgreSQL, PostgreSQL with a read replica, MySQL, MySQL with a replica and FoundationDB. The system, JMAP and SCIM suites pass. authorization.rs skipped fork permissions that guard no registry object; the audit suite checks a plain user is refused instead. |
||
|
|
08f29926d4 |
SQL queries time out; readiness follows the data store
Cluster rehearsal 3: with PostgreSQL paused (docker pause, so its kernel still answered TCP keepalives), requests on connections already checked out hung until it came back, and /healthz/ready stayed 200 through the outage. #41 bounded getting a connection, not using one. Client-side query limits (store::backend::query_timeout). Every operation on a PostgreSQL or MySQL connection now runs under a time limit. A server-side statement_timeout (or MySQL's MAX_EXECUTION_TIME, which covers SELECTs only) can't do this: the server that would enforce it is the one not answering. When an operation runs out, its connection is closed instead of pooled, since a query may still be in flight on it or a transaction open: deadpool's Object::take on PostgreSQL; Conn::disconnect on MySQL, which marks the connection closed before it sends anything, so the pool discards it even when the server never answers. - query, 2 minutes: reads, writes (the whole transaction with its retries), blobs, SQL lookups, search queries and indexing. These take milliseconds; two minutes leaves room for a large blob over a slow link and still ends a hang. - maintenance, 30 minutes: range deletes (account removal, purges), unindexing, purge_store, and creating tables and indexes at startup, which can legitimately run long in one statement. Their existing chunked fallback for server-side statement timeouts is unchanged. - iterate (exports, reindexing, maintenance scans) can run for hours, so the query limit bounds each wait for the database (preparing, the query starting, the next row) rather than the whole scan. The limits are fixed, like the pool timeouts; the DataStore schema has no field for them. Tests set them with Store::with_query_timeouts (test_mode only). Readiness. /healthz/ready answered 200 whenever a data store was configured. It now reads one key from the data store with a 2 s limit and reuses the answer for 2 s, so probes can't load the database; while one probe runs, others get the last answer. The first failed probe of an outage is logged. /healthz/live stays 200: restarting a node doesn't bring its database back, and an orchestrator restarting on failed liveness would restart every node at once. The container HEALTHCHECK already uses /healthz/live. Tests, store::pool_timeout (a proxy that stops forwarding while keeping connections open plays the paused database): - postgres_query_timeout, mysql_query_timeout (new): with four pooled connections open, a read, a scan and a write each fail with "Query timed out" 2.0 s after the pause (2 s test limit); once the proxy forwards again the store answers. With the limits set to an hour (upstream's behavior), the read was still waiting at the test's 20 s limit. - postgres_readiness (new, STORE=PostgreSql): a node's data store goes through the proxy; /healthz/ready is 200, 503 about 4 s after the pause while /healthz/live stays 200, and 200 again about 2 s after it ends. - postgres_pool_timeout, mysql_pool_timeout: pass as before. store::store_tests (PostgreSql, MySql, including the MariaDB statement timeout step) and store::task_locks (PostgreSql) pass; store::search_tests (PostgreSql) fails at the same ordering assertion (query.rs:684) as on main. |
||
|
|
2c684be5c9 |
Registry writes apply to the running settings without ReloadSettings
A 3-node rehearsal found that saving an MtaDeliverySchedule left it
unknown to the queue ("Queue strategy not found") until someone ran
x:Action ReloadSettings; only Directory and Authentication writes
reloaded (DIR-17). The admin UI has to remember a separate reload after
every save, and a script or API client that doesn't gets a server
running stale settings.
x:<Object>/set now reloads the running settings when it created,
updated or destroyed an object they are built from, and broadcasts the
same RegistryChange::Reload over the coordinator as ReloadSettings, so
every node applies it:
- Settings objects (MTA, spam filter, listeners, tracers, Sieve system
scripts, cluster roles, directories, ...: the object types the core,
telemetry, listener and directory builders read) get a full reload.
- Certificates, lookup stores and blocked/allowed IPs get their own
targeted reloads.
- Accounts, domains, roles and other data read as needed, stores (they
take a restart) and applications (their own reload action) get none.
Full reloads are coalesced: a write waits for a reload that started
after it was stored and joins one if it can, so a burst of writes, or
a request with many objects, costs one or two reloads, not one each.
The write itself is never undone. When the reload is refused (build
errors in objects that were working, the rule from the previous
commit), the set response says so in a new x:settingsReload field,
{"applied": false, "description": "Saved, but the running settings
were not reloaded. <object>: <error>"}; {"applied": true} otherwise.
The field is absent when the write needs no reload. The description
helper is shared with ReloadSettings' refusal.
Each reload sends the queue a ReloadSettings event, so the SMTP test
harness's read_event, try_read_event and assert_no_events now pass over
those; expect_reload_settings still waits for one.
system::auto_reload::settings_reload_tests (new): an MtaVirtualQueue
and an MtaDeliverySchedule created over JMAP are in the running
settings with no ReloadSettings, and gone once destroyed; eight
concurrent creates all land; a write whose reload fails is stored and
reported applied: false with the error; a domain write carries no
x:settingsReload. On main the new schedule is missing. The cluster
broadcast test (three nodes, PostgreSQL + NATS) now checks that every
node has a schedule created on node 0 without a reload.
|
||
|
|
999ae12cc7 |
Settings reload: no DNS at build time, don't refuse over old failures
A 3-node rehearsal found every settings reload refused, cluster-wide,
because one node couldn't resolve the Pyzor server:
- PyzorConfig::parse resolved the host while building the settings and
made a failed lookup a build error. It now keeps the host and port and
resolves when a message is checked (an IP address is used as is, a
name is reused for five minutes, the lookup counts against the Pyzor
timeout). A failure there is a Pyzor error for that message.
- A milter's hostname was resolved the same way, with a blocking
to_socket_addrs in async code. An IP address is kept; a name is now
resolved on each connection.
Other build-time I/O is already non-fatal: directories that can't
connect become unavailable with a warning (DIR-21), and the AI model
locality check only warns.
reload_registry swapped the core only when the whole build was free of
errors, while boot runs with whatever built. One failing object thus
refused every later reload, and the running settings went stale. Now a
reload is refused only for errors in objects that built when the
running settings were built (at boot or by the last applied reload):
applying it would lose those. Objects that already failed then are
missing from the running settings anyway, as at boot, so their errors
are logged and returned as known_errors but don't hold the reload back.
Refusing on new errors keeps a bad edit from taking a working object
out of service; the admin gets the error instead.
ReloadSettings now says "Settings were not reloaded." and names the
object and its error ("Tracer with id ...: Only one console tracer is
allowed"), with a count of any further errors. A refused reload after a
directory change logs its errors too.
system::reload::reload_tests (new): with Pyzor enabled on an
unresolvable host, ReloadSettings succeeds (on main it fails with
"Invalid address: failed to lookup address information"); an IP host
needs no lookup; a new build error refuses the reload, names the object
and leaves the running settings unchanged; the same error, once known
from the running settings' build, no longer blocks; once fixed, a new
error there blocks again. smtp::inbound::milter's session test now
names its milter "localhost", so the connect-time lookup is exercised.
|
||
|
|
7c80a12d75 |
Task manager: release task locks on stop, recheck claims held elsewhere
A cluster rehearsal (PostgreSQL + NATS) left index tasks pending well past the one-hour task lock after the node that claimed them was stopped or killed. The exact cause there isn't confirmed; this closes every path found in the task manager that stretches a takeover past the lock, or keeps a task claimed without running it: - A graceful stop never released the locks it held, so every task the node had claimed stayed blocked for an hour. The server now tracks the locks it holds (common::ipc::TaskLocks) and, once the shutdown signal arrives, stops claiming and releases them before exiting. - A node that failed to claim a task (another node held it) set its own local hold for a full lock lifetime from that scan. If the holder claimed it just after the scan began, or ran on a clock ahead, that hold ran out a moment before the lock did and was set for another hour: two hours in all. Such claims are now tried again every five minutes (a twelfth of the lock lifetime), and the task manager wakes up for them: before, a node without a coordinator could sleep up to five minutes past the recheck, or until something else woke it. - A worker that panicked took its task type down on that node for good, while the scan kept claiming that type's tasks and failing to hand them over, re-taking each lock as it expired and so starving every other node of them. Each batch now runs on a task of its own; a panic is logged, the batch's locks are released and the worker carries on. A failed hand-over releases the lock too. - A claimed task the worker couldn't read, or found gone, kept its lock for the hour. It is released. - An IndexDocument task for a file (not indexed) returned no result, which shifted every later result in the batch onto the wrong task in update_tasks. It returns Ignored. Nothing queues such a task today. The lock lifetime stays one hour; it now lives per server so the tests can shorten it. store::task_locks::task_lock_tests plays a second node by writing its locks straight into the in-memory store: tasks it claimed and abandoned run here once its locks expire, including locks that outlive this node's view of them, and a graceful stop hands this node's locks back at once and claims nothing more. It passes on RocksDB, SQLite and PostgreSQL. With the old recheck it fails. |
||
|
|
fee6b74e79 |
Listeners can be stopped one at a time, which LP-2 needs
The legacy-protocols switch has to close the IMAP, POP3 and ManageSieve ports and leave everything else accepting. The server could not do that. Two findings from the source, both now recorded in the spec. A settings reload never closes a port: cache/reload.rs parses the listeners only to collect configuration errors and drops the result, and sockets are bound once at startup through init.servers.spawn in main.rs. And there is only one shutdown signal -- Listeners::spawn makes a single watch channel and hands every listener a clone -- so the one thing the server could do was stop all of them at once, port 25 included. That answers the spec's open question 1, and the answer was neither of the two it offered. So each listener gets its own channel. ListenerControl holds the sending ends keyed by listener id; firing one breaks that accept loop, which drops its TcpListener and closes the socket. The accept loop itself is unchanged -- it already did the right thing, it just had no way to be told about one listener. stop_matching takes a predicate and a keep list, because the inbound listener shares its protocol with submission and telling them apart is the caller's job (LP-3), not this registry's. spawn_with_control is a second method rather than a change to spawn. The registry owns the senders, so a dropped registry would stop every listener at once; the four test callers pass no registry and keep the old shared channel exactly as it was. Whole-server shutdown now fires the per-listener channels too, since the returned sender no longer reaches them. No policy, no JMAP and no screen yet: this is only the mechanism, with seven tests over stopping one, stopping many, sparing port 25 and sparing submission. It closes no port on its own, and it does not touch the host's firewall or any port-forward -- that is LP-20, and stays the operator's. |
||
|
|
6a53d47106 |
Mark the files this fork changed (AGPL section 5(a))
The AGPL asks a modified version to carry prominent notices saying it was modified, and giving a date. Publishing the source is the conveyance that asks for it, so it wants doing before the repository is public rather than at the release. Every upstream file the fork changed now says so in its header, beneath the notice it came with: 164 files, found by diffing against the upstream snapshot branch rather than by guessing, so the list is what actually differs. Files the fork wrote itself already carry their own copyright and need nothing. Upstream's notices are untouched, which its licence requires and which was already true. The README says the same thing in prose, since the obligation is on the work as a whole and not only its Rust files. Builds unchanged: the server and the test binary both compile. |
||
|
|
0755fad51e |
Scale-out storage: IMAP CONDSTORE raises the mark, and the two-node check (ST-7, test 11)
A FETCH with CHANGEDSINCE presents its mod-sequence to the read scope, so a replica must have that change before it answers, as a JMAP sinceState already did. replica_cluster_tests covers test 11: with the replica's replay paused, a write on one node is read back through a second store with its own marks, sharing through Redis. Composite stores nest store futures deeply enough to pass rustc's default query depth once both postgres and redis are compiled in, so the server crates raise their recursion limit. |
||
|
|
9490fc4677 |
AI spam classification: the model's opinion as one bounded spam signal, and the llm_prompt Sieve function (AI-1 to AI-28)
The classifier sends only the subject and text, between unforgeable markers after the operator's prompt, to an OpenAI-compatible endpoint the operator configured; nothing is preset. Its answer maps to an LLM_ tag whose score is clamped (+5.0, -1.0 by default) and can never discard or reject on its own; X-Spam-LLM is sanitized, encoded and folded, and a planted one is removed. Failures, timeouts past the ceiling, a full slot or a paused model leave mail flowing untagged. llm_prompt answers trusted scripts, and accounts holding interactAi within an hourly limit. Redirects aren't followed and no content or secret is logged. The limits live in inbuxa:AiLimits. Acceptance tests 1 and 3 to 21; test 2 as the re-enabled shared llm case, whose setup no longer waits on a rules file from a developer's own path; test 22 written as the ignored ai_compat. |
||
|
|
faedf7a1da |
Versioning: INBUXA's own dated version, with the Stalwart base shown
inbuxa --version, the banner, startup events, OpenTelemetry and the JMAP implementation string read "2026.9.18 (Stalwart 0.16.22)". The base comes from Cargo, which keeps following upstream so version bumps merge cleanly. Received headers and IMAP ID carry the INBUXA version. |
||
|
|
c9c761fab5 |
Quick fixes from the first boot: recovery admin, upsell, warnings
- The recovery administrator (INBUXA_RECOVERY_ADMIN, or STALWART_RECOVERY_ADMIN) is honored only in bootstrap and recovery mode. On a configured server it's ignored with a startup warning. Before, it was a standing full-admin login for as long as the variable stayed set. - The Enterprise upsell error is replaced by "This feature isn't available in INBUXA yet" for the features still to be rebuilt. - Workspace warnings: 25 to 0. cargo fix removed the unused imports. The seven places where Enterprise code used to plug in keep their parameters, each with an inbuxa: comment naming the rebuild that uses it again. The antispam test's mock-server imports are back behind pending-rebuild. |
||
|
|
89665decfa |
Rebrand to INBUXA: product name, protocol greetings, sign-in page, calendar templates, logos, README
One branding module (types::brand!) supplies the name to every protocol greeting, the HTTP realm, the startup banner, Received headers, the user agent, the IMAP ID response and calendar PRODIDs. The sign-in and calendar pages take ihasmail's palette and font and the INBUXA logo, and calendar emails embed the INBUXA lockup. Default calendar and address book names follow. Protocol identifiers (urn:stalwart:jmap, vnd.stalwart Sieve extensions) and upstream copyright notices are unchanged. |
||
|
|
7dae9b29fd |
Import upstream v0.16.22, stripped
Upstream commit: 474dd0229cb20cf513036619781ed97bd8073c3f Enterprise-only files removed or emptied: 63 Enterprise-only snippets removed: 117 in 50 files Dangling module declarations removed: 5 Cargo edits turning enterprise off: 14 Verification: clean Enterprise feature gates left for rebuilt features: 19 in 18 files Produced by tools/fork/strip.py. The full report is in docs/fork/strip-reports/ on main. |