A cluster rehearsal (PostgreSQL + NATS) left index tasks pending well
past the one-hour task lock after the node that claimed them was stopped
or killed. The exact cause there isn't confirmed; this closes every path
found in the task manager that stretches a takeover past the lock, or
keeps a task claimed without running it:
- A graceful stop never released the locks it held, so every task the
node had claimed stayed blocked for an hour. The server now tracks the
locks it holds (common::ipc::TaskLocks) and, once the shutdown signal
arrives, stops claiming and releases them before exiting.
- A node that failed to claim a task (another node held it) set its own
local hold for a full lock lifetime from that scan. If the holder
claimed it just after the scan began, or ran on a clock ahead, that
hold ran out a moment before the lock did and was set for another
hour: two hours in all. Such claims are now tried again every five
minutes (a twelfth of the lock lifetime), and the task manager wakes
up for them: before, a node without a coordinator could sleep up to
five minutes past the recheck, or until something else woke it.
- A worker that panicked took its task type down on that node for good,
while the scan kept claiming that type's tasks and failing to hand them
over, re-taking each lock as it expired and so starving every other
node of them. Each batch now runs on a task of its own; a panic is
logged, the batch's locks are released and the worker carries on. A
failed hand-over releases the lock too.
- A claimed task the worker couldn't read, or found gone, kept its lock
for the hour. It is released.
- An IndexDocument task for a file (not indexed) returned no result,
which shifted every later result in the batch onto the wrong task in
update_tasks. It returns Ignored. Nothing queues such a task today.
The lock lifetime stays one hour; it now lives per server so the tests
can shorten it.
store::task_locks::task_lock_tests plays a second node by writing its
locks straight into the in-memory store: tasks it claimed and abandoned
run here once its locks expire, including locks that outlive this node's
view of them, and a graceful stop hands this node's locks back at once
and claims nothing more. It passes on RocksDB, SQLite and PostgreSQL.
With the old recheck it fails.
These upstream files were changed after the fork marked the files it had
modified, and never got the notice: six by the listener and schema-cache
work on 2026-09-20, two by the name check. Found by diffing against the
upstream snapshot branch, as before.
Five conflicts, resolved:
- crates/common/src/auth/authentication.rs: upstream's get_directory_for_token
and JwtClaims replace extract_jwt_domain; the per-domain directory code
(DIR-1, DIR-5 to DIR-7) is kept, and the token lookup routes through it.
The release's one new Enterprise snippet was the body of
get_directory_for_issuer, which stays returning None: a token naming no
address gets the server default, as DIR-2 specifies and as v0.16.22 did.
- crates/common/src/manager/application.rs: upstream's rewrite of the tests,
with the temp directory names renamed again, and the 5(a) notice the
name-purge change should have added.
- crates/common/src/network/mta.rs: both sides' imports.
- crates/main/Cargo.toml: the AGPL-only license kept, version 0.16.23.
- Cargo.lock: upstream's, with the fork's crates added by Cargo.
The join: the policy decides, features owns the listener objects,
ListenerControl owns the running sockets, and only Server has both.
Server::set_protocol_policy is what a click performs. It applies the locks
to what was asked before storing anything (LP-21), so what is recorded is
what the server allows. Closing removes each listener object and then stops
its socket; opening puts the object back and then spawns it. The order is
the point in both directions -- a socket stopped while its object remains
returns on the next restart, and a socket spawned before its object exists
has nothing to come back to.
saved_listeners is carried over from the stored policy rather than taken
from the request. A client never sets it, and a /set that omitted it would
otherwise lose the listeners still waiting to come back.
Putting a listener back has to bind a fresh socket, so it re-parses from
the registry -- the objects are already back by then -- rather than trying
to revive the saved one. Only main knows which session manager a protocol
wants, so it leaves a spawner behind at startup and spawn_listener is now
shared between that and the initial spawn. Without a spawner a restored
listener is reported as pending a restart rather than promised, which is
what the test servers will see.
A listener that cannot be put back does not stop the others and stays
saved for another try (LP-5).
Still nothing an operator can reach: no JMAP method calls this yet, and no
sign-in is refused. What it does do is close and reopen a port on a
running server, which is the part that did not exist this morning.
The legacy-protocols switch has to close the IMAP, POP3 and ManageSieve
ports and leave everything else accepting. The server could not do that.
Two findings from the source, both now recorded in the spec. A settings
reload never closes a port: cache/reload.rs parses the listeners only to
collect configuration errors and drops the result, and sockets are bound
once at startup through init.servers.spawn in main.rs. And there is only
one shutdown signal -- Listeners::spawn makes a single watch channel and
hands every listener a clone -- so the one thing the server could do was
stop all of them at once, port 25 included. That answers the spec's open
question 1, and the answer was neither of the two it offered.
So each listener gets its own channel. ListenerControl holds the sending
ends keyed by listener id; firing one breaks that accept loop, which drops
its TcpListener and closes the socket. The accept loop itself is unchanged
-- it already did the right thing, it just had no way to be told about one
listener. stop_matching takes a predicate and a keep list, because the
inbound listener shares its protocol with submission and telling them
apart is the caller's job (LP-3), not this registry's.
spawn_with_control is a second method rather than a change to spawn. The
registry owns the senders, so a dropped registry would stop every listener
at once; the four test callers pass no registry and keep the old shared
channel exactly as it was.
Whole-server shutdown now fires the per-listener channels too, since the
returned sender no longer reaches them.
No policy, no JMAP and no screen yet: this is only the mechanism, with
seven tests over stopping one, stopping many, sparing port 25 and sparing
submission. It closes no port on its own, and it does not touch the host's
firewall or any port-forward -- that is LP-20, and stays the operator's.
Upstream commit: 474dd0229cb20cf513036619781ed97bd8073c3f
Enterprise-only files removed or emptied: 63
Enterprise-only snippets removed: 117 in 50 files
Dangling module declarations removed: 5
Cargo edits turning enterprise off: 14
Verification: clean
Enterprise feature gates left for rebuilt features: 19 in 18 files
Produced by tools/fork/strip.py. The full report is in docs/fork/strip-reports/ on main.