Report reschedules keep the task queue readable
ci / fork-checks (pull_request) Successful in 52s
ci / build (pull_request) Successful in 4m0s

Setting deliverAt on an internal DMARC or TLS report wrote the new task
queue row with the report's object type (0x21, 0x6e) instead of the task
type (7, 8), and left the task row at its old due. The task manager's scan
failed on that row with store.data-corruption ("Failed to iterate over task
queue"), and because the error ended the whole scan, every task due after
the row stopped running on every node.

- reschedule_ops writes the new queue row through schedule_task_with_id, so
  it carries the task type and the task row gets the new due. It removes
  the row the task is actually queued under (the task's due, which differs
  from deliverAt once the task has been retried) and any row an earlier
  reschedule left at deliverAt.
- x:DmarcInternalReport/set and x:TlsInternalReport/set lock the report's
  task while they move it, as x:Task/set does, refuse while the report is
  being sent, release the locks however the request ends, and wake the task
  manager.
- The task manager logs a queue row it can't read (id, due, key, value) and
  skips it instead of ending the scan. It then repairs the row from its task:
  the row is rewritten with the task's type, and a row with no task behind
  it is removed. A row holding a report's object type for a report task is
  what the old reschedule wrote: the task is moved to that row's time, as
  the reschedule intended, and its old queue row is removed. Stores that
  already hold such a row recover on their own once it comes due.
- x:Task/query with a type filter skips an unreadable row instead of
  failing.

Test: smtp::reporting::reschedule (RocksDB and PostgreSQL). It fails on
main: x:Task/get shows the old due, and with that check removed, neither
report nor a later task ever runs.
This commit is contained in:
2026-09-24 18:09:44 -07:00
parent b90a7f173e
commit 1a7859a8cc
6 changed files with 602 additions and 33 deletions
+63 -5
View File
@@ -2,6 +2,8 @@
* SPDX-FileCopyrightText: 2020 Stalwart Labs LLC <[email protected]>
*
* SPDX-License-Identifier: AGPL-3.0-only OR LicenseRef-SEL
*
* Modified by Coffey Labs in 2026 for INBUXA.
*/
use crate::{
@@ -15,22 +17,42 @@ use jmap_proto::{error::set::SetError, types::state::State};
use jmap_tools::{Key, Value};
use registry::{
jmap::IntoValue,
schema::prelude::{Object, ObjectInner, ObjectType, Property},
schema::{
prelude::{Object, ObjectInner, ObjectType, Property},
structs::Task,
},
types::{EnumImpl, datetime::UTCDateTime},
};
use services::task_manager::lock::TaskLockManager;
use smtp::reporting::index::{ExternalReportIndex, InternalReportIndex};
use std::str::FromStr;
use store::{
U64_LEN, ValueKey,
registry::{RegistryFilter, RegistryFilterValue, RegistryQuery},
write::{BatchBuilder, RegistryClass, ValueClass, key::KeySerializer},
write::{BatchBuilder, RegistryClass, TaskQueueClass, ValueClass, key::KeySerializer},
};
use trc::AddContext;
use types::id::Id;
pub(crate) async fn report_set(
mut set: RegistrySetResponse<'_>,
set: RegistrySetResponse<'_>,
) -> trc::Result<RegistrySetResponse<'_>> {
// inbuxa: task locks taken to reschedule reports are released however
// the request ends; a held lock is renewed, so a leaked one would keep
// the report's task from ever running
let server = set.server;
let mut locked_tasks = Vec::new();
let result = report_set_locked(set, &mut locked_tasks).await;
for task_id in locked_tasks {
server.remove_index_lock(task_id).await;
}
result
}
async fn report_set_locked<'x>(
mut set: RegistrySetResponse<'x>,
locked_tasks: &mut Vec<u64>,
) -> trc::Result<RegistrySetResponse<'x>> {
let object_id = set.object_type.to_id();
// Reports cannot be created
@@ -89,12 +111,45 @@ pub(crate) async fn report_set(
.get_value::<Object>(ValueKey::from(key.clone()))
.await?
{
// inbuxa: the report's task shares its id. Hold the task
// while its queue rows move, as x:Task/set does, and move the
// row the task is actually queued under
if !set.server.try_lock_task(item_id).await {
set.response.not_updated.append(
id,
SetError::forbidden().with_description(
"The report is being sent and cannot be rescheduled".to_string(),
),
);
continue;
}
locked_tasks.push(item_id);
let queued = set
.server
.store()
.get_value::<Task>(ValueKey::from(ValueClass::TaskQueue(
TaskQueueClass::Task { id: item_id },
)))
.await?;
match &mut report_obj.inner {
ObjectInner::DmarcInternalReport(report) => {
report.reschedule_ops(&mut batch, item_id, report_obj.revision, deliver_at);
report.reschedule_ops(
&mut batch,
item_id,
report_obj.revision,
deliver_at,
queued.as_ref(),
);
}
ObjectInner::TlsInternalReport(report) => {
report.reschedule_ops(&mut batch, item_id, report_obj.revision, deliver_at);
report.reschedule_ops(
&mut batch,
item_id,
report_obj.revision,
deliver_at,
queued.as_ref(),
);
}
_ => {}
}
@@ -156,6 +211,9 @@ pub(crate) async fn report_set(
.write(batch.build_all())
.await
.caused_by(trc::location!())?;
// inbuxa: a rescheduled report may now be due sooner than the task
// manager's next scan
set.server.notify_task_queue();
}
Ok(set)
+4 -9
View File
@@ -463,15 +463,10 @@ pub(crate) async fn task_query(
.set_values(typ.is_some()),
|key, value| {
if let Some(typ) = typ {
let task_type =
TaskType::from_id(value.deserialize_be_u16(0)?).ok_or_else(|| {
trc::StoreEvent::DataCorruption
.into_err()
.ctx(trc::Key::Key, key.to_vec())
.ctx(trc::Key::Value, value.to_vec())
.caused_by(trc::location!())
})?;
if task_type != typ {
// inbuxa: a row whose type can't be read matches no type
// filter; the task manager logs and repairs it
let task_type = value.deserialize_be_u16(0).ok().and_then(TaskType::from_id);
if task_type != Some(typ) {
return Ok(true);
}
}