SQL pools time out; task locks are a renewed five-minute lease
A 3-node rehearsal (PostgreSQL + NATS + Garage) found two ways a crash leaves work stuck: Pool hangs. The PostgreSQL pool (deadpool) was built with no timeouts, so a request waited for a free connection, and for one to be opened or recycled, for as long as it took: forever when the server stopped answering. MySQL's pool (mysql_async) has no wait timeout at all. - PostgreSQL: wait 30 s (or the store's timeout if longer), create the store's timeout or 15 s (it bounds the whole handshake, where tokio-postgres's connect_timeout covers only the TCP connect), recycle 10 s. The pool config is now always set, not only with poolMaxConnections. - MySQL: every connection is taken through MysqlStore::conn(), which gives up after 30 s. - Both: TCP keepalive after 60 s idle, so a server that vanished without closing the connection is noticed in minutes rather than the two-hour system default. The DataStore schema has no pool timeout settings, so these are fixed defaults; the store's own timeout bounds connecting on PostgreSQL. Task locks. A task lock lasted an hour, so after a hard crash the dead node's tasks waited up to an hour and five minutes. The lock is now a five-minute lease: while this node runs a task, the task manager renews its lock every third of the lifetime (InMemoryStore::renew_lock, a compare-and-set on the store backends and SET XX EX on Redis, which leaves a lock that already expired alone). A killed node's tasks run elsewhere within about five minutes plus the claim recheck. A task this node holds isn't handed to a worker again by the scan. store::pool_timeout (new): a local listener that accepts connections and never answers plays a hung server; a PostgreSQL store with a 2 s timeout returns an error in about 4 s, and a MySQL store in 30 s. Without the timeouts both wait for good. store::task_locks gains a task held for 1.5 lock lifetimes: its lease is still held, and released when the task ends.
This commit is contained in:
@@ -401,6 +401,57 @@ impl InMemoryStore {
|
||||
}
|
||||
}
|
||||
|
||||
/// inbuxa: extends a lock this node holds to `duration` seconds from now.
|
||||
/// Returns false when the lock is gone or has expired: it may have been
|
||||
/// taken by someone else since, so it is left alone.
|
||||
pub async fn renew_lock(&self, prefix: u8, key: &[u8], duration: u64) -> trc::Result<bool> {
|
||||
match self {
|
||||
InMemoryStore::Store(store) => {
|
||||
let key = KeyValue::<()>::build_key(prefix, key);
|
||||
let key = ValueClass::InMemory(InMemoryClass::Key(key));
|
||||
let Some(lock_expiry) = store
|
||||
.get_value::<u64>(ValueKey::from(key.clone()))
|
||||
.await
|
||||
.caused_by(trc::location!())?
|
||||
else {
|
||||
return Ok(false);
|
||||
};
|
||||
let now = now();
|
||||
if lock_expiry <= now {
|
||||
return Ok(false);
|
||||
}
|
||||
|
||||
let mut batch = BatchBuilder::new();
|
||||
batch.assert_value(key.clone(), AssertValue::U64(lock_expiry));
|
||||
batch.set(key, (now + duration).serialize());
|
||||
match store.write(batch.build_all()).await {
|
||||
Ok(_) => Ok(true),
|
||||
Err(err) if err.is_assertion_failure() => Ok(false),
|
||||
Err(err) => Err(err
|
||||
.details("Failed to renew lock.")
|
||||
.caused_by(trc::location!())),
|
||||
}
|
||||
}
|
||||
InMemoryStore::Sharded(store) => {
|
||||
Box::pin(
|
||||
store
|
||||
.member(&KeyValue::<()>::build_key(prefix, key))
|
||||
.renew_lock(prefix, key, duration),
|
||||
)
|
||||
.await
|
||||
}
|
||||
#[cfg(feature = "redis")]
|
||||
InMemoryStore::Redis(store) => {
|
||||
store
|
||||
.renew_lock(&KeyValue::<()>::build_key(prefix, key), duration)
|
||||
.await
|
||||
}
|
||||
InMemoryStore::Static(_) | InMemoryStore::Http(_) => {
|
||||
Err(trc::StoreEvent::NotSupported.into_err())
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
pub async fn remove_lock(&self, prefix: u8, key: &[u8]) -> trc::Result<()> {
|
||||
self.key_delete(KeyValue::<()>::build_key(prefix, key))
|
||||
.await
|
||||
|
||||
Reference in New Issue
Block a user