SQL pools time out; task locks are a renewed five-minute lease
A 3-node rehearsal (PostgreSQL + NATS + Garage) found two ways a crash leaves work stuck: Pool hangs. The PostgreSQL pool (deadpool) was built with no timeouts, so a request waited for a free connection, and for one to be opened or recycled, for as long as it took: forever when the server stopped answering. MySQL's pool (mysql_async) has no wait timeout at all. - PostgreSQL: wait 30 s (or the store's timeout if longer), create the store's timeout or 15 s (it bounds the whole handshake, where tokio-postgres's connect_timeout covers only the TCP connect), recycle 10 s. The pool config is now always set, not only with poolMaxConnections. - MySQL: every connection is taken through MysqlStore::conn(), which gives up after 30 s. - Both: TCP keepalive after 60 s idle, so a server that vanished without closing the connection is noticed in minutes rather than the two-hour system default. The DataStore schema has no pool timeout settings, so these are fixed defaults; the store's own timeout bounds connecting on PostgreSQL. Task locks. A task lock lasted an hour, so after a hard crash the dead node's tasks waited up to an hour and five minutes. The lock is now a five-minute lease: while this node runs a task, the task manager renews its lock every third of the lifetime (InMemoryStore::renew_lock, a compare-and-set on the store backends and SET XX EX on Redis, which leaves a lock that already expired alone). A killed node's tasks run elsewhere within about five minutes plus the claim recheck. A task this node holds isn't handed to a worker again by the scan. store::pool_timeout (new): a local listener that accepts connections and never answers plays a hung server; a PostgreSQL store with a 2 s timeout returns an error in about 4 s, and a MySQL store in 30 s. Without the timeouts both wait for good. store::task_locks gains a task held for 1.5 lock lifetimes: its lease is still held, and released when the task ends.
This commit is contained in:
@@ -22,11 +22,34 @@ use crate::{
|
||||
use ::registry::schema::{enums::PostgreSqlRecyclingMethod, structs};
|
||||
use ahash::AHashSet;
|
||||
use deadpool_postgres::{
|
||||
Config, ManagerConfig, Object, Pool, PoolConfig, RecyclingMethod, Runtime,
|
||||
Config, ManagerConfig, Object, Pool, PoolConfig, RecyclingMethod, Runtime, Timeouts,
|
||||
};
|
||||
use std::time::Duration;
|
||||
use tokio_postgres::NoTls;
|
||||
use utils::tls::rustls_client_config;
|
||||
|
||||
/// inbuxa: how long a request waits for a pooled connection.
|
||||
pub(crate) const POOL_WAIT_TIMEOUT: Duration = Duration::from_secs(30);
|
||||
/// inbuxa: how long opening a connection may take when the store sets no
|
||||
/// timeout of its own.
|
||||
pub(crate) const POOL_CREATE_TIMEOUT: Duration = Duration::from_secs(15);
|
||||
/// inbuxa: how long checking a pooled connection before reuse may take.
|
||||
pub(crate) const POOL_RECYCLE_TIMEOUT: Duration = Duration::from_secs(10);
|
||||
/// inbuxa: idle time before TCP keepalive probes start.
|
||||
pub(crate) const POOL_KEEPALIVE_IDLE: Duration = Duration::from_secs(60);
|
||||
|
||||
/// inbuxa: the pool's timeouts. Opening a connection is bounded by the
|
||||
/// store's own timeout when it has one; waiting for one covers at least that
|
||||
/// long, so a slow connect isn't cut short by the wait.
|
||||
pub(crate) fn pool_timeouts(connect_timeout: Option<Duration>) -> Timeouts {
|
||||
let create = connect_timeout.unwrap_or(POOL_CREATE_TIMEOUT);
|
||||
Timeouts {
|
||||
wait: POOL_WAIT_TIMEOUT.max(create).into(),
|
||||
create: create.into(),
|
||||
recycle: POOL_RECYCLE_TIMEOUT.into(),
|
||||
}
|
||||
}
|
||||
|
||||
impl PostgresStore {
|
||||
pub async fn open(config: structs::PostgreSqlStore) -> Result<Store, String> {
|
||||
// inbuxa: ST-15: where the primary is, to tell a replica from it
|
||||
@@ -46,9 +69,20 @@ impl PostgresStore {
|
||||
PostgreSqlRecyclingMethod::Clean => RecyclingMethod::Clean,
|
||||
},
|
||||
});
|
||||
if let Some(max_conn) = config.pool_max_connections {
|
||||
cfg.pool = PoolConfig::new(max_conn as usize).into();
|
||||
}
|
||||
// inbuxa: upstream set no pool timeouts, so a request waited for a
|
||||
// free connection, or for one to be made or recycled, for as long as
|
||||
// it took: forever when the server stopped answering. A worker now
|
||||
// gets an error instead and the task or request is retried.
|
||||
let mut pool = config
|
||||
.pool_max_connections
|
||||
.map(|max_conn| PoolConfig::new(max_conn as usize))
|
||||
.unwrap_or_default();
|
||||
pool.timeouts = pool_timeouts(cfg.connect_timeout);
|
||||
cfg.pool = pool.into();
|
||||
// Notice a server that went away without closing the connection in
|
||||
// minutes rather than the system default of two hours
|
||||
cfg.keepalives = true.into();
|
||||
cfg.keepalives_idle = POOL_KEEPALIVE_IDLE.into();
|
||||
|
||||
let primary_pool = if config.use_tls {
|
||||
cfg.create_pool(
|
||||
|
||||
Reference in New Issue
Block a user