Do not spend login attempts on an outage nobody caused

ihasmail runs in its own container, usually on its own host, so Stalwart being
briefly unreachable is an ordinary Tuesday. Sign-in handled it almost right:
a 401 is invalid_credentials, a timeout is 504 and anything else is 502, none
of which reads as a rejected password.

What it got wrong was the counting. RateLimiter.check() consumes an attempt
when it is called, and it is called before the upstream is contacted; reset()
only runs on success. So every try against an unreachable server burned a
credential attempt, and after ten of them the person was locked out for the
rest of the fifteen-minute window -- including after the server came back. A
thirty-second blip became a quarter-hour lockout, and the second failure was
entirely ihasmail's own doing.

A 401 is a judgement about the password and stays counted. A 502 or 504 is the
upstream failing to answer, says nothing about the credentials, and is now
refunded -- one attempt back, not the key cleared, so a run of real failures
with an outage in the middle still adds up. The old-server refusal refunds too:
those credentials were accepted.

Both guessing keys are refunded, not just the username one. Refunding only
that would not have fixed it -- ten retries still spend the per-address budget,
and behind one office NAT that budget belongs to the whole building, so a
company-wide outage would lock out the company.

Which needs a backstop, because "not counted" must not mean "unlimited": each
attempt still costs an outbound connection that may sit there until
UPSTREAM_TIMEOUT, and an outage is the one moment the endpoint is cheapest to
abuse. So there is a second ceiling per address, twenty times looser and never
refunded. A person retrying will not come near it; something hammering will.

Both messages now say the quiet part -- "This is not a problem with your
password" -- for somebody already worried they have forgotten it.

Closes #239.
This commit is contained in:
2026-09-02 14:08:23 -07:00
parent f6f6ce0b5e
commit 607afeb4ad
4 changed files with 173 additions and 2 deletions
+37
View File
@@ -107,3 +107,40 @@ test("only a PDF blob may be framed, and only by us", async () => {
assert.equal(securityHeadersFor("image/png", true), "DENY");
assert.equal(securityHeadersFor("text/html", true), "DENY");
});
/*
* #239: retrying through an outage must not lock somebody out of the recovery.
*
* STALWART_URL at the top of this file is 127.0.0.1:1 — nothing listens there,
* so every sign-in here is the outage case. Before the fix, the eleventh of
* these came back 429 and stayed 429 for fifteen minutes, outliving whatever
* had actually been wrong.
*/
test("an unreachable upstream does not spend login attempts", async () => {
const app = createApp();
const login = () =>
app.request("/api/auth/login", {
method: "POST",
headers: { "content-type": "application/json", "x-requested-with": "ihasmail" },
body: JSON.stringify({ username: "[email protected]", password: "hunter2" }),
});
// Comfortably past LOGIN_RATE_LIMIT, which defaults to 10.
for (let i = 0; i < 25; i++) {
const res = await login();
assert.notEqual(res.status, 429, `attempt ${i + 1} was rate limited`);
assert.ok(res.status === 502 || res.status === 504, `attempt ${i + 1} said ${res.status}`);
}
});
test("an unreachable upstream says it is not the password", async () => {
const app = createApp();
const res = await app.request("/api/auth/login", {
method: "POST",
headers: { "content-type": "application/json", "x-requested-with": "ihasmail" },
body: JSON.stringify({ username: "[email protected]", password: "hunter2" }),
});
const body = (await res.json()) as { error: string; message: string };
assert.notEqual(body.error, "invalid_credentials");
assert.match(body.message, /not a problem with your password/i);
});