Commit Graph
4 Commits
Author SHA1 Message Date
jcoffey-dev a90df656e7 Describe an installation in a file, and converge to it
One file describes the whole installation: which machine runs what, under
which names. Each machine acts on its own part of it and prints the command
to run on the others, which nobody but their operator runs. Emit, never
execute -- no agent, no console-held credential, no machine reaching
another.

  inbuxa plan  -f topology.json     what would change here; changes nothing
  inbuxa apply -f topology.json     make this machine match it
  inbuxa export                     the file, from what is already here

plan diffs the file against what is installed rather than against what
happens to be running: a container stopped by hand is still installed, and
offering to install it again would be a lie about what is about to happen.
The state that makes that possible -- intent, which the machine itself
cannot tell you -- is /etc/inbuxa/install.json.

Front ends across machines, not a clustered mail server. Two machines each
running one is refused, and the refusal says why: a second node needs a
shared store and cluster configuration, which this does not set up. Two of
the same component on one machine is refused too -- two webmails need two
ports and two names, and the file says neither.

The shrink is the half worth proving, and it found the bug that mattered:
apply ran first boot every time, so the second one asked a configured server
for bootstrap credentials it had stopped accepting, and adding or removing a
front end could not work at all. An installed machine now converges instead:
the deployment is rewritten from the shapes asked for, --remove-orphans
takes away what the file no longer lists, data volumes are left alone, and
the secrets generated the first time are kept rather than rolled.

Two more found the same way:

- Certificates were turned on even where nothing holds port 80. On a machine
  with no proxy the order can only fail, and it stopped the install over it.
  It now says whose job they are instead.
- A converge that restarts the server reported "Done" while it was still
  coming back. It waits.

Twenty-eight checks, on Debian 13 and Fedora 43: install from a file, plan
the same file and be told there is nothing to do, remove the webmail and
watch it go while the mail stays, put it back.
2026-09-24 09:51:16 -07:00
jcoffey-dev e55ff6b2df Offer only what this machine can deliver
The installer knew one distribution: Debian, with docker, on a new enough
release. Everything else it would have offered and then failed at.

Three facts now decide what the matrix offers, and each is read from the
machine rather than assumed:

- The container runtime. Docker where the machine has one, podman on the Red
  Hat family, which ships no docker at all. Both go through the same compose
  plugin: compose speaks the Docker API and podman serves it, so there is one
  compose file and one deployment path, not two. Telling a Fedora operator to
  add Docker's own repository to a machine that already has a container
  runtime would have been the wrong trade.

- glibc. The server binary is downloaded, not built here, and it is linked
  against 2.39. Rocky 9 (2.34), Debian 12 and Ubuntu 22.04 cannot run it, so
  the host shape is refused there with the version found and the container
  shape named as the answer -- rather than installing a file that cannot
  start.

- The operating system itself. This compiles for macOS and Windows because Go
  compiles anything, and on either it would read no os-release, find no
  systemd, and describe a machine that does not exist. It now says what it is
  and exits.

Two bugs the other distributions found, both of which Debian could not have:

- The survey reported the first thing in the way and stopped, so on Fedora it
  asked to start podman.socket, and then -- having done it -- asked for the
  compose plugin. Needs are named now, not described, and reported together.

- apply used the survey taken before dependencies were installed, so on a
  machine that had no runtime at all it installed podman and then reached for
  docker. It re-surveys after resolving, and stops if containers still are
  not usable.

The lab takes DISTRO now: debian13, debian12, ubuntu2404, fedora, rocky9,
arch, each with its own disk and ssh port so several can be up at once. The
cases no longer say "docker" either. install-local passes on Debian 13,
Fedora 43 and Rocky 9 -- 20 checks each, ending with a sign-in to the webmail
the installer put there.
2026-09-22 22:22:45 -07:00
jcoffey-dev 8725d8c11b Obtain certificates, and wait for the one that matters
The public shape now works end to end: real ports, Caddy in front, and both
programs that need certificates getting them from the same CA -- Caddy for
the front ends over TLS-ALPN-01, the mail server for its own names over
HTTP-01, which Caddy forwards on port 80.

Proved in the lab against Pebble, with a DNS stub answering every name with
the machine's own address, so no public name or public CA is involved:
twenty checks, ending with IMAPS and submissions presenting a certificate
for the mail host that verifies against the CA, and the webmail sending
sign-in to the server as the first-party client the server registered.

Two things the test found, both of which would have shipped:

- The proxy fronted four of the server's five names. The server puts
  ua-auto-config in its own certificate too, so its challenge was never
  forwarded, one name failed, and the whole order failed with it -- leaving
  the mail ports on a self-signed certificate while everything else looked
  healthy. The list now matches what the server asks for.

- Nothing waited for the certificate. An order that fails is not retried on
  its own and a restart does not start a new one, so the install declared
  itself finished over a self-signed certificate. It now waits, asks again
  every 45 seconds, and reports the issuer -- or says plainly that the
  server will keep trying once the domain resolves here, which is the
  ordinary case on a first install.

--acme-directory and --acme-ca-root are what let a private CA be used: the
root is added to the server image's own bundle and given to Caddy, because
neither sees the other's trust store.
2026-09-22 19:11:52 -07:00
jcoffey-dev 7141e565ea Install the suite, for real, in containers
The plan now happens. "inbuxa install --local --domain example.test
--install-deps --yes" on a machine with nothing on it ends with a mail
server, a console and a webmail running, an administrator and a first
mailbox created, and the records the domain needs written out.

The sequence is the one ihasmail-oneshot worked out against a running
server, which is why its JMAP client and its Docker handling came across
nearly whole: bring the server up in bootstrap mode with a credential that
lives in an override file for that step only, complete bootstrap, bring the
rest up without it -- so no recovery credential outlives the setup -- exempt
the front ends from the auto-ban, restart for the settings that need it,
create the first account, and write down the password nothing else holds.

New here: three services rather than two. The console is static files that
learn their server's address at start, and the webmail is given the
first-party OAuth client secret that the server is given too.

Twenty checks in the lab, from a bare Debian 13. The two worth having are
the ones that catch an install that looks fine and is not: nothing in the
running server carries a recovery admin any more, and the account the
installer created can sign in to the webmail it installed.

Two bugs the lab caught, both of which would have shipped:

- the private addresses were worked out on a copy of the stack, so the
  server was told the webmail speaks from "", and refused it.
- the console image rewrites index.html when it starts, so a read-only root
  filesystem left it restarting forever. The webmail keeps read_only; the
  console cannot have it until that rewrite moves.
2026-09-22 18:26:31 -07:00