f99082d4ef9b474a7ccc91c7adc7edb80d27d632
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a90df656e7 |
Describe an installation in a file, and converge to it
One file describes the whole installation: which machine runs what, under which names. Each machine acts on its own part of it and prints the command to run on the others, which nobody but their operator runs. Emit, never execute -- no agent, no console-held credential, no machine reaching another. inbuxa plan -f topology.json what would change here; changes nothing inbuxa apply -f topology.json make this machine match it inbuxa export the file, from what is already here plan diffs the file against what is installed rather than against what happens to be running: a container stopped by hand is still installed, and offering to install it again would be a lie about what is about to happen. The state that makes that possible -- intent, which the machine itself cannot tell you -- is /etc/inbuxa/install.json. Front ends across machines, not a clustered mail server. Two machines each running one is refused, and the refusal says why: a second node needs a shared store and cluster configuration, which this does not set up. Two of the same component on one machine is refused too -- two webmails need two ports and two names, and the file says neither. The shrink is the half worth proving, and it found the bug that mattered: apply ran first boot every time, so the second one asked a configured server for bootstrap credentials it had stopped accepting, and adding or removing a front end could not work at all. An installed machine now converges instead: the deployment is rewritten from the shapes asked for, --remove-orphans takes away what the file no longer lists, data volumes are left alone, and the secrets generated the first time are kept rather than rolled. Two more found the same way: - Certificates were turned on even where nothing holds port 80. On a machine with no proxy the order can only fail, and it stopped the install over it. It now says whose job they are instead. - A converge that restarts the server reported "Done" while it was still coming back. It waits. Twenty-eight checks, on Debian 13 and Fedora 43: install from a file, plan the same file and be told there is nothing to do, remove the webmail and watch it go while the mail stays, put it back. |
||
|
|
e55ff6b2df |
Offer only what this machine can deliver
The installer knew one distribution: Debian, with docker, on a new enough release. Everything else it would have offered and then failed at. Three facts now decide what the matrix offers, and each is read from the machine rather than assumed: - The container runtime. Docker where the machine has one, podman on the Red Hat family, which ships no docker at all. Both go through the same compose plugin: compose speaks the Docker API and podman serves it, so there is one compose file and one deployment path, not two. Telling a Fedora operator to add Docker's own repository to a machine that already has a container runtime would have been the wrong trade. - glibc. The server binary is downloaded, not built here, and it is linked against 2.39. Rocky 9 (2.34), Debian 12 and Ubuntu 22.04 cannot run it, so the host shape is refused there with the version found and the container shape named as the answer -- rather than installing a file that cannot start. - The operating system itself. This compiles for macOS and Windows because Go compiles anything, and on either it would read no os-release, find no systemd, and describe a machine that does not exist. It now says what it is and exits. Two bugs the other distributions found, both of which Debian could not have: - The survey reported the first thing in the way and stopped, so on Fedora it asked to start podman.socket, and then -- having done it -- asked for the compose plugin. Needs are named now, not described, and reported together. - apply used the survey taken before dependencies were installed, so on a machine that had no runtime at all it installed podman and then reached for docker. It re-surveys after resolving, and stops if containers still are not usable. The lab takes DISTRO now: debian13, debian12, ubuntu2404, fedora, rocky9, arch, each with its own disk and ssh port so several can be up at once. The cases no longer say "docker" either. install-local passes on Debian 13, Fedora 43 and Rocky 9 -- 20 checks each, ending with a sign-in to the webmail the installer put there. |
||
|
|
8725d8c11b |
Obtain certificates, and wait for the one that matters
The public shape now works end to end: real ports, Caddy in front, and both programs that need certificates getting them from the same CA -- Caddy for the front ends over TLS-ALPN-01, the mail server for its own names over HTTP-01, which Caddy forwards on port 80. Proved in the lab against Pebble, with a DNS stub answering every name with the machine's own address, so no public name or public CA is involved: twenty checks, ending with IMAPS and submissions presenting a certificate for the mail host that verifies against the CA, and the webmail sending sign-in to the server as the first-party client the server registered. Two things the test found, both of which would have shipped: - The proxy fronted four of the server's five names. The server puts ua-auto-config in its own certificate too, so its challenge was never forwarded, one name failed, and the whole order failed with it -- leaving the mail ports on a self-signed certificate while everything else looked healthy. The list now matches what the server asks for. - Nothing waited for the certificate. An order that fails is not retried on its own and a restart does not start a new one, so the install declared itself finished over a self-signed certificate. It now waits, asks again every 45 seconds, and reports the issuer -- or says plainly that the server will keep trying once the domain resolves here, which is the ordinary case on a first install. --acme-directory and --acme-ca-root are what let a private CA be used: the root is added to the server image's own bundle and given to Caddy, because neither sees the other's trust store. |
||
|
|
7141e565ea |
Install the suite, for real, in containers
The plan now happens. "inbuxa install --local --domain example.test --install-deps --yes" on a machine with nothing on it ends with a mail server, a console and a webmail running, an administrator and a first mailbox created, and the records the domain needs written out. The sequence is the one ihasmail-oneshot worked out against a running server, which is why its JMAP client and its Docker handling came across nearly whole: bring the server up in bootstrap mode with a credential that lives in an override file for that step only, complete bootstrap, bring the rest up without it -- so no recovery credential outlives the setup -- exempt the front ends from the auto-ban, restart for the settings that need it, create the first account, and write down the password nothing else holds. New here: three services rather than two. The console is static files that learn their server's address at start, and the webmail is given the first-party OAuth client secret that the server is given too. Twenty checks in the lab, from a bare Debian 13. The two worth having are the ones that catch an install that looks fine and is not: nothing in the running server carries a recovery admin any more, and the account the installer created can sign in to the webmail it installed. Two bugs the lab caught, both of which would have shipped: - the private addresses were worked out on a copy of the stack, so the server was told the webmail speaks from "", and refused it. - the console image rewrites index.html when it starts, so a read-only root filesystem left it restarting forever. The webmail keeps read_only; the console cannot have it until that rewrite moves. |