Merge pull request 'Local AI: the model service production runs, and how to build its model' (#1) from packaging/local-ai into main

Reviewed-on: #1
This commit was merged in pull request #1.
This commit is contained in:
2026-09-24 14:55:32 +00:00
3 changed files with 197 additions and 0 deletions
+67
View File
@@ -0,0 +1,67 @@
# Local AI for spam filtering
The model behind inbuxa's AI spam classifier, run on the same machine as the
mail server, as production runs it. Nothing here is used by the installer
yet: it's the reference for its bare-metal mode, and the record of what runs.
The classifier itself is in inbuxa-server (the `ai-spam-classification`
spec). It's off until an administrator turns it on in inbuxa Admin, under
Settings › Spam Filter › Local AI. The model's opinion is one bounded signal
among many (at most +2 points by default), and a slow or missing model never
holds up mail.
## What runs
| Piece | What |
|---|---|
| Server | llama.cpp `b11160`, the CPU build `llama-b11160-bin-ubuntu-x64.tar.gz` (SHA-256 `48ece242…e435`), in `/opt/llama/b11160` |
| Model | Qwen3 4B Instruct 2507 (Apache-2.0), quantized to Q4_K_M: `/opt/llama/models/qwen3-4b-instruct-2507-Q4_K_M.gguf`, SHA-256 `0f5e5250018ea2e4384e8b440f3a51fd1fff31d6694906e63fca3b5abd8a9a28` |
| Service | [`inbuxa-llm.service`](inbuxa-llm.service): in the mail network namespace, on its loopback `127.0.0.1:8080` only, 4 cores and 8 GB at most, 4 slots of 4096 tokens |
The mail server's model entry (`x:AiModel`) points at
`http://127.0.0.1:8080/v1/chat/completions` with the model name
`qwen3-4b-instruct-2507`, which is the service's `--alias`. The address is
the mail namespace's own loopback, so message text never leaves the machine.
Measured on production (host1, CPU only): about 1.2 s per message once warm,
4 s for the first; about 4.3 GB of memory.
## The model, built from source
Qwen publishes no GGUF of this model, and the popular ones are third-party
conversions. [`build-model.sh`](build-model.sh) makes it from Qwen's official
weights instead: it checks each file against Hugging Face's published
checksums, converts with llama.cpp's own converter, quantizes to Q4_K_M, and
compares the result with production's SHA-256.
```bash
packaging/local-ai/build-model.sh ~/ai-models
```
## Installing it by hand
On a Debian 13 host whose mail server runs in the `mail` namespace
(`/run/netns/mail`, from `mail-netns.service`):
```bash
apt-get install libgomp1 # llama.cpp's OpenMP runtime
mkdir -p /opt/llama/models
curl -fLO https://github.com/ggml-org/llama.cpp/releases/download/b11160/llama-b11160-bin-ubuntu-x64.tar.gz
echo "48ece24283876fc3401b737724008c03cbc4c7ba335b6c1aa2a7b6ce2d49e435 llama-b11160-bin-ubuntu-x64.tar.gz" | sha256sum -c
mkdir -p /opt/llama/b11160 && tar -C /opt/llama/b11160 -xzf llama-b11160-bin-ubuntu-x64.tar.gz
install -m 0644 qwen3-4b-instruct-2507-Q4_K_M.gguf /opt/llama/models/
install -m 0644 packaging/local-ai/inbuxa-llm.service /etc/systemd/system/
systemctl daemon-reload && systemctl enable --now inbuxa-llm
ip netns exec mail curl -s http://127.0.0.1:8080/health # {"status":"ok"}
```
A host without the mail namespace drops `NetworkNamespacePath`,
`After=mail-netns.service` and `Requires=mail-netns.service` from the unit;
the service then listens on the host's loopback, which is where a mail server
on that host reaches it.
## Undoing it
`systemctl disable --now inbuxa-llm`. The classifier then fails fast on every
message and mail flows as before. Turn the classifier off in inbuxa Admin to
stop it trying.
+84
View File
@@ -0,0 +1,84 @@
#!/bin/bash
# SPDX-FileCopyrightText: 2026 Coffey Labs
# SPDX-License-Identifier: AGPL-3.0-or-later
#
# Build the local AI spam classifier's model from Qwen's official weights, so
# the file that runs is one you made and can check, not a third party's.
#
# packaging/local-ai/build-model.sh WORKDIR
#
# Downloads Qwen/Qwen3-4B-Instruct-2507 (Apache-2.0) and checks every file
# against the checksums Hugging Face publishes, converts it to GGUF with
# llama.cpp's own converter, and quantizes it to Q4_K_M, the quantization the
# classifier was calibrated with (inbuxa-server's ai-spam-classification
# spec, "Calibration"). Needs docker, curl, python3 and about 20 GB of disk.
#
# Qwen publishes no GGUF of this model; the popular ones are third-party
# conversions. Built with llama.cpp b11160, the result is byte-identical to
# the model production runs, whose SHA-256 is EXPECTED below.
set -euo pipefail
W="${1:?usage: build-model.sh WORKDIR}"
LLAMA=b11160
LLAMA_BIN_SHA=48ece24283876fc3401b737724008c03cbc4c7ba335b6c1aa2a7b6ce2d49e435
MODEL=Qwen/Qwen3-4B-Instruct-2507
NAME=qwen3-4b-instruct-2507
EXPECTED=0f5e5250018ea2e4384e8b440f3a51fd1fff31d6694906e63fca3b5abd8a9a28
mkdir -p "$W/$NAME" "$W/llama-bin"
cd "$W"
echo "== $MODEL, checked against Hugging Face's checksums"
curl -sf "https://huggingface.co/api/models/$MODEL/tree/main" > tree.json
python3 - "$MODEL" "$NAME" <<'PY'
import hashlib, json, subprocess, sys
model, out = sys.argv[1], sys.argv[2]
for f in json.load(open('tree.json')):
path = f['path']
if f.get('type') != 'file' or path.startswith('.'):
continue
dest = f'{out}/{path}'
subprocess.run(['curl', '-sfL', '-o', dest,
f'https://huggingface.co/{model}/resolve/main/{path}'], check=True)
oid = (f.get('lfs') or {}).get('oid')
if oid:
h = hashlib.sha256()
with open(dest, 'rb') as fh:
for block in iter(lambda: fh.read(1 << 22), b''):
h.update(block)
if h.hexdigest() != oid:
sys.exit(f'checksum mismatch: {path}')
print(f' ok {path}')
PY
echo "== llama.cpp $LLAMA: converter source and quantizer"
curl -sfL -o "llama-src-$LLAMA.tar.gz" "https://github.com/ggml-org/llama.cpp/archive/refs/tags/$LLAMA.tar.gz"
tar -xzf "llama-src-$LLAMA.tar.gz"
curl -sfL -o "llama-$LLAMA-bin-ubuntu-x64.tar.gz" \
"https://github.com/ggml-org/llama.cpp/releases/download/$LLAMA/llama-$LLAMA-bin-ubuntu-x64.tar.gz"
echo "$LLAMA_BIN_SHA llama-$LLAMA-bin-ubuntu-x64.tar.gz" | sha256sum -c
tar -C llama-bin -xzf "llama-$LLAMA-bin-ubuntu-x64.tar.gz"
echo "== convert to GGUF (f16)"
# Runs as the container's root: the converter's dependencies look up the
# user's name, and a host uid with no passwd entry breaks that.
docker run --rm -v "$W":/w -w /w python:3.12-slim sh -c "
python -m venv /tmp/v &&
/tmp/v/bin/pip install -q -r llama.cpp-$LLAMA/requirements/requirements-convert_hf_to_gguf.txt \
--extra-index-url https://download.pytorch.org/whl/cpu >/dev/null &&
/tmp/v/bin/python llama.cpp-$LLAMA/convert_hf_to_gguf.py $NAME --outtype f16 --outfile $NAME-f16.gguf &&
chown $(id -u):$(id -g) $NAME-f16.gguf"
echo "== quantize to Q4_K_M"
Q=$(find llama-bin -name llama-quantize -type f | head -1)
LD_LIBRARY_PATH="$(dirname "$Q")" "$Q" "$NAME-f16.gguf" "$NAME-Q4_K_M.gguf" Q4_K_M >/dev/null
GOT=$(sha256sum "$NAME-Q4_K_M.gguf" | cut -d' ' -f1)
echo "== $W/$NAME-Q4_K_M.gguf"
echo " sha256 $GOT"
if [ "$GOT" = "$EXPECTED" ]; then
echo " matches the model production runs"
else
echo " differs from production's ($EXPECTED): a different llama.cpp or upstream file change" >&2
exit 1
fi
+46
View File
@@ -0,0 +1,46 @@
# SPDX-FileCopyrightText: 2026 Coffey Labs
# SPDX-License-Identifier: AGPL-3.0-or-later
#
# The local language model for inbuxa's AI spam classifier: llama.cpp's
# server with Qwen3 4B Instruct 2507 (Apache-2.0), converted from Qwen's
# official weights and quantized to Q4_K_M. It runs in the mail network
# namespace and listens only on its loopback, 127.0.0.1:8080, where the mail
# server's x:AiModel "local" points. Nothing leaves the machine.
#
# Bounded so it can never starve the mail server: 4 cores, 8 GB. Four slots
# match the classifier's default of four requests in flight (inbuxa:AiLimits
# maxConcurrentCalls), each with a 4096-token context.
[Unit]
Description=inbuxa local AI model (llama.cpp, Qwen3 4B Instruct 2507)
After=mail-netns.service
Requires=mail-netns.service
[Service]
NetworkNamespacePath=/run/netns/mail
Environment=LD_LIBRARY_PATH=/opt/llama/b11160/llama-b11160
ExecStart=/opt/llama/b11160/llama-b11160/llama-server \
--model /opt/llama/models/qwen3-4b-instruct-2507-Q4_K_M.gguf \
--alias qwen3-4b-instruct-2507 \
--host 127.0.0.1 --port 8080 \
--threads 4 --parallel 4 --ctx-size 16384 \
--no-webui
DynamicUser=yes
CPUQuota=400%
MemoryMax=8G
Nice=10
Restart=on-failure
RestartSec=10
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
NoNewPrivileges=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectControlGroups=yes
RestrictSUIDSGID=yes
LockPersonality=yes
CapabilityBoundingSet=
SystemCallArchitectures=native
[Install]
WantedBy=multi-user.target