Merge pull request 'Local AI: the model service production runs, and how to build its model' (#1) from packaging/local-ai into main
Reviewed-on: #1
This commit was merged in pull request #1.
This commit is contained in:
@@ -0,0 +1,67 @@
|
|||||||
|
# Local AI for spam filtering
|
||||||
|
|
||||||
|
The model behind inbuxa's AI spam classifier, run on the same machine as the
|
||||||
|
mail server, as production runs it. Nothing here is used by the installer
|
||||||
|
yet: it's the reference for its bare-metal mode, and the record of what runs.
|
||||||
|
|
||||||
|
The classifier itself is in inbuxa-server (the `ai-spam-classification`
|
||||||
|
spec). It's off until an administrator turns it on in inbuxa Admin, under
|
||||||
|
Settings › Spam Filter › Local AI. The model's opinion is one bounded signal
|
||||||
|
among many (at most +2 points by default), and a slow or missing model never
|
||||||
|
holds up mail.
|
||||||
|
|
||||||
|
## What runs
|
||||||
|
|
||||||
|
| Piece | What |
|
||||||
|
|---|---|
|
||||||
|
| Server | llama.cpp `b11160`, the CPU build `llama-b11160-bin-ubuntu-x64.tar.gz` (SHA-256 `48ece242…e435`), in `/opt/llama/b11160` |
|
||||||
|
| Model | Qwen3 4B Instruct 2507 (Apache-2.0), quantized to Q4_K_M: `/opt/llama/models/qwen3-4b-instruct-2507-Q4_K_M.gguf`, SHA-256 `0f5e5250018ea2e4384e8b440f3a51fd1fff31d6694906e63fca3b5abd8a9a28` |
|
||||||
|
| Service | [`inbuxa-llm.service`](inbuxa-llm.service): in the mail network namespace, on its loopback `127.0.0.1:8080` only, 4 cores and 8 GB at most, 4 slots of 4096 tokens |
|
||||||
|
|
||||||
|
The mail server's model entry (`x:AiModel`) points at
|
||||||
|
`http://127.0.0.1:8080/v1/chat/completions` with the model name
|
||||||
|
`qwen3-4b-instruct-2507`, which is the service's `--alias`. The address is
|
||||||
|
the mail namespace's own loopback, so message text never leaves the machine.
|
||||||
|
|
||||||
|
Measured on production (host1, CPU only): about 1.2 s per message once warm,
|
||||||
|
4 s for the first; about 4.3 GB of memory.
|
||||||
|
|
||||||
|
## The model, built from source
|
||||||
|
|
||||||
|
Qwen publishes no GGUF of this model, and the popular ones are third-party
|
||||||
|
conversions. [`build-model.sh`](build-model.sh) makes it from Qwen's official
|
||||||
|
weights instead: it checks each file against Hugging Face's published
|
||||||
|
checksums, converts with llama.cpp's own converter, quantizes to Q4_K_M, and
|
||||||
|
compares the result with production's SHA-256.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
packaging/local-ai/build-model.sh ~/ai-models
|
||||||
|
```
|
||||||
|
|
||||||
|
## Installing it by hand
|
||||||
|
|
||||||
|
On a Debian 13 host whose mail server runs in the `mail` namespace
|
||||||
|
(`/run/netns/mail`, from `mail-netns.service`):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
apt-get install libgomp1 # llama.cpp's OpenMP runtime
|
||||||
|
mkdir -p /opt/llama/models
|
||||||
|
curl -fLO https://github.com/ggml-org/llama.cpp/releases/download/b11160/llama-b11160-bin-ubuntu-x64.tar.gz
|
||||||
|
echo "48ece24283876fc3401b737724008c03cbc4c7ba335b6c1aa2a7b6ce2d49e435 llama-b11160-bin-ubuntu-x64.tar.gz" | sha256sum -c
|
||||||
|
mkdir -p /opt/llama/b11160 && tar -C /opt/llama/b11160 -xzf llama-b11160-bin-ubuntu-x64.tar.gz
|
||||||
|
install -m 0644 qwen3-4b-instruct-2507-Q4_K_M.gguf /opt/llama/models/
|
||||||
|
install -m 0644 packaging/local-ai/inbuxa-llm.service /etc/systemd/system/
|
||||||
|
systemctl daemon-reload && systemctl enable --now inbuxa-llm
|
||||||
|
ip netns exec mail curl -s http://127.0.0.1:8080/health # {"status":"ok"}
|
||||||
|
```
|
||||||
|
|
||||||
|
A host without the mail namespace drops `NetworkNamespacePath`,
|
||||||
|
`After=mail-netns.service` and `Requires=mail-netns.service` from the unit;
|
||||||
|
the service then listens on the host's loopback, which is where a mail server
|
||||||
|
on that host reaches it.
|
||||||
|
|
||||||
|
## Undoing it
|
||||||
|
|
||||||
|
`systemctl disable --now inbuxa-llm`. The classifier then fails fast on every
|
||||||
|
message and mail flows as before. Turn the classifier off in inbuxa Admin to
|
||||||
|
stop it trying.
|
||||||
Executable
+84
@@ -0,0 +1,84 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# SPDX-FileCopyrightText: 2026 Coffey Labs
|
||||||
|
# SPDX-License-Identifier: AGPL-3.0-or-later
|
||||||
|
#
|
||||||
|
# Build the local AI spam classifier's model from Qwen's official weights, so
|
||||||
|
# the file that runs is one you made and can check, not a third party's.
|
||||||
|
#
|
||||||
|
# packaging/local-ai/build-model.sh WORKDIR
|
||||||
|
#
|
||||||
|
# Downloads Qwen/Qwen3-4B-Instruct-2507 (Apache-2.0) and checks every file
|
||||||
|
# against the checksums Hugging Face publishes, converts it to GGUF with
|
||||||
|
# llama.cpp's own converter, and quantizes it to Q4_K_M, the quantization the
|
||||||
|
# classifier was calibrated with (inbuxa-server's ai-spam-classification
|
||||||
|
# spec, "Calibration"). Needs docker, curl, python3 and about 20 GB of disk.
|
||||||
|
#
|
||||||
|
# Qwen publishes no GGUF of this model; the popular ones are third-party
|
||||||
|
# conversions. Built with llama.cpp b11160, the result is byte-identical to
|
||||||
|
# the model production runs, whose SHA-256 is EXPECTED below.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
W="${1:?usage: build-model.sh WORKDIR}"
|
||||||
|
LLAMA=b11160
|
||||||
|
LLAMA_BIN_SHA=48ece24283876fc3401b737724008c03cbc4c7ba335b6c1aa2a7b6ce2d49e435
|
||||||
|
MODEL=Qwen/Qwen3-4B-Instruct-2507
|
||||||
|
NAME=qwen3-4b-instruct-2507
|
||||||
|
EXPECTED=0f5e5250018ea2e4384e8b440f3a51fd1fff31d6694906e63fca3b5abd8a9a28
|
||||||
|
|
||||||
|
mkdir -p "$W/$NAME" "$W/llama-bin"
|
||||||
|
cd "$W"
|
||||||
|
|
||||||
|
echo "== $MODEL, checked against Hugging Face's checksums"
|
||||||
|
curl -sf "https://huggingface.co/api/models/$MODEL/tree/main" > tree.json
|
||||||
|
python3 - "$MODEL" "$NAME" <<'PY'
|
||||||
|
import hashlib, json, subprocess, sys
|
||||||
|
model, out = sys.argv[1], sys.argv[2]
|
||||||
|
for f in json.load(open('tree.json')):
|
||||||
|
path = f['path']
|
||||||
|
if f.get('type') != 'file' or path.startswith('.'):
|
||||||
|
continue
|
||||||
|
dest = f'{out}/{path}'
|
||||||
|
subprocess.run(['curl', '-sfL', '-o', dest,
|
||||||
|
f'https://huggingface.co/{model}/resolve/main/{path}'], check=True)
|
||||||
|
oid = (f.get('lfs') or {}).get('oid')
|
||||||
|
if oid:
|
||||||
|
h = hashlib.sha256()
|
||||||
|
with open(dest, 'rb') as fh:
|
||||||
|
for block in iter(lambda: fh.read(1 << 22), b''):
|
||||||
|
h.update(block)
|
||||||
|
if h.hexdigest() != oid:
|
||||||
|
sys.exit(f'checksum mismatch: {path}')
|
||||||
|
print(f' ok {path}')
|
||||||
|
PY
|
||||||
|
|
||||||
|
echo "== llama.cpp $LLAMA: converter source and quantizer"
|
||||||
|
curl -sfL -o "llama-src-$LLAMA.tar.gz" "https://github.com/ggml-org/llama.cpp/archive/refs/tags/$LLAMA.tar.gz"
|
||||||
|
tar -xzf "llama-src-$LLAMA.tar.gz"
|
||||||
|
curl -sfL -o "llama-$LLAMA-bin-ubuntu-x64.tar.gz" \
|
||||||
|
"https://github.com/ggml-org/llama.cpp/releases/download/$LLAMA/llama-$LLAMA-bin-ubuntu-x64.tar.gz"
|
||||||
|
echo "$LLAMA_BIN_SHA llama-$LLAMA-bin-ubuntu-x64.tar.gz" | sha256sum -c
|
||||||
|
tar -C llama-bin -xzf "llama-$LLAMA-bin-ubuntu-x64.tar.gz"
|
||||||
|
|
||||||
|
echo "== convert to GGUF (f16)"
|
||||||
|
# Runs as the container's root: the converter's dependencies look up the
|
||||||
|
# user's name, and a host uid with no passwd entry breaks that.
|
||||||
|
docker run --rm -v "$W":/w -w /w python:3.12-slim sh -c "
|
||||||
|
python -m venv /tmp/v &&
|
||||||
|
/tmp/v/bin/pip install -q -r llama.cpp-$LLAMA/requirements/requirements-convert_hf_to_gguf.txt \
|
||||||
|
--extra-index-url https://download.pytorch.org/whl/cpu >/dev/null &&
|
||||||
|
/tmp/v/bin/python llama.cpp-$LLAMA/convert_hf_to_gguf.py $NAME --outtype f16 --outfile $NAME-f16.gguf &&
|
||||||
|
chown $(id -u):$(id -g) $NAME-f16.gguf"
|
||||||
|
|
||||||
|
echo "== quantize to Q4_K_M"
|
||||||
|
Q=$(find llama-bin -name llama-quantize -type f | head -1)
|
||||||
|
LD_LIBRARY_PATH="$(dirname "$Q")" "$Q" "$NAME-f16.gguf" "$NAME-Q4_K_M.gguf" Q4_K_M >/dev/null
|
||||||
|
|
||||||
|
GOT=$(sha256sum "$NAME-Q4_K_M.gguf" | cut -d' ' -f1)
|
||||||
|
echo "== $W/$NAME-Q4_K_M.gguf"
|
||||||
|
echo " sha256 $GOT"
|
||||||
|
if [ "$GOT" = "$EXPECTED" ]; then
|
||||||
|
echo " matches the model production runs"
|
||||||
|
else
|
||||||
|
echo " differs from production's ($EXPECTED): a different llama.cpp or upstream file change" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
@@ -0,0 +1,46 @@
|
|||||||
|
# SPDX-FileCopyrightText: 2026 Coffey Labs
|
||||||
|
# SPDX-License-Identifier: AGPL-3.0-or-later
|
||||||
|
#
|
||||||
|
# The local language model for inbuxa's AI spam classifier: llama.cpp's
|
||||||
|
# server with Qwen3 4B Instruct 2507 (Apache-2.0), converted from Qwen's
|
||||||
|
# official weights and quantized to Q4_K_M. It runs in the mail network
|
||||||
|
# namespace and listens only on its loopback, 127.0.0.1:8080, where the mail
|
||||||
|
# server's x:AiModel "local" points. Nothing leaves the machine.
|
||||||
|
#
|
||||||
|
# Bounded so it can never starve the mail server: 4 cores, 8 GB. Four slots
|
||||||
|
# match the classifier's default of four requests in flight (inbuxa:AiLimits
|
||||||
|
# maxConcurrentCalls), each with a 4096-token context.
|
||||||
|
[Unit]
|
||||||
|
Description=inbuxa local AI model (llama.cpp, Qwen3 4B Instruct 2507)
|
||||||
|
After=mail-netns.service
|
||||||
|
Requires=mail-netns.service
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
NetworkNamespacePath=/run/netns/mail
|
||||||
|
Environment=LD_LIBRARY_PATH=/opt/llama/b11160/llama-b11160
|
||||||
|
ExecStart=/opt/llama/b11160/llama-b11160/llama-server \
|
||||||
|
--model /opt/llama/models/qwen3-4b-instruct-2507-Q4_K_M.gguf \
|
||||||
|
--alias qwen3-4b-instruct-2507 \
|
||||||
|
--host 127.0.0.1 --port 8080 \
|
||||||
|
--threads 4 --parallel 4 --ctx-size 16384 \
|
||||||
|
--no-webui
|
||||||
|
DynamicUser=yes
|
||||||
|
CPUQuota=400%
|
||||||
|
MemoryMax=8G
|
||||||
|
Nice=10
|
||||||
|
Restart=on-failure
|
||||||
|
RestartSec=10
|
||||||
|
ProtectSystem=strict
|
||||||
|
ProtectHome=yes
|
||||||
|
PrivateTmp=yes
|
||||||
|
NoNewPrivileges=yes
|
||||||
|
ProtectKernelTunables=yes
|
||||||
|
ProtectKernelModules=yes
|
||||||
|
ProtectControlGroups=yes
|
||||||
|
RestrictSUIDSGID=yes
|
||||||
|
LockPersonality=yes
|
||||||
|
CapabilityBoundingSet=
|
||||||
|
SystemCallArchitectures=native
|
||||||
|
|
||||||
|
[Install]
|
||||||
|
WantedBy=multi-user.target
|
||||||
Reference in New Issue
Block a user