Local AI: the model service production runs, and how to build its model #1

Merged
jcoffey-dev merged 1 commits from packaging/local-ai into main 2026-09-24 14:55:32 +00:00
Owner

packaging/local-ai/ holds what runs the AI spam classifier's model on
production's mail host, for the installer's bare-metal mode to use later
(nothing in the installer uses it yet):

  • inbuxa-llm.service: llama.cpp's server with Qwen3 4B Instruct 2507
    (Q4_K_M), in the mail network namespace on 127.0.0.1:8080 only, bounded
    to 4 cores and 8 GB. Identical to host1's unit apart from the license
    header.
  • build-model.sh: builds the model from Qwen's official weights, checked
    against Hugging Face's checksums, with llama.cpp b11160's converter and
    quantizer, and compares the result with production's SHA-256. Qwen
    publishes no GGUF of this model, so the file that runs is one we made.
    Two builds from the same weights came out byte-identical.
  • README.md: what runs and where it came from, a by-hand install, and how
    to undo it.
packaging/local-ai/ holds what runs the AI spam classifier's model on production's mail host, for the installer's bare-metal mode to use later (nothing in the installer uses it yet): - inbuxa-llm.service: llama.cpp's server with Qwen3 4B Instruct 2507 (Q4_K_M), in the mail network namespace on 127.0.0.1:8080 only, bounded to 4 cores and 8 GB. Identical to host1's unit apart from the license header. - build-model.sh: builds the model from Qwen's official weights, checked against Hugging Face's checksums, with llama.cpp b11160's converter and quantizer, and compares the result with production's SHA-256. Qwen publishes no GGUF of this model, so the file that runs is one we made. Two builds from the same weights came out byte-identical. - README.md: what runs and where it came from, a by-hand install, and how to undo it.
jcoffey-dev added 1 commit 2026-09-24 14:54:53 +00:00
packaging/local-ai/ holds what runs the AI spam classifier's model on
production's mail host, for the installer's bare-metal mode to use later
(nothing in the installer uses it yet):

- inbuxa-llm.service: llama.cpp's server with Qwen3 4B Instruct 2507
  (Q4_K_M), in the mail network namespace on 127.0.0.1:8080 only, bounded
  to 4 cores and 8 GB. Identical to host1's unit apart from the license
  header.
- build-model.sh: builds the model from Qwen's official weights, checked
  against Hugging Face's checksums, with llama.cpp b11160's converter and
  quantizer, and compares the result with production's SHA-256. Qwen
  publishes no GGUF of this model, so the file that runs is one we made.
  Two builds from the same weights came out byte-identical.
- README.md: what runs and where it came from, a by-hand install, and how
  to undo it.
jcoffey-dev merged commit d4351beb78 into main 2026-09-24 14:55:32 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: inbuxa/inbuxa-installer#1