jcoffey-dev is traveling from Thursday 1 October through Sunday 4 October. Issues and pull requests are welcome, and will get an answer after that. Thanks for your patience.
Personal-data catalog spec, §6 (Phase 3c).
inbuxa:DataInventory/get evaluates the catalog against the server's
live settings and says what this server holds: for each source and each
object that can hold personal data, its categories and whose data it
is, whether it is collected here at all, what bounds its retention (the
live value of the setting that does, or unbounded), whether it leaves
the host and to which endpoints, and a summary. Every host that
receives something is listed once as a candidate processor with what it
receives. Inside a tenant it answers with the tenant's slice and none
of the server's processors. Read-only, with sysComplianceGet.
inbuxa:InventorySnapshot/get is the history: a dated copy of the
evaluated inventory, recorded when it changes -- after a registry write
to an object the inventory reads, after inbuxa's log, audit or AI
settings change, and on the daily clean-up -- and kept as long as the
audit log's records. ids: null lists every snapshot, newest first; the
full inventory only when asked for.
The catalog is embedded and parsed at start (new dependency: toml,
MIT/Apache); the evaluation is a pure function of it and the live
facts, so each configuration is tested without a server. Loopback
endpoints stay on the host; any other configured endpoint leaves it.
Tested: unit tests for the evaluation (a new install's defaults, an
external blob store, a hosted AI endpoint, telemetry off, a tenant's
slice, hosts from URLs, loopback), snapshots, and the fact gathering's
store and duration rules; the compliance system test, extended (the
officer reads the inventory, a plain user is refused, a tenant's
officer sees its slice and no processors, a webhook to another host
becomes a processor and a snapshot names x:WebHook, a retention change
reads through); the system, audit, legal hold and account lock suites;
fork checks. The system suite failed once of three runs with an email
import's blob not found, in antispam.rs; the same happened once in
purge.rs on the previous branch. Nothing here touches uploads; noted
for a separate look.
Personal-data catalog spec, default D1 (settled 2026-09-28): log files
were never deleted. inbuxa:LogSettings.keepForDays says how many days
rotated log files are kept; unset (null) keeps every file, as before,
and a new install sets 30 days.
It is a fork-owned setting, stored under T + l as audit retention is,
not a field on x:TracerLog: that object is also stored inside
x:Bootstrap with a field after it, so a new field would change
x:Bootstrap's stored format. Server-level, with the tracers'
permissions (sysTracerGet, sysTracerUpdate); changes are in the audit
log, before and after.
Log files are local, so every node deletes its own: hourly, and at once
when the setting changes on that node. Only regular files named
<prefix>.<something> in each enabled log tracer's directory, last
changed more than the limit ago, are removed; the file being written is
never that old, and nothing else in the directory is touched. Minimum
one day. The catalog classifies inbuxa:LogSettings and points the log
file's retention at it.
Tested: unit tests for the file rule (only this log's old files; the
current file, other files and directories stay) and a purge on disk;
the system suite, which reads, sets, refuses zero, restores null and
checks the audit records; fork checks.
The schema, which both sides changed, merged as JSON with no conflicts.
The personal-data catalog (#83) gains upstream's new x:DnsServerPowerDns:
nothing personal but its API key, like the other DNS providers.
Phase 2 of the personal-data catalog spec.
resources/privacy/catalog.toml classifies every object in the schema
(316) and inbuxa's own JMAP objects (12): each property that can hold
personal data, with its categories, and for objects that hold any,
whose data it is, where it lives, its scope and what bounds its
retention (a named setting where there is one). Twenty sources that
are no object -- the log file, exporters, webhooks, spam lookups, the
Explain cache, relays and hooks, push, legacy-use records -- carry the
same facts plus the settings that turn them on, whether the data
leaves the host, and the code that writes it. Classifications of
objects that hold data about people are from the spec's source map;
the rest are typed from the schema alone (address, IP, secret).
tools/fork/privacy-check.py fails CI when an object or inbuxa object
has no entry, when a property the schema types as an address, IP or
secret is left to its object's default, when an entry names an
object, property, setting or code path that is gone, or when it uses
a word outside the catalog's vocabulary. --unlisted prints starting
entries. strip.py's report gains "Unclassified in the privacy
catalog": objects and fields new in an import and not classified,
informational like the Enterprise flags.
Tested: 13 unit tests (tools/fork/tests): the check passes on this
tree; fails on an unclassified object, an address hidden behind a
default, a secret in a set or object reference, stale properties,
objects, settings and code paths, an unlisted inbuxa object and a
word outside the vocabulary; --unlisted's entries; and the strip
report on a synthetic import. The check and the tests run in the
fork-checks job.