jcoffey-dev is traveling from Thursday 1 October through Sunday 4 October. Issues and pull requests are welcome, and will get an answer after that. Thanks for your patience.
Phase 2b of the DLP and mail flow rules spec: every identifier in the
§2.3 catalog, each implemented from its issuer's published rules and
tested against published examples.
US (SSN, ITIN, EIN, ABA routing, driver's licenses, MBI, NPI, DEA), UK
(NI number, NHS number, UTR), Canada (SIN), Australia (TFN, Medicare),
the EU (Germany's tax ID and ID card, France's NIR, Spain's DNI/NIE,
Italy's codice fiscale, the Dutch BSN, Belgium's national number,
Poland's PESEL, Sweden's personnummer, Denmark's CPR, Finland's HETU,
Ireland's PPS, Portugal's NIF, Austria's SVNR), Norway, Switzerland,
India (Aadhaar, PAN), China, Japan, Singapore, South Korea, Brazil (CPF,
CNPJ), Mexico (CURP) and South Africa. 49 detectors in all, plus seven
templates named for what they find.
An identifier that is only digits and whose check about one random
number in ten passes counts alone only in its written form
(536-22-1234, 943 476 5919) and as bare digits only beside a word; ABA
routing numbers and NPIs always need one. Spec §2.3 records this.
A test runs every detector over an ordinary business email (order,
invoice and tracking numbers, dates, amounts, an address) and requires
nothing to fire but the contact detectors. 47 unit tests.
Phase 2a of the DLP and mail flow rules spec: pure functions in
crates/features/src/mailflow, nothing wired into the mail path yet.
- Detectors report distinct values found, each either checked by its
published check digit or counted only beside a corroborating word
within 50 characters. This PR adds the region-free ones: payment
cards (issuer prefixes, Luhn), IBAN (registry lengths, mod 97),
SWIFT/BIC, email addresses and phone numbers in bulk, dates of birth,
passport numbers, private keys and published service-token formats.
Regional identifiers follow, a region per PR.
- Word lists (Aho-Corasick, whole words, any case) and patterns (regex
with a compiled-size limit) count occurrences.
- Attachment text: text files with or without a UTF-16 mark, HTML,
DOCX/XLSX/PPTX, ODT/ODS/ODP and ZIP archives one level deep, read
with the zip and quick-xml crates the workspace already has.
Encrypted files, PDF, legacy binary Office files, nested archives
and anything past the limits come back as not inspectable, with why.
21 unit tests, against the networks' test card numbers and the IBAN
registry's own examples among others.