Local AI Privacy · Arabic PII Evaluation
Mirogate Boundary
Inspect what your AI integration sends. A local-first text gateway with explicit policies, request-local placeholders, and Arabic-focused tests that expose missed personal information.
This is experimental, reversible pseudonymization—not a privacy guarantee. Detection can miss data. Only requests routed through the supported proxy contract are covered.
Two questions, two kinds of evidence
What did the detector miss? The diagnostic suite measures exact entity matches, annotated-character coverage, and false positives. What did the application send? Capture-server tests inspect the outgoing JSON after policy enforcement. Neither result substitutes for the other.
What ships in v0.1.0
- A Python library and CLI with a standard-library-only core.
- A pinned local OpenAI Privacy Filter adapter, Unicode-aware rules, and a hybrid mode. Model setup is explicit; failed model loading does not silently fall back to rules.
- Allow, tokenize, and block policies. Detected secrets always block. Default personal-information handling uses opaque placeholders held in a bounded, expiring, request-local vault.
- An authenticated loopback proxy for a text-only, non-streaming subset of
/v1/chat/completions. Tools, media, unknown fields, and unsupported payloads are rejected. - Optional local restoration of same-request placeholders in assistant text. Restoration is off by default, and restored output remains untrusted.
- Content-free decision receipts, English and Arabic onboarding, a local capture demo, and 99 synthetic diagnostic cases.
Published results, including the misses
The corpus has 90 labelled entities, 81 sensitive cases, and 18 negative controls. It is AI-authored development material—not held out, representative, or independently human-reviewed.
| Detector | Exact-entity recall | Character coverage | False-positive controls |
|---|---|---|---|
| Rules only | 56.67% | 64.41% | 1 of 18 |
| Privacy Filter | 61.11% | 81.72% | 3 of 18 |
| Hybrid | 81.11% | 93.11% | 4 of 18 |
Hybrid detection still left seven sensitive cases incompletely covered, including six entirely missed entities. Its exact-entity precision was 67.59%, versus 69.62% for the model alone and 82.26% for rules. Better coverage came with more false positives.
Model scores come from actual Linux CPU inference. Windows Application Control blocked the model runtime on the development host; its security policy was not changed. See the method, hashes, raw reports, and limitations before interpreting the numbers.
The release-source test matrix passed 80 tests on Windows and Linux with Python 3.11 and 3.13. These tests are not an independent security audit.
Try a local-only round trip
With Python 3.11 or later, clone the repository and create a virtual environment. Activate it using the command for your operating system, then run:
git clone --branch v0.1.0 https://github.com/mirogate/mirogate-boundary.git
cd mirogate-boundary
python -m venv .venv
# Activate .venv before the next commands.
python -m pip install -e .
python examples/capture_demo.py
The demo uses synthetic data, explicitly selects rules-only detection, and captures the exact text received by a local mock provider. It verifies local restoration without an external AI call. The separate model setup guide covers the pinned runtime and checkpoint.
Know the supported boundary
Boundary is not a system-wide network filter. It does not inspect applications that bypass it, general MCP traffic, streaming, tool calls, images, PDF/OCR, or arbitrary API schemas. Context can still identify a person after detected values are replaced. No paid-provider endpoint was used for release validation.
Read the threat model and supported request contract. Evaluate synthetic examples from your own application before considering sensitive workloads.
Build on the public work
This project turns part of Mirogate’s Secure AI Engineering Framework into an inspectable implementation. It complements the Arabic/RTL Agent Security Lab and our secure AI engineering work.
Credit to Abdullah Alsaidi’s Maskode for inspiring the exploration, and to the OpenAI Privacy Filter team. Boundary is an independent implementation, not a new foundation model or a claim to the first privacy proxy.
Contribute invented missed-entity or false-positive cases, independently review Arabic annotations, or reproduce the model evaluation. Do not upload customer records or active credentials.