AI Privacy · · By Mirogate

Testing an AI privacy gateway: Arabic PII and outgoing requests

A detector can miss a name. An integration can send an uninspected field. These are different failures, and an AI privacy gateway needs evidence for both.

Mirogate Boundary is an experimental, open-source Python library, CLI, and local text-chat proxy. It turns one part of our Secure AI Engineering Framework into executable checks: detect locally, apply a policy, then inspect the supported outgoing request. The useful question is not whether a demo looks anonymized. It is what the system can miss, reject, or transmit.

Define the boundary before connecting a client

Version 0.1 accepts a small, non-streaming /v1/chat/completions subset. System, developer, user, and assistant messages must contain string text. Tools, images, file references, content arrays, streaming, and unknown fields are rejected, not forwarded unchanged. The model identifier is inspected too. This deliberate contract makes rejection behavior testable; it also means Boundary is not a drop-in replacement for every SDK or coding agent.

The authenticated server binds only to loopback and uses one explicitly configured upstream. A local client token stays separate from the provider credential. Traffic sent directly by another process, telemetry, shell commands, and MCP requests remain outside this boundary. A local proxy is not an operating-system firewall.

Detection, policy, and restoration have separate jobs

The default hybrid detector combines Unicode-aware rules with local OpenAI Privacy Filter inference. Model installation is explicit; a missing model or detector failure does not silently select rules-only mode. Policy can allow, tokenize, or block detected categories. Detected secrets always block. The qualification matters: a missed secret is still missed, even when transport validation fails closed.

Random placeholders refer to an expiring, bounded, in-memory vault shared only within one request. Restoration is off by default and, when enabled, applies only to visible assistant text—not tool arguments. This is reversible pseudonymization, not irreversible anonymization. Context may still identify someone, and restored model output remains untrusted.

Measure Arabic PII without hiding negative results

The release includes 99 Arabic-focused and mixed-language synthetic cases: 90 annotated entities across 81 sensitive cases, plus 18 negative controls. This is AI-authored development material, not a held-out population sample; annotations have not been independently human-reviewed. It supports reproducible diagnostics, not a claim about general Arabic accuracy.

v0.1 detection results on the same synthetic corpus
MetricRulesPrivacy FilterHybrid
Strict entity precision82.26%69.62%67.59%
Strict entity recall56.67%61.11%81.11%
Annotated-character coverage64.41%81.72%93.11%
Fully covered sensitive cases48 / 8166 / 8174 / 81
Negative controls flagged1 / 183 / 184 / 18
Unlabelled-character redaction1.74%7.27%8.92%

Hybrid covered more annotated text but produced more false positives. Seven sensitive cases remained incompletely covered, leaving 121 annotated characters uncovered. Four of 18 negative controls were flagged. These results do not establish leak-free behavior or a privacy guarantee.

Strict precision and recall require the correct entity boundaries and category. Character coverage asks how much labelled text overlaps a prediction, regardless of category. Distinct overlapping proposals count in strict scoring, while masking merges overlaps. Reporting both avoids confusing an overlong mask with a correctly identified entity.

The evaluation report and raw results preserve case IDs, corpus and source hashes, dependency versions, and calibration settings. Actual model inference ran on Linux CPU with pinned weights and the official Viterbi runtime. A separate Windows attempt was blocked by Application Control; its protections were left unchanged. The repository provides the reproducible workflow, including a preliminary smoke-test miss.

Test the bytes that leave, not only the labels

Detection scores describe recognized spans. They do not prove that an application used the proxy or that no other field carried private data. Boundary separately tests supported request handling with local capture servers: transformed payloads, rejected inputs, policy blocks, and restoration. The release passes 80 tests across Windows and Ubuntu with Python 3.11 and 3.13. These tests are neither a production audit nor evidence of paid-provider compatibility.

Content-free receipts record decisions and category counts without original values or transformed text. A hash chain can be checked against an independently stored head. Without that anchor, someone controlling the receipt file can rewrite or truncate it. Receipts also cannot prove that all network traffic passed through Boundary.

Start with a synthetic, local capture

Use Python 3.11 or newer and a fresh virtual environment. The following demo uses limited rules and a local mock provider; it needs no model download, provider key, or external AI call.

git clone --branch v0.1.0 https://github.com/mirogate/mirogate-boundary.git
cd mirogate-boundary
python -m venv .venv
# Linux/macOS: source .venv/bin/activate
# PowerShell: .\.venv\Scripts\Activate.ps1
python -m pip install -e .
python examples/capture_demo.py

Activate the environment using the line appropriate to your shell before installing. The demo displays the transformed text received by the mock provider and checks restoration. For model-backed evaluation, follow the pinned setup instructions. Before connecting an application, read the proxy contract and threat model.

A decision checklist for your integration

Our secure AI engineering work starts with those system boundaries. A detector adds a control; it does not replace data minimization, access control, or an accountable deployment decision.

Make the next contribution a reproducible failure

Credit goes to Abdullah Alsaidi's Maskode for inspiring this exploration, and to OpenAI Privacy Filter for the local model. Boundary is independently implemented and claims no endorsement. Synthetic dialect, transliteration, contextual-name, and false-positive cases are welcome; never submit customer records or real secrets.

Explore the source repository, get the experimental v0.1.0 release, read the Mirogate launch story on Medium, or join the discussion on Mirogate's LinkedIn announcement.