PII redaction for LLMs belongs in the gateway, not in each app
How one gateway handles PII redaction for LLMs for every tool a company runs: the order of the checks, one-way placeholders and an audit log without prompt text.
AFKzona Group · 7 min read
- The short answer
- Personal data leaves through the prompt, not through the model
- The gateway is mostly a question of order
- Injection screening catches disguised attacks and reworded ones
- Redaction makes the provider's copy as small as the task
- Evaluating PII redaction for LLMs: five questions to ask
- What generalises
- Questions
The short answer
- PII redaction for LLMs means removing personal data from a prompt before it leaves the company for a model provider, in a gateway every model call passes through.
- Injection screening, redaction and the budget check run in a fixed order, and a request that fails any of them is refused before the provider sees it.
- Injection screening catches disguised attacks by pattern and reworded ones by meaning, and decoy endpoints feed it real attempts.
- The audit log records who called which model, at what cost and with how many detections, and never the prompt itself.
In May 2025, VDAI, Lithuania's data protection authority, told organisations using AI systems that if the goal can be reached without personal data, that option should be preferred. A policy document can repeat that sentence. Only something in the path of the request can enforce it.
The gateway we built is that enforcement. It sits between a company's applications and the model provider and handles PII redaction for LLMs, injection screening, token budgets and the audit log in one place, so every tool that calls a model works under the same rules.
Personal data leaves through the prompt, not through the model
The exposure happens when someone pastes a complaint, a CV or a bank statement into a tool that calls a model. Whatever is in the request reaches the provider. A gateway puts one checkpoint on that path, so the rule is written once and applied to every application that calls a model, without anyone having to remember it.
Under Article 4(1) of the GDPR, a name and an identification number are both personal data. Article 5(1)(c) asks that data be limited to what is necessary. Summarising a complaint does not need the customer's personal code, so the gateway removes it before the text leaves.
The gateway is mostly a question of order
The work is less in detection than in sequence. Injection screening, redaction and the budget check run in that order before the model call, and a failing check stops the request with an error. A refused request never reaches the model provider: the gateway fails closed by construction, not because someone remembered to set a flag.
Four details carry most of the weight.
Redaction is one-way by design. A match becomes a placeholder such as [EMAIL] or [IBAN], and no table maps it back. There is no store of original values for anyone to leak later, and detection results never hold the raw value either.
Each tool gets the mode it needs. Block refuses a request that contains personal data, redact replaces it and sends the rest, warn counts it and lets it through. A strict preset blocks; a balanced one redacts. In a multi-tenant setup each tenant has its own personal-data mode and types, injection settings and list of allowed models.
The answer is scanned too. The model's reply passes the same detector on its way back, and streamed answers from OpenAI and Anthropic models are held and released whole, so a name split across two chunks is still caught.
The log keeps metadata, not text. Each entry records user, team, model, tokens, cost, duration and the number of detections, and blocked requests are logged with their reason. There is no column for the prompt, raw or redacted. When an injection is blocked, the log keeps at most the first 50 characters of the matched phrase: enough to see the attack, too little to hold a document.
Budgets are monthly token totals per user and team, checked before every call. Set to block, a user over the limit gets an error instead of an answer.
Injection screening catches disguised attacks and reworded ones
Prompt injection is text that tries to override the instructions the model was given. The gateway screens every user message twice over: against known attack phrasings after undoing the usual disguises, and against the meaning of known attacks, so a rewording written from scratch is caught as well as a copied one.
The pattern screen checks 21 attack phrasings after normalising the usual evasions: Cyrillic and Greek look-alike letters, zero-width characters, digits standing in for letters, and words spaced out letter by letter. "ign0re prev1ous instruct1ons" is caught the same way as the plain version. Sensitivity is set per tool, so a translation tool where staff routinely ask the model to play a role runs the 16 high-severity patterns.
The meaning screen compares each message with 25 reference attacks and blocks anything at a cosine similarity of 0.82 or higher. A company can add up to 50 examples of its own, and a separate set of decoy endpoints we built records real attempts against fake AI services and exports them as fresh examples for the reference set.
The gateway also hardens every system prompt with rules against revealing its instructions or obeying instructions placed inside user content. When it adds knowledge-base context, it tells the model to answer only from that context.
Redaction makes the provider's copy as small as the task
Data minimisation is the GDPR principle the gateway puts into practice: the provider receives what the task needs and nothing more. Personal data in Lithuanian text arrives in many shapes, so the detectors are built for them and each company adds its own.
The gateway decides what the provider receives, request by request, and the log proves it afterwards.
IBANs are typed in groups of four, mobiles as +370 or 8 6xx, and names decline: Jonas, Jono, Jonui. The gateway has 14 detector types, from e-mail, phone, IBAN and card numbers checked with the Luhn algorithm to national ID numbers, dates of birth and names written with Lithuanian letters. Each tool runs the set it needs, and company patterns add local formats and internal identifiers: contract numbers, case references, account codes. Every custom pattern is checked for runaway shapes before it runs.
The result is a small, known exposure, and a log that shows which teams and users trigger detections most often.
Evaluating PII redaction for LLMs: five questions to ask
| Question | What good looks like |
|---|---|
| Which mode does each tool run in? | Block, redact or warn, chosen per tool or tenant and written down. |
| Is it configured for your formats? | Patterns for spaced IBANs, +370 and 8 6xx numbers, personal codes and your own identifiers, proven on a test set you own. |
| Is the model's answer scanned as well as the prompt? | Yes, including streamed answers. |
| What does the audit log store? | Counts, costs, models and detections, never prompt text. |
| What happens when a check fails? | The request stops before the provider call, and the refusal is logged with its reason. |
What generalises
Put every model call through one gateway before deciding how clever its detection should be. A rule enforced in one place can be tested, logged and changed once. A rule copied into six applications has to be tested six times, and the seventh application will not have it.
Then configure it with your own data, written the way your staff write. A gateway earns trust when a team can show which identifiers it removes, which tools run in block mode and what the log recorded last month.
If your teams send customer data to model providers and you want that path set up on your own documents, see our AI governance work, ask for a technical review, or book a call.
Common questions
Can I put customer personal data into ChatGPT or another language model?
Everything in a prompt reaches the provider, so the GDPR asks for a legal basis, a processing agreement and a reason the data is needed at all. The data minimisation principle, and Lithuania's supervisory authority VDAI in its 2025 guidance on AI tools, point the same way: if the task works without personal data, remove it before the call. A gateway does that for every tool at once.
What is PII redaction for LLMs?
It is the step that finds personal data in a prompt, such as e-mail addresses, phone numbers, IBANs, personal codes and names, and replaces it with placeholders like [EMAIL] before the text is sent to the model. It works best in a gateway between your applications and the provider, so the rule is enforced once for every tool rather than separately in each.
Which personal data does an LLM gateway detect?
The gateway we built has 14 detector types, from e-mail addresses, phone numbers and IBANs to card numbers checked with the Luhn algorithm, national ID numbers, dates of birth and names written with Lithuanian letters. Each tool runs the set it needs, and a company adds patterns for its own identifiers, such as contract numbers or case references.
How does a gateway detect prompt injection?
It compares user messages with known attack phrasings such as 'ignore previous instructions', after normalising tricks like look-alike letters, invisible characters and spaced-out words. A second screen compares the message's meaning with a set of reference attacks, which catches rewordings no fixed phrase list contains. Decoy endpoints that record real attempts supply fresh examples for that set.
What should an AI audit log store?
Enough to answer who used which model, when, at what cost and with what result, and nothing that turns the log itself into a copy of sensitive documents. Our gateway records user, team, model, tokens, cost, duration and the number of detections for every call, logs blocked requests with their reason, and has no column for prompt text at all.