13/08/2026
You can get the benefit of an LLM on sensitive text without ever handing it the sensitive part, which is important for applications in highly regulated domains.
The personal data in a document is rarely what the model needs. It needs the context: who said what, in what order, about what. So before any call goes out, we run the text through a detection pass that finds the identifying details (names, phone numbers, addresses, dates of birth ...) and swaps each one for a surrogate token. Alice becomes PERSON_1, +32 16 79 20 79 becomes PHONE_1. The sentences around them stay exactly as they were.
The trick is that the swap is reversible, not redaction:
- A keyed cipher maps every real value to a stable surrogate, so the model still reasons about the same person or place consistently across the whole text
- One key per document, reused across every related piece, so the whole set de-identifies and re-identifies together
The model works on the full context with the identities masked. When its answer comes back, we reverse the mapping with the same key, and PERSON_1 becomes Alice again before anyone reads it.
This makes the message masked for the model, meaningful for the client.
One honest caveat: anything the model paraphrases away has no surrogate left to reverse. But for everything it carries through, you get full utility on the surface and no identities underneath.