PII Redaction for AI Prompts: What to Remove and How
Updated · EraseAI
PII redaction means removing or replacing personally identifiable information before text is stored or shared. For AI, it is the single most effective way to get the benefit of a model without handing it someone's identity: the model can still summarize the complaint, draft the reply or find the bug, it just never sees who it is about.
What counts as PII
Treat secrets the same way even though they aren't personal data: API keys, passwords and tokens should never reach an AI tool.
- Direct identifiers: full name, email address, phone number, home address, national ID, passport and driving licence numbers.
- Financial identifiers: payment card numbers, bank account and IBAN numbers.
- Sensitive categories: health information, biometric data, religion, ethnicity, sexual orientation. Under GDPR these are "special category" data with stricter rules.
- Indirect identifiers: date of birth, job title plus employer, customer or patient IDs. One alone may be harmless; combined they can identify a person.
Redaction, masking and pseudonymization
- Redaction removes the value or replaces it with a label: "Call Sarah Lee on +65 8123 4567" becomes "Call [NAME] on [PHONE]". Simple and safe; the model still understands the sentence.
- Masking hides part of a value, like showing only the last four digits of a card.
- Pseudonymization replaces values with consistent stand-ins ("Customer A", "Customer B") so relationships survive. Useful for data analysis; keep the mapping table away from the AI tool.
How to automate it
EraseAI does this at the Send button in ChatGPT, Claude and Gemini: it lists what it found and offers Sanitize & Send, which replaces each item with a label and sends the rest. Developers can call the same detection from their own apps through the EraseAI API.
- Detect with a mix of patterns (card numbers with checksum validation, key formats, phone and ID formats) and context ("my password is …").
- Show the person what was found so they can confirm. Silent rewriting surprises people and hides false positives.
- Replace with readable labels like [EMAIL] or [API KEY] so the prompt still makes sense.
- Check attachments too: text files, PDFs, spreadsheets and screenshots carry as much PII as typed messages.
- Keep the original on the device. The point is that the identity never leaves.
Common mistakes
- Redacting the name but leaving the email address that contains it.
- Forgetting screenshots and PDFs.
- Pasting the redacted text, then the original "for context" in the next message.
- Relying on the AI to redact for you: by then the data has already been sent.
Check every message before it reaches AI
EraseAI stops API keys, passwords, card numbers and personal data in ChatGPT, Claude and Gemini. Free in Chrome, no account needed.
Frequently asked questions
Does redacted data still count as personal data under GDPR?
Properly anonymized data does not, but pseudonymized data still does, because it can be re-identified with the mapping. Redaction that removes identifiers entirely moves text towards anonymization.
Will redaction make AI answers worse?
Rarely. Models answer "[NAME] was charged twice on [CARD]" just as well as the original. Keep labels descriptive so the meaning survives.