LLM Data Security: Risks, Controls and a Practical Checklist

Updated · EraseAI

LLM data security is about protecting information as it flows into, through and out of large language models. Whether you use a chat assistant or build your own AI features, the same questions apply: what data goes in, where it is stored, who can see it, and what can come back out.

How data moves through an LLM

A message you send to an AI assistant typically passes through several stages, and each one is a place data can persist:

  1. Your device and browser, including extensions and clipboard history.
  2. The provider's application, which stores the conversation so you can see it later and may keep it for abuse monitoring.
  3. The model, which processes the text. Depending on the provider and plan, conversations may be used to train or improve future models unless you opt out.
  4. Connected tools: web search, file stores, plugins and agents that can read from or write to other systems.
  5. The output, which may be copied, shared by link, or indexed if a share page is public.

The main risks

  • Sensitive input. People paste secrets, personal data and confidential documents. This is the most common and most preventable risk. See AI data loss prevention.
  • Retention you didn't expect. Deleting a chat in the interface does not always mean the provider deleted it. In 2025 a US court ordered OpenAI to preserve ChatGPT conversation logs, including deleted ones, as part of ongoing litigation.
  • Training and human review. Consumer AI plans may use conversations to improve models, and some providers state that human reviewers may read samples. Business plans usually exclude training by default. Check the current terms of each tool you use.
  • Shared and indexed conversations. Share links have exposed private chats to search engines when they were made public.
  • Prompt injection in documents, web pages and emails that an AI tool reads, which can make it leak data or take actions. See prompt injection explained.
  • Over-broad integrations. An assistant connected to your email, drive or ticketing system can surface data to people who should not see it.

Controls that work

  • Minimize what goes in. Check messages and files before they are sent and redact what the task doesn't need. A PII redaction step keeps the question and drops the identity.
  • Choose the right plan. Use business or API tiers with no-training commitments, retention controls and audit logs for work data.
  • Turn off training and history on personal accounts where the provider allows it.
  • Limit integrations to the data each team actually needs, and review permissions like you would for any other app.
  • Treat outputs as untrusted when they come from content you did not write, and never let an AI agent act on sensitive systems without confirmation.
  • Write it down. A short AI acceptable use policy makes expectations clear.

Checklist

  1. List every AI tool in use, and who uses it.
  2. Decide which are approved, on which plans.
  3. Confirm training, retention and sharing settings for each.
  4. Put an automatic check on what is sent from browsers and phones.
  5. Scan inputs and outputs in AI features you build, before they reach the model.
  6. Restrict connectors and agents to least privilege.
  7. Review findings monthly and update the policy.

Check every message before it reaches AI

EraseAI stops API keys, passwords, card numbers and personal data in ChatGPT, Claude and Gemini. Free in Chrome, no account needed.

Frequently asked questions

Do LLMs remember what I type?

The model itself doesn't keep a memory of one conversation for other users, but the provider stores conversations, and depending on the plan they may be used to train future models. Some assistants also have a memory feature that recalls details across your own chats.

Is the API safer than the chat app?

Usually, for data handling: API and business plans from major providers generally don't train on your data by default and offer retention controls. The data still leaves your environment, so minimizing sensitive input still matters.

Related guides