AGENSPHERE/ JOURNAL
← JOURNAL
ENTERPRISE AI · PART 4 OF 9
J-028NOTE4 MIN READ

Data compliance with LLMs: know where every prompt goes

Every prompt is a data transfer. Customer records, contracts and employee details leave your systems the moment they are pasted into a model. LLM data privacy starts with a gateway that sees, redacts, routes and records every call.

IN SHORT
  • Using an LLM sends data to a processor. Compliance questions follow: what personal or confidential data is sent, to which provider, in which region, how long it is retained, and who can see it.
  • Risk comes less from the model and more from the paths around it: employees pasting data into unapproved tools, prompts and logs stored without retention rules, and retrieval that ignores permissions.
  • A central AI gateway gives one place to classify and redact sensitive data, route requests to approved providers and regions, apply retention, and log every call.
  • Laws such as the GDPR, India's DPDP Act and the EU AI Act raise the bar for knowing and documenting how data is used. Engineering controls make that provable.

This is part 4 of Enterprise AI. The first three parts covered cost, context and ownership. This part covers the question legal and security teams ask first, and often too late: where does our data go when someone uses AI?

This is an engineering note, not legal advice. Work with your legal and privacy teams on what your obligations are.

01Why is every LLM prompt a data privacy question?

When an application or an employee sends a prompt to a hosted model, the content of that prompt leaves your environment and is processed by the provider. If the prompt contains a customer's name and order history, an employee's performance review or a confidential contract, that data has now been shared with a processor.

That brings the usual data protection questions into every AI call:

  • What personal or confidential data is in the prompt?
  • Where is it processed, and does that region meet your residency requirements?
  • How long is it retained by the provider, and is it used for training?
  • Who inside your company and the provider can see prompts and outputs?
  • Can you prove all of the above to an auditor or a regulator?

Enterprise agreements with major providers typically address training and retention terms. That covers the provider. It does not cover the rest of the path.

02Where does data actually leak?

Usually not through the model. Through the paths around it:

Unapproved tools. Employees paste data into whatever AI tool is convenient, including personal accounts with consumer terms. This is the “shadow AI” problem in part 5.

Logs nobody governs. Applications log full prompts and responses for debugging, and those logs end up in observability tools with long retention and broad access.

Retrieval without permissions. A RAG assistant that searches everything can surface HR or finance documents to anyone who asks the right question (see filter by permission before you rank).

Over-sharing in context. Systems send whole records when the task needed two fields, because it was easier.

Outputs that travel. Generated summaries containing personal data get emailed, pasted into tickets or stored in new places, each a new copy to govern.

03What does a compliance-ready architecture look like?

Put one AI gateway between every application and every model provider. Nothing calls a model directly. The gateway is where policy is enforced and evidence is collected.

gateway/policy.py
def handle(request: AIRequest, caller: Caller) -> AIResponse:
    policy = policies.for_app(caller.app)                       # approved providers, regions, data classes
    findings = classify(request.content)                        # PII, financial, health, secrets
    if findings.blocked_for(policy):
        return deny(request, reason=findings.summary())         # e.g. health data to an unapproved region
    redacted, vault_map = redact(request.content, findings, mode=policy.redaction)
    provider = select_provider(request.task, region=policy.region, data_class=findings.max_class)
    response = provider.call(redacted)
    restored = reinsert(response, vault_map) if policy.allow_reinsert else response
    audit.log(caller, request.task, provider, findings, retention=policy.retention)  # no raw content by default
    return restored

The gateway's responsibilities:

  1. Classify content for personal data, financial data, credentials and other sensitive classes.
  2. Minimise and redact. Replace names, emails, account numbers and IDs with placeholders before the call, and restore them in the response where the use case allows.
  3. Route by policy. Send each data class only to approved providers and regions; keep the most sensitive work on private or self-hosted models if required.
  4. Apply retention. Decide what is logged, for how long, and who can read it. Default to metadata, not raw content.
  5. Record evidence. Every call logged with caller, purpose, provider, region and data classes found, so a data protection review is a query, not an investigation.

04What do the regulations change?

Regulation varies by country and sector, but the direction is consistent. Data protection laws such as the EU's GDPR and India's Digital Personal Data Protection Act require a lawful purpose for processing, minimisation, security safeguards, and accountability for processors. The EU AI Act adds obligations based on how AI is used, with stricter requirements for high-risk uses, including documentation and record-keeping.

Common to all of them: you must know what data you process, limit it to what is necessary, protect it, and demonstrate that you did. A gateway with classification, routing and audit logs turns those obligations into things your systems do on every call.

Compliance is not a clause in a vendor contract. It is a property of every path your data takes.

05Where does this fit?

The gateway sits in L0 infrastructure; classification, minimisation and permission-aware retrieval sit in L1 data and knowledge. The same gateway also does routing (part 1) and auditing (part 8), which is why it is the foundation of the owned intelligence layer.

Previous: Unowned intelligence. Next: Shadow AI.

Building something like this?DESCRIBE A SYSTEM →