Beyond Encryption: Protecting Proprietary Data inside the LLM Inference Loop

Beyond Encryption: Protecting Proprietary Data inside the LLM Inference Loop
Beyond Encryption: Protecting Proprietary Data inside the LLM Inference Loop

Ask most security teams how their AI data is protected, and the answer covers two states well: encrypted at rest in storage, and encrypted in transit over the network. Ask what protects it during the third state — the moment it's decrypted, sitting in GPU memory, and actively being processed by a model — and the answer is usually much less confident. That gap is the inference loop, and it's the part of the AI data lifecycle most security architectures haven't caught up to yet.

The State Encryption Doesn't Cover

Encryption at rest protects data sitting in storage. Encryption in transit protects data moving over a network. Neither protects data during computation, because computation requires the data to be decrypted in memory — a model can't run inference on ciphertext it can't read.

This "data in use" state has always existed, but it mattered less when the computation involved was a database query or a simple business logic function. It matters enormously now, because LLM inference means an entire proprietary document, customer record, or trade secret sits decrypted in GPU memory for the full duration of a forward pass, and that memory is a real attack surface with real access paths: a compromised hypervisor, a misconfigured multi-tenant GPU scheduler, a malicious co-tenant on shared infrastructure, or simply an operator with more access than they should have.

What Actually Happens During Inference

When a prompt is sent to a model, the request payload — which may contain the exact proprietary content an organization is most protective of — gets decrypted, tokenized, loaded into GPU memory alongside the model weights, and processed through the full forward pass.

Intermediate activations, attention states, and the eventual output all exist in plaintext in memory during this window. On shared or multi-tenant infrastructure, the question of who else has any access path to that memory space — even briefly, even accidentally — is a real security question, not a theoretical one.

This matters differently depending on what's actually at risk. For consumer-facing chat applications processing low-sensitivity queries, this gap is a minor theoretical concern. For an enterprise running inference over unreleased financial results, unfiled patent language, or protected health information, it's a genuine architectural requirement that "encrypted at rest and in transit" simply doesn't satisfy — the same principle behind why regulated industries increasingly require on-premise or sovereign AI infrastructure in the first place.

Confidential Computing: Closing the Gap

Confidential computing addresses this directly by extending encryption into the compute layer itself, using hardware-based trusted execution environments (TEEs) that keep data encrypted even while it's being processed. Modern implementations — from CPU-level TEEs to GPU-level confidential computing modes on recent accelerator generations — create an isolated, attested execution environment where data is decrypted only inside hardware-enforced boundaries that even the infrastructure operator can't inspect.

The practical effect: even on shared infrastructure, even with a cloud provider or platform operator in the loop, the actual content of a request and the intermediate state of inference remain opaque to everything outside the trusted execution boundary — including, critically, the platform operator's own administrative access.

This is a materially different guarantee than "we have strict internal access controls," because it removes the need to trust that those controls are correctly configured and never bypassed.

Beyond TEEs: The Rest of the Inference-Loop Threat Model

Confidential computing addresses the hardware memory exposure problem, but a complete inference-loop protection strategy has to cover a few more specific risks:

  • Prompt and output logging. The most common real-world leak isn't a sophisticated memory attack — it's a logging or observability pipeline that captures full request and response content by default, for debugging or analytics purposes, and retains it indefinitely without anyone treating it as sensitive data. This is an operational discipline problem more than a hardware one, and it's usually the highest-probability leak path in practice.
  • Model weight protection. For organizations that fine-tune models on proprietary data, the resulting weights can themselves encode sensitive information, and protecting them requires the same rigor as protecting the training data itself — access control, encryption at rest for model artifacts, and monitoring for unauthorized model extraction attempts.
  • Multi-tenant isolation guarantees. On shared inference infrastructure, isolation between tenants needs to be verifiable, not just claimed — attestation mechanisms that let a customer confirm their workload is actually running in an isolated, confidential environment, rather than trusting a provider's description of their architecture.
  • Output-side data minimization. Inference doesn't just consume sensitive data — it can inadvertently reproduce it in outputs, particularly for fine-tuned models that may have memorized training examples. Output filtering and monitoring for unintended verbatim reproduction of sensitive training content is a distinct control from anything protecting the input side.

Where This Fits Into a Broader Data Protection Strategy

Inference-loop protection is one layer in a stack that also includes the access control, logging discipline, and data flow mapping we've covered in the context of self-hosting LLMs for regulated industries.

None of these controls substitutes for the others — confidential computing protects data during processing, but a logging pipeline that captures full plaintext content downstream of that processing defeats the purpose entirely.

A genuinely defensible architecture treats the full lifecycle — at rest, in transit, in use, and downstream in logs and outputs — as one continuous chain, not a set of independently satisfied checkboxes.

Conclusion

"We encrypt everything" is an incomplete claim if it only covers data at rest and in transit, because LLM inference necessarily involves a state where data has to be readable to be processed — and that state has real, exploitable access paths on shared infrastructure.

Confidential computing closes the hardware-level gap; disciplined logging, model weight protection, and output monitoring close the rest. Organizations handling genuinely sensitive data in AI workloads need to ask, specifically, what protects that data during the moments it's actually being reasoned over — not just before and after.

Want to understand what confidential computing and inference-loop protections would look like for your specific workload? Book a call with our team or explore the documentation to get started.

Learn more at