Cohere Adds Encrypted AI Inference It Says Even It Cannot Read

The newly available Model Vault capability aims to protect data while a model processes it, but customers will need to assess exactly what its attestation evidence covers.

By 3 min read
Cohere Adds Encrypted AI Inference It Says Even It Cannot Read
Cohere Adds Encrypted AI Inference It Says Even It Cannot Read

Listen to this story

The audio brief

About 1:38
0:001:38
Read transcript
Cohere says it can now process enterprise AI prompts without being able to read them itself. The capability, called the Encrypted tier, is newly available in Cohere’s single-tenant Model Vault service. It uses confidential computing to protect data not only while it is stored or moving across a network, but also while a model is actively processing it in memory. That matters because ordinary encryption normally has to be removed at the point where the AI runs. Cohere says its design keeps data encrypted outside a protected execution environment, putting Cohere staff, cloud providers, and cluster operators outside that boundary. The system combines a confidential virtual machine, built with Intel TDX or AMD SEV-SNP, with NVIDIA GPUs running in confidential-computing mode. Cohere says memory and the connections between the CPU and GPU are protected, with data decrypted only inside those environments. Customers keep the keys needed to decrypt returned responses. The other half of the pitch is verification. Each inference returns an attestation report covering the hardware, software, and security policies protecting that workload. But the public description does not say whether that check happens when the system boots, for every request, or on another schedule. Cohere plans to open-source the Model Vault serving stack, so customers can examine logging, data export, and leakage risks. Existing Model Vault customers pay no extra, but the key question is whether the evidence and hardware assumptions satisfy their security requirements.

Story brief

3 key points

Cohere has added confidential inference to its single-tenant Model Vault service, extending protection beyond data at rest and in transit to the CPU and GPU memory used during processing. The company says operators cannot read customer prompts, while customers retain keys for returned responses and receive attestation reports for verification. The tier costs existing Model Vault customers no extra, but buyers still...

  1. 01

    Model Vault combines Intel TDX or AMD SEV-SNP confidential VMs with NVIDIA confidential GPUs.

  2. 02

    Attestation reports cover the hardware, software, and security policies protecting each inference workload.

  3. 03

    The published account does not clarify whether attestation is checked at boot, per request, or on another schedule.

Cohere has made confidential-computing support newly available in Model Vault, its single-tenant inference service. The company says the Encrypted tier protects customer data during the usually exposed moment when an AI model reads and processes a prompt, putting Cohere, cloud providers, and other infrastructure operators outside that boundary.

That is a different proposition from ordinary encryption. Data is commonly protected when stored and sent over a network, but it must normally become readable in system memory for a model to run. Cohere says its new design keeps data encrypted outside a protected execution environment, so its staff, the cloud provider, and cluster operators cannot access it.

A managed service with a harder privacy claim

Model Vault already gives a customer a dedicated model deployment rather than a shared API. Confidential computing targets a separate concern: whether the people operating the underlying infrastructure can inspect the workload. Cohere’s approach combines a confidential virtual machine using Intel TDX or AMD SEV-SNP with NVIDIA GPUs in confidential-computing mode.

The GPU portion is important because inference does not stay on a CPU. Cohere says its design protects memory and CPU-to-GPU connections, with data decrypted only where processing must occur inside the protected CPU and GPU environment. Customers retain the keys to decrypt returned responses, according to the company.

The report matters as much as the enclosure

Cohere’s pitch is not simply that customers should trust the boundary. It says every inference returns an attestation report that customers can use to check the hardware, software, and security policies safeguarding a workload. The company says approved hardware and software configuration values are registered with Intel Trust Authority when a confidential VM starts.

What customers can verify—and what remains open

  • Cohere says attestation evidence identifies the hardware, software, and security policies protecting an inference workload.
  • It is not clear from the published account whether verification occurs at boot, on every request, or through another cadence.
  • Cohere plans to open-source the full Model Vault serving stack for independent assessment of logging, export, and leakage risks.

A crowded category, aimed at Cohere’s own buyers

Confidential inference is not a new category. The public account names Tinfoil, Edgeless Systems’ Privatemode, Phala, Maple AI, and Confer as providers offering related privacy-oriented services. Cohere’s distinction is applying the approach to its own enterprise models in its existing dedicated serving product, rather than offering an independent privacy service around other models.

Cohere says the capability carries no added cost for existing Model Vault customers. But price parity does not settle the procurement question. For organizations handling sensitive data, the practical test is whether the attestation evidence, hardware assumptions, and future serving-stack transparency meet their own security requirements.

Sources

  1. venturebeat.comCohere's Model Vault now encrypts AI inference so even Cohere cannot see enterprise customers' data

Loading discussion...

Cohere Adds Encrypted AI Inference It Says Even It Cannot Read | Superpower Daily