Productspublished

Nutanix Enterprise AI 2.8 Adds Agent Token Quotas and MCP Access Controls

Nutanix is putting controls around what agents can reach and consume, but says secure backend infrastructure, API keys and roles remain essential.

By 2 min read
Nutanix Enterprise AI 2.8 Adds Agent Token Quotas and MCP Access Controls

Listen to this story

The audio brief

About 1:43
0:001:43
Read transcript
Nutanix is adding token quotas and access controls for AI agents, giving enterprises a central way to limit what those agents can use—and how much they can consume. Enterprise AI 2.8 is now generally available with an MCP gateway inside Nutanix Agent Gateway, plus a separate Nutanix Cloud Platform MCP Server for controlled access to infrastructure managed by Nutanix software. The practical change is governance. Companies can set token-consumption limits by team, user, or individual agent, and expanded MCP logs can show which data sources and applications an agent accessed, along with the actions it took. But these logs focus on activity at the connection layer; according to Thomas Cornely, Nutanix’s executive vice president of product management, inspecting prompts is not the product’s default purpose. That distinction matters because the gateway is not a complete security boundary. Nutanix says MCP servers are only as secure as the backend infrastructure, API keys, and role-based access controls behind them. The release also broadens private inference. Customers can fine-tune models below eight billion parameters with LoRA, and use multi-GPU, batch, and speculative decoding. Nutanix says speculative decoding can deliver up to two-and-a-half times faster token generation, depending on the model and infrastructure. Private, open-weight models can also shift some workloads from per-token billing toward infrastructure costs. The next constraint to watch is deployment: Nutanix Kubernetes Platform 2.19 is planned to bring Kubeflow, Milvus, and Slurm across virtualized and bare-metal environments, while the underlying servers, keys, and roles still carry the security burden.

Story brief

3 key points

Nutanix’s Enterprise AI 2.8 gives organizations finer control over agent consumption and MCP-connected tools, including quotas by team, user, or agent and logs of accessed data sources, applications, and actions. The release also broadens private inference with LoRA tuning for models under 8 billion parameters and multi-GPU, batch, and speculative decoding options; Nutanix claims up to 2.5× faster generation in...

  1. 01

    MCP activity logs can capture data sources, applications, and actions, while prompt inspection is not the default focus.

  2. 02

    Private Inference shifts some workloads from per-token billing to infrastructure costs using privately deployed open-weight models.

  3. 03

    LoRA fine-tuning is limited to models below 8 billion parameters.

Nutanix has made Enterprise AI 2.8 generally available with an MCP gateway, token tracking and quotas, expanded audit logs, and private-inference features. The release creates a central layer for controlling agent access and consumption, but the systems behind those connections remain the security boundary.

A gate for tools, data and infrastructure

The generally available Model Context Protocol gateway sits in Nutanix Agent Gateway and is designed to govern an agent’s access to tools and data. A separate Nutanix Cloud Platform MCP Server gives agents controlled access to infrastructure managed by Nutanix software. Enterprise AI 2.8 can also track token use and apply consumption quotas at the team, user or agent level.

The logging is aimed at the connections agents use. Nutanix says MCP activity logs can record the data sources and applications an agent accessed, along with its actions. Thomas Cornely, Nutanix’s executive vice president of product management, said inspecting prompts is not the product’s default purpose.

Private inference offers a different cost model

Private Inference adds Low-Rank Adaptation fine-tuning for models with fewer than 8 billion parameters, multi-GPU inference using tensor parallelism, batch inference and speculative decoding. Cornely said basic work such as Python scripting can run on privately deployed open-weight models, where customers pay for infrastructure rather than individual tokens.

Private Inference additions

  • LoRA fine-tuning for models below the 8-billion-parameter threshold.
  • Multi-GPU inference, batch inference and speculative decoding.
  • Nutanix claims speculative decoding can increase token-generation speed by up to 2.5 times, depending on the model and infrastructure configuration.

Kubernetes management is the next release

Nutanix Kubernetes Platform 2.19 is scheduled for an upcoming release and will extend Kubernetes management across virtualized and bare-metal environments. Its application catalog will offer curated deployments of Kubeflow, Milvus and Slurm. NKP Metal is intended to automate operating-system, firmware and container deployment, while NKP on AHV integrates with Nutanix Flow for network isolation.

Sources

  1. siliconangle.comNutanix expands cloud platform with controls for agentic AI - SiliconANGLE