Nutanix Enterprise AI 2.8 Adds Agent Token Quotas and MCP Access Controls
Nutanix is putting controls around what agents can reach and consume, but says secure backend infrastructure, API keys and roles remain essential.
Listen to this story
The audio brief
Story brief
3 key pointsNutanix’s Enterprise AI 2.8 gives organizations finer control over agent consumption and MCP-connected tools, including quotas by team, user, or agent and logs of accessed data sources, applications, and actions. The release also broadens private inference with LoRA tuning for models under 8 billion parameters and multi-GPU, batch, and speculative decoding options; Nutanix claims up to 2.5× faster generation in...
- 01
MCP activity logs can capture data sources, applications, and actions, while prompt inspection is not the default focus.
- 02
Private Inference shifts some workloads from per-token billing to infrastructure costs using privately deployed open-weight models.
- 03
LoRA fine-tuning is limited to models below 8 billion parameters.
Nutanix has made Enterprise AI 2.8 generally available with an MCP gateway, token tracking and quotas, expanded audit logs, and private-inference features. The release creates a central layer for controlling agent access and consumption, but the systems behind those connections remain the security boundary.
A gate for tools, data and infrastructure
The generally available Model Context Protocol gateway sits in Nutanix Agent Gateway and is designed to govern an agent’s access to tools and data. A separate Nutanix Cloud Platform MCP Server gives agents controlled access to infrastructure managed by Nutanix software. Enterprise AI 2.8 can also track token use and apply consumption quotas at the team, user or agent level.
The logging is aimed at the connections agents use. Nutanix says MCP activity logs can record the data sources and applications an agent accessed, along with its actions. Thomas Cornely, Nutanix’s executive vice president of product management, said inspecting prompts is not the product’s default purpose.
Private inference offers a different cost model
Private Inference adds Low-Rank Adaptation fine-tuning for models with fewer than 8 billion parameters, multi-GPU inference using tensor parallelism, batch inference and speculative decoding. Cornely said basic work such as Python scripting can run on privately deployed open-weight models, where customers pay for infrastructure rather than individual tokens.
Private Inference additions
- LoRA fine-tuning for models below the 8-billion-parameter threshold.
- Multi-GPU inference, batch inference and speculative decoding.
- Nutanix claims speculative decoding can increase token-generation speed by up to 2.5 times, depending on the model and infrastructure configuration.
Kubernetes management is the next release
Nutanix Kubernetes Platform 2.19 is scheduled for an upcoming release and will extend Kubernetes management across virtualized and bare-metal environments. Its application catalog will offer curated deployments of Kubeflow, Milvus and Slurm. NKP Metal is intended to automate operating-system, firmware and container deployment, while NKP on AHV integrates with Nutanix Flow for network isolation.
Sources
- siliconangle.comNutanix expands cloud platform with controls for agentic AI - SiliconANGLE