Productspublished

Broadcom Validates Five AI Model Families for VMware’s On-Premises Model Service

The update gives enterprises a path to run models from NVIDIA, Google, NEC, Alibaba Cloud and Z.ai inside their own infrastructure, with new sharing, routing and code-isolation services around them.

By 2 min read
Broadcom Validates Five AI Model Families for VMware’s On-Premises Model Service
Broadcom Validates Five AI Model Families for VMware’s On-Premises Model Service

Listen to this story

The audio brief

About 1:32
0:001:32
Read transcript
Broadcom has validated five AI model families for VMware Cloud Foundation, giving enterprises a supported path to run them inside their own infrastructure and offer them internally as a model-as-a-service. The lineup includes NVIDIA’s Nemotron 3, Google DeepMind’s Gemma 4, NEC’s cotomi, Alibaba Cloud’s Qwen 3.7-Max, and Z.ai’s GLM 5.2. The important context is that these are named validations, not a fixed five-model stack. VCF uses vLLM as its default runtime and supports more than 150 open-source models. It also runs across AMD, Intel, and NVIDIA hardware, so customers can choose the models and infrastructure that fit their applications. VMware AI Factory adds the operating layer around those models. Multi-tenant namespaces let business units share deployments instead of creating separate copies, which VMware says can reduce duplication and GPU waste. An AI Gateway provides one interface for models running on premises or in the cloud, with prompt routing, token and usage limits, and application authorization. There is also a security boundary for agent workflows: virtualized container spaces isolate code generated by agents, while controls govern tool access, agent invocation, and output validation before an action is taken. So the immediate result is a tested VCF route for five additional model families, within a much broader catalog. The constraint is that enterprises still have to decide which models, hardware, and controls fit each workload.

Story brief

3 key points

Broadcom is expanding VMware Cloud Foundation’s on-premises AI path beyond direct model access: customers can deploy five validated families—Nemotron 3, Gemma 4, cotomi, Qwen 3.7-Max and GLM 5.2—and offer them internally as a managed service. The validation is not a fixed stack: VCF uses vLLM by default, supports more than 150 open-source models and runs across AMD, Intel and NVIDIA hardware. VMware AI Factory adds...

  1. 01

    VCF’s default vLLM runtime supports more than 150 open-source models, so the five validations expand a catalog rather than define its limits.

  2. 02

    VMware AI Factory supports multi-tenant namespaces, reducing duplicate deployments and potential GPU waste.

  3. 03

    AI Gateway routes workloads across on-premises and cloud models, with usage limits and application authorization.

Enterprises using VMware Cloud Foundation now have a validated route to run five additional AI model families on premises and offer them internally as a model-as-a-service. Broadcom validated NVIDIA’s Nemotron 3, Google DeepMind’s Gemma 4, NEC’s cotomi, Alibaba Cloud’s Qwen 3.7-Max and Z.ai’s GLM 5.2.

The change is aimed at organizations that want to serve models from their own infrastructure rather than simply give individual teams direct access to a model provider. Broadcom says the validation is intended to let customers deploy the named models on premises and deliver them as a service to internal users.

A wider catalog sits behind the five validations

The five families are named validations within a broader runtime catalog. Broadcom says VCF uses vLLM as its default model runtime and can support more than 150 open-source models. The company also says VCF supports mixed compute across AMD, Intel and NVIDIA hardware, giving customers a choice of CPU and GPU infrastructure.

The services that turn a model into a shared resource

VMware AI Factory supplies the surrounding deployment package. It combines VCF with customer-selected AI software, accelerator architectures and certified AI ReadyNodes from vendors including Cisco, Dell, HPE, Lenovo and Supermicro. VMware says it automates hardware provisioning, software-stack enablement, infrastructure deployment and lifecycle management.

Three services address separate deployment tasks

  • Multi-tenant model sharing uses isolated namespaces so tenants or business units can access shared models without separate deployments. VMware says this is intended to reduce redundant deployments and GPU waste.
  • AI Gateway provides one interface for models hosted on premises and in the cloud, with prompt routing, token and usage limits, and application authorization.
  • Secure virtualized container spaces isolate agent-generated code. A control layer governs agent invocation, tool access and output validation before an action is taken.

Validation narrows the infrastructure question

Broadcom is not presenting the five validations as a single prescribed model stack. Its approach combines a default runtime, support for a larger model catalog and mixed hardware options with services that govern how teams share, route and run AI workloads. The immediate result is a tested VCF path for the five newly named families; enterprises still choose which models and infrastructure fit their applications.

Sources

  1. investors.broadcom.comVMware Cloud Foundation Brings Leading AI Models to the Private AI Cloud | Broadcom Inc.
  2. crn.comVMware Explore 2026: 5 Biggest AI, VCF And Security Launches