Broadcom Validates Five AI Model Families for VMware’s On-Premises Model Service
The update gives enterprises a path to run models from NVIDIA, Google, NEC, Alibaba Cloud and Z.ai inside their own infrastructure, with new sharing, routing and code-isolation services around them.
Listen to this story
The audio brief
Story brief
3 key pointsBroadcom is expanding VMware Cloud Foundation’s on-premises AI path beyond direct model access: customers can deploy five validated families—Nemotron 3, Gemma 4, cotomi, Qwen 3.7-Max and GLM 5.2—and offer them internally as a managed service. The validation is not a fixed stack: VCF uses vLLM by default, supports more than 150 open-source models and runs across AMD, Intel and NVIDIA hardware. VMware AI Factory adds...
- 01
VCF’s default vLLM runtime supports more than 150 open-source models, so the five validations expand a catalog rather than define its limits.
- 02
VMware AI Factory supports multi-tenant namespaces, reducing duplicate deployments and potential GPU waste.
- 03
AI Gateway routes workloads across on-premises and cloud models, with usage limits and application authorization.
Enterprises using VMware Cloud Foundation now have a validated route to run five additional AI model families on premises and offer them internally as a model-as-a-service. Broadcom validated NVIDIA’s Nemotron 3, Google DeepMind’s Gemma 4, NEC’s cotomi, Alibaba Cloud’s Qwen 3.7-Max and Z.ai’s GLM 5.2.
The change is aimed at organizations that want to serve models from their own infrastructure rather than simply give individual teams direct access to a model provider. Broadcom says the validation is intended to let customers deploy the named models on premises and deliver them as a service to internal users.
A wider catalog sits behind the five validations
The five families are named validations within a broader runtime catalog. Broadcom says VCF uses vLLM as its default model runtime and can support more than 150 open-source models. The company also says VCF supports mixed compute across AMD, Intel and NVIDIA hardware, giving customers a choice of CPU and GPU infrastructure.
The services that turn a model into a shared resource
VMware AI Factory supplies the surrounding deployment package. It combines VCF with customer-selected AI software, accelerator architectures and certified AI ReadyNodes from vendors including Cisco, Dell, HPE, Lenovo and Supermicro. VMware says it automates hardware provisioning, software-stack enablement, infrastructure deployment and lifecycle management.
Three services address separate deployment tasks
- Multi-tenant model sharing uses isolated namespaces so tenants or business units can access shared models without separate deployments. VMware says this is intended to reduce redundant deployments and GPU waste.
- AI Gateway provides one interface for models hosted on premises and in the cloud, with prompt routing, token and usage limits, and application authorization.
- Secure virtualized container spaces isolate agent-generated code. A control layer governs agent invocation, tool access and output validation before an action is taken.
Validation narrows the infrastructure question
Broadcom is not presenting the five validations as a single prescribed model stack. Its approach combines a default runtime, support for a larger model catalog and mixed hardware options with services that govern how teams share, route and run AI workloads. The immediate result is a tested VCF path for the five newly named families; enterprises still choose which models and infrastructure fit their applications.
Sources
- investors.broadcom.comVMware Cloud Foundation Brings Leading AI Models to the Private AI Cloud | Broadcom Inc.
- crn.comVMware Explore 2026: 5 Biggest AI, VCF And Security Launches