Kasm Pairs Xeon 6 Workspaces With Local AI, Says Data Stays In-House
The partnership packages local model inference into managed workspaces, but Kasm’s claims on speed, capability, and cost still await deployment-specific proof.
Listen to this story
The audio brief
Story brief
3 key pointsKasm’s updated AI Workspaces package targets enterprises that want model inference inside their own network, using Intel Xeon 6 AMX CPUs and OpenVINO rather than requiring a GPU. The company claims one node can reach subscription cost parity at about 40 provisioned users, while its control plane can also span NPUs and discrete GPUs. Qwen3-Coder-30B-A3B is a cited deployment. The economics and performance remain...
- 01
CPU inference runs in ephemeral containers, keeping sessions and model-serving work within the organization’s perimeter.
- 02
Kasm cites cost parity with per-seat AI subscriptions at roughly 40 provisioned users per node; above that, it claims better economics.
- 03
The control plane is designed to span Xeon AMX CPUs, integrated NPUs, and discrete GPUs.
Kasm Technologies has expanded its Intel partnership to run local large language model inference in Kasm AI Workspaces on Intel Xeon 6 processors with Advanced Matrix Extensions, or AMX. Kasm says the design can keep enterprise data within an organization rather than send it to third-party inference providers.
How the workspace design works
The architecture combines Kasm’s containerized workspace platform with Intel’s OpenVINO toolkit and Xeon 6 processors equipped with AMX. Kasm says it is designed to run inference without a GPU, putting the model-serving work inside an ephemeral, or temporary, workspace container.
That deployment approach is aimed at organizations that want AI workloads handled within their own perimeter. Kasm positions the setup for chat, retrieval-augmented generation, tool calls, code assistance, and autonomous agents; it identifies Qwen3-Coder-30B-A3B as one open-weight model running on the architecture.
Kasm says its architecture reaches cost parity with per-seat AI subscriptions at approximately 40 provisioned users per node, with more favorable economics above that level.
One workspace layer across hardware
CPU inference is only one path in Kasm’s pitch. The company says its workspace control plane can orchestrate AI workspaces across AMX-equipped CPUs, integrated NPUs, and discrete GPUs, then deliver sessions to browsers on endpoints. That gives IT teams one stated management layer for lighter interactive work and workloads needing GPU acceleration.
A GPU-sharing option
- Kasm 1.19 supports SR-IOV bifurcation of Intel Arc Pro cards.
- One physical GPU can serve multiple isolated workspaces as virtual functions.
The commercial case is still company-supplied
Kasm says recent open-weight models can approach leading frontier-model capability on selected enterprise workloads and describes the experience as interactive. Its announcement does not include benchmark measurements, model settings, or workload-specific results for those assertions. The 40-user threshold is likewise a company-supplied estimate, not an independent price comparison.
Sources
- prnewswire.comKasm Technologies Expands Intel Partnership to Deliver Private AI Through Kasm AI Workspaces on Intel Xeon 6 with AMX