Productspublished

Kasm Pairs Xeon 6 Workspaces With Local AI, Says Data Stays In-House

The partnership packages local model inference into managed workspaces, but Kasm’s claims on speed, capability, and cost still await deployment-specific proof.

By 2 min read
Kasm Pairs Xeon 6 Workspaces With Local AI, Says Data Stays In-House

Listen to this story

The audio brief

About 1:34
0:001:34
Read transcript
Kasm is packaging local AI inference into managed workspaces that run on Intel Xeon 6 processors, aiming to keep enterprise data inside the organization instead of sending it to an outside inference provider. The key detail is that Kasm says this can work without a GPU. Its workspace platform uses Intel’s OpenVINO toolkit and Xeon’s Advanced Matrix Extensions, or AMX, to run models inside temporary containers. That means the model-serving session, along with the user’s workspace, stays within the company’s network perimeter. Kasm positions the setup for chat, retrieval-augmented generation, tool calls, coding assistance, and autonomous agents. One cited example is Qwen3-Coder-30B-A3B, an open-weight model. The company also says its control plane can manage workloads across AMX-equipped CPUs, integrated NPUs, and discrete GPUs, giving IT teams one layer for both lighter interactive tasks and more demanding acceleration. There’s a separate sharing feature in Kasm 1.19: Intel Arc Pro GPUs can be divided among isolated workspaces using SR-IOV virtual functions. The business case is still mostly a claim from Kasm. It estimates cost parity with per-seat AI subscriptions at about 40 provisioned users per node, with better economics above that. But the announcement provides no independent benchmarks, model settings, workload results, or pricing analysis. The central question is whether those savings and interactive performance hold in a specific enterprise deployment.

Story brief

3 key points

Kasm’s updated AI Workspaces package targets enterprises that want model inference inside their own network, using Intel Xeon 6 AMX CPUs and OpenVINO rather than requiring a GPU. The company claims one node can reach subscription cost parity at about 40 provisioned users, while its control plane can also span NPUs and discrete GPUs. Qwen3-Coder-30B-A3B is a cited deployment. The economics and performance remain...

  1. 01

    CPU inference runs in ephemeral containers, keeping sessions and model-serving work within the organization’s perimeter.

  2. 02

    Kasm cites cost parity with per-seat AI subscriptions at roughly 40 provisioned users per node; above that, it claims better economics.

  3. 03

    The control plane is designed to span Xeon AMX CPUs, integrated NPUs, and discrete GPUs.

Kasm Technologies has expanded its Intel partnership to run local large language model inference in Kasm AI Workspaces on Intel Xeon 6 processors with Advanced Matrix Extensions, or AMX. Kasm says the design can keep enterprise data within an organization rather than send it to third-party inference providers.

How the workspace design works

The architecture combines Kasm’s containerized workspace platform with Intel’s OpenVINO toolkit and Xeon 6 processors equipped with AMX. Kasm says it is designed to run inference without a GPU, putting the model-serving work inside an ephemeral, or temporary, workspace container.

That deployment approach is aimed at organizations that want AI workloads handled within their own perimeter. Kasm positions the setup for chat, retrieval-augmented generation, tool calls, code assistance, and autonomous agents; it identifies Qwen3-Coder-30B-A3B as one open-weight model running on the architecture.

Kasm’s subscription comparison
Approximately 40Provisioned users per node

Kasm says its architecture reaches cost parity with per-seat AI subscriptions at approximately 40 provisioned users per node, with more favorable economics above that level.

One workspace layer across hardware

CPU inference is only one path in Kasm’s pitch. The company says its workspace control plane can orchestrate AI workspaces across AMX-equipped CPUs, integrated NPUs, and discrete GPUs, then deliver sessions to browsers on endpoints. That gives IT teams one stated management layer for lighter interactive work and workloads needing GPU acceleration.

A GPU-sharing option

  • Kasm 1.19 supports SR-IOV bifurcation of Intel Arc Pro cards.
  • One physical GPU can serve multiple isolated workspaces as virtual functions.

The commercial case is still company-supplied

Kasm says recent open-weight models can approach leading frontier-model capability on selected enterprise workloads and describes the experience as interactive. Its announcement does not include benchmark measurements, model settings, or workload-specific results for those assertions. The 40-user threshold is likewise a company-supplied estimate, not an independent price comparison.

Sources

  1. prnewswire.comKasm Technologies Expands Intel Partnership to Deliver Private AI Through Kasm AI Workspaces on Intel Xeon 6 with AMX