Microsoft Adds Experimental GGUF Model Support to Windows ML Through llama.cpp
Developers can use an OpenAI-compatible endpoint for local prototypes or a new native runtime for finer control. The new Windows ML capabilities remain experimental.
Loading page…
Developers can use an OpenAI-compatible endpoint for local prototypes or a new native runtime for finer control. The new Windows ML capabilities remain experimental.
Listen to this story
Microsoft’s Windows ML update, available October 7, 2026, lets developers run GGUF language models locally through an experimental llama.cpp integration, while keeping GGUF and ONNX behind one Text Generation API. The change gives Windows apps a more direct path to local model inference and can connect speech transcription to generation without sending workloads to the cloud. A separate experimental Windows-native runtime adds lower-level control, but Microsoft advises checking its limitations before using it in production.
Windows ML selects llama.cpp for GGUF models and supports ONNX through the same Text Generation API.
A prototype can use an OpenAI-compatible endpoint, allowing developers to connect the OpenAI SDK to a local Windows ML server.
Speech Recognition transcribes audio with a developer-supplied ONNX Whisper model, whose output can feed local text generation.
Developers can now bring GGUF-format AI models into Microsoft’s Windows ML framework and run them locally through experimental llama.cpp integration. Available October 7, 2026, the update adds a simpler route for text and speech applications, alongside a preview of a Windows-native runtime for developers who need more control over model execution.
Windows ML is Microsoft’s framework for running trained AI models on Windows PCs. It supports workloads across graphics processors, dedicated AI processors and CPUs from AMD, Intel, NVIDIA and Qualcomm. Microsoft says local execution can reduce delays, keep workload data on the device and avoid per-token cloud inference charges.
The new Text Generation API accepts GGUF and ONNX, two model formats, through the same programming interface. Windows ML selects the execution engine automatically, using llama.cpp for GGUF models. Microsoft says developers can take a new GGUF model from Hugging Face and run it through this stack with a few lines of code.
For prototypes, the task APIs also expose an OpenAI-compatible endpoint. Developers can point the OpenAI SDK at a local Windows ML server rather than learn a new client interface. The model runs on the device; the familiar SDK supplies the connection.
Underneath those task interfaces sits the new Windows ML Runtime API. The experimental preview offers direct handling of Windows-native images, video frames, audio buffers and text. Microsoft describes zero-copy paths, which pass data without copying it, instead of requiring developers to write preprocessing and format-conversion code.
Developers can link models in a pipeline and choose whether each stage runs on a CPU, GPU or dedicated AI processor. They can also compile models in advance to help them start faster. The text and speech APIs use this runtime, so apps can switch to more detailed controls when needed.
Existing ONNX Runtime APIs remain fully supported and ship alongside the native path. Microsoft says deeper Windows-native optimizations will arrive through the new runtime over time.
The development-stack context includes official native PyTorch CPU builds for Windows Arm64 and NVIDIA’s CUDA-enabled packages for supported Arm64 hardware. Triton’s Windows distribution supports custom GPU kernels and torch.compile, which generates optimized code. Microsoft’s example trains or develops with those tools, then exports the model to ONNX for deployment through Windows ML.
To try the native runtime, Microsoft directs developers to Windows ML 2.7.2021 Experimental. It advises checking supported scenarios and known limitations before relying on the new capabilities in production.
Loading discussion...
Join the conversation
Explain where automatic choices help and where control matters.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.