Run an LLM inference server on a Mac
Turns a Mac into an LLM inference server run from the menu bar.
Coding / product dossier
Mac LLM server that cuts agent wait times from 90s to 5s
Product brief
oMLX turns your Mac into a full LLM inference server, run from the menu bar. It serves text, vision, OCR, embedding and reranker models with continuous batching, plus a RAM+SSD tiered KV cache that survives restarts, so Claude Code and Cursor respond in about 5s instead of 90s. OpenAI and Anthropic compatible APIs drop straight in. Native Swift, not Electron. Apache 2.0, open source.
Why we selected it
Editorially selected for practical value, novelty, and audience relevance.
Product preview saved with our daily selection
Capability scan
Only capabilities supported by the product information we collected are listed here.
Turns a Mac into an LLM inference server run from the menu bar.
Serves text, vision, OCR, embedding, and reranker models.
Includes continuous batching for model serving.
Uses a RAM-and-SSD tiered KV cache that survives restarts.
Provides APIs described as compatible with OpenAI and Anthropic APIs.
Best-fit use cases
FAQ
It is described as serving text, vision, OCR, embedding, and reranker models.
It runs from the Mac menu bar.
The product description says its RAM-and-SSD tiered KV cache survives restarts.
Its APIs are described as compatible with OpenAI and Anthropic APIs.
Yes. It is described as open source under the Apache 2.0 license.
Selection history