Runs an LLM inference server on a Mac
oMLX turns a Mac into an LLM inference server run from the menu bar.
Coding / product dossier
Runs local text, vision, OCR, embedding, and reranker models as a Mac LLM server.
Product brief
oMLX turns your Mac into a full LLM inference server, run from the menu bar. It serves text, vision, OCR, embedding and reranker models with continuous batching, plus a RAM+SSD tiered KV cache that survives restarts, so Claude Code and Cursor respond in about 5s instead of 90s. OpenAI and Anthropic compatible APIs drop straight in. Native Swift, not Electron. Apache 2.0, open source.
Why we selected it
A focused, open-source local inference server with text, vision, OCR, embedding, and reranking support. OpenAI- and Anthropic-compatible APIs make it a practical way for Mac-based builders to test or run local model
Product preview saved with our daily selection
Capability scan
Only capabilities supported by the product information we collected are listed here.
oMLX turns a Mac into an LLM inference server run from the menu bar.
It serves text, vision, OCR, embedding, and reranker models.
oMLX includes continuous batching.
It uses a RAM-and-SSD tiered KV cache that survives restarts.
oMLX offers OpenAI- and Anthropic-compatible APIs.
Best-fit use cases
FAQ
oMLX is a Mac LLM inference server operated from the menu bar.
It serves text, vision, OCR, embedding, and reranker models.
The product description states that it provides OpenAI- and Anthropic-compatible APIs.
It uses a tiered KV cache across RAM and SSD that survives restarts.
Yes. It is described as open source under the Apache 2.0 license.