A private AI that
lives on your Mac

Chat with it, ask it about your own documents, make images, music, video and 3D, and let it run errands on your computer — free, open source, and fully offline. It runs every major open model, faster than LM Studio, and nothing you type ever leaves the machine.

Download for Mac View on GitHub
macOS 26+
M1 – M5
Free & open source
Star on GitHub
MLX Core app — DiffusionGemma block diffusion running live in the chat interface

Three things that feel like magic

The fastest way to "get" it. Each takes under a minute — and each is something only an AI on your own Mac can do.

One · ⌃Space

Ask anything, from anywhere

Press ⌃Space over any app — your browser, an email, a game — and a little prompt box appears. Ask your question; the answer streams in right there. No tab-switching, no logging in.

It's the fastest quick-question, "what's the word for…" habit you'll pick up. Follow-ups remember the conversation, and Esc tucks it away.

The Quick Launcher: a Spotlight-style prompt panel floating over another app
⌃Space over any app — ask, read, dismiss.
Two · Your files

Chat with your own documents

Drag a folder into the chat and ask about what's inside — a lease, a stack of PDFs, your class notes, a pile of receipts. "Which invoice is unpaid?" · "Summarize this into five bullets." · "Quiz me on chapter 3."

Because nothing leaves your Mac, this is the stuff you'd never paste into an online chatbot — taxes, medical letters, contracts — answered safely at home.

Asking a local model questions about your own documents in the chat window
Attach a folder and ask — your documents stay on your machine.
Three · Voice

Just talk to it, hands-free

Turn on Voice Mode from the menu bar, say "Hey Loki," and speak. It listens, thinks, and answers out loud — perfect while cooking, tidying, or away from the keyboard. Your voice is transcribed on the Mac and never uploaded.

Want it to reply in a familiar voice? Record a short clip and it can even answer in a cloned voice.

Hands-free Voice Mode with an animated listening orb
Hands-free Voice Mode — say the wake word and just talk.

Not just chat — it makes things

Pictures, music, voices, video, even 3D models — all generated on your Mac, with no monthly fee and no watermark. Pick a tab, describe what you want, and go.

Images

Make art from a sentence

Describe a picture — "a watercolor fox in a snowy forest at golden hour" — and get it in seconds. Great for cards, wallpapers, posters, and ideas.

Photo editing

Edit a photo by just asking

Drop in a photo and say "make the sky a sunset" or "remove the background." It keeps the rest of the picture — no design skills needed.

Music

Turn a mood into a track

"Lo-fi study beat, 90 bpm, rainy." Get original background music and jingles you can actually use — nothing to license.

Voice

Any text, read aloud

Turn writing into natural speech — DIY audiobooks, voiceovers, or practice scripts — optionally in a voice you cloned from a few seconds of audio.

Video

Bring a photo to life

Animate a still image or generate a short clip from a prompt — with a soundtrack, or characters that speak your lines.

3D

A photo into a 3D model

Snap or drop a single photo and get a 3D model out — handy for 3D printing, games, and AR tinkering.

The image generation pane creating a picture from a text prompt
Type a description, get a picture — free and local.
The video generation pane creating a clip from a prompt with LTX-Video
Generate a video from a prompt or a photo.

An assistant with hands

Flip on Agent mode and it can tidy files, look things up, and build small things for you — with a safety rail you control.

Agent mode

Give it a chore, watch it work

Ask in plain words: "Organize my Downloads folder by type." · "Rename these photos by the date they were taken." · "Build me a little birthday-countdown web page" — and it opens it in your browser when it's done.

  • You set the leash. Choose how much it's allowed to do; anything riskier pauses and asks first.
  • It stays in its lane. File actions are confined to a folder you pick — or to an isolated Linux VM.
  • Nothing to configure. The tools are built in — just describe the goal.
Agent mode running tools to complete a task, with each step shown
Agent mode shows each step it takes — no black box.
Schedules

Daily jobs, on autopilot

Hand it a recurring task — "every morning at 8, summarize my watched sites" — and it runs on its own, saving a transcript each time.

Telegram

Text your Mac from your phone

Message your own model from anywhere through a private Telegram bot — no cloud AI, and only your chat can reach it.

For the curious

Power tools, when you want them

Coding assistants, a built-in API, and a sandbox are all in there — ignore them until you're ready, then dive in.

The same AI, in your pocket

MLX Chat — Local AI puts this exact open-source engine on your iPhone: chat, voice cloning, image and music generation — all on-device, all offline.

MLX Chat — Local AI

One engine, Mac to iPhone

It's not a companion app that phones home to a server — the same mlx-serve engine behind everything on this page is compiled straight into the app and runs on the phone's GPU. No cloud, no account, and it works in airplane mode.

  • Chat with Gemma 4 running on the phone — or use built-in Apple Intelligence with nothing to download.
  • Clone your voice from an 8-second sample, then have it read anything aloud.
  • Make images and music right on the phone — the same FLUX.2 and ACE-Step models as the Mac app.

Download on the App Store

Needs a recent iPhone — A17 Pro-class or newer, iOS 26+.

MLX Chat on iPhone streaming a local model's answer
MLX Chat on iPhone generating an image on-device

The whole point is that it's yours

A cloud chatbot meters you, remembers you, and needs a signal. This one asks for none of that — because it never leaves your Mac.

$0
forever — no subscription
0
accounts, keys, or sign-ups
100%
runs on your own Mac
✈️
works with no internet

What will you do first?

Start here

New to local AI?

The friendly 5-minute tour — free, private, offline.

Get started →
Create

Edit photos with words

“Make the hair blue” — subject, pose & scene survive.

Deep dive →
Animate

Turn photos into video

Clips with synced sound — even talking characters.

Deep dive →
Compose

Type a vibe, get a song

Original 48 kHz stereo music — lyrics optional.

Deep dive →
Sculpt

Turn photos into 3D models

One picture in, a textured GLB mesh out.

Deep dive →
Speak

Clone any voice

Six seconds of audio, no transcript, all local.

Deep dive →
Code

Run Claude Code — free

Your coding agent, offline, no API key.

Deep dive →
Ask

Summon AI over any app

⌃Space launcher, voice, phone & schedules.

Deep dive →
Unleash

Sandbox your agent

Shell commands hit a Linux VM, not your Mac.

Deep dive →
Compare

Faster than LM Studio

+48% on identical models — keeps your library.

Deep dive →
Swap

Replace Ollama in one line

Your Ollama apps connect unchanged.

Deep dive →
Accelerate

Same answers, 2× faster

Speculative decoding, verified exact.

Deep dive →
Trust

Tool calls that don't break

Small-model mistakes repaired mid-flight.

Deep dive →
Rank

The local LLM tier list

Community S-tiers, filtered to your Mac's memory.

Cast your vote →

Faster than LM Studio.
Every model.

Under the hood, MLX Core is mlx-serve — a native Zig inference server for Apple Silicon with OpenAI- and Anthropic-compatible APIs. Identical 4-bit MLX weights, same machine, same prompts: it wins every cell, and speculative decoding pushes the lead further where it counts. Full LM Studio comparison →

199
tok/s decode · Gemma 4 E2B 4-bit
4,891
tok/s prefill
~200%
faster on Qwen 27B MTP
48%
geomean faster than LM Studio · 18 workloads

Code completion — where the drafter and native MTP shine. Tokens per second — how fast the AI writes; higher is better. Apple M4 Max (128 GB) · identical 4-bit MLX weights · ctx 4096 · temp 0 · LM Studio (MLX runtime) as baseline.

The marquee capability

Run DeepSeek V4 Flash locally on your Mac

The 284-billion-parameter flagship — running on your own machine, no cloud, no API key. If you have a 96 GB+ Apple Silicon Mac, it's one click away in the Model Browser.

  • Built on Salvatore Sanfilippo's antirez/ds4 engine — native Metal kernels, byte-validated against the reference forward.
  • One-click download, served from the same model picker as everything else.
  • Agent mode and MCP tool calling work on DSV4 too — the full toolset is inlined into the prompt.
  • A single self-contained binary — kernel sources are embedded and staged at first launch.
Get MLX Core
284B
parameters · running on your desk
96 GB+
unified memory
0
cloud calls
1
binary

Three ways to draft ahead

Generate multiple tokens per forward pass, verified exactly — so output is identical, just faster. Works on every API surface, streaming or not, tools included: agent loops that echo file contents into edits decode at ~2×. Smart gates keep it on where it pays and step aside where it doesn't.

PLD

Prompt Lookup Decoding

Model-agnostic n-gram drafting from the prompt + generated text. Works on every architecture — Gemma, Qwen, Llama, Mistral, Nemotron-H, LFM2.5 — with nothing extra to download.

up to 2.1× on agent tool loops, echo & RAG
DRAFTER

Gemma 4 assistant drafter

A tiny cross-attention drafter reuses the target model's own K/V cache to propose blocks of tokens. Tuned block sizes per target (E2B → 31B).

up to +65% on Gemma 4 code completion
MTP

Qwen native multi-token prediction

Qwen 3.5/3.6 checkpoints with a trained MTP sidecar draft with the model's own head — 3 tokens per round, a controller that self-tunes depth per request, MoE sidecars included. Auto-loads, zero setup.

up to on code & agent edits (Qwen3.6-27B)
ADAPTIVE

Gates that know when to quit

A prompt-time repetition score disables drafting on novel content; a runtime acceptance gate backs off mid-decode when drafts stop landing. You never pay for speculation that won't pay back.

exact output, zero quality cost

How fast is native MTP? On Qwen3.6-27B 4-bit (M4 Max): code completion 29.0 → 58.4 tok/s (2×), echo/edit loops 58.7 tok/s, free-form +38% — with sidecars like ddalcu/Qwen3.6-27B-4bit-MTP-MLX-Serve. In a head-to-head on the identical checkpoint and prompts, mlx-serve out-decodes the reference MTP runtime at all 8 context rungs from 0.5K to 64K (+11–30%) — and at 64K context it decodes 33.5 tok/s where plain autoregressive does 21.7. Speculative decoding, in depth →

A complete local-AI stack

Everything a private AI setup needs in one Mac app — plus the deep dives above when you want the full story on any of it.

Nothing to install, nothing to configure

One lightweight native app — no dependencies, no setup scripts, no gigabyte runtime. The first answer after launch comes 3.5× faster thanks to eager warmup, and warm conversation turns round-trip in about 0.1 s.

Works with the apps you already use

Claude Code, Cursor, Continue, Raycast, Open WebUI — anything built for OpenAI or Anthropic can talk to your Mac instead of the cloud. Streaming, tool calling, embeddings, and WebSockets included.

Run Claude Code locally →

Your Ollama apps just work

mlx-serve speaks Ollama's language natively, so tools built for it — Raycast, Obsidian, Enchanted, Open WebUI — connect unchanged: same models, faster engine. And mlx-serve run gemma4 downloads, serves, and chats from one command.

The Ollama swap, in depth →

Several chats at once

Multiple conversations decode together through one model — about 1.6× total throughput at 4-way parallel — and every stream stays byte-accurate under load, verified by a 24-hour stress test.

Longer documents, less memory

The working memory a long conversation needs can be compressed up to ~4×, so 16K-token contexts fit on Macs that couldn't hold them before — or you serve more chats in parallel at the same length.

Long chats resume instantly

Re-opening a long conversation used to mean sitting through a full re-read of its history. The prefix cache now persists to SSD, so work the server already did is restored instead of recomputed — across model switches, app relaunches, and reboots. A chat that took 40 seconds to warm up answers in under 3, with byte-identical output. Opt-in from Settings, since it can hold gigabytes.

Watch your server work

Flip on Metrics panel and the server's own homepage grows a live dashboard: decode and prefill tokens/sec with sparklines, requests in flight, time-to-first-token, cache hit rate, GPU load and memory — moving as you generate, including through a long prompt. The same numbers export at a Prometheus /metrics endpoint under vLLM-compatible names, so existing Grafana dashboards work unchanged. Opt-in, with no measurable cost to tokens/sec.

Share it on your network, safely

Serving to other machines? Set an API key and every off-box request — OpenAI, Anthropic, and Ollama APIs, plus the metrics page — has to present it, as a Bearer token, x-api-key, HTTP Basic, or a query parameter. Your own Mac stays trusted and key-free, so local apps and tools carry on untouched.

Built-in agent + MCP

Ten built-in tools — shell, file read/write/edit, search, browse, web search, memory — with a per-tool approval dialog. Connect MCP servers from a curated marketplace, or extend with markdown skills.

Agent Sandbox — a VM, not your Mac

One toggle routes every agent shell command into an isolated Linux VM built on Apple's own Virtualization framework. It boots in under a second, servers the agent starts mirror live to localhost, and a green shield shows when commands run isolated. Let the agent go wild — your files stay untouched.

How the sandbox works →

⌃Space Quick Launcher

A Spotlight-style panel that summons over any app: hit ⌃Space, type, and the answer streams in from your local model — no window shuffling. Follow-ups keep their context, ⌘↩ hands the thread to the chat window, and Esc dismisses while the reply finishes in your sidebar.

The always-on assistant →

One server, every model

A Model Browser built around what you already have: Discover to find new models, My Models for everything this Mac can load — including what LM Studio downloaded and your own folders — and Downloads with a live badge. Finished models stay listed with a one-click Use, and a badge marks whichever one the server is actually serving. Resumable HuggingFace transfers, hot-switching in place, no restart.

Writes whole paragraphs at once

Google's DiffusionGemma composes 256-token blocks in parallel instead of word by word — up to 25 tokens per pass, ~30% faster than the reference implementation — and streams block-by-block into any chat app.

Ask questions about your files

Attach a folder of mixed files — notes, PDFs, exports — and ask in plain language. About 500 files index in ~7 seconds, everything stays in memory, and nothing leaves your Mac.

Telegram bot — model in your pocket

Make a bot, paste its token, flip a switch — and message your local model from your phone, anywhere. No public URL, port-forwarding, or cloud relay; it long-polls over your normal connection, even behind home Wi-Fi. Turn on Agent mode to run tools and edit files from your phone; the bot locks to the first chat that messages it.

Hands-free voice mode

Say “Hey Loki” and just talk. Speech is transcribed on-device with Apple's recognizer — audio never leaves your Mac — and the reply is spoken back in a system voice, with barge-in to interrupt mid-answer. Works from the menu bar with no window open; flip on Agent mode to run tools by voice.

Voice cloning & TTS

Give it 6–8 seconds of any voice and it speaks your text in that voice — zero-shot, no training. Runs fully local on Qwen3-TTS with adjustable speed and expressiveness; the reference audio alone clones the voice, no transcript needed — validated bit-for-bit against the reference.

Voice cloning, in depth →

Image generation & photo editing

Krea-2-Turbo (12.9B, photorealistic) and FLUX.2 run entirely on your Mac — validated pixel-faithful to the reference, any size 256²–2048². And it edits, not just generates: attach a photo, type "make the hair blue", and the subject, pose, and scene survive. Image-to-image variations with a strength slider, runtime style LoRAs, inline generation from chat. Nothing leaves your Mac.

Editing & LoRAs, in depth →

Video generation — with a voice

LTX-Video 2.3 turns a prompt, a photo, or a soundtrack into a 24 fps clip with synced audio, all on-device. Put spoken lines in quotes and your characters talk; attach a real voice clip — or a line spoken by local TTS — and the performance follows it, original audio muxed into the MP4. Animate a photo from the First-frame slot; two-stage presets when you want maximum quality, ~2× faster this release.

Video generation, in depth →

Type a vibe, get a song

New: describe a track — "upbeat synthwave with driving bass" — optionally paste lyrics, and ACE-Step 1.5 Turbo composes original 48 kHz stereo music in 8 diffusion steps, entirely on-device. 10 seconds to 10 minutes; BPM, key, and time signature steerable; instrumental or vocal. Every track saves its exact prompt and settings beside it, so any take is reproducible.

Music generation, in depth →

Photo → 3D model

New: drop in a photo and Hunyuan3D-2.1 — ported natively to Apple Silicon — builds a 3D mesh: the subject is cut out automatically, the shape is generated on-device, and an optional full-PBR paint stage textures it from the same photo. The finished GLB spins in a built-in viewer and opens anywhere glTF does.

Photo → 3D, in depth →

The app updates itself

MLX Core checks the releases page once a day and shows a banner when a new version ships — one click downloads the notarized build, swaps it in place, and relaunches. Models, chats, and settings untouched. And the welcome screen now installs the mlx-serve command onto your PATH with one click.

The CLI, documented →

Agents on a schedule

Define a task once and let it run unattended — once, on an interval, daily, or weekly — even with no window open. Each run is a full agent with tool access (auto-approve or approve per call) and a saved transcript you can read back.

One-click coding agents

Point Claude Code, pi, or OpenCode at your local model from a folder picker — the app wires up the environment and routes every request to mlx-serve. Each agent is handed the context window your model actually serves, not a hardcoded guess, and a turn that spends ten minutes writing one big file streams to completion instead of dying at its client's five-minute timeout. Real coding agents, your code and prompts never leaving your machine.

Claude Code, fully local →

Up and running in seconds

Download the app, pick a model, go. Prefer a terminal? Two commands.

MLX Core
MLX Core
The Mac app — chat, agents, images, video & voice
Get
For developers — Terminal
# Install via Homebrew (or grab the app from Releases)
brew tap ddalcu/mlx-serve https://github.com/ddalcu/mlx-serve
brew install mlx-serve

# One command: download, serve, chat — Ollama-style
mlx-serve run gemma4

# Or serve everything you've pulled, loading on demand
mlx-serve serve
For developers — use the API
curl http://localhost:11234/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": true
  }'

# Point Claude Code at your Mac
export ANTHROPIC_BASE_URL=http://localhost:11234

The latest open models,
one click each

Chat models, image models, video, and voice — downloaded right inside the app, all running on your Mac. If it's open and it matters, it runs here.

Chat & agents

Gemma 4

Google · E2B / E4B / 31B / 26B-A4B MoE · vision

Qwen 3.6

Alibaba · 27B & 35B-A3B · vision · extra-fast MTP builds

DiffusionGemma

Google · writes whole blocks at once · 26B-A4B

Llama 3

Meta · 8B · 70B

Mistral

Mistral AI · 7B · 8x7B

And thousands more

Nemotron-H · LFM2.5 · Qwen 3-Next · Gemma 3 · any GGUF

Create — image, video & voice

Questions, answered

The things people actually ask in HN comments, Discord, and AI search.

Do I need to be technical to use this?

No. MLX Core is a normal menu-bar app: install it, click to download the recommended model — it's sized to your Mac — and start chatting. The advanced bits (agent tools, coding assistants, the developer API) stay out of the way until you go looking for them. There's a friendly 5-minute getting-started tour when you're ready.

Is it really free?

Yes — completely. Open source under the MIT license: no subscription, no account, no per-message cost. You download a model once and use it as much as you like — and the images, music, voice, video, and 3D are free to make too.

Is mlx-serve faster than LM Studio?

Yes, where it counts — +48% geomean across 18 workloads on identical 4-bit MLX weights (Gemma 4 E2B/E4B/31B/26B-A4B-MoE and Qwen 3.6 27B/35B-A3B-MoE, vs LM Studio 0.4.15). The lead comes from speculative decoding: PLD up to 2.1× on echo-heavy work, the Gemma drafter +65% and native Qwen MTP +98% on code, plus a faster MoE decode path (+27% raw on Qwen 35B-A3B).

Does mlx-serve replace LM Studio?

For most use cases, yes. mlx-serve runs the same MLX and GGUF models, exposes an OpenAI-compatible API on the same kind of port, and ships a native menu-bar app instead of an Electron one. It goes deeper on the API surface than LM Studio's newer compatibility endpoints — fuller Anthropic Messages and OpenAI Responses coverage, plus a WebSocket transport — and adds things LM Studio doesn't have: MCP tool calling, agent mode with 10 built-in tools, KV-cache quantization, continuous batching, and the antirez/ds4 engine for DeepSeek V4 Flash.

Does mlx-serve replace Ollama on Mac?

On Apple Silicon, yes — mlx-serve speaks the Ollama API natively (/api/chat, /api/generate, /api/tags, /api/embed, /api/pull…), so Raycast, Obsidian, Enchanted, Open WebUI, and ollama-python/js work unchanged: drop in http://localhost:11234 wherever you had http://localhost:11434. The CLI matches too — mlx-serve run gemma4 downloads, serves, and chats in one command. Underneath, it runs llama.cpp and native MLX with the Mac-specific optimizations Ollama doesn't ship — Metal kernels through mlx-c, speculative decoding, and a shared-prefix KV cache.

Can I run GGUF models on Mac without Python?

Yes. mlx-serve embeds llama.cpp's inference library (libllama) inside the same signed, notarized binary. Point --model at any .gguf file and the server auto-detects the format and routes to the right engine — no pip, no venv, no llama-server to install separately. DeepSeek V4 Flash GGUFs go through the dedicated antirez/ds4 engine instead, also embedded.

Does mlx-serve work with Claude Code?

Yes — natively. mlx-serve implements Anthropic's /v1/messages endpoint including streaming, tool calling, and extended thinking. Point Claude Code at it with ANTHROPIC_BASE_URL=http://localhost:11234. The MLX Core app ships a one-click Launch Claude Code button that wires up the env vars for you.

What about the OpenAI SDK, Continue, Cursor, Open WebUI?

All work — anything that talks the OpenAI chat-completions or Anthropic Messages wire protocol does. mlx-serve also implements the newer OpenAI Responses API (/v1/responses) for clients that want stateful chains via previous_response_id, plus a WebSocket transport on the same endpoint.

Can mlx-serve run DeepSeek V4 Flash locally?

Yes, on 96 GB+ Apple Silicon Macs. Open the MLX Core Model Browser, pick DeepSeek-V4-Flash, hit Download — the server routes the GGUF through the embedded ds4 engine (native Metal kernels, byte-validated against the reference forward). Agent mode and MCP tools work on DSV4 too.

What models are supported?

Native MLX dispatch for Gemma 3/4, Qwen 3 / 3.5 / 3.6 / 3-Next, Llama 3.x, Mistral, Nemotron-H, LFM2.5, and DeepSeek V4 Flash. Anything else as GGUF via embedded llama.cpp — Qwen, Llama, Mistral, Gemma, DeepSeek, Phi, Yi, and thousands more from HuggingFace. On the media side: FLUX.2 and Krea-2-Turbo for images, LTX-Video 2.3 for video, and Qwen3-TTS for speech and voice cloning — all running natively on-device.

Does it support tools / function calling?

Yes, on both API surfaces. The server detects tool-call patterns across architectures (Hermes XML, Gemma 4 <|tool_call>, raw JSON, ChatML), repairs common Qwen 3.5/3.6 escape quirks, and emits OpenAI-style tool_calls deltas in the SSE stream. The MLX Core app ships 10 built-in tools (shell, file I/O, search, browse, web search, memory) and connects to MCP servers from a curated marketplace. Malformed tool-call JSON from small models is repaired at the API layer.

How does it stay this small / fast?

Zig with direct mlx-c FFI — no Python runtime, no Electron, no IPC bridge. The release binary is ~4.5 MB. Eager warmup at boot page-faults weights and pre-compiles decode kernels (first request 3.5× faster). Multi-turn agent loops reuse KV across turns and skip re-prefilling system prompts via a shared-prefix cache that survives interleaved subagent traffic; a Claude Code-sized prompt tokenizes in 8 ms, so a warm agent turn round-trips in ~0.1 s end to end.

Is the inference exact, or quantized output drift?

For greedy decoding (temp=0), mlx-serve is byte-identical to the reference for the first ~30-80 generated tokens, with long-tail divergence inherent to INT4 float-reduction order. For temp > 0, the Leviathan probability-ratio sampler keeps speculative decoding mathematically exact in distribution. Equivalence is pinned by automated tests on every release.

Can mlx-serve generate images, video, and audio locally?

Yes — all on-device, no Python. Image: Krea-2-Turbo (a 12.9B photorealistic model) and FLUX.2 run natively on MLX, validated pixel-faithful to the reference — and it edits photos from a plain instruction ("make the hair blue") while keeping subject and scene, does image-to-image variations with a strength slider, and takes runtime style LoRAs. Video: LTX-Video 2.3 turns a prompt, a photo, or a soundtrack into a clip with synced audio — put spoken lines in quotes and characters talk, lips synced to the voice. Audio: Qwen3-TTS does zero-shot voice cloning from a few seconds of reference audio — no transcript needed. Chat and every media type share one local server and one memory budget: a model loads on demand and unloads when done, so a chat model and a media model can coexist.

Where does my data go?

Nowhere. Everything runs locally on your Mac — no analytics, no telemetry, no cloud calls. The HTTP server binds to 127.0.0.1 by default. Open source under MIT.

How do I install it?

The easiest way is the MLX Core app from GitHub Releases (signed and notarized DMG). Or via Homebrew: brew tap ddalcu/mlx-serve https://github.com/ddalcu/mlx-serve && brew install --cask mlx-core. CLI server alone: brew install mlx-serve.

Have another question? Open an issue · ★ Star the repo if mlx-serve saved you from spinning up another Electron app.