22 September 2026 · Release 26.9.5
Why mlx-serve is so fast: 239 tokens a second on a Mac, no Python
Four 27B chats at 119 tokens a second combined, 56 ms to the first token on a continued chat, and 8-bit KV as fast as bf16 at 32k. What the engine does with each read of the weights.