Tokens per second measured by the people running the models, not by us. Pick your chip to see what to expect, or compare settings to find out which ones are worth turning on.
Every result comes from the same pinned workload: a context ladder (512 / 1k / 2k / 4k / 8k / 16k tokens of synthetic code with a planted constant). Run it yourself from the menu bar in MLX Core.