mlx-serve/Deep dive · Music generation

Type a vibe.
Get a song.

"Upbeat synthwave with driving bass" — and ACE-Step 1.5 XL Turbo composes an original 48 kHz stereo track in 8 diffusion steps, entirely on your Mac. Paste lyrics and it sings them; leave them out for an instrumental. No cloud, no subscription, no usage caps.

Download MLX Core How it works
8 diffusion steps per track 10 s – 10 min per generation 48 kHz stereo output
How it works

Describe. Generate. Keep every take.

Open the Audio pane's Music tab, describe the style — genre, mood, instrumentation — and generate. Want vocals? Paste lyrics (verse/chorus tags steer the arrangement) and pick the vocal language. Want control? Pin the BPM, key, and time signature; leave them free and the model chooses.

Style-prompt starters and original lyric templates ship in the Examples menus, and you can save your own to reuse. Every track lands in a persistent history — play, stop, reveal in Finder — with a settings file beside it recording the exact prompt, lyrics, and parameters, so any take is reproducible.

MLX Core Audio pane, Music tab — style prompt, lyrics, duration and musical steering controls with a generation history
The Music tab: style prompt, optional lyrics, musical steering — and a history of every take.
Under the hood

A music diffusion model, ported natively

ACE-Step 1.5 XL Turbo runs on the same no-Python engine as everything else here: a text encoder conditions a 32-layer diffusion transformer, an audio VAE decodes to waveform, and the turbo schedule needs just 8 steps end to end. The port is validated against the reference implementation with per-component numerical oracles.

One click downloads the converted model (~6.3 GB); it generates comfortably in about 9 GB of memory, loads on demand, and frees itself afterwards — chat and music coexist in one server process.

  • No subscription, no credits — generate all night on your own silicon.
  • Your lyrics stay yours — nothing you write leaves the machine.
  • Works offline — once the model is downloaded, airplane mode is fine.

Music that plugs into everything else

API

One endpoint for developers

POST /v1/audio/music-generations with a prompt, optional lyrics, duration, and musical steering — get a 48 kHz stereo WAV back, with per-step SSE progress. Full API docs.

Video

Score your generated videos

Feed a generated track into audio-to-video and the visuals perform to your soundtrack — beat, mood, and all.

Voice

Sibling tab: cloning & TTS

The Audio pane's Voice tab does zero-shot voice cloning — six seconds of audio, no transcript, bit-for-bit validated.

Reproducible

Every take documented

A .txt sidecar beside each track records the model, prompt, lyrics, and parameters — recreate or iterate on any result later.

Local music generation, answered

What hardware do I need?

An Apple Silicon Mac with 16 GB works well — the 8-bit model generates in about 9 GB of memory and is a ~6.3 GB one-click download.

Can it sing my lyrics?

Yes. Paste lyrics and the vocal performance follows them — structure tags like verse and chorus steer the arrangement, and the vocal language is selectable. Empty lyrics produce an instrumental.

How much control do I get over the music itself?

Style, mood, and instrumentation come from the prompt; BPM (30–300), key/scale, and time signature can be pinned exactly or left to the model. Track length runs from 10 seconds to 10 minutes, and a seed makes any generation repeatable.

Who owns the output?

It's generated on your machine from your prompt — there's no service holding rights over your renders. ACE-Step is an open-weights model; check its license for the fine print on commercial use of outputs.

More deep dives

Your studio. Your silicon. No meter running.

Download MLX Core, grab the music model with one click, and render as many takes as you like.