LTX-Video 2.5, 2.3 and MiniMax-H3 run natively on Apple Silicon — clips with synchronized audio, muxed straight to an mp4 on your desk. Start from a prompt, from your own photo, or from a voice clip your characters perform to.
Type a prompt, pick a quality preset, get a clip with synchronized audio. The one-stage path runs guidance-free by default — ~2× faster per step with a more natural look.
Drop an image into the First-frame slot and the clip begins exactly from it — VAE-encoded on-device and locked as the clean opening frame, then animated forward.
Attach speech or music, and the video is generated against that clip — timing, performance, and lip sync follow the audio, and the original recording lands in the mp4, not a re-synthesis.

Write spoken words in quotes in your prompt — short phrases with acting directions between them — and LTX generates the voice, timed to the picture. Want a specific voice? Attach a real recording, or type a line and have the local Qwen3-TTS speak it — including a voice you cloned from a few seconds of reference audio.
Any WAV/MP3/M4A works, the frame count auto-fits the clip length, and audio guidance on the Quality presets steers toward clean speech — clearer voices, less stray background noise.

The Fast preset runs the distilled one-stage pipeline. The Quality and Super-Quality presets run the full reference two-stage pipeline natively: a guided half-resolution pass on the dev model (CFG + modality guidance — Super adds a second-order sampler), a learned 2× latent upscale, then a distilled refine at full resolution.
Style LoRAs apply here too — the same runtime "filters on steroids" that restyle image generations restyle your clips, zero quality loss on the base weights.
.safetensors, dial the strength.LTX-Video 2.5 ships in a 4-bit pack (36 GB) and an 8-bit quality pack (59 GB) with its own text encoder bundled, so there is no separate download on first use. The 8-bit pack keeps faces, legs and fur that 4-bit loses, and adds a Diffusion decoder toggle: the decoder Lightricks' own published clips use, which denoises the frames for sharper texture and edges. The default canvas and frame ladder are sized to your Mac's memory, and 2.3 keeps working right beside it.
"decoder": "diffusion" over the API).MiniMax-H3 denoises the clip and a stereo soundtrack together in one pass, so the sound is produced with the video rather than dubbed on after. Describe the scene, then what you want to hear after overall_soundscape:. The REF2VA build composes the clip around pictures, clips or audio you attach (<Picture 1>, <Video 1>, <Audio 1> in the prompt), and Turbo renders in 4 steps instead of 30.
chain_windows), LoRAs stack on top.Local video generation is the heaviest thing a consumer Mac can do. Here's the real bill — once.
The app checks free RAM before starting and offers to stop the chat server if it's competing for memory. Outputs land in ~/.mlx-serve/generations/videos/, by date. API users: POST /v1/video/generations with prompt, optional first_frame_image, audio, pipeline, lora_paths (up to 8, stacked), and on LTX 2.5 "decoder": "diffusion".
Really. The full LTX-Video pipelines (2.5 and 2.3) and MiniMax-H3 — diffusion transformers, 3D VAEs, audio decoders — were ported natively to MLX and validated tensor-by-tensor against the reference implementation. Your prompts, photos, and clips never leave the Mac.
The audio clip is encoded on-device and held fixed while the video denoises around it, so motion and lip sync follow the sound. The mp4 gets your original recording at native quality — trimmed to the clip length, never re-synthesized.
Yes — the Speech & sound section chains into the local TTS: type the line, pick (or clone) a voice with Qwen3-TTS, and the generated speech drives the video. All three models run through the same local server.
The ~50 GB snapshot ships both transformer variants (one-stage distilled + two-stage dev, ~11 GB each), the upscaler, and the Gemma text encoder — so every quality preset works offline without re-downloading. The downloader pulls only the files the engine actually reads, not the repo's full ~70 GB. LTX 2.5 travels lighter: 36 GB in 4-bit or 59 GB in 8-bit, text encoder bundled.
Download MLX Core, grab LTX-Video with one click, and turn a prompt, a photo, or a voice memo into a clip — audio and all.