Sogni engineering · New model family · August 10, 2026
Your prompt is now a director
MiniMax H3 and H3 Turbo are live on the Sogni Supernet — the first open-weights video family on our network that renders picture and 32 kHz stereo audio together: scripted dialogue, timestamped cuts, foley, and score, all from one prompt. Here is every workflow, demonstrated, timed, and priced — and what it unlocks that LTX‑2.3 and Wan 2.2 couldn't do.
Every generation above — the wok flames, the rain on the awning, the cut to the close-up at 4.5 seconds, both actors' lines, the vinyl-crackle piano that enters under the second shot — came out of the model in a single pass. No editing, no dubbing, no post.
For current pricing, workflow access, and model IDs, see the MiniMax H3 model page.
What shipped
Eight modes across two speed classes, one prompt contract.
| Selector | Workflow | Steps |
|---|---|---|
minimax-h3 / -t2v | Text-to-video | 20 |
minimax-h3-i2v | First-frame image-to-video | 20 |
minimax-h3-flf2v | First frame → last frame | 20 |
minimax-h3-r2v | Reference-set-to-video (9 img / 3 vid / 3 audio) | 20 |
minimax-h3-turbo | Turbo text-to-video | 4 |
minimax-h3-i2v-turbo | Turbo image-to-video | 4 |
minimax-h3-flf2v-turbo | Turbo first → last frame | 4 |
minimax-h3-r2v-turbo | Turbo reference-set-to-video | 4 |
All modes run at 24 fps on a fixed frame grid from 5.17 to 15.08 seconds, at 768p-class resolutions (1344×768 landscape or portrait). The checkpoint is CFG-distilled at guidance 1 — there is no negative prompt, no scheduler to tune. Turbo is a LightX2V 4-step distillation of the same checkpoint: every mode, same prompt contract, 37.5% of the cost. Reference-to-video was the last mode to get a Turbo build, and it landed after this post first went up — it is in section 04 below.
01 · Text-to-video: a two-hander with cuts
H3's prompt is a structured shot list, not a caption. You write the timeline — [Shot 1], then [Shot 2] At 00:04.500, the camera cuts to… — give every speaking character a stable ID, and put the exact words inside dialogue tags:
integrated_multimodal_description: [Shot 1] … The vendor with a rough,
cheerful voice (S1) slides a steaming bowl across the counter and says:
<d>[English] Last bowl tonight, extra chili for you.</d> [Shot 2] At
00:04.500, the camera cuts to a close-up across the counter. … (S2) says:
<d>[English] Okay, that is worth the rain.</d> [Shot 3] At 00:08.000, …
overall_soundscape: Steady rain patters on the canvas awning and hisses on
hot oil. The wok clangs and scrapes…
non_diegetic_music: A sparse, warm electric-piano motif with soft vinyl
crackle enters under Shot 2…
The standard render is the hero clip at the top of this page. Here it is against H3 Turbo, same brief, same seed:
An honest note on A/B-ing t2v: at 4 steps the denoising trajectory diverges, so the same seed gives you a different take on the same screenplay — Turbo staged the market brighter and more frontal, Standard went moodier and more atmospheric. Both followed the shot list and both delivered the dialogue. For directly comparable A/Bs, use the anchored modes below, where the input frames pin the composition.
02 · Image-to-video: a still that talks
H3 i2v pins your image as the true first frame with a one-line alignment preamble, then animates it — with sound. The source is an original Krea 2 Turbo still, generated on Sogni with a detail-enhancer LoRA stack. We asked the watchmaker to finish the adjustment, set down his tweezers, lift the loupe, and deliver a line over a room full of off-rate ticking clocks.
Both preserve the source identity cleanly — the pores, the silver stubble, the brass clutter. Standard spent its extra steps on subtler facial articulation through the line delivery; Turbo, from the same locked frame, is 6.8× faster and holds up remarkably well at a quarter of the steps.
03 · First → last frame: choreograph to a pose
Give flf2v two anchor images and the model invents the motion — and the audio — that connects them. Both anchors are original stills: the end frame is a Krea 2 Identity Edit of the first (same violinist, same plaza, bow raised, crowd added). H3 has to compose the performance in between, including the diegetic violin line.


04 · Reference-to-video: cast, dress, and place a character
minimax-h3-r2v is a separate checkpoint (now with a Turbo build alongside it) that conditions on a labelled reference set — up to 9 images, 3 videos, and 3 audio clips. Its six-field prompt gives every reference one explicit job. We handed it three unrelated stills:



The braid, the freckles, the gown's cowl neckline, the wet tile and skyline bokeh — composed into one tracking shot of a person who exists in no single source image, walking to the railing and delivering a line. There is no equivalent workflow on LTX‑2.3 or Wan 2.2.
Update: reference-to-video now has a Turbo build too, and it is the most surprising of the family. The clip on the right is the same three references, the same six-field prompt and the same seed at 4 steps instead of 20: same copper braid and freckles, same emerald cowl, same rain-slicked terrace and string lights, and the close-up still holds her identity at full resolution. It renders in 2 m 16 s against roughly 7.5 minutes and costs 48 SPARK instead of 128. The Turbo take reads a little punchier — deeper contrast, hotter bulbs, more saturated skyline — where Standard is the more restrained, naturalistic grade. Casting a character from a reference set is no longer the slow, expensive corner of this model family.
What this unlocks vs LTX-2.3 and Wan 2.2
We ran the same Night Market brief through our previous defaults, prompt-adapted to each model's own contract:
Both are respectable clips — LTX brings real audio and dialogue in its one continuous shot, and the Wan 2.2 t2v baseline moves nicely but renders silent. Neither can cut. The screenplay we gave H3 is simply not expressible in either model's contract:
| Capability | Wan 2.2 t2v | LTX-2.3 | MiniMax H3 |
|---|---|---|---|
| Native audio | silent | ✓ | ✓ 32 kHz stereo, joint |
| Multi-shot cuts in one generation | — | single shot by design | ✓ timestamped [Shot N] |
| Stable speaker IDs across cuts | — | — | ✓ (S1), (S2), (S1,S2) |
| Dialogue carrying across a cut | — | — | ✓ <scenetrans> |
| Off-screen voiceover | — | — | ✓ contract phrase |
| Dialogue languages | — | English-centric | 11 languages |
| First → last frame | — | morph-LoRA path | ✓ native, with audio |
| Reference-set casting | — | — | ✓ 9 img / 3 vid / 3 audio, Standard and Turbo |
| Dedicated score field | — | in-prose only | ✓ non_diegetic_music |
| Max clip length | ≈ 5 s | ≈ 10 s | 15.08 s |
Wan gave you motion. LTX gave you a shot. H3 gives you a scene.
Speed & cost, measured
Wall-clock: one render each on Sogni Supernet fast-network workers, 2026‑08‑09 (varies with the worker you land on). Prices reflect the current network pricing update.
| Render | Wall clock | SPARK | USD | vs Standard |
|---|---|---|---|---|
| Std t2v 10 s | 7 m 29 s | 162.00 | $0.81 | — |
| Turbo t2v 10 s | 2 m 24 s | 60.75 | $0.30 | 3.1× faster, 37.5% cost |
| Std i2v 8 s | 12 m 16 s | 128.00 | $0.64 | — |
| Turbo i2v 8 s | 1 m 48 s | 48.00 | $0.24 | 6.8× faster, 37.5% cost |
| Std flf2v 8 s | 15 m 50 s | 128.00 | $0.64 | — |
| Turbo flf2v 8 s | 4 m 16 s | 48.00 | $0.24 | 3.7× faster, 37.5% cost |
| Std r2v 8 s | ≈ 7.5 m | 128.00 | $0.64 | — |
| Turbo r2v 8 s | 2 m 16 s | 48.00 | $0.24 | 3.3× faster, 37.5% cost |
| LTX-2.3 t2v 10 s | 3 m 23 s | 76.52 | $0.38 | — |
| Wan 2.2 t2v 5 s | 2 m 22 s | 50.77 | $0.25 | — |
Worth pausing on: at the updated prices, a fully scripted 10-second two-hander with dialogue, cuts, and score on H3 Turbo costs 60.75 SPARK ($0.30) — only $0.05 more than the silent five-second Wan 2.2 t2v baseline, and 21% less than the single-shot LTX render.
Notes for prompt writers
- H3 wants its three-field contract —
integrated_multimodal_description,overall_soundscape,non_diegetic_music— not a caption. i2v and flf2v add one exact alignment preamble line pinning your frames to the timeline. - Every sound must be asked for. Audio is generated jointly with the picture; an empty soundscape field is a silent alley.
- There is no negative prompt — the checkpoint is CFG-distilled at guidance 1. Write negative direction into the prose.
- Durations snap to the 24 fps frame grid (
124 + n×17frames): ask for 10 s, get 10.13 s. Keep every cut timestamp inside the snapped duration. - On r2v, give each reference one explicit job — identity, wardrobe, location — and say who wins on conflict. Unassigned references are the top cause of identity drift.
- Turbo defaults to the
er_sdesampler;eulerandsa_solverare available for A/B tests.
Available now. MiniMax H3 and H3 Turbo are already live in Sogni for pay-as-you-go Spark and eligible Sogni Unlimited subscription plans. Explore the models and start creating.
More from Sogni
- MiniMax Music 3 vs ACE-Step 1.5 XL — Which AI Music Model Wins?
- Starting & Ending Frames — Direct the Whole Shot
- Upscale to 8K Without Changing Your Image — RTX Video Super Resolution on Sogni
All reference stills are original Krea 2 Turbo / Krea 2 Identity Edit generations made on Sogni — no extracted video frames were used anywhere in this work. Clips on this page are web-ready encodes (H3 at 1344×768, LTX at 1536×864, Wan at 1280×720) — use the ⛶ View at… button on any clip to open it at its encoded size. Gallery masters keep the original bitrate and audio tracks. Rendered on the Sogni Supernet, 2026-08-09; the Turbo reference-to-video clip added 2026-08-15.