/ Blogs & Newsletters

Sogni engineering · New model family · August 10, 2026

Your prompt is now a director

MiniMax H3 and H3 Turbo are live on the Sogni Supernet — the first open-weights video family on our network that renders picture and 32 kHz stereo audio together: scripted dialogue, timestamped cuts, foley, and score, all from one prompt. Here is every workflow, demonstrated, timed, and priced — and what it unlocks that LTX‑2.3 and Wan 2.2 couldn't do.

H3 Standard t2v · three shots, two speakers 10.13 s @ 24 fps native 1344×768 unmute for dialogue ↗

Every generation above — the wok flames, the rain on the awning, the cut to the close-up at 4.5 seconds, both actors' lines, the vinyl-crackle piano that enters under the second shot — came out of the model in a single pass. No editing, no dubbing, no post.

For current pricing, workflow access, and model IDs, see the MiniMax H3 model page.

What shipped

Eight modes across two speed classes, one prompt contract.

SelectorWorkflowSteps
minimax-h3 / -t2vText-to-video20
minimax-h3-i2vFirst-frame image-to-video20
minimax-h3-flf2vFirst frame → last frame20
minimax-h3-r2vReference-set-to-video (9 img / 3 vid / 3 audio)20
minimax-h3-turboTurbo text-to-video4
minimax-h3-i2v-turboTurbo image-to-video4
minimax-h3-flf2v-turboTurbo first → last frame4
minimax-h3-r2v-turboTurbo reference-set-to-video4

All modes run at 24 fps on a fixed frame grid from 5.17 to 15.08 seconds, at 768p-class resolutions (1344×768 landscape or portrait). The checkpoint is CFG-distilled at guidance 1 — there is no negative prompt, no scheduler to tune. Turbo is a LightX2V 4-step distillation of the same checkpoint: every mode, same prompt contract, 37.5% of the cost. Reference-to-video was the last mode to get a Turbo build, and it landed after this post first went up — it is in section 04 below.

01 · Text-to-video: a two-hander with cuts

H3's prompt is a structured shot list, not a caption. You write the timeline — [Shot 1], then [Shot 2] At 00:04.500, the camera cuts to… — give every speaking character a stable ID, and put the exact words inside dialogue tags:

integrated_multimodal_description: [Shot 1] … The vendor with a rough,
cheerful voice (S1) slides a steaming bowl across the counter and says:
<d>[English] Last bowl tonight, extra chili for you.</d> [Shot 2] At
00:04.500, the camera cuts to a close-up across the counter. … (S2) says:
<d>[English] Okay, that is worth the rain.</d> [Shot 3] At 00:08.000, …

overall_soundscape: Steady rain patters on the canvas awning and hisses on
hot oil. The wok clangs and scrapes…

non_diegetic_music: A sparse, warm electric-piano motif with soft vinyl
crackle enters under Shot 2…

The standard render is the hero clip at the top of this page. Here it is against H3 Turbo, same brief, same seed:

Standard 20 steps 7 m 29 s render 162.00 SPARK
Turbo 4 steps 2 m 24 s render 60.75 SPARK

An honest note on A/B-ing t2v: at 4 steps the denoising trajectory diverges, so the same seed gives you a different take on the same screenplay — Turbo staged the market brighter and more frontal, Standard went moodier and more atmospheric. Both followed the shot list and both delivered the dialogue. For directly comparable A/Bs, use the anchored modes below, where the input frames pin the composition.

02 · Image-to-video: a still that talks

H3 i2v pins your image as the true first frame with a one-line alignment preamble, then animates it — with sound. The source is an original Krea 2 Turbo still, generated on Sogni with a detail-enhancer LoRA stack. We asked the watchmaker to finish the adjustment, set down his tweezers, lift the loupe, and deliver a line over a room full of off-rate ticking clocks.

Krea 2 Turbo still: elderly watchmaker with loupe at a cluttered walnut bench
Source frame — Krea 2 Turbo (krea2_turbo_fp8_scaled), 1344×768, an original still, never a video frame
Standard 12 m 16 s render 128 SPARK 8.00 s
Turbo 1 m 48 s render 48 SPARK 8.00 s

Both preserve the source identity cleanly — the pores, the silver stubble, the brass clutter. Standard spent its extra steps on subtler facial articulation through the line delivery; Turbo, from the same locked frame, is 6.8× faster and holds up remarkably well at a quarter of the steps.

03 · First → last frame: choreograph to a pose

Give flf2v two anchor images and the model invents the motion — and the audio — that connects them. Both anchors are original stills: the end frame is a Krea 2 Identity Edit of the first (same violinist, same plaza, bow raised, crowd added). H3 has to compose the performance in between, including the diegetic violin line.

First frame: violinist poised, bow lowered, empty plaza at dusk
First frame — 0.00 s · Krea 2 Turbo
Last frame: violinist mid-performance, bow raised, small crowd gathered
Last frame — 8.00 s · Krea 2 Identity Edit v1.2
Standard 15 m 50 s render 128 SPARK
Turbo 4 m 16 s render 48 SPARK

04 · Reference-to-video: cast, dress, and place a character

minimax-h3-r2v is a separate checkpoint (now with a Turbo build alongside it) that conditions on a labelled reference set — up to 9 images, 3 videos, and 3 audio clips. Its six-field prompt gives every reference one explicit job. We handed it three unrelated stills:

Identity reference: woman with copper-red braid and freckles in denim jacket
<Picture 1> identity
Wardrobe reference: emerald silk evening gown on a dress form
<Picture 2> wardrobe
Location reference: rain-slicked rooftop bar with string lights and skyline
<Picture 3> location
Standard r2v 20 steps ≈ 7.5 m render 128 SPARK
Turbo r2v 4 steps 2 m 16 s render 48 SPARK

The braid, the freckles, the gown's cowl neckline, the wet tile and skyline bokeh — composed into one tracking shot of a person who exists in no single source image, walking to the railing and delivering a line. There is no equivalent workflow on LTX‑2.3 or Wan 2.2.

Update: reference-to-video now has a Turbo build too, and it is the most surprising of the family. The clip on the right is the same three references, the same six-field prompt and the same seed at 4 steps instead of 20: same copper braid and freckles, same emerald cowl, same rain-slicked terrace and string lights, and the close-up still holds her identity at full resolution. It renders in 2 m 16 s against roughly 7.5 minutes and costs 48 SPARK instead of 128. The Turbo take reads a little punchier — deeper contrast, hotter bulbs, more saturated skyline — where Standard is the more restrained, naturalistic grade. Casting a character from a reference set is no longer the slow, expensive corner of this model family.

What this unlocks vs LTX-2.3 and Wan 2.2

We ran the same Night Market brief through our previous defaults, prompt-adapted to each model's own contract:

LTX-2.3 distilled 10.04 s · 1536×864 web encode 3 m 23 s render 76.52 SPARK
Wan 2.2 t2v 4.97 s · 1280×720 · silent 2 m 22 s render 50.77 SPARK

Both are respectable clips — LTX brings real audio and dialogue in its one continuous shot, and the Wan 2.2 t2v baseline moves nicely but renders silent. Neither can cut. The screenplay we gave H3 is simply not expressible in either model's contract:

CapabilityWan 2.2 t2vLTX-2.3MiniMax H3
Native audiosilent✓ 32 kHz stereo, joint
Multi-shot cuts in one generationsingle shot by design✓ timestamped [Shot N]
Stable speaker IDs across cuts✓ (S1), (S2), (S1,S2)
Dialogue carrying across a cut✓ <scenetrans>
Off-screen voiceover✓ contract phrase
Dialogue languagesEnglish-centric11 languages
First → last framemorph-LoRA path✓ native, with audio
Reference-set casting✓ 9 img / 3 vid / 3 audio, Standard and Turbo
Dedicated score fieldin-prose only✓ non_diegetic_music
Max clip length≈ 5 s≈ 10 s15.08 s

Wan gave you motion. LTX gave you a shot. H3 gives you a scene.

Speed & cost, measured

Wall-clock: one render each on Sogni Supernet fast-network workers, 2026‑08‑09 (varies with the worker you land on). Prices reflect the current network pricing update.

RenderWall clockSPARKUSDvs Standard
Std t2v 10 s7 m 29 s162.00$0.81
Turbo t2v 10 s2 m 24 s60.75$0.303.1× faster, 37.5% cost
Std i2v 8 s12 m 16 s128.00$0.64
Turbo i2v 8 s1 m 48 s48.00$0.246.8× faster, 37.5% cost
Std flf2v 8 s15 m 50 s128.00$0.64
Turbo flf2v 8 s4 m 16 s48.00$0.243.7× faster, 37.5% cost
Std r2v 8 s≈ 7.5 m128.00$0.64
Turbo r2v 8 s2 m 16 s48.00$0.243.3× faster, 37.5% cost
LTX-2.3 t2v 10 s3 m 23 s76.52$0.38
Wan 2.2 t2v 5 s2 m 22 s50.77$0.25

Worth pausing on: at the updated prices, a fully scripted 10-second two-hander with dialogue, cuts, and score on H3 Turbo costs 60.75 SPARK ($0.30) — only $0.05 more than the silent five-second Wan 2.2 t2v baseline, and 21% less than the single-shot LTX render.

Notes for prompt writers

Available now. MiniMax H3 and H3 Turbo are already live in Sogni for pay-as-you-go Spark and eligible Sogni Unlimited subscription plans. Explore the models and start creating.

More from Sogni

All reference stills are original Krea 2 Turbo / Krea 2 Identity Edit generations made on Sogni — no extracted video frames were used anywhere in this work. Clips on this page are web-ready encodes (H3 at 1344×768, LTX at 1536×864, Wan at 1280×720) — use the ⛶ View at… button on any clip to open it at its encoded size. Gallery masters keep the original bitrate and audio tracks. Rendered on the Sogni Supernet, 2026-08-09; the Turbo reference-to-video clip added 2026-08-15.