What FastH3 is
FastH3 Preview v1 is the FastVideo team's four-step distillation of MiniMax H3, built at UC San Diego's Hao AI Lab with Nuva Lab and NVIDIA's FastGen group. Two things make it fast. The model was distilled with data-free DMD2 so a clip needs four transformer passes instead of Standard H3's twenty. And it swaps dense attention for Video Sparse Attention at 90% sparsity, so each of those passes does far less work. Sparse attention pays off most on long clips, which is why a full 15-second render is where the gap opens widest.
On Sogni, FastH3 is now the default Turbo engine for text-to-video, image-to-video, and first-and-last-frame video. It runs Kijai's INT8 conversion of the checkpoint, keeps H3's native 32 kHz stereo audio, works with the H3 LoRAs, and bills 4 Spark per output second at every resolution. The previous LightX2V four-step Turbo stays available behind a switch. Multi-reference video has no FastH3 mode yet.
FastVideo's model card says the preview was distilled for text-to-video, with frame-guided and reference modes still to come. Our test is image-to-video on purpose. If the distilled weights hold a specific face, a spoken line, and a cut when they are anchored to a real still, that is the harder case.
How we timed it
Every variable that could move was pinned. One worker, targeted by its worker NFT so the Supernet could not route us to a bigger card. One source still. One prompt, unchanged across tiers, including a scripted line of dialogue, a mid-clip cut, and a meow. One seed. The maximum H3 length of 362 frames, which is 15.08 seconds at 24 fps.
| Model | MiniMax H3 image-to-video, FL2VA FP8 checkpoint for Standard, Balanced, and Turbo; FastVideo FastH3 4-step Preview v1 (INT8) for FastH3 Turbo |
|---|---|
| Tiers | Standard: 20 steps, res_multistep. Balanced: 8 steps, LightX2V. Turbo: 4 steps, LightX2V. FastH3 Turbo: 4 steps, FastVideo VSA. All at guidance 1, simple scheduler, no LoRAs, audio on. |
| Resolutions | 768 × 1024 and 480 × 640, both the 3:4 portrait presets in Sogni Web |
| Length | 362 frames, 15.08 s, 24 fps, 32 kHz stereo audio generated in the same pass |
| Seed | 20260902 for every clip shown on this page. The timing repeat of each configuration ran with a fresh socket-assigned seed. |
| Worker | mark.and.worker #19, RTX 5090 32 GB, Sogni comfy-worker production image, Windows 11 with Docker Desktop and WSL2 |
| Timing | Wall clock from the worker's job-started event to the completed file, as reported over the Sogni socket. Includes sampling, video and audio VAE decode, encoding, and upload. Excludes model loading and queue wait. |
| Runs | Two per configuration, back to back. The clips shown are the first runs, all on the same seed; the chart uses the second, hot run. 16 clips in total for 1810 Spark ($9.05) at retail. |
The runs were submitted through the Sogni SDK with the in-prompt worker selector, which is available to Premium Spark and SOGNI jobs. The cat's script is Mark's; we added the meow.
The full prompt, exactly as submitted
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Grainy cinematic live-action with warm, diffused bathroom lighting. The video begins exactly from <Picture 1>, preserving the black cat’s face, amber-green eyes, wet fur, surrounding white bubbles, centered close-up composition, and analog film texture. The cat remains almost perfectly still, glances suspiciously from side to side, and then looks directly into the camera. A few bubbles pop softly around his ears. The cat has a relaxed, streetwise adult male voice with a warm Spanish accent (S1). His natural feline mouth and muzzle move subtly in precise synchronization without developing human lips. He begins with a secretive whisper, then speaks quickly and casually: <d>[English] Pssst. Hey, Sogni fam. MiniMax H3 Turbo videos on the network just got two times faster! You didn’t hear that from me, though. I’m just a cat in a bath. You get me, homie?</d> He finishes with a knowing expression and holds eye contact with the camera. [Shot 2] At 00:11.500, the camera cuts to a slightly wider overhead view of the same cat surrounded by the same bath bubbles. After completely finishing his line, he gives the camera one final suspicious glance, then opens his mouth and lets out one short, ordinary cat meow, a real feline meow with no human voice. He slowly closes both eyes and peacefully sinks beneath the bubbles. His face disappears first, followed by his ears, while the foam gently closes over him. The water becomes still and one final bubble pops at exactly 15.00 seconds. Preserve the cat’s identity, realistic feline anatomy, bath environment, lighting, and grainy photographic style across both shots. overall_soundscape: Quiet tiled-bathroom room tone with gentle water movement and tiny bubbles popping around the cat. His opening whisper sounds intimate and close to the microphone, followed by clear Spanish-accented speech. One short, genuine cat meow, then a soft watery burble as he sinks beneath the foam, followed by one isolated bubble pop at the end. non_diegetic_music: A sparse, playful upright-bass pattern plays quietly beneath the scene. The music stops abruptly with the final bubble pop.
The 768p renders
Same still, same prompt, same seed. Hot, Standard needs 11 min 55 s, Balanced 5 min 30 s, LightX2V Turbo 3 min 0 s, and FastH3 Turbo 2 min 6 s; each card shows its own run's time. Play them with sound: the spoken line and the meow are part of the test.
The reference look. Softest grain, most natural fur, the muted warm tone of the source still. Twenty steps buy subtlety, not a different cat.
Open the original fileKeeps the most film grain and the warmest tone of the four. Subtler mouth shapes.
Open the original fileThe LightX2V look: sharp, bright eyes, a wide toothy mouth on the loud syllables.
Open the original fileIdentity holds through both shots. Cleaner, less grainy than the source; the eyes stay sharp and the foam keeps its sparkle.
Open the original fileThe 480p renders
Sogni Web starts H3 at 480p, so this is the size most first drafts render at. Hot, Standard takes 3 min 39 s, Balanced 1 min 46 s, Turbo 1 min 7 s, and FastH3 55 s. The gap narrows here because the fixed decode and upload tail is a bigger share of a small render.
Same character as 768p Standard with less resolution to spend it on. Still the grainiest, warmest clip at this size.
Open the original fileGrain and warmth survive the drop to 480p. The overhead meow reads clearly.
Open the original fileClose to FastH3 480p frame for frame. The cat goes fully under at the end.
Open the original fileSame beats as 768p at a quarter of the pixels. Softer fur, but the cut, the meow, and the sink all land.
Open the original fileWhere the time goes
The socket reports the start of every sampling step, so we can split each hot render into the diffusion passes and everything after them: decoding 362 frames of video and 15 seconds of audio, encoding the MP4, and uploading it. That tail is close to fixed for a given resolution, which is why four-step tiers stop scaling the way step counts suggest they should. The sampling column extends the measured average step length to cover the final step.
| Size | Tier | First run | Hot run | vs Standard | Steps | Sampling | Decode, encode, upload |
|---|---|---|---|---|---|---|---|
| 768p | Standard | 11 min 51 s | 11 min 55 s | baseline | 20 | 11 min 10 s | 45 s |
| 768p | Balanced | 5 min 35 s | 5 min 30 s | 2.2× | 8 | 4 min 46 s | 45 s |
| 768p | Turbo | 3 min 29 s | 3 min 0 s | 4.0× | 4 | 2 min 12 s | 48 s |
| 768p | FastH3 Turbo | 2 min 11 s | 2 min 6 s | 5.7× | 4 | 1 min 19 s | 46 s |
| 480p | Standard | 3 min 43 s | 3 min 39 s | baseline | 20 | 3 min 10 s | 29 s |
| 480p | Balanced | 1 min 44 s | 1 min 46 s | 2.1× | 8 | 1 min 20 s | 26 s |
| 480p | Turbo | 1 min 8 s | 1 min 7 s | 3.3× | 4 | 40 s | 27 s |
| 480p | FastH3 Turbo | 54 s | 55 s | 4.0× | 4 | 26 s | 29 s |
The first run of each configuration is the clip embedded above. It follows a model switch on a worker that was also serving public traffic, so it can carry kernel warm-up. The hot run is the number to plan around.
Warm-up is not a rounding error. A FastH3 480p canary we ran before the matrix, the first FastH3 job after that worker restarted, took 2 min 40 s from job start to finished file. Inside the matrix the same configuration took 54 s. Sparse-attention kernels and the INT8 path pay a one-time cost on a fresh process, which is one more reason a busy worker that keeps FastH3 loaded is the one you want your job to land on.
One honest wrinkle: after the FastH3 pairs, the switch back to the LightX2V checkpoint ran the 32 GB card out of memory on the first Turbo 768p attempt. The worker recovered by unloading everything, but it then reloaded the transformer partly offloaded to system memory, which would have made every later number meaningless. We restarted the worker and re-ran the remaining six configurations from a clean process, so the first Turbo run below followed a cold start.
What each clip costs
FastH3 is the cheapest tier at both sizes as well as the fastest. Retail pay-as-you-go rates before plan discounts, with the exact quote the socket returned for these 15.08-second clips.
| Tier | Rate at 768p / 480p | 15 s at 768p | 15 s at 480p |
|---|---|---|---|
| Standard | 16 / 10 Spark per second | 241 Spark ($1.21) | 151 Spark ($0.75) |
| Balanced | 10 / 6 Spark per second | 151 Spark ($0.75) | 90 Spark ($0.45) |
| Turbo | 6 / 4 Spark per second | 90 Spark ($0.45) | 60 Spark ($0.30) |
| FastH3 Turbo | 4 / 4 Spark per second | 60 Spark ($0.30) | 60 Spark ($0.30) |
Sogni production pricing as of September 2, 2026. Unlimited and Unlimited Pro plans cover eligible H3 jobs without spending Spark.
What to look for
Sound on, and watch the same three beats in every clip: the whispered opening, the cut to the overhead shot at 11.5 seconds, and the meow before the sink. Every tier delivered all three from the same still and seed, which is the headline for anyone who has used earlier four-step distillations. FastH3 did not lose the cat, the shot, or the script.
- Standard is the reference. It keeps the most of the source photograph: soft film grain, the muted warm cast, fur that reads as wet fur. The mouth shapes on the loud syllables are the most restrained of the four.
- Balanced sits closest to Standard in texture and tone, with more of the grain surviving than in either four-step tier. If you like the film look, this is the cheap way to keep it.
- Turbo (LightX2V) trades grain for punch: cleaner foam, brighter eyes, and a wider, toothier mouth when the cat gets loud.
- FastH3 Turbo lands next to Turbo rather than next to Standard. Slightly cleaner and smoother than LightX2V, with sharp eyes and intact whiskers; the film grain is mostly gone. At 480p the two four-step tiers are hard to tell apart in stills.
The cat's own line oversells one number. He says Turbo got two times faster; Sogni's model page says up to 2×. On this card with a full 15-second clip we measured FastH3 at 1.4× the LightX2V Turbo at 768p and a smaller margin at 480p, because a large share of any four-step render is decode and upload that no distillation touches. The 5.7× figure against Standard is the one that changes how you work.
What we cannot score from frames is lip sync and voice, so we will not pretend to. Play the 768p Standard and FastH3 clips back to back with audio and decide whether the difference is worth 11 min 55 s of your GPU's time against 2 min 6 s.
Verdict
For a full 15-second, 768p, image-to-video clip with dialogue on a consumer RTX 5090, FastH3 Turbo is 5.7× faster than Standard and 1.4× faster than the LightX2V Turbo it replaces as Sogni Web's default, and it is the cheapest tier at either size. It kept the cat, the cut, and the script. What it gives up is the film grain and some of the subtlety that twenty steps buy, so Standard stays the right call for a final master where texture matters. For everything before that, drafts, timing passes, script iterations, and most social clips, FastH3 is now where we start.
Sources
- MiniMax H3 with FastH3 Turbo on the Sogni Supernet
- FastVideo FastH3 4-step Preview v1 VSA DataFree model card
- Hao AI Lab: FastVideo FastH3 V1 preview post
- LightX2V MiniMax H3 Turbo distillation
- MiniMax H3 model card and community license
- Sogni client SDK examples, including the MiniMax H3 workflow used here