Sogni / Blogs & Newsletters

FastH3 renders MiniMax H3 5.7× faster. Same cat, same GPU, stopwatch running.

We rendered one 15-second image-to-video prompt through all four MiniMax H3 speed tiers on a single RTX 5090, twice each, at 768p and 480p. FastH3 Turbo, FastVideo's sparse-attention four-step distillation, finished the 768p clip in 2 min 6 s. Standard H3 took 11 min 55 s. Every render is below so you can judge what the speed costs.

A black cat with amber-green eyes peering out of white bath foam, shot on grainy film
The source frame. Every clip on this page started from this one still and this one prompt.

Hot render time for the same 15-second clip on one RTX 5090

Second of two back-to-back runs per tier, measured from the worker's job-started event to the finished MP4 arriving. Model loading and queue time are excluded.

768 × 1024 15.08 s, 362 frames at 24 fps

Standard 11:55baseline
Balanced 5:302.2× faster
Turbo 3:004.0× faster
FastH3 Turbo 2:065.7× faster

480 × 640 15.08 s, 362 frames at 24 fps

Standard 3:39baseline
Balanced 1:462.1× faster
Turbo 1:073.3× faster
FastH3 Turbo 0:554.0× faster

Worker: mark.and.worker #19, an RTX 5090 with 32 GB on the public Sogni Supernet, running the production comfy-worker image. The worker kept serving other people's jobs between ours, which is why we timed from job start rather than from submission.

What FastH3 is

FastH3 Preview v1 is the FastVideo team's four-step distillation of MiniMax H3, built at UC San Diego's Hao AI Lab with Nuva Lab and NVIDIA's FastGen group. Two things make it fast. The model was distilled with data-free DMD2 so a clip needs four transformer passes instead of Standard H3's twenty. And it swaps dense attention for Video Sparse Attention at 90% sparsity, so each of those passes does far less work. Sparse attention pays off most on long clips, which is why a full 15-second render is where the gap opens widest.

On Sogni, FastH3 is now the default Turbo engine for text-to-video, image-to-video, and first-and-last-frame video. It runs Kijai's INT8 conversion of the checkpoint, keeps H3's native 32 kHz stereo audio, works with the H3 LoRAs, and bills 4 Spark per output second at every resolution. The previous LightX2V four-step Turbo stays available behind a switch. Multi-reference video has no FastH3 mode yet.

FastVideo's model card says the preview was distilled for text-to-video, with frame-guided and reference modes still to come. Our test is image-to-video on purpose. If the distilled weights hold a specific face, a spoken line, and a cut when they are anchored to a real still, that is the harder case.

How we timed it

Every variable that could move was pinned. One worker, targeted by its worker NFT so the Supernet could not route us to a bigger card. One source still. One prompt, unchanged across tiers, including a scripted line of dialogue, a mid-clip cut, and a meow. One seed. The maximum H3 length of 362 frames, which is 15.08 seconds at 24 fps.

ModelMiniMax H3 image-to-video, FL2VA FP8 checkpoint for Standard, Balanced, and Turbo; FastVideo FastH3 4-step Preview v1 (INT8) for FastH3 Turbo
TiersStandard: 20 steps, res_multistep. Balanced: 8 steps, LightX2V. Turbo: 4 steps, LightX2V. FastH3 Turbo: 4 steps, FastVideo VSA. All at guidance 1, simple scheduler, no LoRAs, audio on.
Resolutions768 × 1024 and 480 × 640, both the 3:4 portrait presets in Sogni Web
Length362 frames, 15.08 s, 24 fps, 32 kHz stereo audio generated in the same pass
Seed20260902 for every clip shown on this page. The timing repeat of each configuration ran with a fresh socket-assigned seed.
Workermark.and.worker #19, RTX 5090 32 GB, Sogni comfy-worker production image, Windows 11 with Docker Desktop and WSL2
TimingWall clock from the worker's job-started event to the completed file, as reported over the Sogni socket. Includes sampling, video and audio VAE decode, encoding, and upload. Excludes model loading and queue wait.
RunsTwo per configuration, back to back. The clips shown are the first runs, all on the same seed; the chart uses the second, hot run. 16 clips in total for 1810 Spark ($9.05) at retail.

The runs were submitted through the Sogni SDK with the in-prompt worker selector, which is available to Premium Spark and SOGNI jobs. The cat's script is Mark's; we added the meow.

The full prompt, exactly as submitted
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Grainy cinematic live-action with warm, diffused bathroom lighting. The video begins exactly from <Picture 1>, preserving the black cat’s face, amber-green eyes, wet fur, surrounding white bubbles, centered close-up composition, and analog film texture. The cat remains almost perfectly still, glances suspiciously from side to side, and then looks directly into the camera. A few bubbles pop softly around his ears. The cat has a relaxed, streetwise adult male voice with a warm Spanish accent (S1). His natural feline mouth and muzzle move subtly in precise synchronization without developing human lips. He begins with a secretive whisper, then speaks quickly and casually: <d>[English] Pssst. Hey, Sogni fam. MiniMax H3 Turbo videos on the network just got two times faster! You didn’t hear that from me, though. I’m just a cat in a bath. You get me, homie?</d> He finishes with a knowing expression and holds eye contact with the camera.

[Shot 2] At 00:11.500, the camera cuts to a slightly wider overhead view of the same cat surrounded by the same bath bubbles. After completely finishing his line, he gives the camera one final suspicious glance, then opens his mouth and lets out one short, ordinary cat meow, a real feline meow with no human voice. He slowly closes both eyes and peacefully sinks beneath the bubbles. His face disappears first, followed by his ears, while the foam gently closes over him. The water becomes still and one final bubble pops at exactly 15.00 seconds. Preserve the cat’s identity, realistic feline anatomy, bath environment, lighting, and grainy photographic style across both shots.

overall_soundscape: Quiet tiled-bathroom room tone with gentle water movement and tiny bubbles popping around the cat. His opening whisper sounds intimate and close to the microphone, followed by clear Spanish-accented speech. One short, genuine cat meow, then a soft watery burble as he sinks beneath the foam, followed by one isolated bubble pop at the end.

non_diegetic_music: A sparse, playful upright-bass pattern plays quietly beneath the scene. The music stops abruptly with the final bubble pop.

The 768p renders

Same still, same prompt, same seed. Hot, Standard needs 11 min 55 s, Balanced 5 min 30 s, LightX2V Turbo 3 min 0 s, and FastH3 Turbo 2 min 6 s; each card shows its own run's time. Play them with sound: the spoken line and the meow are part of the test.

Standard 11 min 51 s This run, seed 20260902. Hot repeat 11 min 55 s, the baseline. 241 Spark ($1.21) per clip.

The reference look. Softest grain, most natural fur, the muted warm tone of the source still. Twenty steps buy subtlety, not a different cat.

Open the original file
Balanced 5 min 35 s This run, seed 20260902. Hot repeat 5 min 30 s, 2.2× faster than Standard. 151 Spark ($0.75) per clip.

Keeps the most film grain and the warmest tone of the four. Subtler mouth shapes.

Open the original file
Turbo 3 min 29 s This run, seed 20260902. Hot repeat 3 min 0 s, 4.0× faster than Standard. 90 Spark ($0.45) per clip.

The LightX2V look: sharp, bright eyes, a wide toothy mouth on the loud syllables.

Open the original file
FastH3 Turbo 2 min 11 s This run, seed 20260902. Hot repeat 2 min 6 s, 5.7× faster than Standard. 60 Spark ($0.30) per clip.

Identity holds through both shots. Cleaner, less grainy than the source; the eyes stay sharp and the foam keeps its sparkle.

Open the original file

The 480p renders

Sogni Web starts H3 at 480p, so this is the size most first drafts render at. Hot, Standard takes 3 min 39 s, Balanced 1 min 46 s, Turbo 1 min 7 s, and FastH3 55 s. The gap narrows here because the fixed decode and upload tail is a bigger share of a small render.

Standard 3 min 43 s This run, seed 20260902. Hot repeat 3 min 39 s, the baseline. 151 Spark ($0.75) per clip.

Same character as 768p Standard with less resolution to spend it on. Still the grainiest, warmest clip at this size.

Open the original file
Balanced 1 min 44 s This run, seed 20260902. Hot repeat 1 min 46 s, 2.1× faster than Standard. 90 Spark ($0.45) per clip.

Grain and warmth survive the drop to 480p. The overhead meow reads clearly.

Open the original file
Turbo 1 min 8 s This run, seed 20260902. Hot repeat 1 min 7 s, 3.3× faster than Standard. 60 Spark ($0.30) per clip.

Close to FastH3 480p frame for frame. The cat goes fully under at the end.

Open the original file
FastH3 Turbo 54 s This run, seed 20260902. Hot repeat 55 s, 4.0× faster than Standard. 60 Spark ($0.30) per clip.

Same beats as 768p at a quarter of the pixels. Softer fur, but the cut, the meow, and the sink all land.

Open the original file

Where the time goes

The socket reports the start of every sampling step, so we can split each hot render into the diffusion passes and everything after them: decoding 362 frames of video and 15 seconds of audio, encoding the MP4, and uploading it. That tail is close to fixed for a given resolution, which is why four-step tiers stop scaling the way step counts suggest they should. The sampling column extends the measured average step length to cover the final step.

SizeTierFirst runHot runvs StandardStepsSamplingDecode, encode, upload
768p Standard 11 min 51 s 11 min 55 s baseline 20 11 min 10 s 45 s
768p Balanced 5 min 35 s 5 min 30 s 2.2× 8 4 min 46 s 45 s
768p Turbo 3 min 29 s 3 min 0 s 4.0× 4 2 min 12 s 48 s
768p FastH3 Turbo 2 min 11 s 2 min 6 s 5.7× 4 1 min 19 s 46 s
480p Standard 3 min 43 s 3 min 39 s baseline 20 3 min 10 s 29 s
480p Balanced 1 min 44 s 1 min 46 s 2.1× 8 1 min 20 s 26 s
480p Turbo 1 min 8 s 1 min 7 s 3.3× 4 40 s 27 s
480p FastH3 Turbo 54 s 55 s 4.0× 4 26 s 29 s

The first run of each configuration is the clip embedded above. It follows a model switch on a worker that was also serving public traffic, so it can carry kernel warm-up. The hot run is the number to plan around.

Warm-up is not a rounding error. A FastH3 480p canary we ran before the matrix, the first FastH3 job after that worker restarted, took 2 min 40 s from job start to finished file. Inside the matrix the same configuration took 54 s. Sparse-attention kernels and the INT8 path pay a one-time cost on a fresh process, which is one more reason a busy worker that keeps FastH3 loaded is the one you want your job to land on.

One honest wrinkle: after the FastH3 pairs, the switch back to the LightX2V checkpoint ran the 32 GB card out of memory on the first Turbo 768p attempt. The worker recovered by unloading everything, but it then reloaded the transformer partly offloaded to system memory, which would have made every later number meaningless. We restarted the worker and re-ran the remaining six configurations from a clean process, so the first Turbo run below followed a cold start.

What each clip costs

FastH3 is the cheapest tier at both sizes as well as the fastest. Retail pay-as-you-go rates before plan discounts, with the exact quote the socket returned for these 15.08-second clips.

TierRate at 768p / 480p15 s at 768p15 s at 480p
Standard16 / 10 Spark per second241 Spark ($1.21)151 Spark ($0.75)
Balanced10 / 6 Spark per second151 Spark ($0.75)90 Spark ($0.45)
Turbo6 / 4 Spark per second90 Spark ($0.45)60 Spark ($0.30)
FastH3 Turbo4 / 4 Spark per second60 Spark ($0.30)60 Spark ($0.30)

Sogni production pricing as of September 2, 2026. Unlimited and Unlimited Pro plans cover eligible H3 jobs without spending Spark.

What to look for

Sound on, and watch the same three beats in every clip: the whispered opening, the cut to the overhead shot at 11.5 seconds, and the meow before the sink. Every tier delivered all three from the same still and seed, which is the headline for anyone who has used earlier four-step distillations. FastH3 did not lose the cat, the shot, or the script.

The cat's own line oversells one number. He says Turbo got two times faster; Sogni's model page says up to 2×. On this card with a full 15-second clip we measured FastH3 at 1.4× the LightX2V Turbo at 768p and a smaller margin at 480p, because a large share of any four-step render is decode and upload that no distillation touches. The 5.7× figure against Standard is the one that changes how you work.

What we cannot score from frames is lip sync and voice, so we will not pretend to. Play the 768p Standard and FastH3 clips back to back with audio and decide whether the difference is worth 11 min 55 s of your GPU's time against 2 min 6 s.

Verdict

For a full 15-second, 768p, image-to-video clip with dialogue on a consumer RTX 5090, FastH3 Turbo is 5.7× faster than Standard and 1.4× faster than the LightX2V Turbo it replaces as Sogni Web's default, and it is the cheapest tier at either size. It kept the cat, the cut, and the script. What it gives up is the film grain and some of the subtlety that twenty steps buy, so Standard stays the right call for a final master where texture matters. For everything before that, drafts, timing passes, script iterations, and most social clips, FastH3 is now where we start.

Try FastH3 Turbo on Sogni Run a worker on the Supernet

Sources

More from Sogni