New in MiniMax H3 · Keyframes at any frame · September 27, 2026
Sogni launches multi-keyframe MiniMax generations
Until today, MiniMax H3 on Sogni could only be anchored at two moments: the first frame and the last. Now you can pin up to eight keyframes at any frame in between — exact moments you choose — on Image to Video, Sound to Video and Reference to Video. H3 cuts to each keyframe on its frame, and with Sound to Video the character keeps lip-syncing your audio straight through every cut.
That is one generation. We gave H3 a portrait, six seconds of Bhad Bhabie’s “Gucci Flip Flops”, and two more stills of the same woman: a close-up and a three-quarter angle. We pinned the close-up at 2.88 seconds and the three-quarter angle at 4.25 seconds. H3 wrote everything else — the opening shot, both cuts, every mouth shape — and landed on each still exactly on its frame.
What’s new: keyframes anywhere on the timeline
Before: the first frame and the last. Now: any frame.
| Before | Now | |
|---|---|---|
| Where you can pin an image | first frame, last frame | first, last, and any frame in between |
| Keyframes between them | none | up to 8, at exact times you set |
| Image to Video and Sound to Video | start (and end) on your images; H3 invents every shot between | the shots between look like the images you pin |
| Reference to Video | references only, no pinned frames | references plus up to 8 keyframes |
| Lip sync to your audio | yes | yes, straight through every keyframe and cut |
A keyframe is an image pinned to an exact moment. You decide what the scene looks like at that frame; H3 builds the video to arrive on it and generates the motion — and the sound — in between. On Sogni they run on MiniMax H3, an open-weights model that renders picture and sound together and can lip-sync your own audio across the keyframes.
- Up to 8 keyframes per video, at any frame between the first and the last, alongside your first frame (and last frame, on first/last modes).
- On every H3 video mode except text-to-video: Image to Video (first frame, or first and last), Sound to Video (starting picture + audio, first and last + audio, or audio only), and Reference to Video.
- Your prompt still tells the story. H3 never reads a keyframe as a reference image, so the words for each shot should describe what its keyframe shows.
01 · A talking head that cuts like an edit
Here is what went in. The audio is two lines of the verse, beat and all; Sound to Video passes it through untouched and makes the picture perform it.



Uploaded audio · “This a big watch, diamond drippin’ off of the clock / Pull the 6 out, wintertime, droppin’ the top” (Bhad Bhabie, “Gucci Flip Flops”)
Both takes follow the same written shot list and both lip-sync the verse. The difference is who chose the shots. On the right, H3 invents its own close-up and its own second angle. On the left, the close-up is the one we pinned, the cut to the three-quarter angle lands on frame 102 exactly, and her mouth keeps moving with the verse through both pinned frames — no freeze, no pop.
02 · Audio only: keyframes keep your character
Sound to Video can also start from audio alone, with no picture at all. That is exactly where keyframes earn their keep: without them, nothing tells H3 who is performing.
Same audio, same words, same shots. On the right H3 casts someone new — a different woman, with heavier makeup, another sweater and a black chair. On the left the two keyframes pin her in: the woman from our stills, rapping the verse, cutting from the close-up to the three-quarter angle on time. Two stills turned “a voice” into “this person, rapping this”.
03 · First frame, last frame, and everything between
With first and last frames plus audio, keyframes go in the middle. Here the clip starts on the portrait, passes through both keyframes, and ends on a fourth still: the three-quarter angle again, her lips closed as the verse ends.




Keyframes also work like a storyboard. Pin up to eight stills in one clip — one per shot, each at the moment its shot should begin — and describe every shot in the prompt. On Reference to Video they sit alongside your reference images.
Directing with keyframes
What we learned making these clips.
- Describe every keyframe in its shot. H3 uses the still to pin the frame, but reads the story from your words. If the words for a shot describe something different from its still, H3 can flash the still for a single frame.
- Start a new shot at each keyframe that changes the angle. In the clips above, the three-quarter angle cuts exactly on its frame.
- A keyframe that is only a closer view of the same angle may arrive as a push-in. H3 sometimes connects a same-angle close-up with a smooth camera move instead of a cut. On one seed it did so every time, whatever we wrote; on another it cut cleanly. Both are shown below. Want a guaranteed cut? Change the angle.
- Two differently framed or lit stills in one continuous shot will cross-fade. Give them separate shots.
- On Sound to Video, the audio drives the performance and keyframes decide how it looks. Put a keyframe in a pause if you want a closed mouth on that frame; mid-word works too, as the three-quarter angle above shows.
How to use it
In Sogni: open MiniMax H3 and choose Image to Video, Sound to Video, or Reference to Video. Under Keyframes, press Add keyframe, pick a still, and set its time; you can add up to eight. Then describe each keyframe’s shot in your prompt, or let the Screenwriter draft it for you.
With the SDK: pass keyframes as a list of stills and the frame each one pins (24 fps, so frameIndex = seconds × 24), and pass an exact frames count from the H3 grid so every time lands where you put it.
const project = await sogni.projects.create({
type: 'video',
network: 'fast',
modelId: 'minimax-h3-fastvideo-int8_ia2v_turbo', // Sound to Video
positivePrompt: prompt, // describe each keyframe's shot
referenceImage: portrait, // first frame
referenceAudio: audio, // drives the performance
keyframes: [
{ image: closeUp, frameIndex: 69 }, // 2.88 s
{ image: threeQuarter, frameIndex: 102 } // 4.25 s, new angle
],
frames: 158, // 6.58 s on the H3 grid (124 + n×17)
width: 768,
height: 1120,
steps: 4,
guidance: 1,
tokenType: 'spark'
});
Speed, measured
One RTX 5090 Sogni worker, models already loaded.
| Workflow | Keyframes | Render | vs none |
|---|---|---|---|
| Sound to Video, 6.6 s, 768×1120 | 0 | 56.3 s | — |
| 2 | 55.7 s | same | |
| Sound to Video, 8 s, 1344×768 | 0 | 63.8 s | — |
| 1 | 63.3 s | same | |
| 2 | 68.0 s | +7% | |
| 8 | 106.0 s | +66% | |
| Reference to Video Turbo, 8 s, 960×544 | 0 | 49.6 s | — |
| 1 | 50.4 s | +2% | |
| 2 | 53.7 s | +8% | |
| 8 | 57.3 s | +16% |
One or two keyframes cost next to nothing in render time. A full eight adds about two thirds to a FastH3 clip and about a sixth on Reference to Video.
Pricing: your first two keyframes are included. Each keyframe after the second adds a sliver of output time at your job’s normal per-second price — 0.75 seconds on FastH3 and 0.3 seconds on every other tier, the same way we already bill extra inputs like reference video. On an 8-second FastH3 clip that is 3 Spark per extra keyframe: one or two keyframes cost nothing extra, and a full eight take the clip from 32 to 50 Spark.
Available now. Keyframes are live on MiniMax H3 Image to Video, Sound to Video and Reference to Video in Sogni. Explore MiniMax H3 and start creating.
Keyframes in MiniMax H3: quick answers
What are keyframes in MiniMax H3?
A keyframe is an image you pin to an exact moment of the video. MiniMax H3 builds the video to arrive on each keyframe at its frame and generates the motion and sound in between.
Where can I put keyframes?
At any frame between the first and the last frame, not just at the start and end. You can pin up to 8 keyframes per video, in addition to a first frame and, on first/last-frame modes, a last frame.
Which MiniMax H3 modes support keyframes?
Every MiniMax H3 video mode except text-to-video: Image to Video (first frame, or first and last), Sound to Video (starting picture plus audio, first and last plus audio, or audio only) and Reference to Video.
Does lip sync still work with keyframes?
Yes. On Sound to Video your uploaded audio drives the performance and keyframes pin how the video looks at their times; in our tests the character's lips kept following the audio through every pinned frame and cut.
What do keyframes cost?
Your first two keyframes are included. Each keyframe after the second adds output time at your job's per-second rate: 0.75 seconds on FastH3 and 0.3 seconds on every other tier. An 8-second FastH3 clip costs 32 Spark with up to two keyframes and 50 Spark with eight.
Do keyframes make renders slower?
One or two keyframes add almost nothing. Eight keyframes add about two thirds to a FastH3 render and about a sixth on Reference to Video Turbo.
