/ Blogs & Newsletters

New in MiniMax H3 · Keyframes at any frame · September 27, 2026

Sogni launches multi-keyframe MiniMax generations

Until today, MiniMax H3 on Sogni could only be anchored at two moments: the first frame and the last. Now you can pin up to eight keyframes at any frame in between — exact moments you choose — on Image to Video, Sound to Video and Reference to Video. H3 cuts to each keyframe on its frame, and with Sound to Video the character keeps lip-syncing your audio straight through every cut.

FastH3 Sound to Video · starting picture + audio + 2 keyframes 6.58 s @ 24 fps sound on ↗

That is one generation. We gave H3 a portrait, six seconds of Bhad Bhabie’s “Gucci Flip Flops”, and two more stills of the same woman: a close-up and a three-quarter angle. We pinned the close-up at 2.88 seconds and the three-quarter angle at 4.25 seconds. H3 wrote everything else — the opening shot, both cuts, every mouth shape — and landed on each still exactly on its frame.

What’s new: keyframes anywhere on the timeline

Before: the first frame and the last. Now: any frame.

BeforeNow
Where you can pin an imagefirst frame, last framefirst, last, and any frame in between
Keyframes between themnoneup to 8, at exact times you set
Image to Video and Sound to Videostart (and end) on your images; H3 invents every shot betweenthe shots between look like the images you pin
Reference to Videoreferences only, no pinned framesreferences plus up to 8 keyframes
Lip sync to your audioyesyes, straight through every keyframe and cut

A keyframe is an image pinned to an exact moment. You decide what the scene looks like at that frame; H3 builds the video to arrive on it and generates the motion — and the sound — in between. On Sogni they run on MiniMax H3, an open-weights model that renders picture and sound together and can lip-sync your own audio across the keyframes.

01 · A talking head that cuts like an edit

Here is what went in. The audio is two lines of the verse, beat and all; Sound to Video passes it through untouched and makes the picture perform it.

First frame: blonde woman in a beige open-knit sweater holding her ponytail above her head, white studio
Starting picture · 0.00 s
Keyframe: the same woman, closer, facing the camera with her hands on her ponytail
Keyframe · 2.88 s · closer, in a pause
Keyframe: the same woman seen from a three-quarter angle, hands in her lap, mid-word
Keyframe · 4.25 s · new angle, mid-word

Uploaded audio · “This a big watch, diamond drippin’ off of the clock / Pull the 6 out, wintertime, droppin’ the top” (Bhad Bhabie, “Gucci Flip Flops”)

With keyframes 2 keyframes pinned
Without keyframes same prompt

Both takes follow the same written shot list and both lip-sync the verse. The difference is who chose the shots. On the right, H3 invents its own close-up and its own second angle. On the left, the close-up is the one we pinned, the cut to the three-quarter angle lands on frame 102 exactly, and her mouth keeps moving with the verse through both pinned frames — no freeze, no pop.

02 · Audio only: keyframes keep your character

Sound to Video can also start from audio alone, with no picture at all. That is exactly where keyframes earn their keep: without them, nothing tells H3 who is performing.

Audio + 2 keyframes
Audio only no keyframes

Same audio, same words, same shots. On the right H3 casts someone new — a different woman, with heavier makeup, another sweater and a black chair. On the left the two keyframes pin her in: the woman from our stills, rapping the verse, cutting from the close-up to the three-quarter angle on time. Two stills turned “a voice” into “this person, rapping this”.

03 · First frame, last frame, and everything between

With first and last frames plus audio, keyframes go in the middle. Here the clip starts on the portrait, passes through both keyframes, and ends on a fourth still: the three-quarter angle again, her lips closed as the verse ends.

First frame
First frame · 0.00 s
Keyframe at 2.88 s
Keyframe · 2.88 s
Keyframe at 4.25 s
Keyframe · 4.25 s
Last frame: the three-quarter angle with her lips closed in a small smile
Last frame · 6.54 s
FastH3 first + last + audio + 2 keyframes

Keyframes also work like a storyboard. Pin up to eight stills in one clip — one per shot, each at the moment its shot should begin — and describe every shot in the prompt. On Reference to Video they sit alongside your reference images.

Directing with keyframes

What we learned making these clips.

Cuts to the close-up one seed
Pushes in to it another seed

How to use it

In Sogni: open MiniMax H3 and choose Image to Video, Sound to Video, or Reference to Video. Under Keyframes, press Add keyframe, pick a still, and set its time; you can add up to eight. Then describe each keyframe’s shot in your prompt, or let the Screenwriter draft it for you.

With the SDK: pass keyframes as a list of stills and the frame each one pins (24 fps, so frameIndex = seconds × 24), and pass an exact frames count from the H3 grid so every time lands where you put it.

const project = await sogni.projects.create({
  type: 'video',
  network: 'fast',
  modelId: 'minimax-h3-fastvideo-int8_ia2v_turbo', // Sound to Video
  positivePrompt: prompt,          // describe each keyframe's shot
  referenceImage: portrait,        // first frame
  referenceAudio: audio,           // drives the performance
  keyframes: [
    { image: closeUp, frameIndex: 69 },       // 2.88 s
    { image: threeQuarter, frameIndex: 102 }  // 4.25 s, new angle
  ],
  frames: 158,                     // 6.58 s on the H3 grid (124 + n×17)
  width: 768,
  height: 1120,
  steps: 4,
  guidance: 1,
  tokenType: 'spark'
});

Speed, measured

One RTX 5090 Sogni worker, models already loaded.

WorkflowKeyframesRendervs none
Sound to Video, 6.6 s, 768×1120056.3 s—
255.7 ssame
Sound to Video, 8 s, 1344×768063.8 s—
163.3 ssame
268.0 s+7%
8106.0 s+66%
Reference to Video Turbo, 8 s, 960×544049.6 s—
150.4 s+2%
253.7 s+8%
857.3 s+16%

One or two keyframes cost next to nothing in render time. A full eight adds about two thirds to a FastH3 clip and about a sixth on Reference to Video.

Pricing: your first two keyframes are included. Each keyframe after the second adds a sliver of output time at your job’s normal per-second price — 0.75 seconds on FastH3 and 0.3 seconds on every other tier, the same way we already bill extra inputs like reference video. On an 8-second FastH3 clip that is 3 Spark per extra keyframe: one or two keyframes cost nothing extra, and a full eight take the clip from 32 to 50 Spark.

Available now. Keyframes are live on MiniMax H3 Image to Video, Sound to Video and Reference to Video in Sogni. Explore MiniMax H3 and start creating.

Keyframes in MiniMax H3: quick answers

What are keyframes in MiniMax H3?

A keyframe is an image you pin to an exact moment of the video. MiniMax H3 builds the video to arrive on each keyframe at its frame and generates the motion and sound in between.

Where can I put keyframes?

At any frame between the first and the last frame, not just at the start and end. You can pin up to 8 keyframes per video, in addition to a first frame and, on first/last-frame modes, a last frame.

Which MiniMax H3 modes support keyframes?

Every MiniMax H3 video mode except text-to-video: Image to Video (first frame, or first and last), Sound to Video (starting picture plus audio, first and last plus audio, or audio only) and Reference to Video.

Does lip sync still work with keyframes?

Yes. On Sound to Video your uploaded audio drives the performance and keyframes pin how the video looks at their times; in our tests the character's lips kept following the audio through every pinned frame and cut.

What do keyframes cost?

Your first two keyframes are included. Each keyframe after the second adds output time at your job's per-second rate: 0.75 seconds on FastH3 and 0.3 seconds on every other tier. An 8-second FastH3 clip costs 32 Spark with up to two keyframes and 50 Spark with eight.

Do keyframes make renders slower?

One or two keyframes add almost nothing. Eight keyframes add about two thirds to a FastH3 render and about a sixth on Reference to Video Turbo.