/ Blogs & Newsletters

Sogni engineering · New model family · August 16, 2026

H3 directs. LTX‑2.5 brings the camera rig.

MiniMax H3 still wears the director's crown: precise shot lists, timestamped cuts, speaker IDs, and reference casting. LTX‑2.5 answers as the production camera system — native 1920×1088 by default, resolutions up to 4K, and 60 fps at 1080p, synchronized 48 kHz stereo audio, clips up to 20 seconds, stronger continuity across edits, and a full video-to-video control suite. The launch lineup is all-distilled: five models, eleven workflows, one 8‑step fast path at roughly a third of the old 30‑step Dev-tier price. Here is every mode, with each example labelled by its actual input workflow.

LTX-2.5 · T2V text-to-video · three shots, two cuts, one prompt 10.04 s @ 24 fps native 1920×1088 unmute for dialogue ↗

Three shots. A wide of her walking the snowy street, a hard cut inside to the bookshop where she thaws out over a mug, then a second hard cut to a firelit macro of that mug with the steam coming off it. This is a text-to-video (T2V) generation: its only input is one prompt. Two edits, three lighting setups, one pass — and the coat, scarf, and snow on her shoulders carry across every cut. That is the T2V headline: not that LTX invented the cut, but that this native-1080p path holds visual continuity across it while generating the audio jointly, and generating fast.

For current pricing, workflow access, and model IDs, see the LTX‑2.5 22B model page. Full parameter reference, mode-by-mode notes and a sample gallery live in the LTX‑2.5 documentation.

What shipped

Five distilled models, eleven generation paths, one prose contract.

ModelWorkflowsSteps
ltx25-22b-int8_t2v_distilledText-to-video, single-shot or multishot8
ltx25-22b-int8_i2v_distilledImage-to-video · first → last frame8
ltx25-22b-int8_a2v_distilledAudio-to-video8
ltx25-22b-int8_ia2v_distilledImage + audio-to-video8
ltx25-22b-int8_v2v_distilledV2V: canny · depth · pose · detailer · inpaint · outpaint8

Mode labels used below: T2V starts from text only; I2V starts from one image; FLF2V starts from first and last images; and V2V transforms an existing video. Every example carries its mode beside the model name.

LTX‑2.5 is LTX's newest open-weights 22B audio-video transformer, and Sogni ships it on the official distilled path: the INT8‑ConvRot checkpoint, a custom Gemma 4 12B text encoder, separate video and audio VAEs, and the official two-stage pipeline with a 2× latent spatial upscaler. Every public launch mode renders at a fixed 8 steps. Defaults are 1920×1088 at 24 fps, and — as on LTX‑2.3 — the same paths accept resolutions up to 4K, and frame rates up to 60 fps at 1080p. Sogni's current workflows run up to 20 seconds, with 48 kHz stereo audio generated jointly with the picture. Upstream 2.5 also introduces native 4K HDR, RAW/EXR workflows, and optional automatic duration; this article stays focused on the five generation paths currently routed on Sogni.

What LTX‑2.5 actually fixes

This is more than a checkpoint refresh, but several improvements need careful wording.

AreaLTX-2.3LTX-2.5What creators should notice
Shot structureDesigned around one continuous shot; cuts could happen off-spec on a favorable seedNative multishot, trained to hold character, environment, lighting, voice, and style across cutsTwo to four connected shots are now an intended workflow, not a lucky exception
Prompt understandingGemma 3 with the enlarged 2.3 text connectorCustom Gemma 4 12B encoder plus an optional prompt enhancerLonger sequences retain more characters, camera instructions, actions, and lighting details
Decode and detailRebuilt 2.3 latent space and VAENew diffusion video decoder with fidelity allocated by scene complexityCleaner faces, texture, typography, product detail, and fast motion with fewer smeared areas
Distilled fast path8-step distilled checkpointStill 8 steps, but trained to retain more full-model quality, prompt adherence, and motion consistencyThe speed story is quality per step and fewer retries — not a lower step count
DurationExplicit frame countOptional upstream duration predictorThe upstream model can infer clip length from the action; Sogni's current public workflows still expose an explicit duration
Audio when underspecifiedOfficial guidance warns that vague prompts can produce generic ambienceNo documented silence-by-omission change; the 2.5 guide still asks for explicit audio and continuity at every cutWrite no music, no score, no soundtrack when you want only dialogue, room tone, and effects
CompatibilityBroad LoRA and IC-LoRA ecosystemMost upstream 2.3 adapters are expected to work, but must be validatedOn Sogni, Voice ID-LoRA, the transition LoRA, and 10Eros remain explicit 2.3 rollback paths

The audio row matters for anyone who fought an unwanted score in LTX‑2.2 or 2.3. The tempting claim is that a prompt with no sound description now comes back silent. That is not what LTX documents. LTX's own 2.3 audio guide says an underspecified prompt can generate generic ambient sound, while the 2.5 prompt guide tells creators to describe ambient sound, music, speech, or singing and to state what continues or changes at every cut. The practical fix is explicit direction: “Natural diegetic sound only. No music, no score, no soundtrack.” If you want an actually silent result, ask for silence directly and verify the final output; omission is not a mute switch.

01 · Text-to-video (T2V): continuity across the cut

MiniMax H3 got to purposeful multi-shot direction first, and it still gives creators the more precise shot contract. The question for LTX‑2.5 is different: can its T2V path keep a character and scene coherent when plain prose asks it to cut? It generates connected shots in one pass: give the cut a numeric timestamp, re-identify your character on every shot, and say what the audio does while the picture changes. Here is the exact text-only caption behind the film at the top of this page, written to the same contract our built-in prompt enhancer targets:

Wide shot follows a woman in a red wool coat and grey scarf walking
through heavy falling snow past glowing shop windows at night, her breath
clouding as she says brightly: "Almost there." At exactly 3.5 seconds,
a hard cut switches to a medium shot of the same woman inside a warm
lamplit bookshop, snow melting on her shoulders as she cradles a steaming
mug and says: "Worth the walk." At exactly 7.0 seconds, a hard cut
switches to an extreme close-up of the mug in her hands beside a crackling
fire, steam curling off the surface and firelight flickering across her
fingers. Falling snow gives way to a low fire and quiet indoor murmur
across the cuts.

Now the measurement. We handed that caption to LTX‑2.5 and to LTX‑2.3 verbatim — same seed, same 1920×1088, same 241 frames, both on the 8‑step distilled path. The only direction either model got about the edits was those two timestamped sentences:

LTX-2.5 · T2V text-only input · cuts at 3.3 s & 6.0 s 1 m 52 s render 91.83 SPARK
LTX-2.3 · T2V same text-only input · cuts at 4.8 s & 7.3 s 2 m 59 s render 76.52 SPARK

Both models cut. That is worth saying plainly, because LTX‑2.3 is usually described as single-shot only: hand it a timestamped caption and it will, on a good seed, give you real edits — here at 4.8 s and 7.3 s against the 3.5 s and 7.0 s we asked for. LTX‑2.5 lands its first cut within two-tenths of the mark.

The difference that matters shows up across the cuts. LTX‑2.5 treats the three shots as one scene: the same red wool coat, the same grey scarf, the same woman, and snow that melts on her shoulders once she is indoors. LTX‑2.3 re-dresses her every time it cuts — the plain red coat becomes a star-patterned jacket in the bookshop and a polka-dotted sleeve at the fire — and it keeps snowing indoors, flakes drifting past the bookshelves. Continuity across an edit is the whole job of a multi-shot model, and that is the gap.

Three caveats worth knowing. First, treat the timestamp as direction rather than a frame-exact edit point: this take cut at 3.3 s and 6.0 s against the 3.5 s and 7.0 s we asked for. Second, caption discipline decides whether the cut happens at all — over-length captions collapse it, so keep one sentence per shot and stay under about 140 words. Third, and most useful: make the two shots genuinely different places. When we asked for a close-up inside the same room, the model gave us a smooth push-in instead of an edit; a scene change with a real change of space and light is what makes it commit to a cut. Iterating a couple of seeds for a timing-critical edit is normal, and the built-in prompt enhancer in Sogni Studio / Pro writes this structure for you.

02 · Image-to-video (I2V): detail that has to survive motion

This is I2V, not T2V: the generation starts from a single original still, never a frame lifted from a video. It is also the hardest test in this article. A macro beauty close-up gives a video model nowhere to hide — skin is the one texture every viewer knows by heart, and fourteen seconds of speech is long enough for a face to drift, smooth over, or fall apart. Same still, same caption, both models:

Original still: close-up portrait of a woman with freckled skin and glossy lips
Source frame — an original still, 1088×1536
LTX-2.5 · I2V 6 m 00 s render 102.72 SPARK 14.04 s unmute for dialogue ↗
LTX-2.3 · I2V 4 m 15 s render 85.60 SPARK 14.04 s

Both takes do the job the caption asked for: she turns from a three-quarter gaze into the lens, delivers the line with the emphasis written into the prompt, and breaks into a real smile on the last beat — with the audio generated in the same pass, mouth shapes tracking the words. Identity, freckle pattern, lip gloss, and the warm beige backdrop all survive fourteen seconds of continuous motion, which is roughly twice as long as anything else in this article.

Getting there took more attempts on LTX‑2.5 than on LTX‑2.3. Working from a different still of the same subject — a tighter macro crop with her lips pressed shut — the first takes animated the head, held the identity and smiled on cue without ever actually speaking. Later takes from that same still did land, so it is not that one frame simply cannot work; it is that the odds shift with the frame you start from. The still above, with the lips already parted, landed on every take we ran. LTX‑2.3 spoke from either still.

Our honest summary for dialogue on 2.5 today: the source frame decides it, and no amount of prompt engineering rescues a bad one. We ran twenty-eight takes across two stills of the same subject. The normally-framed portrait landed clean lip sync on every single take. The extreme macro crop landed roughly one in four — and stayed at one in four through eight different caption strategies, including splitting the dialogue into fragments with acting beats the way LTX‑2.5’s own guide recommends, naming the language and accent, stating the speech is already underway, and adding the literal phrase “perfect lip sync”. None of them moved it reliably. Changing the frame did.

So: audition your source frame, not your prompt. Run two or three takes; if the mouth is not tracking the words, replace the still rather than rewriting the caption. Watch out for the failure that looks like success — a face moving convincingly but out of time with the audio, which reads as fine until you actually listen. LTX‑2.3 was markedly more dependable on identical material, and text-to-video was dependable throughout; if a shot has to talk and has to land, those are the safer paths today.

Open both at full size and the difference is texture. LTX‑2.5 keeps individual pores, the fine grain across her cheek, sharp edges on every freckle, and the wet specular break on her nose and lower lip. LTX‑2.3 renders the same face softer — the grain smooths toward wax, freckle edges bleed into the skin, and the highlights spread. Sampled at two, five, and nine seconds, that gap is consistent rather than a single unlucky frame. This is the clearest example in the post of what the newer decoder and prompt encoder actually buy you.

One instruction still matters disproportionately here: state the camera on every shot. This caption ends with “the camera remains static with a slow, subtle push-in” — leave that out and the model will invent a move of its own and drift off the subject in the final seconds.

03 · First/last-frame video (FLF2V): two anchored images

This is FLF2V: it begins with two image inputs, not a text-only prompt. On LTX‑2.3, two-image interpolation rode along on the I2V path, with ValiantCat's well-loved transition LoRA available when you wanted real control over the middle. LTX‑2.5 has a first-class first/last-frame template in the official pipeline instead. There is no LTX‑2.5 build of that transition LoRA yet, and the 2.3 weights do not currently attach to a 2.5 job on Sogni — the transition LoRA stays an LTX‑2.3 path for now. Both anchors below are original stills: a Krea 2 Turbo alpine summit at sunrise, and the same summit edited to deep night. Nothing in the frame is allowed to move except the light.

First frame: an alpine summit at golden sunrise above a sea of clouds
First frame — 0.00 s · Krea 2 Turbo
Last frame: the same summit at night under the Milky Way
Last frame — 8.04 s · Qwen Image Edit, same composition
LTX-2.5 · FLF2V first + last image inputs 1 m 36 s render 73.54 SPARK
LTX-2.3 · FLF2V I2V keyframe path · same two images 4 m 12 s render 61.28 SPARK

This is the mode where LTX‑2.5's detail is easiest to see. Hold either clip at full resolution and the granite does not move: every fracture line, every wind-carved cornice, every rope of cloud in the valley sits exactly where the stills put it, while the alpenglow drains off the west faces and the Milky Way comes up behind the summit. Nothing melts, nothing swims — which is the hard part, because interpolating between two fixed anchors is precisely where a video model is most tempted to cross-fade.

The two takes differ in pacing rather than fidelity. LTX‑2.5 spends the whole eight seconds on the transition, so the sky darkens continuously and the stars arrive a few at a time. LTX‑2.3 rushes: it reaches full night by about 3.5 seconds and then holds a nearly static frame for the rest of the clip. Same endpoints, same price bracket — one of them is a shot, the other is a slideshow with a long tail.

04 · Text-to-video (T2V): the MiniMax H3 Turbo benchmark

LTX‑2.5 is not the only open-weights model on the network that can cut. MiniMax H3 got there first, and it still cuts more precisely — its shot list takes explicit timestamps and holds speaker IDs across the edit. This is an H3 Turbo T2V generation: the rainy night-market two-hander from our launch post, created from text alone at H3's native 768p class. It remains the bar for scripted multi-shot direction:

H3 Turbo · T2V text-only input · 4 steps · 1344×768 2 m 24 s render 60.75 SPARK 10.13 s

These two families are complements, not substitutes. H3 remains the director's model: explicitly timestamped cuts, stable speaker IDs across shots, and reference-set casting. LTX‑2.5 is the cinematographer's camera: native 1080p by default — trading up to 4K when you want resolution, or 60 fps at 1080p when you want frame rate — against H3's fixed 768p class, clips up to twenty seconds against fifteen, audio-driven animation, and the full video-to-video control suite:

CapabilityLTX-2.3LTX-2.5H3 Turbo
Native audio✓ 48 kHz stereo✓ 32 kHz stereo
Resolution & frame rate1536-class default; 4K max, 60 fps at 1080p1920×1088 default; 4K max, 60 fps at 1080p1344×768 fixed, 24 fps
Multi-shot cuts in one generationoccasionally, off-spec✓ prose "hard cut" (2–3 shots)✓ timestamped [Shot N]
Wardrobe / identity held across cutsre-dresses the subject✓ held in our runs✓ (S1), (S2) speaker IDs
Reference-set casting✓ Standard and Turbo r2v
First → last framekeyframe path + LoRA✓ native template
Audio-to-video / image+audio
V2V control (canny/depth/pose/detailer/in-out-paint)✓ six modes
Max public clip length20 s20 s15.08 s
Steps per render8 / 30 (Dev)84

One more difference matters more than any row in that table, and it is the one you feel on a deadline. H3 usually just works. Give it a shot list and the first generation is usually a usable generation. LTX‑2.5 often takes several attempts to land the perfect shot. Across the multishot seeds behind this post, roughly half produced edits crisp enough to publish; the rest went soft, drifted at the end, or quietly turned a cut into a push-in. Budget a few takes per shot.

What you buy with those extra takes is a look you cannot get out of H3 — 1920×1088 by default and up to 4K when you want it, against a fixed 768p class, twenty-second clips against fifteen, and the grade and texture you can see in the firelit mug, the wood and window light of the luthier’s bench, and the granite on the summit. LTX also carries the larger ecosystem: six V2V control modes, audio-driven animation, and a deep community bench of adapters and workflows around the 2.x line that has no real equivalent on H3 yet.

H3 directs the scene and lands it on the first take. LTX‑2.5 gives you a camera, a grade and a control rig — and asks you to roll a few more times.

05 · Video-to-video (V2V): controls inherit the source

This is V2V: the input is an existing video, not text alone or a still image. Sogni routes six workflows on the same 8-step model: canny, depth, pose, detailer, masked inpainting, and canvas outpainting. Canny and depth preserve edge or spatial structure; pose transfers motion and adds a still image for the subject's appearance; inpaint limits regeneration to a mask; outpaint extends the canvas.

Here is canny control running over a separate LTX‑2.5 I2V clip of a violin maker at his bench — the same eight seconds, re-rendered against its own edge map:

source source I2V clip 8.04 s · 1920×1088
LTX-2.5 · V2V canny 1 m 51 s render 69.14 SPARK

Every edge the control pass was handed survives: the window and its light, the tool rack, the bench line, the outline of the instrument, and the craftsman's pose and timing beat for beat. Everything the caption was free to invent has changed — the carved spruce becomes a moulded white shell, the wood shavings become drifts of pale dust, the warm amber shop cools to grey-green daylight, and the craftsman himself is recast. Same scene, same blocking, different world.

One limit to note from the same pair: canny carries structure, not performance. Freeze both clips at 5.5 seconds and the source is mid-sentence while the restyle's mouth is closed. Treat a control pass as a way to restage a scene, not to re-skin a specific take of dialogue.

Speed & cost, measured

Wall-clock render time: one render each on Sogni Supernet fast-network workers, 2026‑08‑14. Fast-network GPUs differ, so treat each number as one honest sample, not a benchmark — the same 8-second 1080p LTX‑2.5 job clocked 1 m 20 s on one worker and over 3 m on another during our runs, and on some workers an LTX‑2.3 job of identical shape beats a 2.5 job. The structural facts are what hold: both distilled families run 8 steps, the Dev tier runs 30, H3 Turbo runs 4 much heavier ones. The 2.3 Dev row is priced but unrendered — its wall clock varied too widely across workers for one honest sample. The H3 Turbo sample carries over from our launch-post run (2026‑08‑09). Prices are live network estimates; every LTX‑2.5 render is also covered credit-free under fair use on Sogni Unlimited plans.

RenderResolutionWall clockSPARKUSD
2.5 multishot T2V 10 s (3 shots)1920×10881 m 52 s91.83$0.46
2.3 T2V 10 s (same caption)1920×10882 m 59 s76.52$0.38
2.3 Dev T2V 10 s1920×1088267.83$1.34
H3T T2V 10 s1344×7682 m 24 s60.75$0.30
2.5 I2V 14 s (portrait)1088×15366 m 00 s102.72$0.51
2.3 I2V 14 s (portrait)1088×15364 m 15 s85.60$0.43
2.5 FLF2V 8 s1920×10881 m 36 s73.54$0.37
2.3 FLF2V 8 s1920×10884 m 12 s61.28$0.31
2.5 V2V canny 8 s1920×10241 m 51 s69.14$0.35

Notes for prompt writers

Available now. All five LTX‑2.5 models are live on the Sogni Supernet for pay-as-you-go Spark and included credit-free under fair use on every Sogni Unlimited subscription plan. Explore the model and start creating, or read the full LTX‑2.5 documentation.

Sources and further reading

All source stills are original generations made on Sogni — the luthier and summit from Krea 2 Turbo, and the night summit a Qwen Image Edit of that same frame. No extracted video frames were used as image inputs anywhere in this work. Matched A/B pairs share the same caption, seed, resolution, and frame count. Clips are web-ready encodes: LTX comparisons at 1920×1088 and H3 Turbo at 1344×768, carried over from our MiniMax H3 post. Use the ⛶ View at… button on any clip to open it at its encoded size. Rendered on the Sogni Supernet, 2026-08-14/15.

More from Sogni