Sogni engineering · New model family · August 16, 2026
H3 directs. LTX‑2.5 brings the camera rig.
MiniMax H3 still wears the director's crown: precise shot lists, timestamped cuts, speaker IDs, and reference casting. LTX‑2.5 answers as the production camera system — native 1920×1088 by default, resolutions up to 4K, and 60 fps at 1080p, synchronized 48 kHz stereo audio, clips up to 20 seconds, stronger continuity across edits, and a full video-to-video control suite. The launch lineup is all-distilled: five models, eleven workflows, one 8‑step fast path at roughly a third of the old 30‑step Dev-tier price. Here is every mode, with each example labelled by its actual input workflow.
Three shots. A wide of her walking the snowy street, a hard cut inside to the bookshop where she thaws out over a mug, then a second hard cut to a firelit macro of that mug with the steam coming off it. This is a text-to-video (T2V) generation: its only input is one prompt. Two edits, three lighting setups, one pass — and the coat, scarf, and snow on her shoulders carry across every cut. That is the T2V headline: not that LTX invented the cut, but that this native-1080p path holds visual continuity across it while generating the audio jointly, and generating fast.
For current pricing, workflow access, and model IDs, see the LTX‑2.5 22B model page. Full parameter reference, mode-by-mode notes and a sample gallery live in the LTX‑2.5 documentation.
What shipped
Five distilled models, eleven generation paths, one prose contract.
| Model | Workflows | Steps |
|---|---|---|
ltx25-22b-int8_t2v_distilled | Text-to-video, single-shot or multishot | 8 |
ltx25-22b-int8_i2v_distilled | Image-to-video · first → last frame | 8 |
ltx25-22b-int8_a2v_distilled | Audio-to-video | 8 |
ltx25-22b-int8_ia2v_distilled | Image + audio-to-video | 8 |
ltx25-22b-int8_v2v_distilled | V2V: canny · depth · pose · detailer · inpaint · outpaint | 8 |
Mode labels used below: T2V starts from text only; I2V starts from one image; FLF2V starts from first and last images; and V2V transforms an existing video. Every example carries its mode beside the model name.
LTX‑2.5 is LTX's newest open-weights 22B audio-video transformer, and Sogni ships it on the official distilled path: the INT8‑ConvRot checkpoint, a custom Gemma 4 12B text encoder, separate video and audio VAEs, and the official two-stage pipeline with a 2× latent spatial upscaler. Every public launch mode renders at a fixed 8 steps. Defaults are 1920×1088 at 24 fps, and — as on LTX‑2.3 — the same paths accept resolutions up to 4K, and frame rates up to 60 fps at 1080p. Sogni's current workflows run up to 20 seconds, with 48 kHz stereo audio generated jointly with the picture. Upstream 2.5 also introduces native 4K HDR, RAW/EXR workflows, and optional automatic duration; this article stays focused on the five generation paths currently routed on Sogni.
What LTX‑2.5 actually fixes
This is more than a checkpoint refresh, but several improvements need careful wording.
| Area | LTX-2.3 | LTX-2.5 | What creators should notice |
|---|---|---|---|
| Shot structure | Designed around one continuous shot; cuts could happen off-spec on a favorable seed | Native multishot, trained to hold character, environment, lighting, voice, and style across cuts | Two to four connected shots are now an intended workflow, not a lucky exception |
| Prompt understanding | Gemma 3 with the enlarged 2.3 text connector | Custom Gemma 4 12B encoder plus an optional prompt enhancer | Longer sequences retain more characters, camera instructions, actions, and lighting details |
| Decode and detail | Rebuilt 2.3 latent space and VAE | New diffusion video decoder with fidelity allocated by scene complexity | Cleaner faces, texture, typography, product detail, and fast motion with fewer smeared areas |
| Distilled fast path | 8-step distilled checkpoint | Still 8 steps, but trained to retain more full-model quality, prompt adherence, and motion consistency | The speed story is quality per step and fewer retries — not a lower step count |
| Duration | Explicit frame count | Optional upstream duration predictor | The upstream model can infer clip length from the action; Sogni's current public workflows still expose an explicit duration |
| Audio when underspecified | Official guidance warns that vague prompts can produce generic ambience | No documented silence-by-omission change; the 2.5 guide still asks for explicit audio and continuity at every cut | Write no music, no score, no soundtrack when you want only dialogue, room tone, and effects |
| Compatibility | Broad LoRA and IC-LoRA ecosystem | Most upstream 2.3 adapters are expected to work, but must be validated | On Sogni, Voice ID-LoRA, the transition LoRA, and 10Eros remain explicit 2.3 rollback paths |
The audio row matters for anyone who fought an unwanted score in LTX‑2.2 or 2.3. The tempting claim is that a prompt with no sound description now comes back silent. That is not what LTX documents. LTX's own 2.3 audio guide says an underspecified prompt can generate generic ambient sound, while the 2.5 prompt guide tells creators to describe ambient sound, music, speech, or singing and to state what continues or changes at every cut. The practical fix is explicit direction: “Natural diegetic sound only. No music, no score, no soundtrack.” If you want an actually silent result, ask for silence directly and verify the final output; omission is not a mute switch.
01 · Text-to-video (T2V): continuity across the cut
MiniMax H3 got to purposeful multi-shot direction first, and it still gives creators the more precise shot contract. The question for LTX‑2.5 is different: can its T2V path keep a character and scene coherent when plain prose asks it to cut? It generates connected shots in one pass: give the cut a numeric timestamp, re-identify your character on every shot, and say what the audio does while the picture changes. Here is the exact text-only caption behind the film at the top of this page, written to the same contract our built-in prompt enhancer targets:
Wide shot follows a woman in a red wool coat and grey scarf walking
through heavy falling snow past glowing shop windows at night, her breath
clouding as she says brightly: "Almost there." At exactly 3.5 seconds,
a hard cut switches to a medium shot of the same woman inside a warm
lamplit bookshop, snow melting on her shoulders as she cradles a steaming
mug and says: "Worth the walk." At exactly 7.0 seconds, a hard cut
switches to an extreme close-up of the mug in her hands beside a crackling
fire, steam curling off the surface and firelight flickering across her
fingers. Falling snow gives way to a low fire and quiet indoor murmur
across the cuts.
Now the measurement. We handed that caption to LTX‑2.5 and to LTX‑2.3 verbatim — same seed, same 1920×1088, same 241 frames, both on the 8‑step distilled path. The only direction either model got about the edits was those two timestamped sentences:
Both models cut. That is worth saying plainly, because LTX‑2.3 is usually described as single-shot only: hand it a timestamped caption and it will, on a good seed, give you real edits — here at 4.8 s and 7.3 s against the 3.5 s and 7.0 s we asked for. LTX‑2.5 lands its first cut within two-tenths of the mark.
The difference that matters shows up across the cuts. LTX‑2.5 treats the three shots as one scene: the same red wool coat, the same grey scarf, the same woman, and snow that melts on her shoulders once she is indoors. LTX‑2.3 re-dresses her every time it cuts — the plain red coat becomes a star-patterned jacket in the bookshop and a polka-dotted sleeve at the fire — and it keeps snowing indoors, flakes drifting past the bookshelves. Continuity across an edit is the whole job of a multi-shot model, and that is the gap.
Three caveats worth knowing. First, treat the timestamp as direction rather than a frame-exact edit point: this take cut at 3.3 s and 6.0 s against the 3.5 s and 7.0 s we asked for. Second, caption discipline decides whether the cut happens at all — over-length captions collapse it, so keep one sentence per shot and stay under about 140 words. Third, and most useful: make the two shots genuinely different places. When we asked for a close-up inside the same room, the model gave us a smooth push-in instead of an edit; a scene change with a real change of space and light is what makes it commit to a cut. Iterating a couple of seeds for a timing-critical edit is normal, and the built-in prompt enhancer in Sogni Studio / Pro writes this structure for you.
02 · Image-to-video (I2V): detail that has to survive motion
This is I2V, not T2V: the generation starts from a single original still, never a frame lifted from a video. It is also the hardest test in this article. A macro beauty close-up gives a video model nowhere to hide — skin is the one texture every viewer knows by heart, and fourteen seconds of speech is long enough for a face to drift, smooth over, or fall apart. Same still, same caption, both models:
Both takes do the job the caption asked for: she turns from a three-quarter gaze into the lens, delivers the line with the emphasis written into the prompt, and breaks into a real smile on the last beat — with the audio generated in the same pass, mouth shapes tracking the words. Identity, freckle pattern, lip gloss, and the warm beige backdrop all survive fourteen seconds of continuous motion, which is roughly twice as long as anything else in this article.
Getting there took more attempts on LTX‑2.5 than on LTX‑2.3. Working from a different still of the same subject — a tighter macro crop with her lips pressed shut — the first takes animated the head, held the identity and smiled on cue without ever actually speaking. Later takes from that same still did land, so it is not that one frame simply cannot work; it is that the odds shift with the frame you start from. The still above, with the lips already parted, landed on every take we ran. LTX‑2.3 spoke from either still.
Our honest summary for dialogue on 2.5 today: the source frame decides it, and no amount of prompt engineering rescues a bad one. We ran twenty-eight takes across two stills of the same subject. The normally-framed portrait landed clean lip sync on every single take. The extreme macro crop landed roughly one in four — and stayed at one in four through eight different caption strategies, including splitting the dialogue into fragments with acting beats the way LTX‑2.5’s own guide recommends, naming the language and accent, stating the speech is already underway, and adding the literal phrase “perfect lip sync”. None of them moved it reliably. Changing the frame did.
So: audition your source frame, not your prompt. Run two or three takes; if the mouth is not tracking the words, replace the still rather than rewriting the caption. Watch out for the failure that looks like success — a face moving convincingly but out of time with the audio, which reads as fine until you actually listen. LTX‑2.3 was markedly more dependable on identical material, and text-to-video was dependable throughout; if a shot has to talk and has to land, those are the safer paths today.
Open both at full size and the difference is texture. LTX‑2.5 keeps individual pores, the fine grain across her cheek, sharp edges on every freckle, and the wet specular break on her nose and lower lip. LTX‑2.3 renders the same face softer — the grain smooths toward wax, freckle edges bleed into the skin, and the highlights spread. Sampled at two, five, and nine seconds, that gap is consistent rather than a single unlucky frame. This is the clearest example in the post of what the newer decoder and prompt encoder actually buy you.
One instruction still matters disproportionately here: state the camera on every shot. This caption ends with “the camera remains static with a slow, subtle push-in” — leave that out and the model will invent a move of its own and drift off the subject in the final seconds.
03 · First/last-frame video (FLF2V): two anchored images
This is FLF2V: it begins with two image inputs, not a text-only prompt. On LTX‑2.3, two-image interpolation rode along on the I2V path, with ValiantCat's well-loved transition LoRA available when you wanted real control over the middle. LTX‑2.5 has a first-class first/last-frame template in the official pipeline instead. There is no LTX‑2.5 build of that transition LoRA yet, and the 2.3 weights do not currently attach to a 2.5 job on Sogni — the transition LoRA stays an LTX‑2.3 path for now. Both anchors below are original stills: a Krea 2 Turbo alpine summit at sunrise, and the same summit edited to deep night. Nothing in the frame is allowed to move except the light.


This is the mode where LTX‑2.5's detail is easiest to see. Hold either clip at full resolution and the granite does not move: every fracture line, every wind-carved cornice, every rope of cloud in the valley sits exactly where the stills put it, while the alpenglow drains off the west faces and the Milky Way comes up behind the summit. Nothing melts, nothing swims — which is the hard part, because interpolating between two fixed anchors is precisely where a video model is most tempted to cross-fade.
The two takes differ in pacing rather than fidelity. LTX‑2.5 spends the whole eight seconds on the transition, so the sky darkens continuously and the stars arrive a few at a time. LTX‑2.3 rushes: it reaches full night by about 3.5 seconds and then holds a nearly static frame for the rest of the clip. Same endpoints, same price bracket — one of them is a shot, the other is a slideshow with a long tail.
04 · Text-to-video (T2V): the MiniMax H3 Turbo benchmark
LTX‑2.5 is not the only open-weights model on the network that can cut. MiniMax H3 got there first, and it still cuts more precisely — its shot list takes explicit timestamps and holds speaker IDs across the edit. This is an H3 Turbo T2V generation: the rainy night-market two-hander from our launch post, created from text alone at H3's native 768p class. It remains the bar for scripted multi-shot direction:
These two families are complements, not substitutes. H3 remains the director's model: explicitly timestamped cuts, stable speaker IDs across shots, and reference-set casting. LTX‑2.5 is the cinematographer's camera: native 1080p by default — trading up to 4K when you want resolution, or 60 fps at 1080p when you want frame rate — against H3's fixed 768p class, clips up to twenty seconds against fifteen, audio-driven animation, and the full video-to-video control suite:
| Capability | LTX-2.3 | LTX-2.5 | H3 Turbo |
|---|---|---|---|
| Native audio | ✓ | ✓ 48 kHz stereo | ✓ 32 kHz stereo |
| Resolution & frame rate | 1536-class default; 4K max, 60 fps at 1080p | 1920×1088 default; 4K max, 60 fps at 1080p | 1344×768 fixed, 24 fps |
| Multi-shot cuts in one generation | occasionally, off-spec | ✓ prose "hard cut" (2–3 shots) | ✓ timestamped [Shot N] |
| Wardrobe / identity held across cuts | re-dresses the subject | ✓ held in our runs | ✓ (S1), (S2) speaker IDs |
| Reference-set casting | — | — | ✓ Standard and Turbo r2v |
| First → last frame | keyframe path + LoRA | ✓ native template | ✓ |
| Audio-to-video / image+audio | ✓ | ✓ | — |
| V2V control (canny/depth/pose/detailer/in-out-paint) | ✓ | ✓ six modes | — |
| Max public clip length | 20 s | 20 s | 15.08 s |
| Steps per render | 8 / 30 (Dev) | 8 | 4 |
One more difference matters more than any row in that table, and it is the one you feel on a deadline. H3 usually just works. Give it a shot list and the first generation is usually a usable generation. LTX‑2.5 often takes several attempts to land the perfect shot. Across the multishot seeds behind this post, roughly half produced edits crisp enough to publish; the rest went soft, drifted at the end, or quietly turned a cut into a push-in. Budget a few takes per shot.
What you buy with those extra takes is a look you cannot get out of H3 — 1920×1088 by default and up to 4K when you want it, against a fixed 768p class, twenty-second clips against fifteen, and the grade and texture you can see in the firelit mug, the wood and window light of the luthier’s bench, and the granite on the summit. LTX also carries the larger ecosystem: six V2V control modes, audio-driven animation, and a deep community bench of adapters and workflows around the 2.x line that has no real equivalent on H3 yet.
H3 directs the scene and lands it on the first take. LTX‑2.5 gives you a camera, a grade and a control rig — and asks you to roll a few more times.
05 · Video-to-video (V2V): controls inherit the source
This is V2V: the input is an existing video, not text alone or a still image. Sogni routes six workflows on the same 8-step model: canny, depth, pose, detailer, masked inpainting, and canvas outpainting. Canny and depth preserve edge or spatial structure; pose transfers motion and adds a still image for the subject's appearance; inpaint limits regeneration to a mask; outpaint extends the canvas.
Here is canny control running over a separate LTX‑2.5 I2V clip of a violin maker at his bench — the same eight seconds, re-rendered against its own edge map:
Every edge the control pass was handed survives: the window and its light, the tool rack, the bench line, the outline of the instrument, and the craftsman's pose and timing beat for beat. Everything the caption was free to invent has changed — the carved spruce becomes a moulded white shell, the wood shavings become drifts of pale dust, the warm amber shop cools to grey-green daylight, and the craftsman himself is recast. Same scene, same blocking, different world.
One limit to note from the same pair: canny carries structure, not performance. Freeze both clips at 5.5 seconds and the source is mid-sentence while the restyle's mouth is closed. Treat a control pass as a way to restage a scene, not to re-skin a specific take of dialogue.
Speed & cost, measured
Wall-clock render time: one render each on Sogni Supernet fast-network workers, 2026‑08‑14. Fast-network GPUs differ, so treat each number as one honest sample, not a benchmark — the same 8-second 1080p LTX‑2.5 job clocked 1 m 20 s on one worker and over 3 m on another during our runs, and on some workers an LTX‑2.3 job of identical shape beats a 2.5 job. The structural facts are what hold: both distilled families run 8 steps, the Dev tier runs 30, H3 Turbo runs 4 much heavier ones. The 2.3 Dev row is priced but unrendered — its wall clock varied too widely across workers for one honest sample. The H3 Turbo sample carries over from our launch-post run (2026‑08‑09). Prices are live network estimates; every LTX‑2.5 render is also covered credit-free under fair use on Sogni Unlimited plans.
| Render | Resolution | Wall clock | SPARK | USD |
|---|---|---|---|---|
| 2.5 multishot T2V 10 s (3 shots) | 1920×1088 | 1 m 52 s | 91.83 | $0.46 |
| 2.3 T2V 10 s (same caption) | 1920×1088 | 2 m 59 s | 76.52 | $0.38 |
| 2.3 Dev T2V 10 s | 1920×1088 | — | 267.83 | $1.34 |
| H3T T2V 10 s | 1344×768 | 2 m 24 s | 60.75 | $0.30 |
| 2.5 I2V 14 s (portrait) | 1088×1536 | 6 m 00 s | 102.72 | $0.51 |
| 2.3 I2V 14 s (portrait) | 1088×1536 | 4 m 15 s | 85.60 | $0.43 |
| 2.5 FLF2V 8 s | 1920×1088 | 1 m 36 s | 73.54 | $0.37 |
| 2.3 FLF2V 8 s | 1920×1088 | 4 m 12 s | 61.28 | $0.31 |
| 2.5 V2V canny 8 s | 1920×1024 | 1 m 51 s | 69.14 | $0.35 |
Notes for prompt writers
- LTX wants one flowing prose paragraph — shot scale, scene, character, action, camera, and audio in present tense. Do not rely on
[Shot N]markup alone: name every cut in natural language. The release workflow supplies its own negative prompt, while creative exclusions such as no music still belong in your positive direction. - Write dialogue verbatim in quotes with the speaker and delivery: she says in a calm, seasoned voice: "…". Budget about 3 spoken words per second.
- Omitting sound design is not a mute switch. LTX‑2.3's official audio guide warns that underspecified prompts can produce generic ambience, and 2.5 still expects explicit audio direction. To avoid the unwanted score that often annoyed 2.2/2.3 users, write: “Natural diegetic sound only. No music, no score, no soundtrack.” State actual silence explicitly when that is what you need.
- For multishot, give every shot one sentence with an explicit shot type, timestamp each cut ("At exactly 3.5 seconds, a hard cut switches to…"), re-identify the character on every shot, and say what the audio does across the cut. Two things decide whether you get a cut at all: total caption length (past roughly 140 words the edits collapse) and how different the two shots are — ask for a close-up in the same room and the model will give you a smooth push-in instead of an edit, while a real change of place and light makes it commit. Expect cuts to land within about a second of the timestamp rather than on the frame.
- Durations sit on a frame grid of
1 + 8nframes at 24 fps; Sogni's current public workflows expose clips up to 20 seconds. Upstream 2.5 also ships an optional duration predictor, but an explicit duration is easier to compare and budget. - On I2V,
startingImageStrengthtrades first-frame fidelity against motion freedom — higher gives the model more licence to depart from your still. The FLF2V template exposes an equivalent control per anchor frame. - For I2V dialogue, audition the source frame rather than the prompt. Across twenty-eight takes, one still landed clean lip sync every time and another landed about one in four — and stayed there through eight caption strategies, including LTX‑2.5’s own documented dialogue structure. If the mouth is not tracking the words after two or three takes, change the still. Always verify with sound on: the common failure is a face that moves convincingly but out of time, which looks correct until you listen.
- V2V canny, depth and pose extract structure straight from your source video through one Union IC‑LoRA — no reference image required — and share a single
strengthcontrol (0.3–1.0, default 0.85) with the detailer. Inpaint adds a mask image; outpaint adds a canvas position. V2V output dimensions must be multiples of 128, not the usual 64. - State the camera on every shot. The official caption recipe expects explicit camera motion, and if you want none you have to say so — "the camera remains static". Leave it out of an anchored I2V shot and the model tends to invent a push-in that drifts off your subject in the last second.
Available now. All five LTX‑2.5 models are live on the Sogni Supernet for pay-as-you-go Spark and included credit-free under fair use on every Sogni Unlimited subscription plan. Explore the model and start creating, or read the full LTX‑2.5 documentation.
Sources and further reading
- LTX‑2.5 official release overview
- LTX‑2.5 official model card and checkpoint notes
- LTX‑2.5 open-source documentation: what changed
- Official LTX prompting and multishot guide
- Official LTX‑2.3 audio guide
- LTX‑2.3 release notes
All source stills are original generations made on Sogni — the luthier and summit from Krea 2 Turbo, and the night summit a Qwen Image Edit of that same frame. No extracted video frames were used as image inputs anywhere in this work. Matched A/B pairs share the same caption, seed, resolution, and frame count. Clips are web-ready encodes: LTX comparisons at 1920×1088 and H3 Turbo at 1344×768, carried over from our MiniMax H3 post. Use the ⛶ View at… button on any clip to open it at its encoded size. Rendered on the Sogni Supernet, 2026-08-14/15.