Skip to content
Chris Perkles
AI Tools19 min

MiniMax H3 vs LTX-2.5: 126 Clips, Same Prompts, Two Reels

63 shots, identical prompts, identical seeds, two models, one RTX 5090. The result is two finished 61-second showreels. Here is where the models diverge, where the automation crowns the wrong winner, and why the licence decides everything in the end.

MiniMax H3 vs LTX-2.5: 126 Clips, Same Prompts, Two Reels

Back in March I wrote about LTX 2.3 and what local video generation on an RTX 5090 actually delivers. Two things have changed since. LTX is at 2.5 now, and MiniMax H3 has arrived as a serious alternative.

So I ran the only test that means anything. Not two models each showing off with their favourite prompts, but 63 identical shots, identical seeds, identical resolution, identical step count, the same card, two nights back to back. What came out are two finished showreels of 61.5 seconds each, cut by the same script, on the same beat, to the same music.

The result is not a winner. It is a map of how differently two models read the same set of instructions.

The setup

Both runs went through ComfyUI on an RTX 5090, driven by a Python script that works the shot list, scores the clips, cuts the reel and delivers the result to Dropbox.

MiniMax H3LTX-2.5 distilled (nvfp4)
Resolution1024×5761024×576
Frames124 @ 24 fps (5.17 s)129 @ 25 fps (5.16 s)
Steps2020
Samplerres_multistepeuler
Shots rendered6363
Total GPU time3.87 h2.67 h
Median per clip220 s105 s
Native audioyesyes

The shot list is built in six sections, like a real reel: cold_open, build, peak, breakdown, rebuild, hero. On top of that, nine categories, from FPV drone flights through waterfall macros and automotive to Super 8 home movie material. Of the 63 rendered shots, 45 make the cut. The rest is deliberate overproduction so the curation step has something to choose from.

Every prompt follows the same schema: one line of scene description, a timecode block for the camera move, a line of camera characteristics, a negative line against text and watermarks, an audio line. One of them looks like this:

One continuous five-second shot, no cuts. A dense modern city core at blue hour,
glass towers lit from within, cool blue sky against warm interior light,
anamorphic feel. [0s-5s] The camera rises vertically between two towers, glass
facades sliding down past frame, city grid opening up below. Smooth vertical
drone ascent, perfectly level horizon. No text, no logos, no subtitles, no
watermarks. Audio: distant city hum, wind between buildings, no music.

Same text, same seed, two models. What happens after that is the actual test.

The two reels

Both run 61.5 seconds, 45 clips, cut to 118 BPM (one beat is 0.508 s), with eight moments where the model's native audio punches briefly through the music.

MiniMax H3:

LTX-2.5:

At reel length both hold together surprisingly well. The difference does not live in the overall impression, it lives in individual shots, and there it is dramatic.

Where the models part ways

For the comparison I built contact sheets per shot from both runs, three frames per model out of the same clip, H3 on top, LTX-2.5 below. Plus stacked moving-image comparisons for the shots where the motion is the argument. Nine examples, from the first section of the reel to the last.

Cold open: over the edge

The first shot of the reel is the hardest test in the whole set, because it prescribes a timeline: "[0s-1s] The FPV camera hangs at the cliff lip, still. [1s-5s] It tips forward and plunges down the rock face at high speed, wall rushing past inches away, sea rising fast to fill frame." Plus the image description: "first light, dark basalt rock, cold grey-blue sea far below".

Cold open up close: H3 warms the scene into a sunset, LTX holds the cold morning lightCold open up close: H3 warms the scene into a sunset, LTX holds the cold morning light

Two deviations from H3 at once. First the colour temperature: "first light" and "cold grey-blue sea" turn into a warm sunset with an orange cloud deck. Second the distance: instead of a wall rushing past inches away, the camera stays high above the cliff and rolls over the edge, more landscape shot than plunge.

LTX holds the cold morning light, the grey-blue water and the proximity to the rock. The camera clings to the stone, tips forward, and the sea genuinely fills the frame.

The colour temperature drift in H3 is the more interesting half, because it looks familiar from image generation. Models tip into warm "editorial" mood under uncertainty, because warm backlight is overrepresented in training data. The countermeasure is the same one as there: write time of day and colour temperature into the prompt explicitly and repeatedly, not just once at the top.

Super 8: literal or atmospheric

The prompt asks for "authentic Super 8mm home-movie footage, heavy warm grain, soft halation, slight gate weave, faded amber cast, 1970s emulsion look", plus a young woman laughing in summer light.

H3 on top, LTX-2.5 below: the same Super 8 prompt, two completely different readingsH3 on top, LTX-2.5 below: the same Super 8 prompt, two completely different readings

H3 delivers a clean widescreen shot with a warm colour cast. The image is sharp, digital, technically flawless, and "Super 8" has been read as a grade, not as a medium. LTX-2.5 renders the actual film frame: ragged frame edges, gate weave, the smaller image area with black surrounding it, plus halation blooming off the highlights.

That has a concrete consequence for post. The cutting script carries a Super 8 grade for the breakdown section that adds grain, vignette and a colour shift. For the LTX run the script switched that grade off by itself, with the log line Super-8 grade off (model supplies its own). For the H3 run the grade was necessary, otherwise the section would have sat there as a style break.

Both are defensible. If you want the retro look as a mood and you want to keep control of the grade, H3 is the better starting point. If you want the material to look like it came out of a projector, LTX saves you a whole step.

Super 8, round two: what the look costs

The same conflict on a second shot, and here the price becomes visible. The prompt: "children run through a garden sprinkler in hard summer sun, water catching light, pure unposed joy."

Children in a sprinkler: H3 stays legible, LTX delivers the more convincing film frame and loses the facesChildren in a sprinkler: H3 stays legible, LTX delivers the more convincing film frame and loses the faces

H3 gives you four children in bright summer light, faces clearly readable, the fan of water cleanly drawn against the backlight, a light vignette in the corners. That is usable material: you can see who is in frame and what is happening.

LTX builds the complete film frame with rounded gate corners, a heavy vignette and contrast tipped towards sepia. It looks considerably more like 1975. It also costs two of three frames their legibility, because the spray closes up against the backlight and the children disappear into the darkness of the vignette.

That is the honest trade-off behind the whole Super 8 question: authenticity and legibility pull in opposite directions. For a two-second emotional cutaway, LTX is the better call. For anything where a person has to stay recognisable, H3 plus a grade in post is the safer route.

Camera moves: does the motion actually get executed?

The skyscraper shot above asks for a vertical rise between two towers, with the "city grid opening up below".

Vertical drone ascent: H3 stays between the towers, LTX climbs through to the revealVertical drone ascent: H3 stays between the towers, LTX climbs through to the reveal

H3 stays between the facades for the full five seconds. There is movement, but it is conservative, and the promised reveal never arrives. LTX actually climbs through, and in the last third the complete city grid opens up beneath the camera. The shot tells the story the prompt describes.

This is not a one-off. The motion score the curation script computes reads 7.63 (H3) against 10.51 (LTX) for this shot. The same pattern runs through the entire peak section: LTX executes described camera moves more completely, H3 starts them and then caps them.

FPV through forest: where more motion is not more quality

The counterexample arrives one section later. The FPV shot through pine forest carries the highest motion score in the entire set, in both runs.

FPV through pine forest: H3 holds the forest structure, LTX delivers more perceived speed with less substanceFPV through pine forest: H3 holds the forest structure, LTX delivers more perceived speed with less substance

H3 scores 17.50, LTX 21.44. By the metric, LTX wins clearly. Look at the frames, and the verdict flips.

With H3 the trunks stay trunks across the whole run. The forest floor is readable, sun breaks through the canopy, the motion blur sits at the edges where it belongs, and the scene has a spatial order you can still follow after five seconds. With LTX the speed is noticeably higher, but the foliage smears radially into mush, the trunks dissolve into streaks, and by frame three the forest is more of a texture than a place.

Which puts the earlier claim more precisely: LTX executes described motion more consistently, but past a certain speed it pays for that with structural breakdown. H3 travels slower and holds together instead. For a one-second beat in the cut that is irrelevant. For a shot that stays on screen for three seconds it is not.

Optics: when the description is lens language

The waterfall macro prompt is deliberately technical: "extreme close detail of white water hitting rock, individual droplets suspended, backlit spray, very shallow focus."

Waterfall macro: H3 stays wide and uniformly sharp, LTX delivers individual droplets and real bokehWaterfall macro: H3 stays wide and uniformly sharp, LTX delivers individual droplets and real bokeh

This is the widest gap in the whole test. H3 delivers a solid but wide waterfall shot with everything in focus. There is no trace of "very shallow focus", and none of "individual droplets suspended" either. LTX gives you exactly that: single droplets hanging in the air, real bokeh in the background, backlight in the spray, a focal plane that looks like it came off a macro lens.

If your prompt style is built from lens descriptions, and a DP's prompt style is pretty much exactly that, LTX-2.5 understands more of it.

Geometry: when one word gets read wrong

The automotive shot asks for "a dark blue sports saloon carving a hairpin in alpine light", with an aerial chase from behind and above.

Automotive shot: H3 builds a mountain switchback, LTX a complete 360 degree loopAutomotive shot: H3 builds a mountain switchback, LTX a complete 360 degree loop

H3 reads "hairpin" the way a camera operator would: a tight bend on a mountain road, camera behind the car, guardrail, drop beyond it. It looks like a real drone plate.

LTX reads "hairpin" geometrically and builds a full 360 degree loop into the mountain, like a spiral ramp. The shot is more spectacular, but it is not what was ordered, and it is unusable as a plate in an automotive spot. On top of that the camera sits wide and static instead of chasing.

That is the flipside of better prompt execution. LTX takes individual words more literally, and when the literal reading is wrong, the whole shot is wrong.

Scale and headcount: animals from above

The zebra shot is a good test of whether a model can hold many moving objects together at once.

Zebra herd from above: H3 holds around sixty animals, LTX shows a dozen at higher detailZebra herd from above: H3 holds around sixty animals, LTX shows a dozen at higher detail

H3 renders a dense herd of around sixty animals at high altitude, with clean stripe patterns and shadow casting that holds up across the five seconds. It looks like a genuine wildlife drone shot. LTX drops much lower, shows maybe a dozen animals at higher detail resolution, but with visible anatomy problems and stripe patterns that do not carry everywhere.

Where many objects have to stay plausible in motion at the same time, H3 is clearly ahead.

The hero shot: who wins when stillness is the brief

The last shot of the reel runs four seconds and is the only one where the prompt explicitly asks for calm: "smooth stabilised aerial motion, no jitter, no drift", ending on "a stable held panorama".

Hero shot: H3 holds the panorama, LTX pushes across a rock ridgeHero shot: H3 holds the panorama, LTX pushes across a rock ridge

H3 does exactly that. A held panorama over an alpine valley, sun in the haze layer, minimal drift, an image a logo could land on. LTX instead pushes across a rock ridge, with foreground travelling through frame.

Taken on its own the LTX shot looks better. It is still the wrong shot, because a hero ending needs stillness for the cut to land. And this is precisely where the automation becomes the problem.

Why the automated curation crowns the wrong winner

The curation script scores every clip with a motion score, essentially the average frame change across the runtime, and keeps the best N per section. That works surprisingly well for 62 of 63 shots. For hero_01 it does not.

H3 scores 0.29 for its correctly held panorama. LTX scores 2.31 for the travelling shot that ignores the instruction. Inside the section the score wins, so the shot that disobeyed the direction wins.

The lesson is not "automation is useless". It is more precise than that: a motion metric is a proxy metric, and proxies break exactly where the intent contradicts the metric. In five of six sections more motion genuinely is better. In the hero section it is not. The fix is cheap, the hero section gets an inverted score, but you have to notice it first, and to notice it you have to look at the rejected clips rather than only the finished reel.

A second result from the same curation step is more surprising. Of the 45 selected shots, 39 are identical across both runs. Only six positions differ. Two completely different models, the same selection logic, and the selection agrees 87 percent of the time. In this setup, what turns out well is decided far more by the prompt than by the model.

Speed, and why the average lies

H3 is the more predictable machine. 63 clips, median 220 seconds, minimum 215, maximum 250. That is under 16 percent spread across an entire overnight run. You can plan a night around it.

LTX-2.5 is twice as fast in the normal case, median 105 seconds. But the average sits at 153 seconds, and that gap has a cause: seven consecutive clips fell out of line, from build_08 to peak_car_03, with render times between 255 and 680 seconds. Eleven minutes for a five-second clip, at identical settings to the clip before it.

After that the run caught itself and went back to 100 to 115 seconds for the remaining 46 clips, with no intervention. That smells like memory pressure building across several passes and then clearing. For an unattended run it means: budget with the median, but size the window for the worst case, otherwise you find half a reel in the morning.

The obvious question: does the extra time buy anything? No. build_09 is the slowest clip of the whole run, 631 seconds on LTX against 220 on H3.

The most expensive clip of the run: 631 seconds on LTX, 220 on H3, no visible difference in the resultThe most expensive clip of the run: 631 seconds on LTX, 220 on H3, no visible difference in the result

Both fly through the sea stacks towards the low sun, both hold the geometry, LTX has the slightly nicer anamorphic flare. That is a difference of taste, not of quality. The six extra minutes are pure lost time, not better sampling. Which is why I read the outlier as a resource problem rather than a property of the model.

Bottom line: 3.87 hours of GPU time for H3, 2.67 for LTX, for 63 clips each. LTX is faster, but not as reliably fast as the median suggests.

The part that has nothing to do with image quality

The run was originally meant to complete in a single night: H3 first, then automatically update ComfyUI, then LTX-2.5. At 01:32 the log says:

[01:32:57] ===== updating ComfyUI for LTX-2.5 =====
fatal: Not possible to fast-forward, aborting.
[01:32:57] git pull FAILED — staying on dec5d945, skipping LTX

A local commit in the ComfyUI directory, a git pull --ff-only that was not prepared for it, and half the plan is gone. The script behaved correctly, it logged the rollback point, aborted the update attempt, restarted ComfyUI on the old revision and exited cleanly. The LTX run was then pulled through by hand the next morning, which is why the two reels carry dates from 15 and 16 August.

The point is not the failure. The point is that in multi-stage overnight runs the expensive mistake almost never sits in the model, it sits in the glue between the stages. An update step between two renders is a single point of failure, and when it tips over you do not lose compute time, you lose a calendar day.

Native audio

Both models generate audio alongside the image. 12 of the 63 clips are marked keep_audio in the shot plan, mostly the Super 8 clips and the hero shot, and the cutting script punches that model audio briefly through the music at eight points.

The effect is good, but limited. What works is texture: projector rattle, wind at altitude, water impact, rotor rush. What does not work is anything that needs timing. An engine note that should match a gear change on screen, or hoofbeats that should sit in sync with an animal's stride, does not come out reliably from either model. For a reel with continuous music, native audio is a welcome seasoning. For a spot where sound design carries the piece, it replaces nothing.

The licence decides everything

And now the part that matters more to any company in Austria, Germany or anywhere else in the EU than every image comparison above.

The MiniMax H3 licence lists the EU as an Excluded Territory. Which means everything in this comparison that came out of H3 is a personal experiment and not a usable asset. Not for client projects, not for advertising, not for commercial delivery. That is exactly why the H3 reel above carries AI-EXPERIMENT in its filename.

LTX-2.5 does not have that problem. It is an open model, it runs locally, not a single frame leaves the machine, and the usage terms are settled.

That tips the whole comparison. H3 wins on herds, on stability and on the more plannable render time. LTX wins on optics prompts, on camera moves and on speed. But if you sit in the EU and produce for clients, the question of the better image is academic, because one of the two models simply is not available for that purpose.

That is the lesson missing from every model comparison, and the one that hits first in practice: licence status and territory are a first-order selection criterion, not a detail for legal to handle at the end. A model you are not allowed to ship is not a candidate for commercial work, no matter how well it renders.

What changes on the next run

1. Score the hero section inversely. Less motion should win there, not more. One line in the curation script.

2. Cap camera movement explicitly where it is not wanted. "Stable held panorama" is not enough for LTX. It needs a hard negative, something along the lines of "no dolly, no push, no reveal, locked frame after 1s".

3. Spell out ambiguous craft terms. "Hairpin" has two readings, and LTX reliably picks the geometric one. "Tight switchback turn on a mountain road, guardrail on the outside" leaves no room.

4. Separate grading instructions from the medium. Write "Super 8" and LTX hands you the film frame with it. If only the colour mood is meant, that needs to be written that way: "warm faded colour grade, fine grain, no film borders, no gate weave".

5. Take update steps out of the render run. ComfyUI gets updated and verified before the run, not in the middle of it. An overnight run should render, nothing else.

Conclusion

Two models, the same 63 prompts, the same seeds, the same card, two finished reels. What remains:

LTX-2.5 executes described camera moves more completely, understands lens and focus language considerably better, delivers style instructions like Super 8 as an actual medium rather than a grade, and renders twice as fast in the normal case. In exchange it reads individual words more literally than you would like, and its render time is less dependable than the median suggests.

MiniMax H3 is the calmer machine. It holds many moving objects together plausibly, executes instructions to stay still correctly, and its render time varies by less than 16 percent across an overnight run. In exchange it understates camera moves and largely ignores focus descriptions.

And together they show that the selection logic crowned the same material 87 percent of the time. In this setup, the prompt stack carries more weight than the choice of model.

For commercial work in the EU, only one of the two survives anyway, and that is decided by the licence, not by the image.


Want to bring AI-generated footage into your production? I help with model selection, prompt stack, local setup, and the question of what you are actually allowed to ship at the end. Let's talk →

MiniMax H3LTX-2.5AI VideoRTX 5090ComfyUIVideo Production
Share
Chris Perkles

Chris Perkles

AI consulting, automation and training from Salzburg. Founder of Skyline Medien and AgencyFlow — his own agency now runs on a fraction of its former resources.

Related Articles
MiniMax H3 vs LTX-2.5: 126 Clips, Same Prompts, Two Reels | Chris Perkles Blog