The model everyone's timeline picked up in August
Every few months one video model becomes the thing every creator timeline is testing. Since August, that model is MiniMax H3 — also known by its lineage name, Hailuo 3.0. The clips give away why: sound that lands on the action because it was generated with the picture, characters that hold their design across a 15-second multi-shot take, and 2K output that survives a full-screen watch.
Underneath the demos, H3 is a genuinely different design: an omni-modal model that treats images, video clips, and audio as one conversation, plus an open-weight release that made it the most studied video model of the season. And since late August there are two of them — the original H3 and H3 Max, a near real-time tier that changed what iteration feels like.
This is the plain-language intro: what H3 is, what each tier does, what to make with it, and how to use it without wasting a credit.
What H3 actually is under the hood
H3 is a video generation model built around one idea: everything is a reference. Where most models take a prompt and maybe one image, H3 accepts a working set —
- up to 9 reference images for subjects, outfits, and style
- up to 3 reference video clips for motion and pacing
- up to 3 audio clips to steer sound and rhythm (each attached to a visual)
— and you cite them in the prompt by order, the way you would brief an animator with a mood board. The model composes a clip of 4 to 15 seconds that follows the motion you referenced, keeps the subjects you pinned, and returns native 32 kHz stereo audio generated jointly with the picture. Ambience, foley, even lip-synced dialogue arrive in the same pass, already in sync — there is no separate dubbing step to schedule or pay for.
Resolution works in two stages: the base model renders at 768p, and a hosted regeneration stage lifts the final result to 2K (2560×1440). That split matters for one practical reason: MiniMax open-sourced the base weights, so researchers and toolmakers can build on H3 — but the 2K finishing stage stays hosted. Self-hosting is real, and capped at 768p.
Around the core generation, the family covers precise video editing and audio-video continuation — extending an existing clip with picture and sound together — which is what makes it usable for sequences, not just loops.
H3 Max: the fast tier that beat its own source model
On August 27, a second tier appeared: H3 Max, post-trained on H3's open weights by fal in partnership with MiniMax, and tuned for prompt adherence, aesthetics, and raw throughput.
The headline is speed that changes your working rhythm: officially, a 5-second clip in under 3 seconds; in everyday use, a 10-second clip comes back in roughly the time it takes to play it. Pricing lands around $0.05 per second at 480p and $0.08 at 768p, and within days of launch the tier gained the same mixed reference generation as the base model — images, clips, and audio, with the first few reference images free in the metered pricing.
The surprise is that the fast tier is not the rough tier. On public image-to-video leaderboards, H3 Max currently ranks first — above the base H3 it was trained from. Its one real ceiling is resolution: the 2K stage cannot be inherited from open weights, so Max tops out at 768p, structurally.
The clean way to hold the family in your head:
| MiniMax H3 | H3 Max | |
|---|---|---|
| Best at | Finished shots | Iteration speed |
| Resolution | Up to 2K hosted | 480p / 768p |
| References | 9 images / 3 clips / 3 audio | Same, added days after launch |
| Native audio | Yes | Yes |
| Rhythm | Queue it, review it | Closer to regenerating an image |
For the deeper workflow logic of splitting drafts and finals across the two tiers, see the H3 Max workflow guide.
What people are shipping with it
The demo reels cluster around the jobs H3's design favors:
Anime PVs and music-driven openings. Native audio generated with the picture means beat-synced cuts without an editor pass — describe the track's energy, reference an audio clip, and hits land on motion. This is the use case where H3 first took off among anime creators.
Multi-shot character pieces. A single 15-second generation can cut between locations while one character's face, outfit, and proportions hold — the hard problem of animated storytelling, handled inside one request when the references are good.
Dialogue clips. Lip-synced speech with room tone and ambience in one pass, in multiple languages. Short character moments that used to need three tools now need one prompt.
Style-locked sequences. Give it a visual language — a palette, linework, a period texture — and it holds that look across every cut in the clip, which is why stylized and retro work features so heavily in H3 showcases.
How to use H3 without wasting a credit
H3 is available as a generation model in ArcLoop — the model page covers modes and settings — and the difference between demo-reel results and frustrating ones is almost always the reference discipline around it, not the prompt wording.
Three habits carry most of the weight:
Feed it stable references. Nine image slots are an invitation to consistency, and a trap if the nine images disagree about who your character is. Keep the character as an asset with its canonical references bound to it, and pull from that one source for every generation. Referencing the asset with @ in shot descriptions is more reliable than re-describing the character each time — and it means your tenth H3 clip uses the same identity as your first.
One clear action per clip. H3's 15 seconds reward a shot brief with a beginning, a beat, and an end — not five ideas fighting for the same frames. Write the camera move, the action, and the sound you want, in order.
Generate from a storyboard, not a prompt box. In ArcLoop, open your episode and hit Generate Shots to break it into shot cards, then attach your assets and run generations per card — or batch them from the AI Chat Panel with a plain instruction like "Generate videos for Shots 1, 2, and 3." Every H3 clip lands attached to the shot it belongs to, so comparing takes and rebuilding the cut stays organized instead of scattering across downloads.
Iterate cheap, finish once. Explore timing and blocking on the Max tier or at lower settings, and spend the 2K pass on shots that have already earned their place in the cut — the same cost-control logic that applies to every tiered model family. If you are weighing H3 against the other finishing-grade engine in its class, the H3 vs Seedance 2.0 comparison walks that decision properly.
FAQ
Is MiniMax H3 the same thing as Hailuo 3.0?
Yes — H3 is the international name for the model line previously known as Hailuo. Some leaderboards also list the Max tier under an internal name, MiniMax H3 Turbo.
Is MiniMax H3 free to use?
The base model's weights are open, and hosted access is pay-per-use. The Max tier currently offers a small number of free daily generations on its host's tool page. For project work inside ArcLoop, H3 runs on your ArcLoop credits with settings depending on the generation mode.
What is the difference between H3 and H3 Max?
Max is a post-trained speed tier: near real-time generation at up to 768p, currently first on image-to-video leaderboards, with the same reference modes and native audio. The base H3 keeps the hosted 2K output stage. Draft on Max, finish on H3 is the split most workflows land on.
Does H3 really generate audio with the video?
Yes — audio is predicted together with the picture rather than added afterward, which is why sound effects land on impacts and dialogue syncs without an editing pass. You can describe the soundscape in the same prompt as the shot.
How long can an H3 video be?
4 to 15 seconds per generation, with audio-video continuation available to extend an existing clip. Longer pieces are built shot by shot — which is where working from a storyboard with persistent character assets beats prompting clips one at a time.
The model is new; the discipline is not
H3 earned its moment: mixed references, native audio, open weights, and now a Max tier fast enough to think with. But every capability on that list pays off in proportion to how stable your inputs are — and that part is workflow, not model. Build your characters as assets, brief shots like a director, iterate on the cheap tier, and spend 2K where the audience will actually look. Open a project, pin your first character, and give H3 something worth recognizing.





