Get Started

MiniMax H3 Max: What the Speed Tier Changes for Anime Creators

A new fast tier of the model behind this season's anime PV wave — and the honest workflow split between drafting fast and finishing at 2K.

Start Creating
MiniMax H3 Max: What the Speed Tier Changes for Anime Creators

Where this guide starts

MiniMax H3 spent August pulling anime creators toward it: an omni-modal video model that generates up to 15-second clips with native stereo audio, reaches 2K on the hosted version, and — the part that matters most for character work — accepts mixed references, so a clip can be steered by images, footage, and audio together. On August 27, a second name appeared next to it: MiniMax H3 Max, a speed-tuned tier post-trained on H3 in partnership with fal.ai. Search interest spiked immediately, and with it a lot of confusion about which one you are supposed to want.

Short version: H3 Max is not a bigger H3. It is a dramatically faster, cheaper H3 with a 768p ceiling — and within that ceiling it is no downgrade: on public image-to-video leaderboards it currently ranks first, ahead of the base model it was trained from. What you give up is exactly one thing — the 2K pass. This guide sorts out what that means for anime and OC work specifically.

H3 Max is not a bigger H3

H3 Max is a post-trained variant of H3 optimized for generation speed and cost:

MiniMax H3MiniMax H3 Max
Resolution768p open weights; up to 2K hosted480p / 768p
ModesText-to-video, image-to-video, reference generation from mixed images, clips, and audioText-to-video, image-to-video, and — added within days of launch — the same mixed reference generation (images, clips, audio)
AudioNative stereo, generated with the pictureNative stereo, generated with the picture
SpeedMinutes-scale for long clipsNear real-time: officially, a 5-second clip in under 3 seconds; early users report 10-second clips generating in roughly 10 seconds
TuningCapability breadthPrompt adherence, aesthetics, and throughput
List price (hosted)Higher per secondAround $0.05/s at 480p, $0.08/s at 768p; reference inputs metered by token with the first few images free

One more fact that reframes the launch: this is not a "fast but worse" tier. H3 Max currently sits first on public image-to-video leaderboards — Design Arena scores it above the base H3 it was post-trained from — because the post-training targeted prompt adherence and aesthetics before speed. Within 768p, Max is arguably the better model.

The resolution ceiling is structural, not a toggle waiting to be enabled. H3 renders at 768p and reaches 2K through a separate regeneration stage; only the base model's weights are open, so a model post-trained from those weights has no 2K stage to inherit. H3 Max at 768p is not "2K coming later" — it is the design.

What both tiers share is the property that made H3 interesting in the first place: audio is predicted together with the video, not scored on top afterward. Ambient sound, impacts, and musical hits land on the motion because they were generated with it.

What near-real-time and cheap actually unlock

Near-real-time, cheap, 768p clips change one production stage completely: iteration. Blocking a cut, testing whether an action reads, trying four timing variants of the same beat — at a nickel per second, with a clip back before you have finished rereading your own prompt, you can afford to be wrong constantly. That is a different creative mode, not just a cheaper one: when a variant costs seconds, you explore instead of committing.

And because reference generation shipped on Max days after launch, iteration no longer means abandoning identity control. You can pin your OC with reference images — the first few are free in the metered pricing — and draft on-model, steering with clips and audio the same way the full model does.

What Max still cannot do is the finishing resolution: the 2K stage belongs to H3 proper, structurally. So the honest split is:

  • Iterate on Max — motion tests, timing variants, layout drafts, reference-pinned identity checks, all at 768p speed.
  • Finish on the full model — the locked cut, re-rendered at 2K with the same references, with the native audio you will actually ship.

If you are comparing H3 against other finishing-grade models instead, that is a different question — the H3 vs Seedance 2 comparison covers that head-to-head.

How to run the draft-then-finish split in ArcLoop

A model tier is only useful inside a workflow that keeps your characters and cut stable while you switch between fast and final. In ArcLoop, that structure is the project itself:

Anchor identity before any generation. Create your characters as assets and bind their reference images. Draft passes are where identity drift runs wildest — a cheap fast clip that mutates your OC's face teaches you nothing about the shot. Referencing the character with @ in every shot description keeps even throwaway drafts on-model, because an asset reference is more reliable than re-describing the character from scratch each time.

Draft at the storyboard, not in a prompt box. Generate shots from your episode's storyboard so every fast pass lands attached to the shot card it belongs to. Four timing variants of Shot 7 stay with Shot 7 — comparable, replaceable, and not lost in a downloads folder.

Promote, don't redo. When a draft variant wins, the shot card already holds its description, its asset references, and its place in the cut. Rerun that shot on MiniMax H3 with full settings for the final pass — same shot, same references, higher tier — instead of rebuilding the prompt from memory.

Cut with drafts, replace with finals. Assemble the episode timeline while shots are still 768p drafts; pacing problems show up at any resolution, and they are cheaper to fix before you have spent finishing-tier credits on shots the edit will kill. Then swap finals in shot by shot. This is the same cost-control logic that applies to every tiered model family: spend where the audience looks, save where you are still deciding.

Ways to waste the speed advantage

  • Finishing on the speed tier. A 768p PV upscaled after the fact is competing against native-2K work this season. Max is a drafting tool; let it be one.
  • Drafting without identity anchors. Fast iterations with an unpinned character produce feedback about noise, not about your shot. Assets first, drafts second.
  • Planning around last week's spec sheet. Reference generation was "coming soon" at Max's launch and shipped within days; most comparison posts online still say it is missing. This family moves weekly — verify capabilities on the endpoint itself before committing a project, in either direction.
  • Comparing tiers on price per second alone. A draft you regenerate eight times at $0.08/s can cost more than one considered pass on the full model. Price the loop, not the clip.
  • Letting one hot model reshuffle a working pipeline. New tiers keep arriving from every model family. If your project structure separates drafting from finishing, each new release slots into one of those two roles in an afternoon — that is the durable skill, not any single model's spec sheet.

FAQ

Is MiniMax H3 Max better than H3?

At 768p, often yes — it leads image-to-video leaderboards above its own base model, and it now matches H3's mode list including mixed reference generation. What H3 keeps exclusively is the 2K output stage. "Better" depends on whether this pass is for iterating or for shipping.

How fast is it really?

The official claim is a 5-second clip in under 3 seconds; early users report 10-second clips coming back in about 10 seconds. Either way it is a different working rhythm — closer to regenerating an image than to queueing a video.

Can I use MiniMax H3 inside ArcLoop?

Yes — H3 is available as a generation model in ArcLoop, with settings depending on the selected generation mode. See the MiniMax H3 model page for what it does best in anime work: PVs, openings, and music-driven sequences where native audio carries the clip. H3 Max is in ArcLoop as its own model too: it takes image, video, and audio references and returns a 5-second 480p clip in around 3 seconds — see the MiniMax H3 Max model page.

Does H3 Max generate audio too?

Yes. The audio-with-video property comes from the shared base model, so speed-tier drafts still land sound on motion — useful for testing whether a beat-synced idea works before finishing it.

My characters need to stay consistent across shots. Which tier?

Both tiers now accept reference images, so identity can be pinned in drafts and finals alike. The discipline that actually decides consistency is upstream: keep the character's canonical references bound to an asset, and feed every generation — fast or final — from that same source instead of from whatever image was closest.

The tier is the tactic, not the strategy

H3 Max's launch is genuinely good news: the drafting half of your loop just got faster and cheaper. But speed tiers reward creators whose projects are structured to exploit them — characters locked as assets, shots living on storyboard cards, an edit that accepts drafts and swaps in finals. Set that structure up once, and every fast new model that ships becomes an upgrade to your iteration speed instead of a threat to your consistency. Open your project, pin your characters, and let the cheap tier be cheap.

Draft loose, finish locked

Build your characters as assets in ArcLoop, iterate rough passes cheaply, and render final shots with MiniMax H3's full 2K and native audio when the cut is decided.

Create Now

Discover More

GPT Image 2.5: Keep One Character Consistent Across Every Shot

GPT Image 2.5: Keep One Character Consistent Across Every Shot

How Many Seconds Does This Shot Get?

How Many Seconds Does This Shot Get?

GPT Image 2.5: Transparent Assets and UI Resource Guide

GPT Image 2.5: Transparent Assets and UI Resource Guide

Dawn, Noon, Dusk, Night: How to Light One Scene in Four Temperatures

Dawn, Noon, Dusk, Night: How to Light One Scene in Four Temperatures