BYTEDANCE Seedance 2.5 30-second one-take clips, 50 reference inputs, released July 31, 2026 | MINIMAX H3 (Hailuo 3.0) Native 2K with stereo audio in one pass, released July 31, 2026 | ALIBABA Wan 3.0 30-second clips from text, images, or a PDF, gated beta August 6, 2026 |
The last ten days of my life have mostly been rendered queues. On July 31, ByteDance pushed Seedance 2.5 live through Jimeng and Doubao, and MiniMax shipped H3 on Hailuo the very same day. Nobody believes that timing was an accident. Then, before anyone finished arguing about those two, Alibaba opened the Wan 3.0 beta to applicants on August 6 and threw a third contender into the ring.
I've been generating with all three since launch week: product shots, dialogue scenes, a night market tracking shot I kept reusing as my control prompt, and one cursed attempt at turning a pitch deck into a video. My Hailuo credit balance is in ruins. My opinions are earned.
Here's what actually separates them once the launch hype fades, and which one deserves your money right now.
THE QUICK VERDICT • MiniMax H3 takes the crown, narrowly. Verified 2K output, the best native audio here, top-three Artificial Analysis placements within days of launch, and roughly a dollar per clip. • Seedance 2.5 has the highest ceiling. Nothing else pairs 30-second single takes with 50 reference inputs. It's the pick for long-form, brand-controlled work. • Wan 3.0 is the wildcard. Document-to-video is genuinely new and per-second pricing is the most transparent, but it's a beta with real rough edges. |
The tale of the tape
Before any opinions, the launch specs side by side. These come from vendor documentation and platform pages, so treat them as claims that independent benchmarks are still verifying.
| Seedance 2.5 | MiniMax H3 | Wan 3.0 | |
| Developer | ByteDance (Seed team) | MiniMax | Alibaba (Tongyi Lab) |
| Status | Released July 31, 2026 | Released July 31, 2026 | Application-gated public beta, August 6, 2026 |
| Max single pass | 30 seconds, with a 180-second Ultra-Long beta on Jimeng | 15 seconds, extendable to around 30 | 30 seconds, plus a separate extension tool |
| Resolution | Up to 4K on Dreamina's marketing, though the native 4K pipeline was announced for Seedance 2.0 | Native 2K (2560×1440) at 24fps, no upscale pass | Priced tiers at 480p, 720p, and 1080p in beta |
| Native audio | Yes, 10+ languages, with lip sync and beat matching driven by a reference track | Yes, stereo dialogue, effects, and room tone in one pass | Yes, though Alibaba openly says audio quality still lags |
| Reference inputs | Up to 50 files: 30 images, 10 videos, 10 audio clips | Up to 9 images, 3 videos, 3 audio clips, capped at 12 files | Text, image, audio, video, plus PDFs, slides, spreadsheets, and live URLs |
| Editing | Region edits, green-screen swaps, white-model shot blocking | Plain-language instruction edits, ranked #1 for editing on Artificial Analysis | Editing folded into the main model rather than a separate one |
| Access | Jimeng, Dreamina, Doubao Pro; API on Volcano Engine Ark from August 7 | Hailuo app, MiniMax Hub, MiniMax Open Platform API, third-party hosts | Qwen Cloud and Alibaba Cloud Model Studio, full API rollout pending |
| Pricing signal | No free API quota; needs about a $30 balance or an existing resource package | Roughly $1 for a 15-second 2K clip on the basic tier, per early access partners | $0.05/sec at 480p, $0.10/sec at 720p, $0.20/sec at 1080p |
| Open weights | No, fully proprietary | Shipped August 3; license excludes the US, EU, UK, and South Korea | Apache 2.0 pledged, not yet shipped, community skeptical |
Sources: vendor launch documentation and platform pages, August 10, 2026. Beta figures can and probably will change.
Meet the three contenders
Seedance 2.5
ByteDance / The heavyweight

ByteDance previewed this at the Volcano Engine FORCE conference on June 23, skipped versions 2.1 through 2.4 to signal a generational jump, then shipped it on July 31. Pedigree matters here: Seedance 2.0 topped the Artificial Analysis Video Arena for both text-to-video and image-to-video earlier in 2026. The 2.5 pitch is control at length: one continuous 30-second take, guided by up to 50 reference files.
Where it wins
• 30-second native single takes, with a 180-second beta mode on Jimeng
• 50 multimodal references keep characters, wardrobe, and products locked
• White-model blocking and green-screen swaps feel like real production tools
Where it hurts
• No free API quota, and you need about $30 in your account just to switch it on
• 4K claims are muddled between the 2.0 and 2.5 announcements
• Full access outside China is still catching up to the Jimeng rollout
MiniMax H3
MiniMax / The all-rounder

Officially MiniMax H3, known to almost everyone as Hailuo 3.0, this was unveiled at WAIC on July 17 and released July 31. MiniMax built an omni-modal system rather than a plain video model: it reads text, images, video, and audio as one context, then returns a clip with stereo sound already in it. Within days it held #1 in video editing, #2 in text-to-video, and #3 in image-to-video on the Artificial Analysis leaderboards.
Where it wins
• Native 2K straight out of the model, not an upscale pass
• Dialogue, sound effects, and ambience generated with the picture
• Plain-language edits: swap a character, change the lighting, fix a line
Where it hurts
• 15-second cap on a single pass, half of what the other two offer
• Open-weight license excludes the US, EU, UK, and South Korea
• 12-file reference cap looks small next to Seedance's 50
Wan 3.0
Alibaba / The wildcard

Wan 3.0 entered an application-gated public beta on August 6 through Alibaba Cloud Model Studio, jumping from the 2-to-15-second range of March's Wan 2.7 to full 30-second single takes. The feature nobody else has: it accepts documents. Feed it a PDF, a slide deck, a spreadsheet, or a live URL and it builds a video from the contents. Alibaba also collapsed four separate Wan models into one, and it even proposes its own clip length from your prompt.
Where it wins
• Document-to-video is a genuinely new input category
• Cleanest pricing of the three, billed per second per resolution
• 30-second takes plus a dedicated extension tool
Where it hurts
• Alibaba itself admits audio and on-screen text accuracy trail the field
• Tops out at 1080p in the current beta price list
• The Apache 2.0 promise rings hollow after Wan 2.5 and 2.6 never opened
Seven rounds, head to head
Spec tables only get you so far. Here's how the three stack up on the things that decide whether a clip is usable.
ROUND 1
Clip length and one-take continuity
Seedance 2.5 and Wan 3.0 both generate 30 seconds in one pass, which sounds like a tie until you check the edges. Seedance shipped first, adds multi-turn extension, and has a 180-second Ultra-Long beta already visible inside Jimeng. My night market tracking shot held its subject, the signage, and the light for the full half minute. Wan matches the ceiling and its intelligent duration feature is clever, but it's a week old and still in beta. H3's 15-second cap puts it a distant third.
Winner: Seedance 2.5
ROUND 2
Resolution and picture quality
This one is messier than the marketing suggests. H3's native 2K at 24fps is confirmed, consistent, and genuinely sharp, with no separate upscaler softening the edges. Seedance's 4K story is complicated: Dreamina advertises 4K output, but the native 4K pipeline was formally announced as part of the Seedance 2.0 update at the same event. Wan 3.0 currently prices out at 1080p. On verified, shipping resolution, H3 holds the high ground. If Seedance's 4K survives independent testing, this round flips.
Winner: MiniMax H3, for now
ROUND 3
Native audio
All three generate sound with the picture. The gap is in how good that sound is. H3 produces stereo dialogue, effects, and room tone in one pass, and the first cafe scene I generated came back with background chatter so believable I replayed it to be sure. Seedance counters with range: more than ten languages, plus a reference track that can drive pacing, beat matching, and lip sync by itself. Wan brings up the rear by Alibaba's own admission, which I respect them for saying out loud.
Winner: MiniMax H3
ROUND 4
Reference control
This is Seedance's home turf and it isn't close. Fifty reference files in one generation, split as 30 images, 10 video clips, and 10 audio tracks, is the largest capacity any commercial video model has claimed. Load it with character sheets, wardrobe shots, product angles, and background plates, and it holds them all across the clip. H3's 12-file cap covers most single-scene jobs, but a campaign with multiple characters and strict product accuracy hits that wall fast. Wan's system is broad rather than deep, and that breadth gets its own round next.
Winner: Seedance 2.5
ROUND 5
Inputs beyond images and text
Wan 3.0 did something nobody asked for and plenty of people needed. It reads structured documents: PDFs, slide decks, spreadsheets, even a live URL, and turns them into video. I fed it an eleven-slide product deck and got back a coherent 30-second explainer that followed the deck's actual structure, not a vague remix of it. It needed a cleanup pass, but as a starting point for training content or internal comms, that's a different product category. Neither rival accepts anything like this, and it's the strongest reason to keep Wan on your radar through the beta.
Winner: Wan 3.0
ROUND 6
Editing after the render
Regenerating a whole clip because one element is wrong is the most expensive habit in AI video, so editing matters more than it sounds. H3 takes plain-sentence instructions: replace a person, remove an object, change the lighting, rewrite a line of dialogue, and it currently sits at #1 in video editing on the Artificial Analysis leaderboard. Seedance answers with region-level edits that change part of a frame without touching the rest, plus green-screen swaps, a more surgical, production-style approach. Both are strong. The independent ranking breaks the tie.
Winner: MiniMax H3
ROUND 7
Price and getting in the door
Wan has the most transparent rate card, a full 30-second clip running $1.50 at 480p or $6 at 1080p, but the beta itself is application-gated. H3 is the best value and the easiest door: around $1 for a 15-second 2K clip per early partners, live in the Hailuo app, the API, and third-party hosts from day one, and its open weights hit Hugging Face on August 3. Seedance has no free API quota and needs a roughly $30 balance before it switches on. Great model, unfriendly front door.
Winner: MiniMax H3
The scorecard
| Round | Seedance 2.5 | MiniMax H3 | Wan 3.0 |
| Length and continuity | Win | Close second | |
| Resolution | Pending 4K verification | Win | |
| Native audio | Close second | Win | |
| Reference control | Win | ||
| Unusual inputs | Win | ||
| Editing | Close second | Win | |
| Price and access | Win | Close second | |
| Total | 2 | 4 | 1 |
Match the model to the job
Scorecards flatten nuance. For most working creators, the right model depends on what's actually sitting on the production calendar this month.
Short social ads with dialogue Pick: MiniMax H3 2K picture, synced stereo audio, and a per-clip cost low enough to iterate freely. | A 30-second one-take brand film Pick: Seedance 2.5 The only model where 30 seconds of continuity and 50 locked references coexist. |
Turning a deck or PDF into video Pick: Wan 3.0 Document-to-video is exclusive to Wan. Nothing else in this fight reads a spreadsheet. | Character-driven, stylized clips Pick: MiniMax H3 The Hailuo line built its name on expressive character motion, and H3 inherits it. |
High-volume 480p drafting Pick: Wan 3.0 At $0.05 per second, throwaway drafts cost $1.50 per 30-second take. | Multi-character brand campaigns Pick: Seedance 2.5 Character sheets, wardrobe, and product refs all stay consistent in one generation. |
Final verdict
After a week and a half of living inside these three models, here's where I've landed, and it surprised me.
I expected Seedance 2.5 to run away with this. On paper it's the most capable video model shipped to date: the longest takes, the deepest reference system, production tools like white-model blocking borrowed from a real film set. When I needed one specific clip to be exactly right, a continuous take with a consistent character and a real product in frame, Seedance was the model I trusted. If your work looks like that, ignore my scorecard and go get it.
But the crown goes to the model I kept reopening without thinking, and that was H3. Almost every session ended the same way: a usable 2K clip with sound that didn't need replacing, the one wrong detail fixed with a sentence instead of a re-roll, and about a dollar spent doing it. It ranks where it should on the independent leaderboards, it's available through more doors than either rival, and it never made me check my balance before letting me work. That's what winning looks like in practice.
Wan 3.0 finishes third and I still wouldn't count it out. Document-to-video solved a real problem for me in one afternoon, the pricing is the fairest here, and the gated beta is four days old. If Alibaba fixes the audio and actually ships those Apache 2.0 weights, this ranking gets rewritten. I'm just not betting production work on promises the Wan team has broken twice before.
THE CROWN, AUGUST 2026 MiniMax H3 takes it, on points Not by knockout. Seedance 2.5 wins on ceiling and Wan 3.0 on ideas, but H3 delivers verified quality, real sound, and low cost to most people, today. Rematch when Seedance's 4K holds up or Wan's weights ship. |
Comments