MiniMax's open multimodal video model

UnrestrictedMiniMaxH3

Direct with text, images, and reference audio. Generate up to 15 seconds at 2K with native stereo sound. Real-person photos are welcome.

3,482,705+ happy users

15s
Maximum length
2K
Resolution
Stereo
Native audio
Yes
Real people
Unrestricted on Easy-Peasy.AI

Your Face Is an Input, Not a Rejection Reason

Animate your portrait, a creator photo, a team headshot, or customer-approved material. Easy-Peasy.AI does not block an ordinary H3 job simply because a recognisable person appears in the input.

Unrestricted does not mean unmoderated. Consent, likeness rights, our Terms of Use, and MiniMax safety filters still apply.

Real portraits accepted

Use selfies, headshots, creator photos, and team images you have permission to use.

Creator images accepted

Use approved creator and customer images as first frames or character references.

One clear workflow

There is no face-safe mode to find and no separate setting to enable before you generate.

Your rights still matter

Only upload likenesses and source material you own or are authorized to use.

Three Ways to Create

One Model, Every Starting Point

Start with an idea, a frame, or a full set of multimodal references. H3 uses the same general-purpose model across all three workflows.

T2V

Text to Video

Describe the scene, action, camera, dialogue, and sound. H3 turns the prompt into a coherent clip with native stereo audio already inside it.

  • 4–15 seconds at 768P or 2K
  • Six landscape, square, and vertical aspect ratios
  • Native dialogue, effects, music, and ambience
I2V

Image to Video

Animate a still as the first frame, or upload both a first and last frame to control where the shot begins and exactly where it lands.

  • First-frame animation from any supported image
  • Optional last frame for controlled transitions
  • Output framing follows the source automatically
R2V

Reference to Video

Combine character, product, and style images with reference audio, then explain each file’s role in natural language. H3 resolves the relationships in one generation.

  • Up to 9 reference images
  • Up to 3 reference audio clips, 15 seconds total
  • Natural-language roles for every uploaded reference
Official Video Examples

See MiniMax H3 in Action

Every clip below comes from the official MiniMax H3 release. Press play with sound to hear the native stereo output.

View the official MiniMax launch release ↗

Full-Modality Direction

A camera move from video, a character from an image, and a voice from audio — combined in one prompt

Native Stereo Sound

Dialogue, effects, ambience, and spatial placement generated with the picture

Film Opening Titles

A complete cinematic title sequence with coherent art direction and typography

Product & Game UI

Interface motion, branded details, and scene transitions held together across a full take

Animated Poster

Portrait-format graphic design turned into a controlled motion piece

Advertising & E-commerce

A vertical commercial with product staging, readable brand treatment, and native sound

Multimodal Context

Tell H3 How the References Relate

Most video tools treat character, product, style, and voice as separate jobs. H3 accepts them as one context. Describe the relationship in natural language and the model decides how to fuse them into the result.

Example direction

Keep the person from Image 1 consistent, feature the product from Image 2, and match the singing voice to Audio 1.

9

Reference images

Characters, products, locations, style, and composition

3

Reference audio clips

Voice timbre, singing, music, and sound direction

One natural-language prompt connects every input →
Native 2K

Detail Regenerated, Not Merely Upscaled

H3 uses in-context regeneration for its 2K tier. The base model looks back at the original prompt and reference context while rebuilding detail—useful for small type, product texture, and brand elements that a conventional upscaler would have to guess.

  • 2K at every duration from 4–15 seconds
  • 768P tier for lower-cost iteration
  • Six output ratios for ads, feeds, and widescreen work
Native Stereo

The Soundtrack Is Part of the Generation

Speech, effects, music, and ambience are jointly modeled with the visual sequence. That gives H3 a shared timeline for mouth movement, impacts, environmental sound, and camera motion—and a real left/right stereo field.

  • Dialogue and singing generated in-scene
  • Action-matched foley and environmental sound
  • Reference audio can guide voice timbre and music
Workflow

How It Works

01

Choose an H3 mode

Open the AI Video Generator and select MiniMax H3, MiniMax H3 Image, or MiniMax H3 Reference depending on what you want to direct.

02

Add your context

Start from text, a first frame, first and last frames, or a mix of reference images and audio. Ordinary photos of real people are accepted.

03

Direct the relationships

Tell H3 what to borrow from each input: a face from Image 1, a product from Image 2, or voice timbre from Audio 1. Natural language connects it all.

04

Set the finish and render

Choose any whole-second duration from 4 to 15 and render at efficient 768P or detail-first 2K. Stereo audio is included automatically.

Prompt Guide

Prompt the Context, Not Just the Scene

The best H3 prompts explain what to create and what role each source should play. These four habits make multimodal direction much more predictable.

Give each reference one clear job

H3 understands relationships between files, but the relationship still needs to be explicit. Name the source and say exactly what should transfer.

Example prompt

Use Image 1 for the presenter’s face and clothing. Keep the product from Image 2 unchanged in every shot. Match the speaking voice to Audio 1, but generate new words from the dialogue below.

Direct sound as part of the shot

Audio is not an afterthought. Describe dialogue, foley, music, ambience, and where they sit in the stereo field alongside the visual action.

Example prompt

She says “we made it” under her breath. Wind moves from left to right as the camera circles. A low engine hum stays centred; no music until the final two seconds.

Use the last frame as a destination

A last frame works best when the prompt explains the journey between the two images instead of merely describing the destination again.

Example prompt

Begin on the empty desk in the first frame. Components assemble in precise mechanical steps as the camera pushes closer, ending on the completed watch in the supplied last frame.

Write visible text exactly

H3 was released with a focus on text and brand rendering. Put required words in quotation marks, state where they appear, and keep the layout direction simple.

Example prompt

A clean product end card. The headline reads “MOVE LIGHTER” in large white condensed type, centred above the shoe. The brand mark from Image 2 remains unchanged in the lower-right corner.

Specifications

MiniMax H3 Technical Details

Maximum duration
15s
Any whole second from 4–15
Resolution
2K
768P or 2K output tiers
Audio
Stereo
Generated natively with every clip
Aspect ratios
6
21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16
Prompt length
4,000
Characters in the Easy-Peasy generator
Reference images
9
JPG, PNG, WEBP, HEIC, and HEIF
Reference audio
3
WAV or MP3, 15 seconds combined
First & last frame
Yes
Controlled start and destination
Real-person inputs
Accepted
No rejection merely for containing a face
768P price
3 cr/s
12 credits for the default 4 seconds
2K price
5 cr/s
20 credits for 4 seconds
Generation Upgrade

MiniMax H3 vs Hailuo 2.3

H3 moves beyond separate text and image generators into one multimodal system—with longer clips, native stereo, reference control, and 2K output.

FeatureMiniMax H3Hailuo 2.3
Maximum duration4–15 seconds6 or 10 seconds
Maximum resolution2K1080P
Native audioStereo audio on every renderSilent video output
Input contextText + image + reference audioText or first-frame image
Reference controlSubject, motion, style, and voice in one modelSeparate text and image workflows
First & last framesSupportedLimited to older dedicated workflows
Output framingSix aspect ratios plus adaptive reference mode16:9 in the Easy-Peasy catalogue
Commercial-Ready Use Cases

What You Can Create

Advertising

Build complete vertical or landscape spots with product action, readable messaging, dialogue, effects, and a finished soundtrack.

E-commerce

Keep the product anchored with image references while the camera, environment, and sound turn a still catalogue shot into a launch asset.

Game & Product UI

Animate interfaces, title screens, and product interactions while preserving the visual language supplied in your references.

Film Titles & Pre-vis

Prototype opening sequences, camera paths, and scene transitions in one coherent take before production begins.

Voice-Led Characters

Use a person or character image with reference audio to guide voice timbre, singing, delivery, and synchronized performance.

Animated Posters

Turn campaign art into a vertical motion piece with controlled typography, depth, particles, music, and stereo sound design.

Testimonials

Our Trustpilot score

Read the comments that people have made on public platforms.

Jeff Wilson

By far the best compilation AI utility

By far the best compilation AI utility out there. Period. We have tried multiple platforms that would work for us - a small business - and it was a bit overwhelming two years ago. Easy-Peasy has been a godsend. So many freaking tools and no fewer than 10 (of the 200) are ones we use at least on a weekly basis. Ironically, their human support is one of the most appreciated elements.

Vivian Daze avatar

Vivian Daze

Great services

Great services, and lovely and attentive customer service. Would reccommend

Cody Crabb

Great service, Best AI tool

Kia Tiow | COGUE avatar

Kia Tiow | COGUE

Easy-peasy is an easy to pick up and awesome AI software. I love it.

Easy-Peasy AI

Dezign121 avatar

Dezign121

Easy-Peasy.AI is very user friendly & a lot more to explore not only in the lesson, very helpful assistant in both work & personal

Blanca Marin avatar

Blanca Marin

Very Good AI…

I really liked the easy-peasy AI service. It's a great tool, but I also had to ask them for a refund because I chose the wrong plan, and they were very quick to respond to my request. They looked into my case and I got my money back without any problems. Of course, I will continue to use it for my projects.

Easy-Peasy AI

Muhammad Md Rahim avatar

Muhammad Md Rahim

Very awesome and practical AI tool that can help improve our productivity and the quality of our work!

Ken Poon avatar

Ken Poon

Very useful and true to “Easy Peasy” on using AI. One platform on all the tools required for business needs. Marianna sharing is super informative on how to use easy-peasy.ai

Dennis Cambronero avatar

Dennis Cambronero

I have been using the tool for +1 year…

I have been using the tool for +1 year and I have to say it is amazing. I use it to create images, bots and access every LLM in one single place. I highly recommend it.

Yeah Likes avatar

Yeah Likes

Practical and Powerful Software which is all In one and versatile. Just finished their workshop and got practical and valuable tips and tricks.

FAQ

Frequently Asked Questions

Everything you need to know about MiniMax H3

What is MiniMax H3?

MiniMax H3 is a general-purpose multimodal generation model released by MiniMax on July 31, 2026. It understands text, images, video, and audio together and generates up to 15 seconds of 2K video with native stereo sound. MiniMax designed it for controlled commercial content across advertising, branding, e-commerce, UI/UX, gaming, and more.

What does “unrestricted” mean for MiniMax H3?

It means Easy-Peasy.AI does not reject an otherwise ordinary image merely because it contains a recognisable real person. You can animate your own portrait, creator photos, team headshots, and customer-approved material. You do not need a special toggle or a separate workflow.

Does unrestricted mean there are no safety rules?

No. Unrestricted refers specifically to ordinary real-person inputs, not to prohibited content. You must have the right and consent to use uploaded likenesses. Deceptive impersonation, unlawful content, sexual content involving real people or minors, and other uses barred by our Terms of Use remain prohibited. MiniMax safety filters also still apply.

Does MiniMax H3 generate audio?

Yes. Every H3 render includes native stereo audio generated together with the picture. The model jointly handles speech, sound effects, music, and ambience instead of attaching a separate soundtrack after the video is complete. H3 does not expose a silent-generation switch, but you can mute or replace the audio later in the editor.

How long can a MiniMax H3 video be?

Choose any whole number of seconds from 4 to 15. Both 768P and 2K tiers support the full duration range. Longer stories can be built from multiple H3 shots in the Easy-Peasy.AI video editor.

What can I use as reference material?

H3 Reference on Easy-Peasy.AI accepts up to 9 reference images and 3 reference audio clips. Audio clips must each be 2–15 seconds long and can total no more than 15 seconds. Reference audio must be paired with at least one reference image.

Can MiniMax H3 use first and last frames?

Yes. Upload a first frame to animate a still, then optionally add a last frame to define the destination of the shot. First/last-frame mode and reference mode are separate: one generation cannot mix frame-control images with reference images or audio.

How much does MiniMax H3 cost on Easy-Peasy.AI?

Pricing is per generated second. The 768P tier costs 3 credits per second and the 2K tier costs 5 credits per second. The default 4-second 768P render costs 12 credits. Native stereo audio and reference audio do not add a separate charge.

Can I use MiniMax H3 videos commercially?

Yes. You can use videos generated on Easy-Peasy.AI for ads, social posts, product pages, client work, games, and other commercial projects, subject to our Terms of Use and your rights to any uploaded source material.