Chinese AI company MiniMax has introduced H3, a multimodal generation model that can work with text, images, video and audio in the same request. The company says the model can create clips up to 15 seconds long in 2K resolution with native stereo sound, while also supporting editing and motion transfer from reference videos.
The launch places MiniMax in more direct competition with Chinese rivals ByteDance and Kuaishou, which have also recently advanced their video-generation offerings. Reuters reports that MiniMax is targeting commercial work such as advertising, e-commerce, product design and games, where creators increasingly need tools that can combine several types of source material.
MiniMax says users can describe relationships between supplied materials in natural language. For example, a creator could provide one video as a camera-motion reference, an image as a character reference and an audio file for vocals. The company says H3 can use these inputs to generate a new video that combines the requested elements.
One model for generation and editing
H3 reflects a broader shift in video AI. Earlier systems often split work into narrowly defined products or modes, such as text-to-video, image-to-video, character reference, video editing or sound generation. MiniMax says it trained H3 to handle these tasks in a more unified way.
According to the company, this approach also covers audio generation. Rather than treating speech, music and sound effects as separate systems, H3 jointly models them and produces stereo audio alongside video. The model also supports what MiniMax calls generalized reference and editing. In practice, this means users can provide different combinations of images, video and audio, then explain the desired change in words.
For content teams, the promise is a workflow with fewer handoffs between specialist tools. A campaign creator could potentially use an existing product image, a motion reference and a music sample in one generation request. However, the company’s performance claims have not been independently verified in the supplied materials.
MiniMax says H3 is particularly strong at following detailed instructions, rendering text and brand elements accurately, and transferring motion between videos. These are important capabilities for commercial production, where visual consistency and legible on-screen text can matter as much as cinematic quality.
Open weights planned
MiniMax says it plans to release H3’s model weights in the coming days, subject to applicable laws and regulations. Model weights are the files that allow developers to run, adapt and fine-tune an AI system themselves. If released as planned, H3 would add to the growing number of Chinese AI models using an open-weight approach.
The company says it designed H3 with compatibility for a range of hardware in mind, including Chinese-made chips. That focus comes as Chinese AI developers seek to reduce dependence on US semiconductor technology.
MiniMax also claims a cost advantage. It says 2K video generation costs less than one-third of comparable mainstream products on a per-second basis. At 768p, it says pricing is below half that of mainstream models producing 720p video. The company did not provide a like-for-like list of competitors or detailed pricing comparisons in its announcement.
Technically, MiniMax attributes H3’s efficiency to a new video and audio tokenizer, called H3-VAE, and an architecture that separates parts of the workload for understanding inputs from those for generating output. It says this design increased training throughput by almost 30 percent and enables 2K generation without a separate conventional upscaling model.
The company plans to publish a technical report later. It also says future H-series models will focus on stronger multimodal understanding, larger model scale and improved visual detail.
Sources
- MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities – MiniMax Research
- China’s MiniMax releases H3 video model – Reuters
Stay up to date
AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox: