Black Forest Labs FLUX 3 is a multimodal video and image model

Black Forest Labs has introduced FLUX 3, a new foundation model designed to generate images and video with synchronized audio, while also serving as a base for robotics systems. The company is initially limiting access to selected users and partners through an Early Access programme.

FLUX 3 Video can create clips up to 20 seconds long from a text prompt, according to Black Forest Labs. It can also animate still images, transform existing video, continue video and audio inputs, and generate transitions between defined keyframes. Every video output includes native audio generation.

The Freiburg-based company says FLUX 3 differs from separate image, video and audio tools because it learns from all three media types within one architecture. The aim is to build a shared representation of the physical world, including objects, movement, sound and cause and effect. Language supplies instructions and higher level concepts.

One model for media and action

Black Forest Labs calls its underlying approach Self Flow. It combines multimodal generation and understanding in the same model, rather than connecting specialised systems after training. The company says it scaled its computing resources and training data to train FLUX 3 on images, video and audio simultaneously.

For content teams, the most immediate applications are video production and editing. The company lists text to video, image to video, video to video, keyframe based generation, multilingual dialogue, animated typography and a broad range of styles and aspect ratios among the model’s capabilities. It also says users can chain clips into longer sequences, using visual references to maintain character consistency across scenes.

Black Forest Labs says early testing suggests FLUX 3 is particularly capable of rendering facial expressions, matching sounds to physical events and generating content in multiple languages. Its image model is also expected to improve prompt following and text rendering compared with earlier FLUX releases. However, FLUX 3 Image is not yet available. The company plans to open an early access phase in the coming weeks.

The company has published preliminary preference tests based on 10 second, 720p text to video clips with audio. It says FLUX 3 was preferred over Luma Ray 3.2 in 93 percent of comparisons and Runway Gen 4.5 in 77 percent. It reported narrower results against Gemini Omni Flash and Seedance 2.0, where FLUX 3 received a 52 percent preference rate.

Those figures should be treated cautiously. Black Forest Labs describes the evaluation as preliminary, and VentureBeat reports that the company has not released the underlying methodology, sample sizes, rater counts or pricing. The published results also describe an earlier model candidate rather than necessarily the version entering Early Access.

Robotics is part of the strategy

FLUX 3 also forms the basis of FLUX mimic, a video action model developed with Mimic Robotics. The partners say a model that has learned motion and physical change from video may require less robot specific training data for manipulation tasks. Mimic Robotics claims some tasks can be adapted with as little as 30 minutes of robot data, though independent validation has not been published.

Black Forest Labs plans four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and FLUX 3 Dev. Video and Action are available only through approved Early Access applications. APIs, public pricing and service commitments have not been announced. The planned FLUX 3 Dev release is intended to provide open weights for a multimodal backbone, but the company has not disclosed a timeline, licence, model size or hardware requirements.

Sources

Stay up to date

AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox:

More info …

About the author

Related posts:

Advertisement

×