Skip to main content

Black Forest Labs unifies image, video, audio and action in FLUX 3

By Steven Van ·

FLUX 3 Video is in early access now, jointly trained with audio and, per the lab, extendable to robotics action prediction.

Black Forest Labs, the team whose FLUX models set the bar for open image generation, just made a much bigger move. FLUX 3 is one multi-modal model for image, video, audio, and action prediction, jointly trained in a single unified architecture, with creations the lab says are truer to life across every style. FLUX 3 Video is available now in early access.

From best-in-class image model to unified media model

FLUX earned its reputation on images, some of the most-used open and commercial image models available. FLUX 3 is the leap from a great image model to a single model that spans modalities. Instead of a separate image model, video model, and audio model stitched together, it's one architecture trained jointly across all of them. That matters because joint training lets the modalities reinforce each other: a model that understands how the world looks, sounds, and moves in one shared representation can produce results that are more coherent across image, video, and audio than a pipeline of specialists.

The action-prediction twist

The genuinely surprising piece is action prediction. Black Forest Labs says the same unified model can be extended to predict actions for robotics, and points to work with mimic and Audi. That reframes FLUX 3 from a media generator into something closer to a world model, one that doesn't just render scenes but reasons about what happens next in them. For creators, the near-term payoff is better generative media; for the field, it's a signal that the line between generative media models and embodied AI is blurring.

Why it matters for creators now

The immediate, usable news is FLUX 3 Video in early access. Given FLUX's track record on image fidelity and prompt adherence, a video model built on the same lineage, and jointly trained with audio, is worth testing against the current crop. "Truer to life in every kind of style" is the claim; the test is whether it holds across the specific looks you actually work in.

Try it

FLUX 3 Video early access is available now via Black Forest Labs. If you already use FLUX for images, this is the natural moment to see whether one unified model can carry your work from still to motion, with sound, in a single system.

Black Forest Labs
Black Forest Labs
The lab behind FLUX — frontier models for image, and now video, audio, and action in one multi-modal architecture.
View Black Forest Labs →

Sources: Black Forest Labs on X, Black Forest Labs.

Read Black Forest Labs unifies image, video, audio and action in FLUX 3 on Creators Toolbox