Black Forest Labs unifies image, video, audio and action in FLUX 3
By Steven Van ·
FLUX 3 Video enters early access, with the same architecture extendable to action prediction for robotics.
FLUX 3 from Black Forest Labs is a single multi-modal model covering image, video, audio, and action prediction, trained jointly in one unified architecture rather than as separate models stitched together. Black Forest Labs says the same architecture can be extended to predict actions for robotics.
FLUX 3 Video is now in early access.
Black Forest Labs
The lab behind FLUX — frontier open and commercial models for image, and now video, audio, and action prediction in one multi-modal architecture.
View Black Forest Labs →