FLUX 3 makes image, video and audio from one model
Black Forest Labs — the lab behind the open-weight Flux image models — has unveiled FLUX 3, and it''s a bigger swing than a new image model. FLUX 3 generates image, video, audio, and even action-prediction from a single set of weights, rather than stitching separate models behind one interface the way most of the field still does.
The pitch is a unified architecture: one network jointly trained across modalities, which the lab frames as "a breakthrough in control, realism, and world understanding." In practice it generates video with native audio, edits images, renders readable text, and can predict robot actions — the last hinting at ambitions well beyond content creation and into physical AI.
The rollout, and why it matters
Access is staged. Video and Action opened in gated early access on day one (via API and private weights to selected partners); image generation lands "in the coming weeks"; and the open-weight Dev release comes last, planned for later in 2026. No pricing has been announced yet.
That open-weight promise is the important part. If FLUX 3 Dev ships as planned, it would be the first publicly available open-weight model that jointly generates video, audio, and images from one architecture — the kind of thing that lets teams self-host a multimodal pipeline instead of renting four separate APIs. For anyone who chose Flux for its openness, this is the roadmap getting a lot more ambitious. We track new model drops in the Latest in AI feed.