FLUX 3: Black Forest Labs' Multimodal AI Model
Black Forest Labs unveiled FLUX 3, a multimodal frontier model that generates image, video, audio, and robot actions from one network. What it does and who it's for.
The company behind one of the most-used open image models is trying to collapse an entire toolbox into a single network. On July 23, 2026, Black Forest Labs (BFL) unveiled FLUX 3, which it describes as a multimodal frontier model trained to understand and generate images, combined audio-and-video clips, and even robot actions — all from one architecture. The launch tagline framed the ambition plainly: “a breakthrough in control, realism, and world understanding — one multimodal model generating image, video, audio and action.”
The release marks a step-change for a startup best known for text-to-image. Where earlier FLUX models produced still images, FLUX 3 is pitched as a general “visual intelligence” system that treats pixels, sound, and physical action as facets of the same problem.
What FLUX 3 does
FLUX 3 generates across four modalities from a single prompt or input:
- Image — still-image generation and editing, the category BFL built its reputation on.
- Video — combined audio-and-video clips of up to 20 seconds generated together, rather than silent footage that needs a separate soundtrack.
- Audio — sound produced jointly with video, so speech, effects, and ambience are aligned to the visuals.
- Action — an extension of the same backbone into robotic vision and action prediction, aimed at controlling physical systems.
The pitch behind bundling these into one network is coherence: a model that learns a shared representation of the visual world should, in theory, keep a character’s appearance, a scene’s lighting, and an object’s physics consistent as it moves between generating a frame, a clip, and the audio that accompanies it. That “world understanding” framing is BFL’s answer to the fragmentation of today’s stack, where teams stitch together one model for images, another for video, and a third for voice.
Like earlier FLUX releases and other systems in the category — including Meta’s Muse Image — FLUX 3 builds on the diffusion lineage its founders helped invent. What is new is the scope: rather than one task done well, BFL is claiming a single multimodal model that spans generation and, with the action head, embodied control.
A phased, gated rollout
FLUX 3 is not shipping all at once, and it is not fully open on day one. BFL laid out a staged release:
- FLUX 3 Video is available now through gated early access.
- FLUX 3 Image is rolling out “in the coming weeks.”
- FLUX 3 Action / FLUX-mimic, the robotics variant, is offered as early access to selected research and commercial robotics partners rather than the general public.
- An open-weight FLUX 3 Dev backbone is planned for later.
That sequencing matters. BFL’s earlier models earned their following in part because open weights let developers fine-tune and self-host — a continuation of the open-source diffusion community the founders came from. FLUX 3 launches the other way around: the frontier capabilities arrive first behind gates and early access, with the downloadable Dev backbone promised for a later date. A placeholder page for the model briefly surfaced at BFL’s site on July 21, two days before the formal reveal.
The gated approach mirrors how the video-generation leaders have handled their most capable systems, releasing through controlled access before wider availability — a nod both to compute costs and to the safety questions that come with realistic synthetic audio and video.
Who’s behind Black Forest Labs
Black Forest Labs was founded in 2024 by Robin Rombach, Andreas Blattmann, and Patrick Esser — researchers who were central to the latent diffusion work that powered Stable Diffusion during their time at Stability AI. That pedigree is why FLUX has been treated from the start as a continuation of the open-weight image-model lineage rather than a newcomer.
The company has raised aggressively to fund the jump into video and multimodal training. In December 2025, BFL closed a $300 million Series B at a $3.25 billion post-money valuation, co-led by Salesforce Ventures and Anjney Midha of Andreessen Horowitz, with participation from a16z, NVIDIA, General Catalyst, and Temasek. The Series B disclosure also revealed a previously unannounced Series A led by a16z, bringing total funding above $450 million. NVIDIA’s presence on the cap table is notable given how much accelerated compute a jointly trained image-video-audio-action model demands.
The competitive picture
FLUX 3 pushes BFL directly into the most contested corner of generative AI. On video, the benchmark competitors are the flagship systems from Google and OpenAI, which defined text-to-video quality and set expectations for clip length, motion, and realism. On images, BFL now competes not only with those labs but with platform players such as Meta, which has been folding image and video generation directly into consumer apps.
BFL’s differentiator is the “one model, many modalities” thesis. Rivals largely ship specialized systems; FLUX 3 bets that a unified backbone — one that can also drive robotics through the FLUX-mimic action head — will generalize better and be cheaper to maintain than a fleet of task-specific models. That is a strong claim, and one the phased, partly-gated launch makes hard to independently verify at release: full public benchmarks await the wider image rollout and the eventual open-weight Dev backbone.
What it means
FLUX 3 is Black Forest Labs’ bid to matter beyond images, and the framing is telling. By packaging video, audio, and robot action into a single “visual intelligence” model, BFL is arguing that the future of generative media is not a stack of specialized tools but one world model that can render a scene, voice it, and act in it. If that thesis holds, it pressures every lab shipping single-purpose generators to justify why their model does only one thing.
The winners to watch are developers and studios who want a coherent image-to-video-to-audio pipeline without gluing three vendors together, and NVIDIA, whose hardware such training and inference runs on. The robotics angle is the wildcard: an action head that shares a backbone with a top-tier video model could give humanoid and industrial-robot builders a cheaper path to perception and control than training from scratch.
The risks are equally real. Twenty-second clips with synchronized, generated audio raise the bar for convincing synthetic media, which is why the gated rollout and delayed open weights read as much as caution as strategy. And BFL’s open-source reputation cuts both ways: shipping the frontier behind gates first, with the downloadable Dev model “later,” is a departure that some of its community will notice.
What to watch next: whether FLUX 3 Image ships on schedule in the coming weeks with public benchmarks against the leading image and video models; when the open-weight Dev backbone actually lands and how capable it is relative to the gated tiers; and whether any of the robotics partners using FLUX-mimic show real-world results that validate the single-backbone bet. For now, FLUX 3 is a bold architectural statement from a well-funded team — and a test of whether one model really can do it all.
Tagged
Keep reading
Chisato · · 6 min read Alibaba Wan-Animate-2: Open-Source Real-Time AI Animation
Alibaba's Tongyi Lab open-sourced Wan-Animate-2, a character-animation model that streams at 24fps under Apache 2.0. What it does and why it matters.
Chisato · · 6 min read Seedance 2.5: ByteDance's 30-Second AI Video Model
ByteDance opened public API access to Seedance 2.5, a model that generates 30-second single-shot clips with native audio. What it does and why it matters.
Chisato · · 6 min read Meta Muse Image: Superintelligence Labs' First Model
Meta launched Muse Image, its first in-house AI image model, across Instagram and WhatsApp — with an invisible watermark and an immediate privacy backlash.