The FLUX image weights already run inside Adobe Photoshop, Picsart, and half the AI design tools creators pay for every month. On August 4, 2026, the lab that made those weights, Black Forest Labs, shipped its first video model: FLUX 3 Video, up to 20 seconds of 1080p footage with native, synchronized sound. That pedigree is the reason this launch matters more than the dozen other video models that dropped this summer, and it is also the reason to read the fine print before you rebuild anything around it.
Here is the honest state of FLUX 3 Video for a working creator: the idea behind it is genuinely new, and the thing you can actually use today is a gated early-access model with no published price and only the company’s own benchmarks to go on. Both of those are true at once. Let me separate what is real from what is marketing.
Who Black Forest Labs actually is
If the name does not register, the work does. Founders Robin Rombach, Andreas Blattmann, and Patrick Esser built VQGAN, latent diffusion, and Stable Diffusion, the research that kicked off the entire open-source image-generation wave. They spun out Black Forest Labs in the summer of 2024, and the company now runs a roughly 100-person team split between Freiburg, Germany and San Francisco. It is valued at about $3.25 billion and has raised more than $450 million from a16z, Salesforce Ventures, Nvidia, Adobe Ventures, and Canva.
You have probably used their output without knowing it. FLUX.1 powers Photoshop’s Generative Fill and image tools, Picsart, and a long list of platforms that quietly license the weights. FLUX.2 Dev, released in late 2025, was a 32-billion-parameter open-weight model that a lot of the Flux versus Midjourney comparisons ran on. This is not a startup trying to get noticed. It is the lab that most of your image tools already depend on, moving into video for the first time.
The one-model bet is the story, not the 20 seconds
Every video lab this year has a spec sheet. What separates FLUX 3 is the architecture underneath it. Instead of stitching together a separate image model, a separate video model, and a separate audio model, Black Forest Labs trained a single set of weights on all three modalities at once, then extended that same backbone to predict robot actions.
Rombach put the reasoning plainly: “You can’t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames.” The claim is that a model which learns motion, sound, and physical cause-and-effect together builds a more coherent internal picture of how the world behaves than one trained on frozen frames.
You can see the intent in how audio was handled. Video prediction ate more than 95% of the training compute, and audio made up less than half a percent of the tokens in a clip with sound. Despite that tiny share, audio was treated as a first-class training signal from the start, not bolted on afterward. Black Forest Labs says that shows up in timing: footsteps landing on the right frame, an object’s impact matching the moment it hits a surface, a door’s sound arriving exactly when it closes. If you have ever synced a foley track by hand, you know that timing is the whole game, and it is where most auto-generated audio still falls apart.
What FLUX 3 Video can do for a creator
Set the philosophy aside and here is the working spec. A single generation runs up to 20 seconds, roughly double the eight-to-ten-second ceiling that most tools still enforce. That is long enough for a full dialogue beat or a continuous camera move without a cut. Native resolution covers 720p and 1080p, with 2K and 4K listed as coming soon.
The input options are broad:
- Text-to-video, image-to-video, and video-to-video. Start from a prompt, a still, or existing footage.
- Keyframes. Set a first frame, a last frame, or several keyframes and let the model fill the transitions.
- Video continuation. Extend an existing clip and its audio by up to four seconds, which is how you chain shots into something longer than the 20-second cap.
- Multiple shots. Generate several scenes and camera angles inside one output.
- Draft mode. Fast, cheaper previews so you are not paying full price to find out a prompt does not work.
Audio comes with the video at no extra charge: dialogue, sound effects, and ambient noise. Dialogue is multilingual, with lip-sync across English, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, Punjabi, and more. For a creator dubbing a channel into other languages, native lip-synced dialogue inside the same generation is a real workflow saver, the kind of thing that used to mean a separate pass in another tool entirely.
The benchmarks, and why to hold them loosely
Black Forest Labs published head-to-head numbers, and they read well. In its own tests on 10-second, 720p clips with audio, FLUX 3 won most comparisons against Luma Ray3.2 and Runway Gen-4.5, and landed at roughly a coin flip against Gemini Omni Flash and Seedance 2.0, the two strongest models in the field right now.
Read that sentence again with the qualifiers intact. These are first-party numbers, labeled by the company as a preliminary evaluation of an early FLUX 3 candidate, not the shipping model. There is no independent arena data yet. I have sat through enough vendor demos in 20-plus years running IT operations to know the pattern: the benchmark that ships with the announcement is the one the vendor chose because it flatters the product. That does not make it fake. It makes it a starting hypothesis you verify on your own footage before it earns a line in your budget. Every one of those competitors is a model your audience’s other favorite creators are already shipping with, and the way FLUX 3 handles your specific prompts is the only test that counts.
The catch: gated access and no real price
This is where the enthusiasm meets procurement reality. FLUX 3 Video launched into gated early access. You request approval; you do not just sign up and generate. It is reachable through partners such as fal and invideo’s agent, but there is no open public API yet, and the general API is promised only for “later in 2026.”
More important for anyone running a business: there is no official price. Any per-second or per-clip figure floating around right now, including the roughly $0.17 per second (about $3.40 for a full 20-second clip) that some resellers have posted, is a reseller’s number, not one Black Forest Labs has published. You cannot build a content pipeline on a cost you have to guess at. When I evaluate a vendor for a fractional COO client, an unpublished price is not a detail to work around, it is a reason to keep the tool in the “watch” column and off the critical path until the number is real.
The open-weight release lands in the same “later this year” bucket. FLUX 3 Dev, an open-weight multimodal backbone you could run yourself, follows the same pattern the company used with FLUX.1 Dev and FLUX.2 Dev. If it ships the way those did, it becomes the most interesting part of this whole story for creators who want local, private, no-metered-cost generation. But it has not shipped, and BFL has not published the license terms, parameter count, or hardware requirements. Promised is not the same as available.
The robot arm is not a distraction
One footnote worth keeping, because it explains why any of this is credible. The same FLUX 3 backbone drives FLUX-mimic, an action-prediction system built with the Swiss firm Mimic Robotics. Their CTO, Elvis Nava, said a robot that used to need 30-plus hours of repeated demonstrations to learn a task can now pick it up from as little as 30 minutes of data, “because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves.”
For a creator, teaching a robot arm is irrelevant. As proof that the model actually learned physics rather than surface-level pixel patterns, it is the most convincing evidence in the launch. A model that can predict how a physical action unfolds is a model more likely to get the timing of your footsteps and your closing doors right.
Where this fits in your stack today
My read, from someone who has watched a lot of promising tools arrive before they were ready to depend on: FLUX 3 Video is worth your attention and not yet worth your workflow. Request access if you can get it, and test the audio timing and the 20-second clips against whatever you run now, whether that is a leader like Kling or Veo or a keyframe-first tool like Luma Ray3.2. If you are choosing a video tool from scratch, our full AI video guide covers where each one actually fits. Treat any output as a preview of where this goes, not a production dependency, until the price is published, the general API opens, and independent benchmarks exist.
The larger signal is the one to actually plan around. The lab whose weights already sit inside the image tools you pay for is now training image, video, and audio as one system, and open-weighting the result on its usual schedule. Whether or not FLUX 3 Video wins the summer, that approach is going to shape the tools you use next, the same way FLUX.1 quietly shaped the ones you use now. Keep a slot open for it, keep your current pipeline running, and let someone else’s budget prove the benchmarks first.
Recent Posts
Lottie Creator 2.0 Turns a Text Prompt Into a Shippable Web Animation
LottieFiles' Lottie Creator 2.0 puts a full motion studio in the browser, builds animations from a plain-language prompt with Motion Copilot, and exports interactive files that load faster than a...
Grok Voice Think Fast 2.0 Talks Back in Seven-Tenths of a Second
xAI's new speech-to-speech model answers in 0.70 seconds and reasons while it talks. A practical look at what real-time voice AI does for creators, and where it still falls short.
