Click on the Edit Content button to edit/add the content.

FLUX 3 Just Broke Out, and It Could Be the AI Model That Dethrones Seedance

Black Forest Labs has officially revealed FLUX 3, and this is much bigger than another image-generation upgrade. The new multimodal model can generate video with native audio, work from image and video references, create multilingual dialogue, edit images, and even provide the foundation for robots that operate in real factories.

Right now, FLUX 3 Video is entering early access. FLUX 3 Image, FLUX 3 Dev, API access, private weights, and action-prediction tools are planned to arrive across the coming weeks and months.

FLUX 3 latest update

DetailCurrent status
FLUX 3 announcementConfirmed
FLUX 3 VideoEarly access
Maximum single-generation lengthUp to 20 seconds
Native audioSupported
Text-to-videoSupported
Image-to-videoSupported
Video-to-videoSupported
FLUX 3 ImageEarly access planned in the coming weeks
FLUX 3 Dev open weightsPlanned
Public pricingNot announced
Full public release dateNot announced

Black Forest Labs announced FLUX 3 on July 23, 2026. The company describes it as a unified multimodal foundation model trained across images, video, and audio instead of treating each format as an isolated system.

What is FLUX 3?

FLUX 3 is a new multimodal generative AI model from Black Forest Labs, the company behind the original FLUX image models.

But calling it an image generator would seriously undersell what Black Forest Labs is building.

FLUX 3 is designed to process and generate several types of information through one underlying architecture:

  • Images
  • Video
  • Audio
  • Language instructions
  • Action prediction

The idea is that images, motion, sound, and physical actions are not separate problems. They are different ways of observing the same world.

A video shows how an object moves. Audio helps reveal what caused an event. Images preserve visual structure. Language connects all of that information to a user’s instructions.

FLUX 3 attempts to learn those relationships together, allowing one model to understand not only what something looks like, but also how it moves, sounds, and behaves.

FLUX 3 is not just FLUX 2 with better image quality

FLUX 1 and FLUX 2 were mainly known as image-generation models. FLUX 3 expands the platform into video, audio, image editing, and physical AI.

That is the real story here.

Instead of releasing one model for images and a separate model for video, Black Forest Labs is building these capabilities on the same multimodal backbone. The company says it trained FLUX 3 across video, images, and audio simultaneously, using a method called Self-Flow.

For creators, this could eventually mean fewer disconnected tools.

You could give the model a reference image, generate a video from it, maintain the same character across multiple scenes, add synchronized sound, change the visual style, and continue the sequence without rebuilding everything from scratch.

That sounds ambitious, and parts of the system are still in development. But the direction is clear. FLUX is no longer positioning itself as only an image-model family.

FLUX 3 Video can generate up to 20 seconds with native audio

The first major release is FLUX 3 Video.

According to Black Forest Labs, the model can generate videos with synchronized native audio lasting up to 20 seconds in a single generation. It supports a surprisingly wide set of workflows.

FLUX 3 Video capabilities

  • Text-to-video generation
  • Image-to-video animation
  • Image references for style and character appearance
  • Video-to-video generation
  • Video and audio continuation
  • Keyframe-to-video transitions
  • Native sound effects
  • Multilingual dialogue
  • Multiple aspect ratios and visual styles
  • Animated typography and motion designs
  • Multi-shot sequences created by chaining clips

Video-to-video could become one of the more useful additions. FLUX 3 can take important elements from a source clip, such as a character, and move them into a different environment or visual context.

Black Forest Labs also says visual references can help maintain character consistency across longer sequences made from several individual clips.

That matters because character drift is still one of the most annoying problems in AI filmmaking. A face can look correct in one shot, then turn into a slightly different person in the next.

FLUX 3 will still need proper independent testing, but it is clearly targeting that problem.

Did FLUX 3 already beat Seedance, Kling and Runway?

Black Forest Labs has published early comparison results, and this is where the announcement becomes much more aggressive.

In preliminary company-run evaluations, BFL says FLUX 3 was preferred over:

FLUX 3 already beat Seedance
  • Seedance 2.0 in 52% of comparisons
  • Gemini Omni Flash in 52%
  • Kling v3 Pro in 60%
  • Grok Imagine Video in up to 69%
  • Runway Gen-4.5 in 77%
  • Luma Ray 3.2 in 93%

The tests used 10-second, 720p text-to-video generations with audio

Those percentages look impressive, especially for a model that is still in early access. But there is an important catch.

These are preliminary results published by Black Forest Labs, not a complete independent benchmark. The company also says the model and its evaluation system are still under development.

So no, we cannot confidently say FLUX 3 has already defeated Seedance, Kling, or Runway.

What we can say is that Black Forest Labs is openly positioning FLUX 3 against the biggest AI video models from day one. It is not entering this market quietly.

FLUX 3 Image is also coming

Video is getting most of the attention, but FLUX 3 Image may still become a major release for existing FLUX users.

Black Forest Labs says the new image model can generate and edit visuals across different styles, resolutions, and aspect ratios. It reportedly performs better than previous FLUX versions when handling complicated prompts and generating accurate text.

Text generation is especially interesting.

FLUX 3 Image

AI image models have improved massively at creating readable words, but longer sentences, unusual fonts, and multilingual typography can still fall apart quickly. Black Forest Labs claims FLUX 3 can render more accurate text in several languages.

FLUX 3 Image is not broadly available yet. The company plans to open an early-access period in the coming weeks before its wider release.

FLUX 3 could eventually power robots too

Here is where things get strange.

The same FLUX 3 backbone being used for video generation is also being adapted for robot action prediction.

Black Forest Labs partnered with mimic robotics to create FLUX-mimic, a video-action model built on FLUX 3. The system has been tested and deployed on industrial tasks involving Audi production environments.

The model is not simply generating a video of a robot moving an object. It uses the visual information it has learned to help predict which physical action should happen next.

Black Forest Labs says the system has been tested on tasks including:

  • Placing parts into organized trays
  • Inserting electronic units into fixtures
  • Assembling components
  • Handling flexible seals and cables
  • Recovering after unsuccessful grasp attempts

The company’s argument is that a model capable of generating believable video already needs to learn something about motion, weight, contact, and cause and effect. That same understanding can then be adapted to control physical machines.

This part of FLUX 3 will not matter immediately to most creators. Still, it shows that Black Forest Labs is thinking far beyond another prompt-based video tool.

When will FLUX 3 fully release?

Black Forest Labs has not announced one universal public release date for every FLUX 3 model.

Instead, the company is planning a staged rollout over the next few weeks and months.

The roadmap currently includes:

FLUX 3 Video

Video and audio generation and editing through APIs and private weight access. Early access requests are already open.

FLUX 3 Image

Image generation and editing through APIs and private weight access. Early access is expected to begin in the coming weeks.

FLUX 3 Action and FLUX-mimic

Action prediction offered through selected commercial and research partnerships.

FLUX 3 Dev

An open-weight multimodal backbone intended to support image, video, audio, and action-prediction development.

The planned open-weight model could be especially important. FLUX gained a large part of its popularity because developers could run, modify, and integrate certain versions outside a completely closed platform.

However, Black Forest Labs has not yet confirmed the hardware requirements, licensing terms, exact model sizes, pricing, or final release date for FLUX 3 Dev.

Can you use FLUX 3 right now?

You can currently request early access to FLUX 3 Video, but that does not mean every user will immediately receive access.

The broader FLUX 3 product page still describes the model as coming soon, while the announcement confirms that the video model has entered an early-access phase.

There is no confirmed free public version yet.

There is also no final public pricing table for FLUX 3 Video, FLUX 3 Image, or FLUX 3 Dev. Those details should become clearer as the company expands access.

Is FLUX 3 really a Seedance killer?

It is too early to call FLUX 3 a Seedance killer.

Seedance 2.0 already has significant attention, an established user base, and strong search demand. FLUX 3 has just entered early access, and most users have not independently tested it yet.

That said, FLUX 3 has several things that could make it a serious competitor:

  • Video generation with native audio
  • Up to 20-second single generations
  • Image, video, and audio references
  • Video-to-video workflows
  • Multilingual dialogue
  • Strong typography ambitions
  • Consistent characters across chained clips
  • A planned open-weight multimodal model
  • A shared foundation for creative tools and physical AI

The biggest advantage may not be one individual feature. It may be the fact that Black Forest Labs is trying to connect all of them inside one model.

Frequently asked questions

What is FLUX 3?

FLUX 3 is a multimodal foundation model from Black Forest Labs. It is being developed for image generation, video with native audio, editing, action prediction, and physical AI applications.

Is FLUX 3 available now?

FLUX 3 Video is available through a limited early-access program. FLUX 3 Image, FLUX 3 Dev, and other capabilities are expected to roll out later.

How long can FLUX 3 videos be?

Black Forest Labs says FLUX 3 can generate videos with native audio lasting up to 20 seconds in one generation. Individual clips can also be chained into longer multi-shot sequences.

Does FLUX 3 generate audio?

Yes. FLUX 3 Video supports native audio generation, including sound effects and multilingual dialogue.

Will FLUX 3 be open source?

Black Forest Labs plans to release FLUX 3 Dev as an open-weight multimodal backbone. The final license, model requirements, and release date have not been announced.

Is FLUX 3 better than Seedance 2.0?

BFL’s preliminary evaluations slightly preferred FLUX 3 in 52% of comparisons with Seedance 2.0. Independent testing is still needed before making a final judgment.

Bottom line

FLUX 3 looks like the moment Black Forest Labs stops being known mainly for AI images.

The new model is moving into video, native audio, editing, longer sequences, image consistency, open weights, and even robot control. That is a much larger bet than simply making FLUX 2 sharper or more prompt-accurate.

Honestly, the early claims are exciting, but they are still early claims. FLUX 3 needs wider access, independent comparisons, transparent pricing, and proper real-world testing before anyone crowns it the new king of AI video.

Still, Seedance, Kling, Runway, and the rest of the market now have another serious competitor to watch.

Leave a Reply

Your email address will not be published. Required fields are marked *