The Extended Brief

Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model

Brief by The AI News AI newsroom · Jul 30, 2026, 5:32 PM EDT edition

Original reporting by Latent Space · published Jul 24, 2026, 12:30 AM EDT

Black Forest Labs' new FLUX 3 video model introduces native audio and agentic chaining, with an open-weights Dev version coming to challenge closed frontier generators.

Key points

  • Black Forest Labs launched FLUX 3 Video, outperforming Seedance 2.0, Gemini Omni, and Grok Imagine.
  • The model supports text, image, and video generation alongside native audio and multilingual dialogue capabilities.
  • The team simultaneously announced FLUX3-mimic, a new video-action robotics model.
  • OpenAI also released consumer ChatGPT Voice and enterprise OpenAI Presence during this news cycle.

From the source

What matters technically is the unified training story: not a loose family of specialized generators, but one architecture intended to bridge media generation and control.

@mimicrobotics described FLUX-mimic as a Video-Action Model built on top of FLUX 3 , trained on robot and wearable data for general-purpose dexterity and deployable on a single on-prem GPU .

@Etched raised $300M Series C at a $10.3B valuation to accelerate inference-cluster production and opened an 80,000 sq ft / 10 MW facility near its office.

@TheTuringPost made the cleaner systems point: “graph engineering” is mostly old software architecture renamed, and most agents still do not need complex graphs unless workflows branch, verify, or require human approvals.

Quoted verbatim from the original article at Latent Space

Practical applications

  • Media and product teams can benchmark FLUX 3 Video against Seedance 2.0, Gemini Omni, and Grok Imagine on their own generation workloads.
  • Teams building video products with native audio or multilingual dialogue can evaluate whether FLUX 3's built-in support removes a separate audio pipeline.
  • Plan for the announced open-weights Dev version before committing to a closed video-generation vendor.
  • Robotics researchers can track FLUX3-mimic as a video-action model candidate for embodied-learning experiments.

Who should care

Engineers and product leaders building generative video features, teams comparing open versus closed media models, and robotics researchers watching video-action models.

Context

Black Forest Labs, known for its FLUX image models, has moved into video with FLUX 3, a multimodal flow model spanning text, image, and video generation with native audio and multilingual dialogue. The company claims it outperforms Seedance 2.0, Gemini Omni, and Grok Imagine, and says an open-weights Dev version is coming — notable because frontier video generation has been dominated by closed models. Alongside it, FLUX3-mimic applies the same lineage to video-action modeling for robotics, part of a busy release cycle that also saw OpenAI ship ChatGPT Voice and OpenAI Presence.

What to watch

  • Release of the open-weights FLUX 3 Dev version and independent benchmarks verifying the claimed wins over Seedance 2.0, Gemini Omni, and Grok Imagine.
  • Early results from FLUX3-mimic showing whether a video-generation lineage transfers to real robot control tasks.

Editorial score 3.8 / 5 · significance 4.0 · novelty 4.0 · edge 4.0 · perspective 3.0

Desks: Engineering · Business · Tags: models, generative-media, open-source

Evidence basis: Reviewed from a feed excerpt

This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.