AUTO-UPDATED

Flux 3 X Mimic: The Next Generation of Video-Action Models

Black Forest Labs has introduced FLUX 3, a multimodal foundation model that powers FLUX-mimic, a new video-action system enabling robots to perform complex manipulation tasks in industrial environments.

Key Points

  • FLUX 3 is a multimodal model trained jointly on images, video, and audio to develop a deep understanding of physical world dynamics.
  • The model powers FLUX-mimic, a collaboration with mimic robotics that enables robots to handle flexible materials and complex assembly.
  • Audi is currently testing and deploying FLUX-mimic robots for production and logistics tasks, including handling cables and electronic control units.
  • The system achieves human-like reaction times, with the backbone processing inputs in under 80ms on a single NVIDIA RTX 5090 GPU.
  • FLUX-mimic demonstrates high sample efficiency, requiring significantly less demonstration data than traditional vision-language-action models to learn new tasks.

Why it Matters

By unifying generative content creation and physical robot control within a single foundation model, Black Forest Labs is shifting the paradigm for industrial automation. This approach allows robots to adapt to complex, variable tasks without the need for costly, manual re-engineering, potentially expanding the scope of flexible automation in manufacturing.
Bfl.ai Published by Black Forest Labs
Read original