The new FLUX 3 multimodal foundation model has entered early access, offering unified generation capabilities for video, audio, and images while supporting advanced physical AI and action prediction.
Key Points
- FLUX 3 utilizes a unified architecture to jointly process and generate images, videos, and audio simultaneously.
- The model supports diverse video features, including text-to-video, image-to-video, and native audio generation for clips up to 20 seconds.
- Early benchmarks show FLUX 3 outperforming competitors like Runway Gen-4.5 and Luma Ray 3.2 in human preference evaluations.
- The company is partnering with mimic robotics to develop FLUX-mimic, a specialized model for dexterous robot manipulation and physical AI.
- Future releases will include API access for video, image, and action prediction, alongside open-weight access for the FLUX 3 Dev model.