Flux 3
FLUX 3 is a multimodal foundation model that learns from images, videos, and audio within a unified architecture. It aims to create a comprehensive representation of the world by understanding object interactions, movements, and corresponding sounds.