Flux 3
Flux 3
FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence. FLUX 3 is now available in Early Access. FLUX 3 - 现实世界模型:迈向作为视觉智能基石的多模态流模型。FLUX 3 现已开放抢先体验。
FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound. FLUX 3 是我们全新的多模态基础模型。它在统一的架构内联合学习图像、视频和音频,因为模型需要学习的并非这些元素中的任何单一孤立部分。相反,模型必须学习对世界的表征:物体如何结合、事物如何运动,以及事件听起来是什么样的。
No single modality provides a complete description. Each is a projection of the same underlying reality, captured by different sensors, each of which loses some information in the process. Images capture spatial structures and relationships at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. Language links these perceptions to goals, abstractions, and instructions. 没有任何单一模态能提供完整的描述。每种模态都是同一底层现实的投影,由不同的传感器捕获,而每个过程都会丢失部分信息。图像捕捉特定时间点的空间结构和关系;视频恢复了时间维度,揭示了时间动态和物理定律;音频揭示了仅靠视觉无法检测到的机械现象与声学之间的因果关系;语言则将这些感知与目标、抽象概念和指令联系起来。
Learn from one and you get a good model of that projection. Learn from all of them at once and their mutual constraints tell you more: the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate and start being evidence about one underlying reality. 只学习一种模态,你只能得到该投影的良好模型。同时学习所有模态,它们之间的相互约束会告诉你更多信息:声音必须与撞击相匹配,运动必须遵循质量规律,未来必须承接过去。模态不再是孤立的,而是成为了关于同一底层现实的证据。
FLUX 3 is our first model built entirely on that principle, and a checkpoint on our mission to develop real-world visual intelligence: models that perceive, predict, and act across physical and digital environments. Early results in content creation and physical AI suggest it is the right path. FLUX 3 是我们首个完全基于该原则构建的模型,也是我们开发现实世界视觉智能使命中的一个里程碑:即在物理和数字环境中感知、预测并行动的模型。在内容创作和物理 AI 领域的早期成果表明,这是一条正确的道路。
FLUX 3: One model, multiple capabilities. FLUX 3:一个模型,多种能力。
FLUX 3 builds on Self-Flow, our approach for efficiently aligning multimodal generation and understanding within the same underlying architecture. Based on this approach, we significantly scaled up compute and data resources to train FLUX 3 across video, images, and audio at the same time. FLUX 3 基于 Self-Flow 构建,这是我们用于在同一底层架构内高效对齐多模态生成与理解的方法。基于此方法,我们大幅扩展了计算和数据资源,以同时跨视频、图像和音频训练 FLUX 3。
Capabilities & Early Evaluations 能力与早期评估
As a result, FLUX 3 is capable of mixing modalities and generating images and video+audio jointly; both from pure text prompts as well as when providing input references such as images and video. We are highlighting a few of the model’s key capabilities below. 因此,FLUX 3 能够混合模态并联合生成图像和视频+音频;无论是通过纯文本提示,还是在提供图像和视频等输入参考时均可实现。我们在下方重点介绍了该模型的一些关键能力。
Video 视频
FLUX 3 can create highly diverse videos with audio up to 20 seconds in length in a single generation. Its core capabilities include the following (all outputs come with native audio generation): FLUX 3 可以在单次生成中创建长达 20 秒、包含音频且高度多样化的视频。其核心能力包括以下内容(所有输出均自带原生音频生成):
- Text-to-video generation. (文本生成视频。)
- Image-to-video generation, either continuing from a starting frame (“animation”) or using images as visual references. (图像生成视频,既可以从起始帧延续(“动画”),也可以使用图像作为视觉参考。)
- Video-to-video generation from a reference clip, carrying central elements of a source video - for instance the same character - into a new scene or context. (基于参考片段的视频生成视频,将源视频的核心元素——例如同一个角色——带入新的场景或语境中。)
- Generative video-audio continuation from input video and audio. (基于输入视频和音频的生成式视频-音频延续。)
- Keyframe-to-video generation for controlled transitions between defined moments. (关键帧生成视频,用于在定义的时刻之间进行受控过渡。)
- Multilingual dialogue. (多语言对话。)
- A broad range of visual styles and aspect ratios, extending far beyond conventional cinematic output. (广泛的视觉风格和纵横比,远超传统的电影输出。)
- Agentic chaining of individual clips into longer, multi-shot sequences. (将单个片段智能链接成更长的多镜头序列。)
- High style diversity — FLUX 3 Video easily handles ranges of styles from candid camcorder footage to animation and cinematics. (高风格多样性——FLUX 3 Video 可以轻松处理从随手拍的摄像机素材到动画和电影级的各种风格。)
- Strong typography generation and animated designs. (强大的排版生成和动画设计。)
Evaluations are early and we expect further improvements. As the model and the harness around it are still in development, these results are preliminary, and we expect further improvements during the early access phase. 评估尚处于早期阶段,我们期待进一步的改进。由于模型及其配套工具仍在开发中,这些结果仅为初步数据,我们预计在抢先体验阶段会有进一步的提升。
While still in development, FLUX 3 Video is already particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities. Furthermore, these capabilities can be combined to create sequences lasting several minutes, where visual references help ensure that the characters remain consistent across all scenes. 尽管仍处于开发阶段,FLUX 3 Video 在捕捉人类面部表情、将声音与物理事件关联以及多语言能力方面已经表现得非常出色。此外,这些能力可以结合起来创建持续数分钟的序列,其中视觉参考有助于确保角色在所有场景中保持一致。
Image 图像
FLUX 3 can synthesize and edit images in a wide variety of styles, aspect ratios, and resolutions. In preliminary evaluations conducted during midtraining, FLUX 3 already shows a significant improvement over earlier versions of FLUX: its ability to handle complex prompts and text generation has improved significantly. FLUX 3 可以合成和编辑各种风格、纵横比和分辨率的图像。在训练中期进行的初步评估中,FLUX 3 已经显示出比早期 FLUX 版本显著的进步:其处理复杂提示词和文本生成的能力有了显著提高。
Action 行动
FLUX 3’s world understanding extends to action prediction. We have taken two routes to it: integrating native action prediction into FLUX 3 directly, scaling up our initial work in Self-Flow; and using the pretrained video backbone as a dynamics-aware foundation that specialized action models can be finetuned from with limited task-specific data. FLUX 3 对世界的理解延伸到了行动预测。我们采取了两种途径:一是将原生行动预测直接集成到 FLUX 3 中,扩展我们在 Self-Flow 方面的初步工作;二是利用预训练的视频骨干网络作为具备动态感知的基础,通过有限的任务特定数据对专门的行动模型进行微调。