Moonworks Lunara: Modeling Artistic Intelligence

Moonworks Lunara: Modeling Artistic Intelligence

Moonworks Lunara:艺术智能建模

Abstract: We formulate \emph{Artistic Intelligence} as exploration driven world realization, leaving space for creative possibility while preserving the semantic, artistic, and compositional structure that must remain true. Moonworks Lunara, a text-to-image model, implements this framework with a novel Diffusion Mixture Transformer architecture.

摘要: 我们将“艺术智能”(Artistic Intelligence)定义为一种由探索驱动的世界实现过程,它在保留必须遵循的语义、艺术和构图结构的同时,为创造性可能性留出了空间。Moonworks Lunara 是一款文本生成图像模型,它通过一种新颖的扩散混合 Transformer(Diffusion Mixture Transformer)架构实现了这一框架。

A new training algorithm iteratively evolves the data distribution through informative sample acquisition and targeted injection of human-created art. We benchmark Lunara against seven image-generation models, including FLUX.2-Klein-4B, Qwen-Image (20B), and GPT-Image-1-Mini.

一种新的训练算法通过信息样本采集和针对性地注入人类创作的艺术作品,迭代地演化数据分布。我们将 Lunara 与七种图像生成模型进行了基准测试,包括 FLUX.2-Klein-4B、Qwen-Image (20B) 和 GPT-Image-1-Mini。

With GPT-5.6 Sol as evaluator, Lunara ranks first in \emph{Aesthetic Quality (8.473 vs. 8.457 GPT-Image-1-mini)}, second in \emph{Emotional Resonance}, and remains competitive in \emph{Content Integrity}. A blind human evaluation over the same evaluation set corroborates the automated metrics, ranking Lunara first. It also stays among the strongest models under conventional measures including CLIPScore and LAION Aesthetic Predictor.

在以 GPT-5.6 Sol 作为评估器的情况下,Lunara 在“审美质量”(8.473 vs. 8.457 GPT-Image-1-mini)方面排名第一,在“情感共鸣”方面排名第二,并在“内容完整性”方面保持竞争力。针对同一评估集的盲测人类评估证实了自动化指标的结果,将 Lunara 排在首位。在包括 CLIPScore 和 LAION 审美预测器在内的传统衡量标准下,它也保持在最强模型之列。

On GenEval, Lunara achieves competitive performance against a broader set of 16 models, including GPT Image 2 and Seedream 4.0. These results place Lunara at the frontier with Artistic Intelligence while maintaining a sub-10B active-parameter footprint and sub-10-second inference latency.

在 GenEval 测试中,Lunara 在与包括 GPT Image 2 和 Seedream 4.0 在内的更广泛的 16 个模型对比中表现出竞争力。这些结果使 Lunara 处于艺术智能的前沿,同时保持了低于 100 亿(sub-10B)的活跃参数规模和低于 10 秒的推理延迟。

Lunara advances the general visual intelligence frontier by shifting the question from whether models can get images right to how deeply they can interpret meaning and realize it as imaginative, expressive worlds.

Lunara 通过将研究重点从“模型能否生成正确的图像”转向“模型能多深刻地解读意义并将其实现为富有想象力和表现力的世界”,推动了通用视觉智能的前沿发展。