MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video An open-weights omni-modal video model with real stereo sound and 2K output — this powerful model is greatly optimized in ComfyUI and can run locally on a 3060.

ComfyUI 实现 MiniMax H3 首日支持:开放权重、原生音频与 2K 视频 这是一款具备真实立体声和 2K 输出能力的开放权重全模态视频模型——该强大模型已在 ComfyUI 中进行了深度优化,甚至可以在 3060 显卡上本地运行。


MiniMax H3 dropped today with open weights, and it’s natively supported in ComfyUI as of this morning. Day zero.

MiniMax H3 于今日发布并开放权重,且从今天早上起已在 ComfyUI 中获得原生支持。实现“首日即用”。

This is a next-generation open-weights video model. Feed it text, images, video, or audio and it generates video with real stereo sound, up to 2K, up to 15 seconds a clip. It is MiniMax’s third-generation video model, following Hailuo 01 and Hailuo 02, and the first the company has released with open weights.

这是一款新一代开放权重视频模型。只需输入文本、图像、视频或音频,它就能生成带有真实立体声、最高 2K 分辨率、单段时长达 15 秒的视频。这是 MiniMax 继 Hailuo 01 和 Hailuo 02 之后的第三代视频模型,也是该公司首个发布开放权重的模型。

Model Highlights / 模型亮点

  • Text-to-video — prompt only. 文生视频 — 仅需提示词。
  • Image-to-video — bring an image to life. 图生视频 — 让静态图像动起来。
  • First-and-last-frame — control the opening frame, the closing frame, or both, and let the model fill in the rest. 首尾帧控制 — 控制起始帧、结束帧或两者,让模型填充中间内容。
  • Reference-to-video — supply reference images, video, or audio and carry a subject, a motion, or a voice through the clip. 参考视频生成 — 提供参考图像、视频或音频,将主体、动作或声音贯穿整个片段。

Output runs to 2K and up to 15 seconds. Audio is generated with the video in the same pass, in stereo, not bolted on afterward.

输出分辨率最高可达 2K,时长最长 15 秒。音频与视频在同一过程中同步生成,且为立体声,而非后期合成。

Multimodal context understanding / 多模态上下文理解

This is the capability MiniMax leads with, and it’s what collapses five separate tasks into one model. Real work rarely draws on one modality. H3 takes images, audio, and video together and resolves them against a prompt that explains how they relate. Describe the relationship between your inputs and the shot you want, and the model handles the cross-modal work itself.

这是 MiniMax 的核心领先能力,它将五个独立的任务整合进了一个模型中。实际工作很少只涉及单一模态。H3 可以同时处理图像、音频和视频,并根据解释它们之间关系的提示词进行解析。你只需描述输入内容与目标镜头之间的关系,模型便会自动处理跨模态的复杂工作。

Native stereo audio / 原生立体声音频

Audio is a property of the model, not a post-process. Every audio output is native stereo.

音频是模型的一种属性,而非后期处理。每一个音频输出都是原生的立体声。

Editing and motion transfer / 编辑与动作迁移

Motion transfer is the one that matters most for graph work. A reference video can supply movement — a camera move, a performance, a cutting rhythm — while the subject and style come from elsewhere. Combined with in-place editing, that means iterating on a shot.

动作迁移是图形化工作流中最关键的部分。参考视频可以提供运动轨迹——如运镜、表演或剪辑节奏——而主体和风格则可以来自其他来源。结合原位编辑功能,这意味着你可以对镜头进行反复迭代。