Strike a Chord! Modal Kinetic Typography
Strike a Chord! Modal Kinetic Typography
奏响和弦!模态动力学排版
Abstract: We introduce modal kinetic typography, which animates a vector glyph to express a semantic concept while keeping it legible. Our key idea is to build motion from the glyph’s natural vibration modes. Specifically, a finite-element eigenproblem assembled from the vector outline yields the glyph’s softest modes, for the whole letter and for each of its parts, allowing it to bend. The problem’s zero-energy solutions, i.e., rigid translations and rotations, are applied in closed form to each part, allowing parts to also move as blocks.
摘要: 我们引入了模态动力学排版(Modal Kinetic Typography),它通过对矢量字形进行动画处理来表达语义概念,同时保持其可读性。我们的核心思想是基于字形的自然振动模态来构建运动。具体而言,通过从矢量轮廓构建的有限元特征值问题,我们得出了整个字母及其各个部分的“最软”模态,从而使其能够弯曲。该问题的零能量解(即刚性平移和旋转)以闭式形式应用于每个部分,使各部分也能作为整体块进行移动。
To animate the glyph, a frozen video diffusion model supervises only the modes’ amplitudes and phases. Our modal approach addresses two weaknesses of prior work. Free-form point optimization under video score distillation (SDS) moves each point and frame independently along noisy gradients, tearing the outline and causing jitter. In contrast, our modes are smooth along the outline and driven by a few whole-cycle harmonics, which restricts these gradients to smooth, seamlessly looping motion.
为了对字形进行动画处理,一个冻结的视频扩散模型仅负责监督模态的振幅和相位。我们的模态方法解决了先前工作中的两个弱点。在视频分数蒸馏(SDS)下的自由形式点优化会沿着噪声梯度独立移动每个点和帧,导致轮廓撕裂并产生抖动。相比之下,我们的模态沿着轮廓是平滑的,并由少数全周期谐波驱动,这限制了这些梯度,从而产生平滑且无缝循环的运动。
On the other hand, structured alternatives rely on skeletons or keypoints from category-specific priors, whereas our modes come from the glyph itself; the only prior is a list naming each letter’s moving parts, generated once for the whole alphabet by a language model. In modal kinetic typography, shape and motion are disentangled by construction: a single base outline is sculpted toward the concept, and the modal drive cannot alter it, so a letter can also be animated without being reshaped.
另一方面,结构化的替代方案依赖于特定类别的先验骨架或关键点,而我们的模态直接源自字形本身;唯一的先验信息是一个列出每个字母运动部件的列表,该列表由语言模型为整个字母表生成一次。在模态动力学排版中,形状和运动在构建时即已解耦:单一的基础轮廓被雕刻以贴合概念,而模态驱动不会改变它,因此字母可以在不改变形状的情况下进行动画处理。
Our method produces more articulated and smoother motion than Dynamic Typography and AniClipart at comparable or better concept alignment, with less glyph tearing than Dynamic Typography, and is preferred by human raters, including in a frozen-shape setting where motion alone must carry the concept. Our results were also preferred over Astra (GPT-6) by human raters.
与 Dynamic Typography 和 AniClipart 相比,我们的方法在概念对齐程度相当或更好的情况下,产生了更具表现力且更平滑的运动,且字形撕裂现象少于 Dynamic Typography。在人类评估中,包括在仅靠运动来传达概念的“固定形状”设置下,我们的方法更受青睐。此外,人类评估者也更倾向于我们的结果,而非 Astra (GPT-6)。