MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models
Computer Science > Computer Vision and Pattern Recognition arXiv:2610.08830 (cs) [Submitted on 27 Sep 2026] Title: MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models Authors: Pengcheng Zheng, Chaoning Zhang, Jiaxin Yan, Sihan Cao, Jianwei Zhang, Xudong Wang, Jiaquan Zhang, Jewon Lee, Tae-Ho Kim, Yang Yang, Heng Tao Shen.
计算机科学 > 计算机视觉与模式识别 arXiv:2610.08830 (cs) [提交于 2026 年 9 月 27 日] 标题:MoR-MLLM:用于高效多模态大语言模型的递归混合架构 作者:Pengcheng Zheng, Chaoning Zhang, Jiaxin Yan, Sihan Cao, Jianwei Zhang, Xudong Wang, Jiaquan Zhang, Jewon Lee, Tae-Ho Kim, Yang Yang, Heng Tao Shen。
Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across vision and language tasks. However, their massive computational and memory demands hinder real-world deployment. While recent efforts reduce costs by employing lightweight language backbones, existing paradigms remain computation-dense due to their static sparsity and depth allocation, which cannot adapt to the semantic complexity of each token.
摘要:多模态大语言模型(MLLMs)在视觉和语言任务中展现出了卓越的推理能力。然而,它们巨大的计算和内存需求阻碍了实际部署。尽管近期的研究通过采用轻量级语言主干来降低成本,但现有的范式由于其静态稀疏性和深度分配,仍然属于计算密集型,无法适应每个标记(token)的语义复杂性。
To this end, we propose MoR-MLLM, a computation-sparse MLLM based on the recent Mixture-of-Recursions (MoR) framework. MoR-MLLM introduces adaptive per-token recursion, allowing the model to dynamically adjust its recursive depth and allocate more computation to visually or linguistically challenging tokens while skipping redundant operations for simpler ones.
为此,我们提出了 MoR-MLLM,这是一种基于近期“递归混合”(Mixture-of-Recursions, MoR)框架的计算稀疏型 MLLM。MoR-MLLM 引入了自适应的逐标记递归,允许模型动态调整其递归深度,并为视觉或语言上具有挑战性的标记分配更多计算资源,同时跳过简单标记的冗余操作。
To stabilize the training of recursive sparsity in multimodal settings, we further design a three-stage MoR-Tuning strategy and an entropy-regularized loss to encourage diverse routing distributions. Extensive experiments show that compared with recent advanced tiny MLLMs, our proposed MoR-MLLM can greatly reduce the training memory and computation complexity while retaining high performance on various vision-language tasks.
为了稳定多模态环境下递归稀疏性的训练,我们进一步设计了三阶段的 MoR 微调策略和熵正则化损失函数,以鼓励多样化的路由分布。大量实验表明,与近期先进的微型 MLLM 相比,我们提出的 MoR-MLLM 可以在保持各种视觉-语言任务高性能的同时,大幅降低训练内存和计算复杂度。