Signal A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training
Summary
The paper explores Multimodal Unsupervised Continual Post-Training (MU-CPT), a task enabling deployed multimodal large language models (MLLMs) to continually evolve from streaming unlabeled data. The authors find that existing unsupervised post-training methods optimize target tokens uniformly, overlooking their heterogeneous visual dependence (VD), and show that token-level VD is crucial for MU-CPT. They propose a Visual Dependence-Aware (VDA) framework with two components: Visually Constrained Optimal Transport, which formulates VD structural distortion during new-task learning as an optimal transport problem to mitigate cross-modal forgetting, and Visually Modulated Adaptation, which exploits VD heterogeneity to promote new-task learning. The method aims to simultaneously maintain old-task stability and new-task plasticity. The paper was submitted to arXiv's cs.AI category on August 26, 2026.
Classification
Evidence 1
- A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training arXiv (cs.AI) 2026-08-26 accessed 2026-08-29T13:47:38+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-13f9cda1c094
