Hacker News· gmays·· 4 小时前AI 评分35
UniEvo-VL:面向多模态模型自改进的自蒸馏训练
UniEvo-VL: Self-Distillation Training for Multimodal Model Self-Improvement
AI 导读
UniEvo-VL 是一种面向多模态模型自进化的自蒸馏训练框架,让单一多模态模型在测试时同时充当教师和学生,基于开源 Qwen-image-2512 将 GenEval 分数从 0.747 提升到 0.808、GenEval2 Soft-TIFA 从 32.97 提升到 35.53。
正文
Authors:Fang Wu, Da Xing, Yanjie Huang, Junxi Wang, Ji Wang, Hejia Geng, Guancheng Wan, Bowen Zuo, Xiaomin Li, Shixiang Tang, Xinyu Xiang, Zehong Wang, Shiyi Du, Peng Xia, Shuangjia Zheng, Yining Hong, Li Erran Li, Jure Leskovec, Yejin Choi
Abstract:Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute. Instead of relying on a separate, often larger, teacher, we leverage their self-critiques as privileged information and ask a single multimodal model to act as both teacher and student with different contexts. The student only sees the vanilla question, while the teacher conditions on the privileged critique. Then training minimizes the per-state divergence between their denoising diffusion distributions over the student's own sampling trajectories. Experiments demonstrate that UniEvo-VL improves the image generation capabilities of multimodal models, while maintaining their sensitivity to additional reflection information. Specifically, we build on top of the open-source Qwen-image-2512 and observe a significant performance gain from 0.747 to 0.808 on GenEval and from 32.97 to 35.53 on GenEval2 Soft-TIFA. Moreover, attempts with more powerful external critics (e.g., GPT5.6-Luna) show that multimodal models with strong judge capabilities can anticipate a higher self-evolving ceiling. Last but not least, mixed text-rendering outcomes show that our self-improvements may not be uniform across different tasks. Our study aims to shed light on the current hot recursive self-improvement research line to enhance the user experience when using multimodal models without external supervision or guidance.
| Subjects: | Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.38721 [cs.AI] |
| (or arXiv:2609.38721v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38721 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Fang Wu [view email]
[v1]
Wed, 30 Sep 2026 00:54:09 UTC (42,272 KB)
来源:Hacker News · arxiv.org