Will Multimodal Models Be Dazzled by Multi-Image Visual Puzzles?2026年5月5日·Zhi Zhu,YaoQi Fan,Zhe Chen,Yue Cao,Yangzhou LiuTong Lu· 0 分钟阅读时长 引用 URL类型会议文章出版物Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)最近更新于 2026年5月5日AuthorsTong Lu南京大学← VMonarch: Efficient Video Diffusion Transformers with Structured Attention 2026年5月5日Arbitrary Generative Video Interpolation 2026年4月24日 →