Less is More: Vision Representation Compression for Efficient Video Generation with Large Language Models
Published in AAAI Conference on Artificial Intelligence (AAAI), 2026
We compress visual token representations and extend next-token prediction to next-sequence prediction for substantially faster autoregressive video generation.
Recommended citation: Yucheng Zhou, Jihai Zhang, Guanjie Chen, Jianbing Shen, and Yu Cheng. Less is More: Vision Representation Compression for Efficient Video Generation with Large Language Models. AAAI, 2026.
Download Paper
