Less is More: Vision Representation Compression for Efficient Video Generation with Large Language Models

Published in AAAI Conference on Artificial Intelligence (AAAI), 2026

We compress visual token representations and extend next-token prediction to next-sequence prediction for substantially faster autoregressive video generation.

Recommended citation: Yucheng Zhou, Jihai Zhang, Guanjie Chen, Jianbing Shen, and Yu Cheng. Less is More: Vision Representation Compression for Efficient Video Generation with Large Language Models. AAAI, 2026.
Download Paper