Language Model Beats Diffusion - Tokenizer is key to visual generation
Lijun Yu, Jose Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A Ross, Lu Jiang
OpenReview ground truth
TL;DR — With a good visual tokenizer, language model style transformer outperforms diffusion models on image and video generation.
Abstract
While Large Language Models (LLMs) are the dominant models for generative tasks in language, they do not perform as well as diffusion models on image and video generation. To effectively use LLMs for visual generation, one crucial component is the visual tokenizer that maps pixel-space inputs to discrete tokens appropriate for LLM learning. In this paper, we introduce \modelname{}, a video tokenizer designed to generate concise and expressive tokens for both videos and images using a common token vocabulary. Equipped with this new tokenizer, we show that LLMs outperform diffusion models on standard image and video generation benchmarks including ImageNet and Kinetics. In addition, we demonstrate that our tokenizer surpasses the previously top-performing video tokenizer on two more tasks: (1) video compression comparable to the next-generation video codec (VCC) according to human evaluations, and (2) learning effective representations for action recognition tasks.
Author context
Most prolific author: 7 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 67% of matchups.
- ▼ lost to Achieving Minimax Optimal Sample Complexit… ×14
- ▲ beat Optimal Sketching for Residual Error Estim… ×14
- ▲ beat Certifying LLM Safety against Adversarial … ×8
- ▼ lost to Robust agents learn causal world models ×6
- ▼ lost to Beyond Weisfeiler-Lehman: A Quantitative F… ×6
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)