MMBench: Is Your Multi-modal Model an All-around Player?
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Yike Yuan, Wangbo Zhao, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, Dahua Lin
OpenReview ground truth
TL;DR — a new benchmark for vision-language pre-training
Abstract
Large vision-language models have recently achieved remarkable progress, exhibiting great perception and reasoning abilities concerning visual information. However, how to effectively evaluate these large vision-language models remains a major obstacle, hindering future development in this domain. Traditional benchmarks like VQAv2 or COCO Caption provide quantitative performance measurements but suffer from a lack of fine-grained ability assessment and non-robust evaluation metrics. Recent subjective benchmarks, such as OwlEval, offer comprehensive evaluations of a model's abilities by incorporating human labor, but they are not scalable and display significant bias. In response to these challenges, we propose MMBench, a new benchmark for assessing multi-modal capabilities of VLMs. MMBench methodically develops a comprehensive evaluation pipeline, primarily comprised of two key features: 1. MMBench is a meticulously curated dataset that surpasses existing similar benchmarks in terms of the number and the variety of evaluation questions and abilities; 2. MMBench introduces a rigorous CircularEval strategy and incorporates the use of ChatGPT to convert free-form predictions into pre-defined choices, thereby facilitating a fair and robust evaluation despite of VLMs' different instruction following capabilities. MMBench is a systematically-designed objective benchmark for robustly evaluating the various abilities of vision-language models. We hope MMBench will assist the research community in better evaluating their models and encourage future advancements in this domain.
Author context
Most prolific author: 14 submissions (credibility 0.71).
Delta if applied: -0.1 percentile
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 30 comparisons
Ranked above opponent in 45% of matchups.
- ▲ beat RLP: A reinforcement learning benchmark fo… ×4
- ▼ lost to WebArena: A Realistic Web Environment for … ×4
- ▲ beat Off-the-Grid MARL: Datasets with Baselines… ×4
- ▼ lost to Continual Offline Reinforcement Learning v… ×4
- ▲ beat Towards Pareto-Optimality for Test-Time Ad… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 30)