SPFormer: Enhancing Vision Transformer with Superpixel Representation
Jieru Mei, Liang-Chieh Chen, Alan Yuille, Cihang Xie
OpenReview ground truth
Abstract
In this study, we present a novel approach to enhance Vision Transformers by leveraging superpixel representation. Unlike the traditional Vision Transformer, which uniformly partitions images into non-overlapping patches of fixed size, our superpixel approach divides an image into distinct, irregular regions, each designed to cluster pixels based on shared semantics for better capturing intricate image details. Note that this superpixel clustering is also applicable at the intermediate feature level. The resulting model, denoted as SPFormer, can be trained end-to-end, and empirically demonstrates superior performance across a range of benchmarks. Additionally, SPFormer provides better interpretability through the visualization of its learned superpixels, and exhibits strong robustness against challenging testing conditions like rotation and occlusion.
Author context
Most prolific author: 11 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 44% of matchups.
- ▼ lost to PF-LRM: Pose-Free Large Reconstruction Mod… ×6
- ▲ beat Exploring High-Order Message-Passing in Gr… ×6
- ▼ lost to Listen to Motion: Robustly Learning Correl… ×4
- ▼ lost to Exploiting Implicit Rigidity Constraints v… ×4
- ▼ lost to A Simple Romance Between Multi-Exit Vision… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)