PapersWithELO
← ICLR 2024 leaderboard

SPFormer: Enhancing Vision Transformer with Superpixel Representation

Jieru Mei, Liang-Chieh Chen, Alan Yuille, Cihang Xie

representation learningVision TransformerSuperpixel
25.90100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
20.30100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
33.10100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

Abstract

In this study, we present a novel approach to enhance Vision Transformers by leveraging superpixel representation. Unlike the traditional Vision Transformer, which uniformly partitions images into non-overlapping patches of fixed size, our superpixel approach divides an image into distinct, irregular regions, each designed to cluster pixels based on shared semantics for better capturing intricate image details. Note that this superpixel clustering is also applicable at the intermediate feature level. The resulting model, denoted as SPFormer, can be trained end-to-end, and empirically demonstrates superior performance across a range of benchmarks. Additionally, SPFormer provides better interpretability through the visualization of its learned superpixels, and exhibits strong robustness against challenging testing conditions like rotation and occlusion.

Author context

Most prolific author: 11 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 44% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)