CAT-Seg: Cost Aggregation for Open-vocabulary Semantic Segmentation
Seokju Cho, Heeseong Shin, Sunghwan Hong, Anurag Arnab, Paul Hongsuck Seo, Seungryong Kim
OpenReview ground truth
Abstract
In this paper, we reinterpret the challenge of open-vocabulary semantic segmentation, where each pixel in an image is labeled with a wide range of text descriptions, as a correspondence problem focusing on the optimal text matching for each pixel. Addressing the limitations of conventional region-to-text matching approaches, we introduce a novel framework, CAT-Seg, grounded on the principles of cost aggregation methods in visual correspondence tasks. This framework refines the initial matching scores between dense image and text embeddings, leveraging a Transformer-based module for cost aggregation, further enhanced with embedding guidance. Notably, by operating on cosine similarity instead of manipulating embeddings directly, our approach enables the end-to-end fine-tuning of the CLIP model for pixel-level tasks, while yielding superior zero-shot capabilities. Empirical evaluations show our method's superior performance, achieving state-of-the-art results across open-vocabulary benchmarks, practical computational efficiency, and robustness for various domains, underscoring its potential for a wide range of open-vocabulary semantic segmentation applications.
Author context
Most prolific author: 9 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 56% of matchups.
- ▼ lost to PF-LRM: Pose-Free Large Reconstruction Mod… ×4
- ▼ lost to Diffusion Models for Open-Vocabulary Segme… ×4
- ▲ beat Towards Precise Prediction Uncertainty in … ×4
- ▲ beat TransCues: Boundary and Reflection-empower… ×4
- ▲ beat Musketeer: Joint Training/Inference for Mu… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)