PapersWithELO
← ICLR 2024 leaderboard

TransCues: Boundary and Reflection-empowered Pyramid Vision Transformer for Semantic Transparent Object Segmentation

Tuan-Anh Vu, Nguyen Truong Hai, Ziqiang Zheng, Binh-Son Hua, Qing Guo, Ivor Tsang, Sai-Kit Yeung

representation learningsemantic segmentationtransparent object segmentationpyramidal vision transformer
79.70100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
72.60100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
80.50100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.42)

OpenReview ground truth

Rejected

TL;DR — We present a novel pyramidal transformer architecture with two object cues which significantly advanced semantic transparent object segmentation and demonstrated remarkable performance across a variety of benchmark datasets

Abstract

Although glass is a prevalent material in everyday life, most semantic segmentation methods struggle to distinguish it from opaque materials. We propose $\textbf{TransCues}$, a pyramidal transformer encoder-decoder architecture to segment transparent objects from a color image. To distinguish between glass and non-glass regions, our transformer architecture is based on two important visual cues that involve boundary and reflection feature learning, respectively. We implement this idea by introducing a Boundary Feature Enhancement (BFE) module paired with a boundary loss and a Reflection Feature Enhancement (RFE) module that decomposes reflections into foreground and background layers. We empirically show that these two modules can be used together effectively, leading to improved overall performance on various benchmark datasets. In addition to binary segmentation of glass and mirror objects, we further demonstrate that our method works well for generic semantic segmentation for both glass and non-glass labels. Our method outperforms the state-of-the-art methods by a large margin on diverse datasets, achieving $\textbf{+4.2}$\% mIoU on Trans10K-v2, $\textbf{+5.6}$\% mIoU on MSD, $\textbf{+10.1}$\% mIoU on RGBD-Mirror, $\textbf{+13.1}$\% mIoU on TROSD, and $\textbf{+8.3}$\% mIoU on Stanford2D3D, demonstrate the effectiveness and efficiency of our method.

Author context

Most prolific author: 11 submissions (credibility 0.97).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 38 comparisons

Ranked above opponent in 57% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)