Referring Expression Matters: Multi-referring Feature Aggregation for Referring Video Object Segmentation
YAN LI, Qiong Wang, Jianwei Zheng, Cong Bai, Lu Zhang
OpenReview ground truth
Abstract
Referring Video Object Segmentation aims to segment object instances referred to by natural language referring expressions in a video sequence. This interaction style is quite simple and flexible, being capable of producing high quality segmentation masks. However, the referring expression variation occurs due to the randomness of expressions provided by users, making the existing state-of-the-art models still face the problem of wrongly identifying the referred object. To address this issue, we present a novel referring video object segmentation network fed with multiple referring expressions. Specifically, a simple but effective neural expression generation module is proposed to map the features of multiple referring expressions to complementary features with less redundancy. This interaction of multiple referring expressions not only is beneficial to identify the referred object but also speeds up the training convergence. We make evaluations of the proposed method on the popular referring video object segmentation datasets, and experimental results demonstrate that our method outperforms the state-of-the-arts by a significant margin in terms of segmentation quality and achieves considerable gains in terms of training convergence speed. Our code and pre-trained models will be available.
Author context
Most prolific author: 1 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 28 comparisons
Ranked above opponent in 44% of matchups.
- ▲ beat TeLLMe what you see: Using LLMs to Explain… ×6
- ▼ lost to EZ-CLIP: EFFICIENT ZERO-SHOT VIDEO ACTION … ×4
- ▼ lost to Motion PointNet: Solving Dynamic Capture i… ×4
- ▼ lost to Amortising the Gap between Pre-training an… ×4
- ▼ lost to Soft Contrastive Learning for Time Series ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 28)