Unsupervised open-vocabulary action recognition with an autoregressive model
Adrian Bulat, Enrique Sanchez, Brais Martinez, Georgios Tzimiropoulos
OpenReview ground truth
Abstract
Current works on zero/few- shot action recognition are largely based on contrastive approaches trained in a supervised manner to select an action class out of a predefined set. Instead, in this work, we propose a new paradigm for zero-shot action recognition based on autoregressive generation of a free-form action-specific caption describing the action occurring in the video. To this end, we propose to adapt an image-based pre-trained autoregressive Vision & Language (V&L) Model for action recognition. We firstly show that direct fine-tuning of an autoregressive model using the action classes suffers from severe overfitting. To alleviate this, we then introduce an unsupervised learning framework consisting of two key components: (a) an unsupervised method for adapting the autoregressive model to action/video data by means of pseudo-caption generation and self-training without using any action-specific labels; (b) a retrieval component for discovering a diverse set of pseudo-captions for each video. In the process, we show that both components are necessary to obtain high accuracy. Our model results in predictions that are fine-grained, interpretable, and naturally open-vocabulary. Importantly, when evaluated for zero- and few-shot action recognition, our approach matches or even outperforms contrastive learning-based methods.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 49% of matchups.
- ▼ lost to Emu: Generative Pretraining in Multimodali… ×6
- ▼ lost to SAN: Inducing Metrizability of GAN with Di… ×4
- ▲ beat Detecting Language Model Attacks With Perp… ×4
- ▼ lost to Stay on Topic with Classifier-Free Guidanc… ×4
- ▲ beat Pay attention to cycle for spatio-temporal… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)