Associative Transformer is a Sparse Representation Learner
Yuwei Sun, Hideya Ochiai, Zhirong Wu, Stephen Lin, Ryota Kanai
OpenReview ground truth
TL;DR — We propose the Associative Transformer (AiT) building upon recent neuroscience studies of the Global Workspace Theory and associative memory.
Abstract
Emerging from the monolithic pairwise attention mechanism in conventional Transformer models, there is a growing interest in leveraging sparse interactions that align more closely with biological principles. Approaches including the Set Transformer and the Perceiver employ cross-attention consolidated with a latent space that forms an attention bottleneck with limited capacity. Building upon recent neuroscience studies of the Global Workspace Theory and associative memory, we propose the Associative Transformer (AiT). AiT induces low-rank explicit memory that serves as both priors to guide bottleneck attention in shared workspace and attractors within associative memory of a Hopfield network. Through joint end-to-end training, these priors naturally develop module specialization, each contributing a distinct inductive bias to form attention bottlenecks. A bottleneck can foster competition of inputs for information writing into the memory. We show that AiT is a sparse representation learner, learning distinct priors through the bottlenecks that are complexity-invariant to input quantities and dimensions. AiT demonstrates its superiority over methods such as the Set Transformer, Vision Transformer, and Coordination in various vision tasks.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 53% of matchups.
- ▲ beat ImAD: An End-to-End Method for Unsupervise… ×6
- ▲ beat A New Type of Associative Memory Network w… ×4
- ▼ lost to STanHop: Sparse Tandem Hopfield Model for … ×4
- ▲ beat Gaussian Process-Based Corruption-resilien… ×4
- ▲ beat The Cyclical Chaos And Its Equilibrium ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)