PapersWithELO
← ICLR 2024 leaderboard

Associative Transformer is a Sparse Representation Learner

Yuwei Sun, Hideya Ochiai, Zhirong Wu, Stephen Lin, Ryota Kanai

self/semi-supervised learningglobal workspace theoryattention mechanismassociative memorylatent bottleneck
40.90100
Fused
band ≈ ±15 pct pts (from σ = 0.31)
39.50100
Mimo
band ≈ ±22 pct pts (from σ = 0.44)
58.30100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.43)

OpenReview ground truth

Rejected

TL;DR — We propose the Associative Transformer (AiT) building upon recent neuroscience studies of the Global Workspace Theory and associative memory.

Abstract

Emerging from the monolithic pairwise attention mechanism in conventional Transformer models, there is a growing interest in leveraging sparse interactions that align more closely with biological principles. Approaches including the Set Transformer and the Perceiver employ cross-attention consolidated with a latent space that forms an attention bottleneck with limited capacity. Building upon recent neuroscience studies of the Global Workspace Theory and associative memory, we propose the Associative Transformer (AiT). AiT induces low-rank explicit memory that serves as both priors to guide bottleneck attention in shared workspace and attractors within associative memory of a Hopfield network. Through joint end-to-end training, these priors naturally develop module specialization, each contributing a distinct inductive bias to form attention bottlenecks. A bottleneck can foster competition of inputs for information writing into the memory. We show that AiT is a sparse representation learner, learning distinct priors through the bottlenecks that are complexity-invariant to input quantities and dimensions. AiT demonstrates its superiority over methods such as the Set Transformer, Vision Transformer, and Coordination in various vision tasks.

Author context

Most prolific author: 3 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)