ICLR 2024 Analytics
fused ranking · n = 730
Validity — a filter, not a chooser
AUC 0.753 (95% CI 0.713–0.794), fused ranking. Bottom-decile bars are red when below the base accept rate (31.1%). “Not accepted” = rejected + withdrawn
Acceptance rate by decile
Decision distribution by percentile
Threshold analysis
| Cut | Not-accepted rate | vs base reject rate |
|---|---|---|
| Bottom 25% | 90.2% rejected | +21.3 pts |
| Bottom 50% | 85.5% rejected | +16.6 pts |
| Bottom 75% | 78.1% rejected | +9.2 pts |
Percentile spread (accepted vs rejected)
| p5 | p25 | median | p75 | p95 | |
|---|---|---|---|---|---|
| Accepted | 21 | 52 | 73 | 88 | 98 |
| Rejected | 4 | 19 | 40 | 64 | 89 |
A rejected paper's median percentile is far below an accepted one's — but the distributions overlap at the top, which is why we filter rather than choose.
Reliability — internal precision
Fraction of observed θ spread that exceeds estimation noise, within each judge's own fit. NOTE: precise ratings are not necessarily correct ratings — a random tournament is internally precise too. Cross-judge agreement is the real validity evidence.
≈0 ⇒ noise · ≳0.3 ⇒ real signal shared by independent judges
Convergence
Spearman ρ of each round's ranking against the final ranking, and mean absolute percentile change (averaged across both judges' runs). Stable for filtering by round ~4; boundary refinement through round 7.
Judge agreement & noise scale
Fitted noise scale ŝ = 0.31; cross-judge agreement 74%. A θ gap of 0.5 ⇒ the better paper wins ≈ 83% of comparisons.
| Round | Swiss τ (within-band) | Cross-band τ | Batch θ spread |
|---|---|---|---|
| 1 | 0.550 | 0.551 | 1.45 |
| 2 | 0.416 | 0.580 | 1.13 |
| 3 | 0.291 | 0.649 | 0.85 |
| 4 | 0.232 | 0.680 | 0.69 |
| 5 | 0.213 | 0.711 | 0.60 |
| 6 | 0.141 | 0.728 | 0.53 |
| 7 | 0.132 | 0.737 | 0.47 |
Cross-band τ stays high while Swiss τ collapses: batches of similar-quality papers are genuinely harder to rank — that's judge noise at the boundary, not judge failure.
Failure cases — where we were wrong
Published deliberately: a credible filter shows its misses. Both lists link to full paper pages.
Worst-ranked but Accepted (we said skip — wrong)
- A Data-Driven Measure of Relative Uncertainty for Misclassification Detection
- OWL: A Large Language Model for IT Operations
- Simplicial Representation Learning with Neural $k$-Forms
- Unraveling the Enigma of Double Descent: An In-depth Analysis through the Lens of Learned Feature Space
- State Representation Learning Using an Unbalanced Atlas
- Generative Modeling with Phase Stochastic Bridge
- Large Content And Behavior Models To Understand, Simulate, And Optimize Content And Behavior
- Time-Varying Propensity Score to Bridge the Gap between the Past and Present
Best-ranked but Rejected (we said read — wrong)
- Achieving Minimax Optimal Sample Complexity of Offline Reinforcement Learning: A DRO-Based Approach
- Certifying LLM Safety against Adversarial Prompting
- Scalabale AI Safety via Doubly-Efficient Debate
- Provably Doubly Accelerated Federated Learning: The First Theoretically Successful Combination of Local Training and Communication Compression
- Generative Marginalization Models
- Dataset Distillation via Adversarial Prediction Matching
- Last-Iterate Convergence Properties of Regret-Matching Algorithms in Games
- Non-Asymptotic Analysis for Single-Loop (Natural) Actor-Critic with Compatible Function Approximation