PapersWithELO
← ICLR 2024 leaderboard

BOWLL: A DECEPTIVELY SIMPLE OPEN WORLD LIFELONG LEARNER

Roshni Ramanna Kamath, Rupert Mitchell, Subarnaduti Paul, Kristian Kersting, Martin Mundt

transfer & meta learningOpen World LearningLifelong LearningContinual LearningActive LearningBenchmark Baseline
19.30100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
18.40100
Mimo
band ≈ ±20 pct pts (from σ = 0.39)
21.80100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

TL;DR — we introduce the first cohesive baseline for lifelong learning in an open world setting

Abstract

The quest to improve scalar performance numbers on predetermined benchmarks seems to be deeply engraved in deep learning. However, the real world is seldom carefully curated and applications are seldom limited to excelling on test sets. A practical system is generally required to recognize novel concepts, refrain from actively including uninformative data, and retain previously acquired knowledge throughout its lifetime. Despite these key elements being rigorously researched individually, the study of their conjunction, open world lifelong learning, is only a recent trend. To accelerate this multifaceted field’s exploration, we introduce its first monolithic and much-needed baseline. Leveraging the ubiquitous use of batch normalization across deep neural networks, we propose a deceptively simple yet highly effective way to repurpose standard models for open world lifelong learning. Through extensive empirical evaluation, we highlight why our approach should serve as a future standard for models that are able to effectively maintain their knowledge, selectively focus on informative data, and accelerate future learning.

Author context

Most prolific author: 9 submissions (credibility 0.97).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 41% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)