PapersWithELO
← ICLR 2024 leaderboard

OmniInput: A Model-centric Evaluation Framework through Output Distribution

Weitang Liu, Ying Wai Li, Tianle Wang, Yi-Zhuang You, Jingbo Shang

general MLcomprehensive model evaluationoutput distribution
18.20100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
22.40100
Mimo
band ≈ ±20 pct pts (from σ = 0.39)
12.60100
DeepSeek
band ≈ ±23 pct pts (from σ = 0.46)

OpenReview ground truth

Rejected

TL;DR — Comprehensive model evaluation, entire input space.

Abstract

We propose a novel model-centric evaluation framework, OmniInput, to evaluate the quality of an AI/ML model's predictions on all possible inputs (including human-unrecognizable ones), which is crucial for AI safety and reliability. Unlike traditional data-centric evaluation based on pre-defined test sets, the test set in OmniInput is self-constructed by the model itself and the model quality is evaluated by investigating its output distribution. We employ an efficient sampler to obtain representative inputs and the output distribution of the trained model, which, after selective annotation, can be used to estimate the model's precision and recall at different output values and a comprehensive precision-recall curve. Our experiments demonstrate that OmniInput enables a more fine-grained comparison between models, especially when their performance is almost the same on pre-defined test sets, leading to new findings and insights for how to train more robust, generalizable models.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 34 comparisons

Ranked above opponent in 49% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)