PapersWithELO
← ICLR 2024 leaderboard

Relevance-based embeddings for efficient relevance retrieval

Kirill Sergeevich Shevkunov, Andrey Ploskonosov, Liudmila Prokhorenkova

self/semi-supervised learningInformation searchRelevance searchNearest neighbor searchRelevance-based embeddingsRecommendation systems
57.20100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
51.60100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
58.40100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.42)

OpenReview ground truth

Rejected

Abstract

In many machine learning applications, the most relevant items for a particular query should be efficiently extracted. The relevance function is typically an expensive neural similarity model making the exhaustive search infeasible. A typical solution to this problem is to train another model that separately embeds queries and items to a vector space, where similarity is defined via the dot product or cosine similarity. This allows one to search the most relevant objects through fast approximate nearest neighbors search at the cost of some reduction in quality. To compensate for this reduction, the found candidates are then re-ranked by the expensive similarity model. In this paper, we propose an alternative approach that utilizes the relevances of the expensive model to make relevance-based embeddings. We show both theoretically and empirically that describing each query by its relevance for a set of support items creates a powerful query representation. Additionally, we investigate several strategies for selecting these support items and show that additional significant improvements can be obtained. Our experiments on diverse datasets show improved performance over existing approaches.

Author context

Most prolific author: 3 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)