PapersWithELO
← ICLR 2024 leaderboard

Distributional Structured Pruning by Lower bounding the Total Variation Distance using Witness functions

Chaitanya Murti, Chiranjib Bhattacharyya

representation learningPruningStructured PruningTotal Variation Distance
65.40100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
65.80100
Mimo
band ≈ ±18 pct pts (from σ = 0.37)
66.30100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

TL;DR — Discrimination-based structured pruning using novel lower bounds on the TV distance.

Abstract

Recent literature introduced the notion of distributional structured pruning (DSP) in Deep Neural Networks by retaining discriminative filters that can effectively differentiate between classes. Crucial to DSP is the ability to estimate the discriminative ability of a filter, which is defined by the minimum pairwise Total Variation (TV) distance between the class-conditional feature distributions. Since the computation of TV distance is generally intractable, existing literature assumes the class-conditional feature distributions are Gaussian, thereby enabling the use of the tractable Hellinger lower bound to estimate discriminative ability. However, the Gaussian assumption is not only restrictive but also does not typically hold. In this work, we address this gap by deriving a lower bound on TV Distance which depends only on the moments of witness functions. Using linear witness functions, the bound establishes new relationships between the TV Distance and well-known discriminant-based classifiers, such as Fisher Discriminants and Minimax Probability machines. The lower bounds are used to produce a variety of pruning algorithms called WitnessPrune by varying the choice of witness function. We empirically show that we can achieve up to 7\% greater accuracy for similar sparsity in hard-to-prune layers using a polynomial witness function as compared to the state-of-the-art.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 40)