VTruST : Controllable value function based subset selection for Data-Centric Trustworthy AI
Soumi Das, Shubhadip Nag, Shreyyash Sharma, Suparna Bhattacharya, Sourangshu Bhattacharya
OpenReview ground truth
Abstract
Trustworthy AI is crucial to the widespread adoption of AI in high-stakes applications with explainability, fairness, and robustness being some of the key trustworthiness metrics. Data-Centric AI (DCAI) aims to construct high-quality datasets for efficient training of trustworthy models. In this work, we propose a controllable framework for data-centric trustworthy AI (DCTAI)- VTruST, that allows users to control the trade-offs between the different trustworthiness metrics of the constructed training datasets. A key challenge in implementing an efficient DCTAI framework is to design an online value-function-based training data subset selection algorithm. We pose the training data valuation and subset selection problem as an online sparse approximation formulation, where the $\textit{features}$ for each training datapoint is obtained in an online manner through an iterative training algorithm. We propose a novel online version of the OMP algorithm for solving this problem. We also derive conditions on the data matrix, that guarantee the exact recovery of the sparse solution. We demonstrate the generality and effectiveness of our approach by designing data-driven value functions for the above trustworthiness metrics. Experimental results show that VTruST outperforms the state-of-the-art baselines for fair learning as well as robust training, on standard fair and robust datasets. We also demonstrate that VTruST can provide effective tradeoffs between different trustworthiness metrics through pareto optimal fronts. Finally, we show that the data valuation generated by VTruST can provide effective data-centric explanations for different trustworthiness metrics.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 48% of matchups.
- ▲ beat Backdoor Attack for Federated Learning wit… ×6
- ▼ lost to Set Learning for Accurate and Calibrated M… ×4
- ▼ lost to Learning Polynomial Problems with $SL(2, \… ×4
- ▼ lost to From Malicious to Marvelous: The Art of Ad… ×4
- ▲ beat Closing the gap on tabular data with Fouri… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)