Preprint

Identifying Biased Subgroups in Ranking and Classification

Eliana PastorPolytechnic University of TurinLuca de AlfaroUNIVERSITY OF CALIFORNIA (SANTA CRUZ)Elena BaralisPolytechnic University of Turin

arXiv (Cornell University)repository2021en

ABI

Annotatsiya

When analyzing the behavior of machine learning algorithms, it is important to identify specific data subgroups for which the considered algorithm shows different performance with respect to the entire dataset. The intervention of domain experts is normally required to identify relevant attributes that define these subgroups. We introduce the notion of divergence to measure this performance difference and we exploit it in the context of (i) classification models and (ii) ranking applications to automatically detect data subgroups showing a significant deviation in their behavior. Furthermore, we quantify the contribution of all attributes in the data subgroup to the divergent behavior by means of Shapley values, thus allowing the identification of the most impacting attributes.

Mavzular

Imbalanced Data Classification Techniques Data Mining Algorithms and Applications Bayesian Modeling and Causal Inference

Identifikatorlar

DOI: 10.48550/arxiv.2108.07450

Iqtiboslar va manbalar

0 ta iqtibos21 ta foydalanilgan manba

Koʻrsatkichlar — AkademScholar · Tez orada