Options
Estimating The Risk Of Individual Discrimination Of Classifiers
Journal
Lecture Notes in Computer Science
Advances in Knowledge Discovery and Data Mining
Date Issued
2023-01-01
Author(s)
WoS ID
WOS:001281030800039
Abstract
Data owners are increasingly liable for the potential harm caused by using their data on underprivileged communities. Stakeholders seek to identify data characteristics that lead to biased algorithms against specific demographic groups, such as race, gender, age, or religion. We focus on identifying feature subsets of datasets where the ground truth response function from features to observed outcomes differs across demographic groups. To achieve this, we propose FORESEE, a decision tree-based algorithm that generates a score indicating the likelihood of an individual’s response varying with sensitive attributes. Our approach enables us to identify individuals most likely to be misclassified by various classifiers, including Random Forest, Logistic Regression, Support Vector Machine, Multi-Layer Perceptron, and k-Nearest Neighbors. The advantage of our approach is that it allows stakeholders to identify risky samples that may contribute to discrimination and use FORESEE to estimate the risk of upcoming samples.
OCDE Subjects
Quartile (Date Issued)
SQ
License
acceso restringido