Evaluation of Data Quality Disparity and Implications for Fair Machine Learning
Authors
Mohit Sharma*, Pratik Mishra*, Sandeep Hans, Abhijnan Chakraborty, Vijay Arya
Mohit Sharma and Pratik Mishra contributed equally to this work.
Abstract
The performance of machine learning (ML) models heavily depends on the quality of the data they are trained on. While prior work often treats data quality as uniform across a dataset, we investigate whether it varies across different population subgroups within a dataset and examine its implications, a phenomenon we refer to as Data Quality Disparity (DQD). Our analysis reveals that many real-world datasets inherently exhibit DQD with underrepresented or marginalized groups frequently associated with lower-quality data compared to more privileged subgroups...
Read More