Publications
Evaluation of Data Quality Disparity and Implications for Fair Machine Learning
Thirty-third IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER-26)
Mohit Sharma*, Pratik Mishra*, Sandeep Hans, Abhijnan Chakraborty, Vijay Arya
Mohit Sharma and Pratik Mishra contributed equally to this work.
The performance of machine learning (ML) models heavily depends on the quality of the data they are trained on. While prior work often treats data quality as uniform across a dataset, we investigate whether it varies across different population subgroups within a dataset and examine its implications, a phenomenon we refer to as Data Quality Disparity (DQD). Our analysis reveals that many real-world datasets inherently exhibit DQD with underrepresented or marginalized groups frequently associated with lower-quality data compared to more privileged subgroups...
Read More