III: Small: Collaborative Research: Fair Data Mining with Insights from Data and Model
INSTITUTION
Purdue University, IN
PRINCIPAL INVESTIGATOR
Jingwen Yan
FUNDING
$597K
YEAR
2024
MOONBASE SCORE
Still being scored
LOADING MOONBASE SCORE
Abstract
The growing utilization of data mining and machine learning systems in critical domains has raised concerns about the potential amplification of societal biases and discrimination. Large-scale pre-trained models, such as Generative Pre-trained Transformer (GPT), also confront fairness issues, further intensifying the need to address societal biases in automated decision-making algorithms. However, the resource-intensive nature of obtaining annotated data for fair algorithms, and the significant computational challenges in debiasing large-scale models, presents substantial obstacles to achieving algorithmic fairness. In response, this project aims to pioneer fundamental research in fair algorithmic decision-making while alleviating the heavy demands on data annotation and computational resources. Ultimately, it will facilitate the development, adoption, and evaluation of fair artificial intelligence (AI) systems by humans. This project will result in novel algorithms and software, fostering broader study in real-world applications such as promoting health equity in Alzheimer's disease research. Moreover, this project prioritizes education and diversity by providing training opportunities for underrepresented minority students, engaging them in cutting-edge computational research. The research objective of this project is to improve fair decision-making in a more efficient and flexible manner. It addresses three fundamental research challenges: 1) establishing a theoretically grounded framework for learning and evaluating fair representations using widely accessible unlabeled data; 2) learning unsupervised fair representation applicable to various downstream tasks for improved flexibility; 3) exploring an efficient strategy tailored for transformers to achieve fairness in large-scale pre-trained models without the need of retraining, thereby enhancing the trade-off between fairness and accuracy while concurrently improving computational and GPU memory efficiency. One important application of this project lies in health equity, particularly in addressing biased predictions in disease studies. By integrating rigorous theoretical analysis with emerging application studies, this research project contributes to the advancement of more equitable and effective AI for societal benefits. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
Are you the primary organization running this research?
The two tools below are built for the principal investigator & host institution behind this project.