Moonbase
← Back to Awards
Directorate for Mathematical and Physical SciencesNSF · NSFNSF

Efficient Data Removal Under High-dimensional Asymptotics with Applications to Risk Estimation and Machine Unlearning

Mohammad Ali Maleki·Columbia University, NY·2025–2028·ACTIVE
Donate

INSTITUTION

Columbia University, NY

PRINCIPAL INVESTIGATOR

Mohammad Ali Maleki

FUNDING

$240K

YEAR

2025

MOONBASE SCORE

Still being scored

LOADING MOONBASE SCORE

Abstract

The rapid expansion of machine learning (ML) and artificial intelligence (AI) has created an urgent need for tools that enhance trust, and control over these technologies. This project aims to address that need by developing methods that help explain why a model makes certain predictions, improve model tuning, and detect harmful or misleading data—such as intentionally corrupted inputs introduced by adversaries—before they compromise the model’s reliability. A promising strategy for tackling these challenges is to analyze the impact of individual data points or subsets of data by estimating how their removal affects the model’s behavior. This research supports the national interest by advancing scientific understanding of AI, improving the robustness of decision-making systems, and contributing to the development of technologies that align with privacy protections. The project investigates whether it is possible to develop computationally efficient algorithms that approximate the output of a model trained without a given subset of data, without having to retrain the model from scratch. This question is particularly challenging in high-dimensional settings, where the number of features is large relative to the sample size. The project focuses on designing data removal methods that are both scalable and theoretically sound in these regimes. The resulting algorithms will be evaluated in two important application areas: risk estimation and machine unlearning. Through this work, the project aims to lay the foundation for practical tools that improve model interpretability and accountability in complex learning systems. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.

Directorate for Mathematical and Physical SciencesArtificial Intelligence (AI)Machine Learning TheorySTATISTICSSTATISTICSthroughunderstandingdimensionalstrategyefficientintelligencemakingsoundworthyreflectshavingadversariesmeritimportantcomplexresultingurgentwithoutaccountabilityunlearning

Are you the primary organization running this research?

The two tools below are built for the principal investigator & host institution behind this project.