University of Copenhagen Explores the Role of Memorization in Fair AI Models

University of Copenhagen Explores the Role of Memorization in Fair AI Models

Presented at the Symposium on Foundations of Responsible Computing (FORC 2025), researchers from the University of Copenhagen’s Department of Computer Science (KU) explore how memorization, typically linked to overfitting, can, under certain conditions, be leveraged to improve fairness metrics in AI classification models. Their paper, titled When Can Memorization Improve Fairness?, analyzes how key group fairness metrics, statistical parity, equal opportunity, and equalized odds, can be influenced, and in some cases superficially optimized, by memorizing targeted subpopulations.

 

This work, led by Bob Pepin, Christian Igel, and Raghavendra Selvan, bridges fairness theory with practical implications for AI systems, especially those deployed in sensitive domains like hiring, healthcare, and public policy.

Fairness through the Lens of Memorization

The study focuses on a counterintuitive phenomenon: memorization, typically considered undesirable in machine learning, can be used to reduce perceived unfairness in classification tasks. By perfectly predicting outcomes for a select group of examples (the "memorized" subset), a classifier can improve group-level fairness metrics. However, this improvement can be superficial or even misleading, a phenomenon the authors label as gaming the metric.

The researchers derive exact mathematical expressions for how memorization affects the three most widely used group fairness metrics. They also provide necessary and sufficient conditions under which memorization can completely eliminate observed fairness gaps.

Explicit Prescriptions for Bias Removal

Key contributions from the paper include:

  • Theoretical Characterization: Closed-form formulas link the fairness bias of a memorizing classifier to the population and memorized subset's composition, along with the classifier’s original bias.
  • Optimal Memorization Strategies: Through systems of linear equations, the authors define exactly which label and group distributions within the memorized dataset can neutralize each fairness metric.
  • Bounds on Required Memorization: Upper and lower limits are provided for how much of the dataset needs to be memorized to fully compensate for existing bias, crucial for real-world feasibility assessments.

Practical Implications: From Dataset Compression to Fair Procurement

The paper outlines diverse scenarios where these results can apply:

  • Dataset Compression: In nearest neighbor models, where only a subset of points is retained, the analysis helps in selecting a subset that optimally balances fairness.
  • Model Development: When fine-tuning prompts for large language models on a development set, the framework quantifies how fairness bias might creep in through selective memorization.
  • Fairness-Constrained Procurement: The results can help public institutions set rigorous standards to prevent vendors from superficially gaming fairness metrics using specialized models on specific subgroups.

Guarding Against Metric Manipulation in Fair AI

While the findings offer a roadmap for fairness improvements, they also underscore the risk of hidden subgroup bias. A classifier that appears fair by standard metrics might still create disparities within protected groups, especially if fairness is achieved through overfitting small, strategically chosen populations.

The work calls for caution in deploying fairness metrics without complementary individual-level assessments, echoing broader concerns from the fairness-in-AI community. It also opens doors to more transparent model design, where fairness is not just a statistical artifact but a meaningful outcome of the training process.

Logo ALMA

SustainML is among these nine innovative projects dedicated to creating a sustainable ML framework for Green AI.

Coordinator Office Address

Plaza de la Encina 10-11, Núcleo 4, 2ª Pl.
28760 Tres cantos - Madrid (España)

  • X
  • LinkedIn

EN-Funded_by_the_EU-POSThis project has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement No 101070408.