Fairness in Machine Learning for Healthcare

Sean Davis, MD, PhD

November 5, 2024

You are designing an image search engine (that will be powered by AI, of course), akin to Google Images.

On a technical level, that’s a piece of cake. You’re a great machine learning engineer, and this is basic stuff!

But say you live in a world where 90 percent of CEOs are male.

Sort of like this world.

What to do?

Should you design your search engine so that it accurately mirrors that reality, yielding images of man after man after man when a user types in “CEO”?


Or, since that risks reinforcing gender stereotypes that help keep women out of the C- suite, should you create a search engine that deliberately shows a more balanced mix, even if it’s a mix that does not reflects reality as it is today?

What is the problem here?

What is Bias?

Social bias

Definition: Systematic favoritism or discrimination.

Example: A hiring algorithm consistently rejects women is socially biased.

Statistical bias

Definition: Systematic error in estimation.

Example: A weather app that always overestimates the probability of rain is statistically biased.

If there’s a predictable difference between two groups on average, then these two definitions will be at odds.

If you design your search engine to make statistically unbiased predictions about the gender breakdown among CEOs, then it will necessarily be biased in the social sense of the word.

And if you design it not to have its predictions correlate with gender, it will necessarily be biased in the statistical sense.

Fairness

What is Fairness?

  • At least 20 different definitions of fairness!

What do the big players do?

What do the big players do?

Types of Fairness

  • Procedural Fairness
  • Distributive Fairness
  • Representational Fairness

Another though experiment

You are a bank officer charged with approving loans. You use an AI system to help you make decisions. That system is trained to predict who will repay a loan and who will not. The FICO score is a key factor in the system’s predictions.

Procedural Fairness

An algorithm is procedurally fair if the procedure is uses to make decisions is fair.

By this measure, your algorithm is fair. It uses a standard, well-established procedure to make its recommendations (the FICO score) and applies that procedure uniformly to all applicants.

Distributive Fairness

Another conception of fairness, known as distributive fairness, says that an algorithm is fair if it leads to fair outcomes. By this measure, your algorithm is failing, because its recommendations have a disparate impact on one racial group versus another.

Another thought experiment

Representational Fairness

Breaches of representational fairness occur when systems reinforce the subordination of some groups along the lines of identity whether because the systems explicitly denigrate a group, stereotype a group, or fail to recognize a group and therefore render it invisible.

Kate Crawford

Interaction

Introduction

  • Machine Learning in Healthcare
  • Importance of Fairness
  • Overview of the Presentation

Bias and Fairness

Types of Harm

  • Harms of Allocation (resources) Allocation of resources, opportunities, or information in a way that systematically disadvantages certain groups.

  • Harms of Representation (identity) Representation of individuals or groups in a way that perpetuates stereotypes or biases.

Background: ML in Healthcare

  • Diagnosis assistance
  • Treatment planning
  • Risk prediction
  • Resource allocation
  • Drug discovery

Why Fairness Matters in Healthcare ML

  • Equal access to quality healthcare
  • Avoiding perpetuation of historical biases
  • Legal and ethical considerations
  • Building trust in ML systems

Sources of Bias in Healthcare ML

  1. Data Collection Bias
  2. Sampling Bias
  3. Label Bias
  4. Feature Selection Bias
  5. Algorithmic Bias

Bias in Healthcare ML

Bias mitigation strategies

Other Fairness-Aware ML Techniques

  • Reweighting
  • Adversarial Debiasing
  • Fair Representation Learning
  • Post-processing Methods

Implementing Fairness in Healthcare ML

  1. Collect diverse and representative data
  2. Choose appropriate fairness metric for the context
  3. Implement fairness-aware algorithms
  4. Regularly audit model performance across groups
  5. Involve diverse stakeholders in development and deployment

Ethical Considerations

  • Transparency and explainability
  • Informed consent for data usage
  • Privacy concerns
  • Potential for unintended consequences

Conclusion

  • Fairness in healthcare ML is crucial but complex
  • No one-size-fits-all solution
  • Balancing act: fairness, accuracy, and domain-specific considerations
  • Ongoing research and ethical considerations are key
  • Collaboration across disciplines is essential!

Questions?