PyCon DE 2026

No, you can't 'eval' your way to fairness

Laura Summers

Laura Summers
About

Laura Summers

Laura is a very technical designer™️, working at Pydantic as Lead Design Engineer. Her side projects include Sweet Summer Child Score (summerchild.dev) and Ethics Litmus Tests (ethical-litmus.site). Laura is passionate about feminism, digital rights and designing for privacy. She speaks, writes and runs workshops at the intersection of design and technology.

  1. Fairness as experience, not state
  2. The gap between the metric and the harm
  3. What genuine intervention looks like
Section one

Fairness as experience, not state

Activity

What's your preferred film format?

  • 70mm / analog
  • IMAX / big digital
Activity

Are you a night owl or early bird?

  • night owl
  • early bird
Activity

Hands up by decade

  • 20s
  • 30s
  • 40s+
«
something was decided about you
on your behalf
without your knowledge or consent
»

What fairness actually is

fairness is felt before it is calculated
The cats
The problem with proxies

Same box. Different reality.

  • American ☑
  • Woman ☑
  • Black ☑
  • candidate for intervention ☑
  • white house vs group house ☐
«
fairness is not a state of the world. it's an experience of it.
»

The gap between the metric and the harm

Harms of allocation

when a system withholds or unfairly distributes opportunities, resources or penalties.

Harms of representation

when a system reinforces the subordination of a group through how it depicts or describes people.

Barocas, Crawford, Shapiro & Wallach, 2017 · Crawford, NeurIPS 2017

Harms of allocation

Low risk

an automated rewards system at a doctor's surgery. algorithm de-prioritises the treat for patients outside the target age range. a child doesn't get a candy. upsetting but recoverable.

High risk

COMPAS. twice as likely to falsely flag Black defendants as future offenders. people stayed in prison longer based on a score nobody could explain. the time does not come back. (Angwin et al., ProPublica 2016)

Harms of representation

Low risk

a face filter that puts bunny ears on you. cute if your face is detected. the effect breaks for darker skin tones. annoying. a minor indignity. nobody's life is at stake.

High risk

a teenage girl searches for mathematicians and sees almost entirely white men. she updates her sense of what's possible. she changes course. the erasure compounds.

so how well do our tools actually measure this?
Fairness metrics & evals

LLM evals with fairness-adjacent metrics

DeepEval (BiasMetric, ToxicityMetric)

Giskard + Phare

Inspect Evals / UK AISI (StereoSet, BOLD, BBQ)

EleutherAI LM Eval Harness (WinoGender, CrowS-Pairs)

Azure AI Evaluation SDK (HateUnfairnessEvaluator)

Venn diagram of fairness and eval library landscape
An unreliable narrator judging its own reliability
The gap

The four jobs

A
B
C
D
  • A Understand the world

  • B Understand foundation models

  • C Understand your system

  • D Intervene

Scales of justice meme: left pan (heavy) labelled 'All of law, philosophy, religion, academia, culture and art on justice and fairness' — right pan (light) labelled 'Me and my lil' Python script'
Section three

What genuine intervention looks like

Design Justice book cover
Design justice

Community-led practices to build the worlds we need

Sasha Costanza-Chock (they/them), MIT Press, 2020. Freely available online.

Design justice network

designjustice.org

founded 2015 at the Allied Media Conference. 10 principles centring people normally marginalised by design. available in multiple languages.

principle #2 We center the voices of those who are directly impacted by the outcomes of the design process.

principle #5 We see the role of the designer (data scientist) as a facilitator rather than an expert.

principle #6 We believe that everyone is an expert based on their own lived experience, and that we all have unique and brilliant contributions to bring to a design process.

Critical tech literacy

The ecosystem

Qualitative + quantitative

A metric is only as good as the hypothesis behind it

  • study the domain — read, follow your curiosity
  • talk to people affected by the system
  • use scanning tools and design activities
  • then measure what you've learned to look for
Case study

"Let's make it fair."

  • two colleagues. one concern. one meeting. everyone leaves satisfied.
  • the ML engineer goes off to tune the model.
  • the product lead goes off to watch the dashboard.
  • neither checks what the other means by "fair."

Six months later

The engineer

tuned the model to achieve equalized odds: the true positive rate and false positive rate are equal across both groups.

"the model works equally well for both groups. it's fair."

The product lead

checked the dashboard. Group A candidates: 52% advanced. Group B candidates: 35% advanced.

"this isn't fair."

choosing a metric is choosing a value system.
The point
installing a library is not a moral position. it is the absence of one.
Fairness washing
«
First we should critique
(trouble, queer, or denormalize)
»

Sasha Costanza-Chock, Design Justice, p. 57

Product: can you intervene further upstream?

Problem statement

does the business model create incentives to find harm, or to ignore it?

Build ↔ iterate loop

monitoring, feedback loops, learning from real use.

01
02
03
04

Design ↔ build loop

what are you encoding before anyone uses it?

Recourse

can the person affected challenge, override, or opt out?

Codesign

Design with, not for

  • bring people affected by the system into the design of it
  • not as test subjects — as co-designers with standing
  • "nothing about us without us" — disability rights movement, 1980s–90s
  • you will be bad at this at first. start anyway.
Consent-based intervention

Offer. Listen. Adapt. Stop when asked.

  • make the intervention visible — never act silently on someone's behalf
  • try once — if they say no, stop
  • if they say yes, learn — offer more things like it
  • never assume you know the answer next time
Consent-based intervention
1

"Yes, offer me more"

the system learns your preferences. surfaces relevant options. builds a picture of what help looks like for you specifically.

2

"No, leave me alone"

the system stops. does not try again. does not assume you'll change your mind. respects the preference as stated.

«
Do you identify as a member of a disadvantaged group? If yes, we'd like to adjust how this system works in your favour. Would you like us to do that?
»

In conclusion…

«
Fairness is not a metric
»
«
Now, go and make some good trouble
»

Thank you

QR code — link to slides