BackPlaybook
GrowthUP Partners

When you're designing an AI evaluation or hiring process

The Move

Insist on side-by-side evaluation, not one-at-a-time. And know that blinding is not a free win — the largest RCT found it can backfire.

The Script

"Before we finalize this evaluation process, I want to flag the evidence on format. Bohnet et al. showed that one-at-a-time evaluation lets stereotypes drive 65% of decisions. Side-by-side, that effect vanishes. And the Australian BETA trial found that de-identifying applications actually reduced women's shortlisting odds. Let me share the research."

Your AI Prompt

I'm designing [an evaluation rubric / a hiring process / an AI tool selection process]. Review the criteria below for bias triggers: [paste rubric or criteria]. Flag where one-at-a-time (separate) evaluation introduces stereotype risk. Recommend specific changes to enable joint (side-by-side) evaluation. Also flag any criteria where blinding could backfire based on the Hiscox 2017 finding.

Why This Works

Bohnet, van Geen & Bazerman (Management Science 2016) found that separate evaluation lets stereotypes drive 65% of decisions — joint evaluation eliminates the effect. But Hiscox et al. (Australian BETA 2017), the largest RCT on blind hiring (2,100+ public servants, ~33,600 assessments), found de-identifying applications reduced women's shortlisting odds by 2.9%. The format matters more than blinding. Wilson & Caliskan (2024) showed AI screening systems favored female-associated names in only 11.1% of comparisons. These tools help you design around the bias, not just name it.

Research Behind This Move

When Performance Trumps Gender Bias: Joint vs. Separate Evaluation

Bohnet, van Geen & Bazerman · Management Science 62(5) · 2016

Going Blind to See More Clearly: Unconscious Bias in Australian Public Service Shortlisting

Hiscox, Oliver, Ridgway, Aranda-Jan, Bedi & Genovese (Australian BETA) · Australian Government Behavioural Economics Team RCT · 2017

Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval

Wilson & Caliskan · AIES 2024, 1578–1590 (arXiv:2407.20371) · October 2024