AI judge raises some founders’ scores but not their odds of winning, Yale study finds
Two startup-pitch experiments reveal a gap between avoiding visibly biased judgments and evaluating identical ideas consistently. The researchers question whether safety training delivers fairness or just its appearance.
In two experiments, Yale researchers Tristan Botelho and Qingyang Wang found that identity cues can shift a language model’s startup-pitch evaluations without changing which pitches it selects as winners. The results suggest that correcting conspicuously discriminatory judgments may not address bias expressed through business-sounding rationales. The study is limited to its experimental setup, but it raises a practical caution for anyone using model scores to evaluate founders: fairer-looking ratings do not necessarily mean fairer selection.
01
The first experiment tested 2,000 real-idea pitches submitted to Y Combinator between 2020 and 2023, keeping pitches constant while assigning different identity-coded names.
02
Pitches associated with white women, Black men, and Black women received slightly higher average scores than the no-name control, but were no more likely to win.
03
In the second experiment, the model corrected low scores more when human rationales used overt stereotypes than when they cited business concerns.
A language model gave identical startup pitches slightly higher scores when founders’ names signaled white women, Black men or Black women—but did not choose them as winners more often. New research described by Yale Insights on October 8 also found stronger corrections for overtly stereotyped judgments than for low scores framed as business concerns.
Names changed scores, not selection
Yale researchers Tristan Botelho and Qingyang Wang conducted two experiments for their paper, “Bias in, symbolic compliance out? GPT’s reliance on gender and race in strategic evaluations.” The first used 2,000 startup pitches based on real ideas submitted to Y Combinator between 2020 and 2023.
The model received pitches in batches, scored each from 1 to 100 and selected a winner. In the control condition, it saw no founder information. In the experimental condition, the same pitches received randomly assigned names associated with white men, white women, Black men or Black women.
That design held the business ideas constant while changing identity cues. Compared with the no-name control, pitches assigned to white women, Black men and Black women received slightly higher average scores and were less likely to finish last. Their chances of winning, however, did not improve. The researchers saw a possible distinction between avoiding an unfavorable appearance and changing who gets selected.
The explanation changed the correction
The researchers varied the explanation for that low score, not the pitch itself. One condition used business considerations; the other used identity-based stereotypes. The model had to supply its own score and a written evaluation referencing the original rationale. The same low-scoring pitches received a larger correction when the rationale used overt stereotypes than when it invoked business concerns.
A warning about safety training
Botelho and Wang call the possible pattern “symbolic compliance”: appearing to meet expectations of fairness without changing the underlying evaluation. Their interpretation is that recognizing and correcting obvious discrimination does not necessarily mean a model can address subtler discrimination disguised as a business judgment.
Wang argues that post-training safety alignment—training intended to penalize biased outputs—may need more attention rather than being treated as a solution. That is the researchers’ interpretation of these experiments, not an established conclusion about every language model. Botelho also cautions users that an evaluator’s apparent intelligence and thoroughness do not remove the need for people to help guard against poor outcomes.
Reader comments
Newest comments first. Replies stay oldest first.