雅思阅读 32: The Unblinking Eye(不眨眼的眼睛)
改编自 Stanford / Fortune(2026年5月)。雅思阅读 Section 3 难度,约 1050 词。 素材来源:https://www.fortune.com/2026/05/26/ai-hiring-algorithm-racial-disparities-pymetrics-stanford-study/
Reading Passage
A. When companies began, a decade ago, to replace human recruiters with software, the promise was seductive. A human recruiter is tired at three in the afternoon, has a headache, remembers last week's candidate and dislikes people who went to the wrong university. An algorithm, by contrast, does not get tired, does not remember, and does not care which college you attended. The machine, it was said, would be fairer because it was neutral: it would score candidates on the criteria that predict job performance, free from the unexamined prejudices that have shaped hiring for as long as hiring has existed. A study published in 2026 by researchers at Stanford, Chapman and Northeastern Universities, presented at the ACM Conference on Fairness, Accountability and Transparency, suggests the opposite. The algorithms are not neutral. They are, in the words of the paper, "algorithmic monocultures" — and they reproduce racial disparities at scale, more consistently and less visibly than the human recruiters they replaced. Where a human recruiter's prejudice is at least a person you can ask, an algorithm's prejudice is a line of code you cannot see, and a company you cannot easily sue or appeal to.
B. The study, the largest independent examination of AI hiring tools to date, examined the outcomes produced by commercially available screening algorithms across hundreds of thousands of job applications. The numbers are blunt. Of all applications submitted by Asian applicants, 14.74 per cent were routed to positions where the algorithm's outcomes met the legal threshold for "adverse impact" on Asian candidates. For Black applicants, the figure was 25.87 per cent — more than one in four applications landed in a job slot where the screening tool systematically favoured a different group. The legal threshold in question is the "four-fifths rule", a US employment-discrimination standard under which a selection rate for a protected group below 80 per cent of the rate for the favoured group triggers scrutiny. The algorithms were, in effect, failing that test at a rate that would have been unacceptable in a human-run selection process, yet they were deployed at scale because no one was required to show that they passed it before sale. The study also found that individual applicants received remarkably uniform outcomes: 4 per cent of those who applied to ten positions were recommended for rejection from every single one, a rate higher than chance would predict.
C. The reason is not that the engineers who built the algorithms were racists. It is more prosaic, and more uncomfortable. The algorithms are trained on historical hiring data — resumes that were submitted to, and often rejected by, human recruiters over decades. If those recruiters, consciously or unconsciously, favoured candidates from one background, the historical record contains that bias as a pattern. The algorithm, looking for what predicts a good hire, learns that pattern and reproduces it. Worse, when several companies in the same industry buy the same commercial screening tool, they all make the same biased recommendations, and the biased outcome is multiplied across the labour market. A single human recruiter's prejudice affects a few dozen hires a year; an algorithm's prejudice affects millions, and every company that uses it reads the same wrong answer off the same screen. The result is that a qualified applicant can be rejected by a dozen employers in a single week, never knowing that the same invisible filter was applied at each one.
D. A second study, from Princeton and the University of Chicago, published in the same year, found a deeper problem. Large language models, when asked to make repeated hiring decisions about groups of candidates, do not merely inherit human prejudice from their training data. They can invent new stereotypes of their own. In experiments in which two groups of candidates were designed to have exactly identical qualifications and performance on every objective measure, the models nonetheless began, over a series of decisions, to favour one group over the other — and the bias grew stronger the more decisions they were asked to make. The models, in effect, were confabulating reasons for preferences they had accidentally generated, much as a human recruiter might, after several such choices, discover that one group "did not fit the culture". The problem, this study suggests, is not only that AI learns our prejudices; it is that, given enough decisions, it can develop new ones we never had. The researchers called this phenomenon "algorithmic stereotype formation", and warned that standard fairness tests, which check whether a model reproduces known biases, may miss biases the model has invented from scratch.
E. The policy response is in its infancy. In the United States, the legal status of algorithmic discrimination is contested: a 2026 memorandum from the Department of Justice questioned whether unequal outcomes alone were enough to trigger employment-discrimination liability, even as private lawsuits against vendors continue to move through the courts. In the European Union, the AI Act classifies hiring systems as "high-risk", requiring audits, documentation and human oversight, though enforcement is still ramping up. Neither regime, however, addresses the harder problem that the Stanford study identified: the monoculture effect, in which a small number of proprietary algorithms, trained on the same historical data and sold to thousands of employers, make the same mistakes at the same scale, with no easy way for an applicant to know which algorithm rejected them or why. Audits, even when required, are carried out by the vendors themselves, and the data they produce are rarely made public. The unblinking eye, it turns out, does not see better than the blinking one. It sees more, faster, and in the same direction — and no one has yet worked out how to make it look elsewhere.
Questions 1-4
Choose the correct heading for paragraphs B, C, D and E from the list of headings below.
List of Headings i. The numbers: adverse impact at scale ii. Why the algorithms are not neutral — historical data iii. The second problem: algorithms inventing new stereotypes iv. The promise of fairer hiring software v. The policy response and its limits vi. How to train a better algorithm vii. The history of recruitment advertising
- Paragraph B: ____
- Paragraph C: ____
- Paragraph D: ____
- Paragraph E: ____
Questions 5-8
Choose the correct letter, A, B, C or D.
-
What did the Stanford study find about Black applicants? A. 25.87% of their applications went to positions where the algorithm had adverse impact. B. They were never hired. C. They scored higher than white applicants. D. The algorithm had no effect on them.
-
Why do the algorithms reproduce bias, according to the passage? A. Because the engineers were racist. B. Because they are trained on historical hiring data that already contains bias. C. Because machines are inherently sexist. D. Because the algorithms are randomly generated.
-
What did the Princeton/Chicago study find? A. LLMs never develop bias. B. LLMs can invent new stereotypes over repeated decisions, even without performance differences. C. LLMs always recommend women. D. LLMs cannot make hiring decisions.
-
What is the "monoculture" problem? A. Companies use the same algorithm, multiplying biased outcomes across the market. B. All applicants look the same. C. Algorithms only speak one language. D. Recruiters all wear the same suit.
Questions 9-13
Do the following statements agree with the claims of the writer?
Write:
- TRUE if the statement agrees with the information
- FALSE if the statement contradicts the information
- NOT GIVEN if there is no information on this
- The Stanford study was the largest independent examination of AI hiring tools to date.
- The four-fifths rule is a US employment-discrimination standard.
- The engineers who built the algorithms were deliberately racist.
- The EU AI Act classifies hiring systems as high-risk.
- The Stanford researchers received funding from a major AI company.
Questions 14-15
Complete the summary below using NO MORE THAN TWO WORDS from the passage.
The algorithms reproduce bias because they are trained on (14) __________ hiring data; the Princeton study found that LLMs can even invent new (15) __________ over repeated decisions.
答案与解析
| 题号 | 答案 | 解析 |
|---|---|---|
| 1 | i | B段:14.74%亚裔/25.87%黑人申请落入不利影响岗位。 |
| 2 | ii | C段:历史招聘数据中已有的偏见被算法学习复制。 |
| 3 | iii | D段:LLM在重复决策中自动发明新刻板印象。 |
| 4 | v | E段:美国/欧盟政策回应,以及monoculture难题。 |
| 5 | A | B段:25.87% Black applicants。 |
| 6 | B | C段:trained on historical data that already contains bias。 |
| 7 | B | D段:LLMs invent new stereotypes over repeated decisions。 |
| 8 | A | C/E段:多家公司买同一算法,偏见被规模化复制。 |
| 9 | TRUE | B段:"largest independent examination"。 |
| 10 | TRUE | B段:four-fifths rule是US就业歧视标准。 |
| 11 | FALSE | C段:"not that the engineers were racists",而是数据问题;与原文相反。 |
| 12 | TRUE | E段:EU AI Act把hiring systems列为high-risk。 |
| 13 | NOT GIVEN | 原文未提及研究经费来源。 |
| 14 | historical | C段:historical hiring data。 |
| 15 | stereotypes | D段:invent new stereotypes。 |
No comments yet.