Human–AI Interaction
Generating the Modal Worker
Whose race and gender do language models represent when they generate occupational personas?

Cross-model audit
The study audits race and gender in occupational personas generated by four language models across 41 occupations, using U.S. workforce data as a benchmark.
What we learned
Under the study’s persona-generation method, generated personas showed less demographic variation than the U.S. workforce benchmark, concentrating occupations around dominant race and gender profiles. For example, GPT-4 assigned Hispanic ethnicity to all 10,000 housekeeper text personas, compared with 51.9% of U.S. maids and housekeeping cleaners in the study’s 2023 BLS benchmark.
Paper authors
Ilona van der Linden, Sahana Kumar, Arnav Dixit, Aadi Sudan, Smruthi Danda, Julianna Dietrich, David C. Anastasiu, and Kai Lukoff.