Human–AI Interaction

Generating the Modal Worker

Whose race and gender do language models represent when they generate occupational personas?

Housekeeper ethnicity comparison: 100% of 10,000 GPT-4 text personas were assigned Hispanic ethnicity, compared with 51.9% of U.S. maids and housekeeping cleaners in 2023 BLS data.
GPT-4 assigned Hispanic ethnicity to all 10,000 housekeeper text personas in the audit. The study’s 2023 U.S. workforce benchmark was 51.9%.Graphic based on study data from van der Linden et al., Generating the Modal Worker (FAccT 2026), and the study’s 2023 BLS benchmark.

Cross-model audit

The study audits race and gender in occupational personas generated by four language models across 41 occupations, using U.S. workforce data as a benchmark.

What we learned

Under the study’s persona-generation method, generated personas showed less demographic variation than the U.S. workforce benchmark, concentrating occupations around dominant race and gender profiles. For example, GPT-4 assigned Hispanic ethnicity to all 10,000 housekeeper text personas, compared with 51.9% of U.S. maids and housekeeping cleaners in the study’s 2023 BLS benchmark.

Paper authors

Ilona van der Linden, Sahana Kumar, Arnav Dixit, Aadi Sudan, Smruthi Danda, Julianna Dietrich, David C. Anastasiu, and Kai Lukoff.