← Daniel Griffin

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Hanna Wallach, Meera Desai, A. Feder Cooper, Angelina Wang, Chad Atalla, Solon Barocas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P. Alex Dow, Jean Garcia-Gathright, Alexandra Olteanu, Nicholas Pangakis, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Vaughan, Matthew Vogel, Hannah Washington, Abigail Z. Jacobs · paper · 2025

01

Excerpt

The measurement tasks involved in evaluating generative AI systems lack sufficient scientific rigor, leading to a tangle of sloppy tests and apples-to-oranges comparisons.

02

Summary

ICML 2025 position paper arguing that generative AI evaluation is fundamentally a social-science measurement problem, and presenting a four-level framework grounded in measurement theory for constructs related to GenAI capabilities, behaviors, and impacts.

03

Why it matters

The foundational pushback against treating any evaluation number as self-evidencing. If the measurement instrument doesn't validly pick out the construct (reasoning, helpfulness, safety, legal competence), a high score is not a capability claim.

04

Source

05

Notes

Sets up measurement and construct validity as prior to evaluation. A benchmark score is a claim about a construct, and the validity of that claim depends on whether the instrument actually measures the construct. The paper argues that most GenAI evaluation skips this step, producing a tangle of sloppy tests and apples-to-oranges comparisons. The authors import a four-level framework from social-science measurement theory and apply it to GenAI. The argument is explicitly not that better metrics solve the problem — it is that capability claims depend on validity work that is social, interpretive, and institutional. Placed against verifiable-reward framings (Royzen; Expanding RLVR), the tension is direct: • Verifiable-reward: the reward is verifiable when the outcome is checkable. • Measurement-validity: checkability of an outcome does not imply the outcome measures the construct you care about. The "verifiable" in verifiable reward is doing more work than it admits. Both can be true at once. A narrow technical task (theorem proved, test suite passed) may have near-trivial validity. A broad capability claim (legal reasoning, medical judgment, general agentic competence) almost never does. The library preserves this disagreement structurally — entries can score high on rewardrichness while scoring low on inputlegibility and unknown on validity. Related entries • Measurement to Meaning (Salaudeen et al. 2025) — the validity-centered framework applied. • Royzen: Standard Signal — poster case for reward richness.