Abstract:
AI evaluation rarely declares the quantity its numbers estimate. That omission shows up as claims that don't survive audit. Agent configurations that clear the subtasks fail the workflows those subtasks compose. Separate harness variance from model variance and the leaderboard leader changes. A flagship coding benchmark was withdrawn by its own creator, its central claim never validated.
The failures localize to two regimes that call for opposite fixes. When the trouble is sampling error, the machinery works: item-level modeling and adaptive selection give reliable comparisons from fewer items. Specification error is harder, because the same fixes measure the wrong thing more precisely, and spreading evidence across benchmarks entrenches shared error instead of averaging it away. Where scores carry stakes, specification error becomes adversarial. Systems optimize against the instrument, and validity becomes an incentive property. The open problems are experimental: how long a claim holds, and which interventions would identify what a score means.
Speaker Biography:
Sanmi Koyejo is an Associate Professor in the Department of Computer Science and a research scientist at Meta. Koyejo leads the Stanford Trustworthy AI Research (STAIR) lab, which develops measurement-theoretic foundations for trustworthy AI systems, spanning AI evaluation science, algorithmic accountability, and privacy-preserving machine learning. Koyejo has received the Presidential Early Career Award for Scientists and Engineers (PECASE), the Skip Ellis Early Career Award, the Alfred P. Sloan Research Fellowship, the NSF CAREER Award, and multiple outstanding paper awards at flagship venues, including NeurIPS and ACL. He has delivered keynote presentations at major conferences, including ECCV and FAccT. He serves in key leadership roles, including as Board President of Black in AI, on the Board of Directors of the Neural Information Processing Systems Foundation, and in other leadership positions in professional organizations that advance AI research and broaden participation in the field.