Stanford HAI Professor, commonsense AI researcher
Yejin Choi
Profile
Yejin Choi is the researcher who keeps asking the question the rest of the field would rather skip past: does a large language model actually understand anything, or is it an extraordinarily gifted mimic? As the Dieter Schwarz Foundation HAI Professor of Computer Science and a Senior Fellow at Stanford HAI, Choi has spent her career probing the gap between what AI systems can say and what they actually know — and she has built the benchmarks, datasets, and demos that make that gap impossible to ignore. For developers weaned on the “just scale it” narrative, her work is the necessary counterweight: proof that fluency and comprehension are not the same thing.
Choi’s path is unusual. Before academia she spent years as a software engineer, and she has spoken candidly about arriving at NLP research later and more circuitously than her peers — which may be why her instincts run against fashion. She built her reputation over a decade at the University of Washington and the Allen Institute for AI (AI2), where she led some of the most cited work in commonsense reasoning. In 2024 she left Seattle for California, joining NVIDIA as a Senior Director of AI Research before landing at Stanford in early 2025. Along the way she collected a 2022 MacArthur “genius” Fellowship — awarded specifically for teaching machines common sense — and appeared on TIME’s 100 Most Influential People in AI in both 2023 and 2025.
Her fingerprints are on much of the infrastructure the field uses to measure reasoning. ATOMIC and COMET turned commonsense knowledge into something a neural network could generate rather than just look up. Benchmarks like HellaSwag and WinoGrande were designed to be adversarially hard for models that were quietly cheating on easier tests — and they became standard yardsticks precisely because they exposed shortcuts. Grover demonstrated that the best detector of AI-generated fake news was the same kind of model that produced it. And Delphi, her most provocative project, asked whether a system could learn to make moral judgments at all. Each was less a product than an argument, delivered in code.
What makes Choi matter to anyone building with AI today is her intellectual honesty about limits. She is not a doomer and not a hype merchant; she’s an empiricist who cheerfully shows you a state-of-the-art model failing at a puzzle a six-year-old would solve. Like Gary Marcus, she doubts that scale alone delivers understanding — but unlike him, she spends her time constructing the experiments and smaller, norm-trained systems that could actually close the gap. In a research culture increasingly organized around parameter counts, Choi insists on asking what the numbers are actually measuring. That skepticism, grounded in real benchmarks, is what keeps the field honest.
Key Articles & Papers
COMET: Commonsense Transformers for Automatic Knowledge Graph Construction Defending Against Neural Fake News (Grover) HellaSwag: Can a Machine Really Finish Your Sentence? WinoGrande: An Adversarial Winograd Schema Challenge at Scale (COMET-)ATOMIC 2020: On Symbolic and Neural Commonsense Knowledge Graphs Can Machines Learn Morality? The Delphi Experiment The Curious Case of Commonsense Intelligence (Daedalus)Videos
Controversies
Choi’s Delphi demo (2021) drew the sharpest reaction of her career. Released as an interactive tool that would render moral verdicts on user-typed scenarios, it quickly produced offensive and inconsistent judgments when probed adversarially, and critics — in the press and in academia — argued that framing ethics as a data-labeling classification problem was both technically shaky and philosophically naive. Choi’s team responded by relabeling Delphi as a research prototype for modeling people’s moral judgments rather than an oracle of right and wrong, and later published a peer-reviewed account in Nature Machine Intelligence. Whether one reads the episode as overreach or as valuable stress-testing of a genuinely hard question, it remains the clearest example of Choi pushing an idea far enough to draw real fire — which is arguably the point of the work.
Spotify Podcasts
YouTube