CEO, World Labs; Stanford HAI co-director
Fei-Fei Li
Profile
If you want to understand why deep learning went from academic curiosity to civilizational force, the story runs straight through Fei-Fei Li. The popular history credits the 2012 AlexNet moment — the convolutional network from Geoffrey Hinton’s lab, built by Ilya Sutskever and Alex Krizhevsky, that crushed the competition and convinced the field that neural nets plus GPUs plus data actually worked. But the benchmark it won was Li’s. Starting in 2007, against the prevailing wisdom that better algorithms were the bottleneck, she bet that the missing ingredient was data at scale. The result was ImageNet: 14 million hand-labeled images across 20,000-plus categories, assembled with WordNet’s taxonomy and an army of Amazon Mechanical Turk annotators. It was unglamorous, expensive, and initially dismissed — and it turned out to be the proving ground the whole revolution needed. That’s the lesson worth internalizing: sometimes the highest-leverage contribution isn’t a clever model, it’s the dataset and the benchmark everyone else optimizes against.
Li’s career is a rare full-stack tour of AI’s institutions. Born in Beijing, she immigrated to New Jersey at 15, learned English while helping run her parents’ dry-cleaning business, studied physics at Princeton, and took her PhD in electrical engineering at Caltech. She joined Stanford in 2009, later directed the Stanford AI Lab (SAIL), and in 2019 co-founded the Stanford Institute for Human-Centered AI (HAI), which she still co-directs. Two of the field’s most influential builders — Andrej Karpathy and Jim Fan — came through her lab. Between 2017 and 2018 she took a high-profile sabbatical as chief scientist of AI/ML at Google Cloud, an episode that ended in controversy (see below) and shaped her subsequent, deliberate emphasis on the human stakes of the technology.
Today she’s CEO and co-founder of World Labs, the spatial-intelligence startup she launched in 2023 and which has now raised roughly $1.23 billion from backers including NVIDIA, AMD, and Andreessen Horowitz. Her thesis is that language models, for all their power, are effectively blind: they manipulate text but have no grounded model of 3D space, physics, distance, or how objects behave. World Labs’ flagship, Marble, generates explorable, persistent 3D worlds from a text prompt, a single image, or a video clip. Li frames spatial intelligence not as a competitor to LLMs but as the missing physical layer — the thing robots, AR, simulation, and embodied agents will need to actually act in the world rather than just describe it.
For a developer learning AI, Li matters on two axes. First, she’s living proof that the data-and-benchmark layer is often where paradigm shifts are actually seeded — a useful counterweight to model-worship. Second, she’s making an explicit, well-capitalized bet that the next frontier is 3D and embodied rather than purely linguistic. Whether or not “world models” become as central as she predicts, her track record of calling the field’s inflection points early — and building the infrastructure to prove it — makes her someone worth reading closely rather than skimming.
Books
Key Articles & Papers
ImageNet: A Large-Scale Hierarchical Image Database ImageNet Large Scale Visual Recognition Challenge (ILSVRC) Deep Visual-Semantic Alignments for Generating Image Descriptions Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations Spatial Intelligence Is AI's Next FrontierVideos
Controversies
Project Maven (Google, 2018). While chief scientist at Google Cloud, Li was involved in Project Maven, a Pentagon contract applying Google’s image-recognition AI to drone footage. Leaked internal emails showed her advising colleagues to avoid publicizing the “AI” framing of the deal for fear of a backlash over “weaponized” AI. When the project became public, thousands of Google employees protested, Google declined to renew the contract and published a set of AI principles, and Li returned to Stanford. The episode is frequently cited as a formative test of tech’s ethics-versus-defense tensions; Li has since spoken about it as part of why she champions human-centered AI, though critics note the gap between her private caution about optics and her public advocacy. (DataCenterKnowledge, The Intercept)
Spotify Podcasts
YouTube