CEO of Ineffable Intelligence, AlphaGo architect
David Silver
Profile
David Silver is the reason “reinforcement learning” stopped being a niche academic subfield and became the thing everyone in AI now argues about. He led AlphaGo, the system whose March 2016 win over Lee Sedol was the last unambiguous “oh” moment before ChatGPT — and then, more importantly, he threw away the human data and built AlphaZero, which learned Go, chess and shogi from nothing but the rules and self-play, and beat everything. That second result is the one that actually shaped his career and, arguably, the current frontier of the field.
The biography is unusually non-linear for a top researcher. Cambridge, 1997, where he shared a college with Demis Hassabis; then a decade out of academia co-founding the video game studio Elixir Studios as CTO; then back to a PhD at the University of Alberta under the Sutton–Barto RL tradition, finishing in 2009; then DeepMind from 2013 to 2026, where he ran the reinforcement learning group. Along the way: the DQN Atari paper, MuZero (planning with a learned model of the environment, so you don’t even need the rules), and co-leading AlphaStar with Oriol Vinyals. He picked up the 2019 ACM Prize in Computing and a Royal Society fellowship in 2021, and he’s still a professor at University College London.
For developers, the most durable thing he ever produced is not a model — it’s his UCL reinforcement learning lecture course, filmed in 2015 and still, a decade on, the single best free introduction to RL on the internet. If you have read Sutton & Barto and bounced off it, watch Silver’s ten lectures instead; the MDP formalism, Bellman equations, TD learning and policy gradients are laid out with a clarity that most modern RLHF/RLVR tutorials simply assume you already have. Anyone doing post-training work today is applying, often without knowing it, machinery Silver taught a generation on a whiteboard.
Now he’s put his thesis on the line. In late 2025 he founded Ineffable Intelligence in London, leaving DeepMind in January 2026; in April 2026 it came out of stealth with a $1.1 billion seed round at a $5.1 billion valuation — reportedly Europe’s largest ever seed — co-led by Sequoia and Lightspeed, with Nvidia, Google, Index and the UK’s sovereign fund alongside. The pitch is a deliberate anti-LLM bet: no pre-training, no human corpus, a “superlearner” that discovers knowledge from its own experience, stated with the mission of making “first contact with superintelligence.” It is either the most important contrarian position in AI or an extremely expensive rerun of the argument that games are not the world. Silver has said any personal proceeds go to high-impact charities. Worth watching either way, because if he’s right, everything you’re currently building on top of next-token prediction has a ceiling.
Key Articles & Papers
Mastering the Game of Go with Deep Neural Networks and Tree Search Mastering the Game of Go without Human Knowledge Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm Human-level Control Through Deep Reinforcement Learning Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model Deterministic Policy Gradient Algorithms Continuous Control with Deep Reinforcement Learning Grandmaster level in StarCraft II using multi-agent reinforcement learning Reward is Enough Welcome to the Era of ExperienceVideos
Controversies
“Reward is Enough.” Silver’s 2021 paper with Satinder Singh, Doina Precup and Richard Sutton argued that reward maximisation alone could account for all intelligence, natural and artificial. It drew sustained pushback: Scalar reward is not enough argued for explicitly multi-objective formulations and flagged alignment risks in single-scalar optimisation, and Reward is not enough challenged the claim across perception, language and social intelligence. A common criticism is that “reward” is defined loosely enough to be unfalsifiable. Silver’s position is best read as a research bet rather than a proven claim — and Ineffable Intelligence is that bet with a billion dollars behind it.
The valuation. A $1.1B seed at $5.1B for a pre-product lab founded months earlier is, by any historical standard, extraordinary — and the honest technical objection is that RL’s greatest hits came in domains with clean reward signals and cheap simulation. Open-ended real-world tasks have neither. Silver’s counter is that this is exactly the problem worth solving. Reasonable people disagree, loudly.
Spotify Podcasts
YouTube