PrometheusRoot
Blog Links Prometheans 100+ AI Books AI Companies Why are you here?
← Prometheans 100+
×
David Silver
builder
Researcher
Website Wikipedia
deepmindalphagoalphazeroreinforcement-learning

Related

pioneer Demis Hassabis
← Prometheans 100+ David Silver

CEO of Ineffable Intelligence, AlphaGo architect

David Silver

CEO and Founder — Ineffable Intelligence Professor — University College London Principal Research Scientist — Google DeepMind
Listen — profile
0:00 / 3:14

Profile

David Silver is the reason “reinforcement learning” stopped being a niche academic subfield and became the thing everyone in AI now argues about. He led AlphaGo, the system whose March 2016 win over Lee Sedol was the last unambiguous “oh” moment before ChatGPT — and then, more importantly, he threw away the human data and built AlphaZero, which learned Go, chess and shogi from nothing but the rules and self-play, and beat everything. That second result is the one that actually shaped his career and, arguably, the current frontier of the field.

The biography is unusually non-linear for a top researcher. Cambridge, 1997, where he shared a college with Demis Hassabis; then a decade out of academia co-founding the video game studio Elixir Studios as CTO; then back to a PhD at the University of Alberta under the Sutton–Barto RL tradition, finishing in 2009; then DeepMind from 2013 to 2026, where he ran the reinforcement learning group. Along the way: the DQN Atari paper, MuZero (planning with a learned model of the environment, so you don’t even need the rules), and co-leading AlphaStar with Oriol Vinyals. He picked up the 2019 ACM Prize in Computing and a Royal Society fellowship in 2021, and he’s still a professor at University College London.

For developers, the most durable thing he ever produced is not a model — it’s his UCL reinforcement learning lecture course, filmed in 2015 and still, a decade on, the single best free introduction to RL on the internet. If you have read Sutton & Barto and bounced off it, watch Silver’s ten lectures instead; the MDP formalism, Bellman equations, TD learning and policy gradients are laid out with a clarity that most modern RLHF/RLVR tutorials simply assume you already have. Anyone doing post-training work today is applying, often without knowing it, machinery Silver taught a generation on a whiteboard.

Now he’s put his thesis on the line. In late 2025 he founded Ineffable Intelligence in London, leaving DeepMind in January 2026; in April 2026 it came out of stealth with a $1.1 billion seed round at a $5.1 billion valuation — reportedly Europe’s largest ever seed — co-led by Sequoia and Lightspeed, with Nvidia, Google, Index and the UK’s sovereign fund alongside. The pitch is a deliberate anti-LLM bet: no pre-training, no human corpus, a “superlearner” that discovers knowledge from its own experience, stated with the mission of making “first contact with superintelligence.” It is either the most important contrarian position in AI or an extremely expensive rerun of the argument that games are not the world. Silver has said any personal proceeds go to high-impact charities. Worth watching either way, because if he’s right, everything you’re currently building on top of next-token prediction has a ceiling.

Key Articles & Papers

Mastering the Game of Go with Deep Neural Networks and Tree Search 2016 — The AlphaGo paper — policy and value networks plus Monte Carlo tree search. The template for combining learned intuition with explicit search. Mastering the Game of Go without Human Knowledge 2017 — AlphaGo Zero: same problem, zero human games, better result. The paper that made 'human data is a crutch' a serious position. Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm 2017 — AlphaZero — one algorithm, three games, no domain knowledge. Generality, not just strength. Human-level Control Through Deep Reinforcement Learning 2015 — DQN. Experience replay and target networks made deep RL stable enough to work at all — still the standard starting point. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model 2020 — MuZero learns its own model of the environment, so planning no longer requires knowing the rules. Deterministic Policy Gradient Algorithms 2014 — The DPG result underpinning DDPG and much of continuous-control RL in robotics. Continuous Control with Deep Reinforcement Learning 2015 — DDPG — deep RL applied to continuous action spaces; a workhorse baseline for years. Grandmaster level in StarCraft II using multi-agent reinforcement learning 2019 — AlphaStar: imperfect information, long horizons, huge action spaces — RL leaving the comfort of board games. Reward is Enough 2021 — The manifesto: reward maximisation alone may be sufficient for general intelligence. Widely cited, widely disputed — read the rebuttals too. Welcome to the Era of Experience 2025 — With Richard Sutton: human data is running out, and the next phase must come from agents learning in the world. This is the intellectual charter for Ineffable Intelligence.

Videos

YouTube video
YouTube video
YouTube video
YouTube video
YouTube video
YouTube video

Controversies

“Reward is Enough.” Silver’s 2021 paper with Satinder Singh, Doina Precup and Richard Sutton argued that reward maximisation alone could account for all intelligence, natural and artificial. It drew sustained pushback: Scalar reward is not enough argued for explicitly multi-objective formulations and flagged alignment risks in single-scalar optimisation, and Reward is not enough challenged the claim across perception, language and social intelligence. A common criticism is that “reward” is defined loosely enough to be unfalsifiable. Silver’s position is best read as a research bet rather than a proven claim — and Ineffable Intelligence is that bet with a billion dollars behind it.

The valuation. A $1.1B seed at $5.1B for a pre-product lab founded months earlier is, by any historical standard, extraordinary — and the honest technical objection is that RL’s greatest hits came in domains with clean reward signals and cheap simulation. Open-ended real-world tasks have neither. Silver’s counter is that this is exactly the problem worth solving. Reasonable people disagree, loudly.

Spotify Podcasts

Nach LLMs kommt das World Model: 1,1 Mrd. $ Pre-Seed für Ineffable Intelligence – Olaf Jacobi (Basic Ventures)
Nach LLMs kommt das World Model: 1,1 Mrd. $ Pre-Seed für Ineffable Intelligence – Olaf Jacobi (Basic Ventures)
Startup Insider
2026
David Silver's Ineffable Intelligence Raises $1.1B Seed
David Silver's Ineffable Intelligence Raises $1.1B Seed
Inferred
2026
[4/28 04:00] David Silver Ineffable Intelligence $1.1B / Talkie-1930 13B Vintage LLM
[4/28 04:00] David Silver Ineffable Intelligence $1.1B / Talkie-1930 13B Vintage LLM
AI News Flash
2026
【4/28 13時】David Silver Ineffable Intelligence $1.1B・Talkie-1930 13B Vintage LLM
【4/28 13時】David Silver Ineffable Intelligence $1.1B・Talkie-1930 13B Vintage LLM
AIニュース速報
2026
【2/22 22時】David Silver DeepMind退社 Ineffable Intelligence設立・SaaSpocalypse ソフトウェア株暴落とM&A加速
【2/22 22時】David Silver DeepMind退社 Ineffable Intelligence設立・SaaSpocalypse ソフトウェア株暴落とM&A加速
AIニュース速報
2026
[2/22 13:00] David Silver leaves DeepMind to found Ineffable Intelligence / SaaSpocalypse deepens as software stocks lose $1T+
[2/22 13:00] David Silver leaves DeepMind to found Ineffable Intelligence / SaaSpocalypse deepens as software stocks lose $1T+
AI News Flash
2026
David Silver: Ineffable Intelligence
David Silver: Ineffable Intelligence
A30.FM Podcast
2026
Is Human Data Enough? With David Silver
Is Human Data Enough? With David Silver
Google DeepMind: The Podcast
2025
David Silver @ RCL 2024
David Silver @ RCL 2024
TalkRL: The Reinforcement Learning Podcast
2024
#86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning
#86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning
Lex Fridman Podcast
2020

YouTube

YouTube video
2020

Related People

pioneer Demis Hassabis
© 2026 PrometheusRoot