Anthropic member of technical staff, scaling reinforcement learning
Sholto Douglas
Profile
Sholto Douglas is one of the most listened-to insider voices explaining how frontier AI models are actually built — not from the outside as a commentator, but from inside the rooms where Gemini and Claude were trained. He is a Member of Technical Staff at Anthropic, where he works on scaling reinforcement learning: the training regime that has, over the past two years, turned language models from autocomplete engines into systems that can reason through competition math, win at competitive programming, and act autonomously across long agentic tasks. If you want to understand why models got dramatically better at coding and reasoning since 2024, Douglas is close to the source of that shift.
His path is unusual. Douglas studied Mechatronic (Space) Engineering at the University of Sydney and was a near-Olympic-level fencer — he placed 21st at the 2017 World Championships — before pivoting into AI. He joined Google DeepMind in 2021 as a research engineer, worked on large-scale language modeling, and became a central figure on the Gemini team, co-leading inference infrastructure and helping shape how those models run efficiently on TPUs at massive scale. He was the person who kicked off and wrote the first version of How to Scale Your Model, DeepMind’s now-widely-cited open textbook on the systems engineering of training and serving large models. He later moved to Anthropic to focus on scaling RL.
What makes Douglas essential for developers isn’t just his résumé — it’s that he can explain the machinery clearly. His long-form conversations with fellow Anthropic researcher Trenton Bricken on the Dwarkesh Podcast, hosted by Dwarkesh Patel, are widely regarded as some of the best available “context dumps” on how modern models are trained, what reinforcement learning actually buys you, how far the current paradigm can scale, and where the bottlenecks are. He is candid about uncertainty in a field full of hype, and specific where most public commentary stays vague.
For someone learning AI today, Douglas is worth following because he sits at the intersection of two things most explanations separate: the low-level systems reality (parallelism, inference cost, hardware constraints) and the high-level capability story (why RL on verifiable tasks is unlocking reasoning and agency). He is bullish on AI progress and on AI coding in particular — he talks openly about the path toward AI “coworkers” — but he grounds that optimism in mechanics rather than vibes. That combination of engineering depth and clear explanation is exactly what makes him a rising, must-listen figure.
Key Articles & Papers
How to Scale Your Model: A Systems View of LLMs on TPUsVideos
Spotify Podcasts
YouTube