UC Berkeley professor, deep RL for robotic control
Sergey Levine
Profile
Sergey Levine is, more than almost anyone else working today, the person who turned “robots that learn” from a research aspiration into an engineering discipline. An Associate Professor of EECS at UC Berkeley, where he runs the Robotic AI & Learning (RAIL) lab within Berkeley AI Research, he has spent a decade building the algorithmic toolkit that most robot-learning systems now quietly depend on. If you have read a paper about a robot arm learning to grasp from raw pixels, an off-policy actor-critic that actually trains stably, or a “vision-language-action” model, there is a good chance Levine’s name is on the citation trail — or on the paper itself.
His research arc is unusually coherent. After a PhD at Stanford (advised by Vladlen Koltun) and a postdoc at Berkeley with Pieter Abbeel, Levine planted a flag on a hard idea: that perception and control should be learned together, end-to-end, rather than bolted together by hand. His 2016 “End-to-End Training of Deep Visuomotor Policies” was a genuine turning point — a single neural network mapping camera pixels to motor torques. From there he became the connective tissue of modern deep RL: Soft Actor-Critic (SAC), co-developed with Tuomas Haarnoja and Abbeel, remains one of the default continuous-control algorithms because it’s the rare method that just works; QT-Opt showed that self-supervised RL could scale to hundreds of thousands of real grasps; and his Conservative Q-Learning and offline-RL agenda made it respectable to learn policies from logged data instead of endless live trial-and-error.
Levine also matters because of who he trains and what he teaches. His CS285: Deep Reinforcement Learning course is, for a large fraction of practitioners, the way they actually learned deep RL — the lectures are free on YouTube and the notes are widely treated as canonical. His collaborators and students — including Chelsea Finn and Karol Hausman — form a good chunk of the current robot-learning establishment. Through his long affiliation with Google (and Google DeepMind), he was central to the Robotics Transformer line — RT-1 and RT-2 — that reframed robot control as a sequence-modeling problem and showed web-scale knowledge could transfer into physical action.
Since 2024 his center of gravity has shifted to industry: he co-founded Physical Intelligence (with Hausman, Finn, and others), a startup building general-purpose robotic foundation models. Its π0 (“pi-zero”) model — a vision-language model fused with a diffusion-based action expert — and its successors (π0.5, π0.6, π0.7) are among the most credible attempts at a single “robot brain” that folds laundry, makes coffee, and assembles boxes without task-specific engineering. The company has raised over a billion dollars at a multibillion-dollar valuation. For a developer, Levine is worth studying because he sits exactly at the seam where the LLM playbook — pretraining, scaling, foundation models — is being ported into the messy, unforgiving physical world, and he’s been more clear-eyed than most about what does and doesn’t transfer.
Key Articles & Papers
End-to-End Training of Deep Visuomotor Policies Soft Actor-Critic: Off-Policy Maximum Entropy Deep RL with a Stochastic Actor QT-Opt: Scalable Deep RL for Vision-Based Robotic Manipulation Conservative Q-Learning for Offline Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems RT-1: Robotics Transformer for Real-World Control at Scale RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control π0: A Vision-Language-Action Flow Model for General Robot Control Offline RL and Large Language Models (Substack)Videos
Spotify Podcasts
YouTube