Physical Intelligence CEO, foundation models for robotics
Karol Hausman
Profile
Karol Hausman is the co-founder and CEO of Physical Intelligence, the San Francisco startup trying to do for robots what GPT did for text: build a single foundation model — a “brain” — that can drive many different machines through many different tasks, rather than hand-coding one robot for one job. If you believe robotics has been stuck not on motors and actuators but on intelligence, Hausman is one of the people betting his career on that thesis, and he has raised more than a billion dollars to test it. Physical Intelligence reportedly reached a valuation around $5.6 billion by late 2025, barely a year after founding — a sign of how much capital is chasing the “physical AI” story.
Hausman’s path is a classic robotics-learning pedigree. He earned a Ph.D. in computer science at USC (advised by Gaurav Sukhatme), with earlier degrees from the Technical University of Munich and Warsaw University of Technology, then spent years at Google Brain and Google DeepMind as a staff research scientist. That’s where he did the work most developers will recognize him for: helping connect large language and vision-language models to actual robot hardware. He was a contributor on SayCan (using an LLM as a high-level planner grounded by what a robot can physically do), PaLM-E (an embodied multimodal model), and the influential RT-2, which showed a vision-language model could be turned into a vision-language-action model that transfers web knowledge directly into robot control. He remains an adjunct professor at Stanford, where he co-taught the deep reinforcement learning course CS 224R.
In 2024 he left DeepMind to co-found Physical Intelligence alongside heavyweights from the same world — including Sergey Levine and Chelsea Finn, longtime collaborators in robot learning. The company’s flagship results are the π (pi) models: π0, a vision-language-action flow model for general robot control, and π0.5, which pushed toward genuine open-world generalization — a mobile robot cleaning kitchens and bedrooms in homes it had never seen in training. That “never seen before” detail is the whole point: the field is littered with demos that work only in the exact room they were recorded in, and Physical Intelligence’s pitch is that co-training on heterogeneous data (multiple robots, web data, high-level semantic prediction) breaks that curse the way scale broke it for language.
For developers, Hausman matters because he sits at the frontier where the transformer playbook meets the messy physical world. The open question he’s staking everything on — does the scaling hypothesis hold for action, not just tokens? — is one of the most consequential unanswered bets in AI. He’s refreshingly candid that robots still fail at tasks a toddler finds trivial, and that no one has yet proven foundation models fully close that gap. Watch this space: if VLA models scale the way LLMs did, Hausman will have been early; if physical intelligence turns out to need more than data and compute, his work will have mapped exactly where the wall is.
Key Articles & Papers
π0: A Vision-Language-Action Flow Model for General Robot Control π0.5: A Vision-Language-Action Model with Open-World Generalization RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan) PaLM-E: An Embodied Multimodal Language Model Open X-Embodiment: Robotic Learning Datasets and RT-X Models Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement LearningVideos
Spotify Podcasts
YouTube