PrometheusRoot
Blog Links Prometheans 100+ AI Books AI Companies Why are you here?
← Prometheans 100+
×
Jim Fan
builder
Researcher
X / Twitter
nvidiaembodied-aiagentsvoyagerrobotics

Related

pioneer Jensen Huang
← Prometheans 100+ Jim Fan

NVIDIA Director of AI, Project GR00T and embodied agent research

Jim Fan

Director of AI & Distinguished Scientist, Project GR00T Co-Lead — NVIDIA First Intern — OpenAI
Listen — profile
0:00 / 4:12

Profile

Linxi “Jim” Fan is NVIDIA’s Director of AI and a Distinguished Research Scientist, co-founder of the GEAR lab (Generalist Embodied Agent Research) alongside Yuke Zhu, and co-lead of Project GR00T — the company’s moonshot to build a foundation model for humanoid robots. His CV reads like a tour of every place AI happened: Columbia valedictorian, Stanford PhD under Fei-Fei Li, stints at Baidu AI Labs and MILA, and — the line he’s most known for — OpenAI’s very first intern, where he worked on World of Bits with Andrej Karpathy back when “web agents” sounded absurd. That last detail matters more than trivia: Fan has been chasing the same idea for a decade, which is agents that act in open-ended environments rather than models that answer questions.

The through-line from that decade is worth understanding if you’re building agents today. MineDojo (2022) argued that you could ground an agent in internet-scale knowledge — YouTube videos, wiki pages, forum posts — instead of a hand-crafted reward. Voyager (2023) took the next step and is still the cleanest demo of the pattern most agent frameworks converged on: GPT-4 writing JavaScript, running it, reading the error, rewriting it, and committing the working version to a persistent skill library it can retrieve later. No gradient updates, no fine-tuning — lifelong learning as code plus memory. Then Eureka closed the loop the other way: an LLM writing reward functions, evaluated at massive parallelism in GPU simulation, beating human-expert rewards on 83% of 29 environments. If you’ve ever wondered what the actual research lineage of “agentic loop with tools and memory” is, it runs through these papers.

Since 2024 that agenda has moved from Minecraft to motors. GR00T N1, released open-weight at GTC 2025, is a 2B-parameter vision-language-action model with a deliberately dual-system design — a VLM that reasons slowly about the scene and instruction, and a diffusion transformer that emits joint actions at real-time frequency. The bottleneck it exists to solve is data: there is no internet of robot trajectories, so GEAR manufactures one. DreamGen fine-tunes video world models to hallucinate photorealistic robot footage, then recovers pseudo-actions with an inverse-dynamics model — “neural trajectories.” A robot trained on a single pick-and-place demo picked up 22 new behaviors this way. The subsequent GR00T N1.5 was reportedly built in 36 hours on synthetic data that would have taken months to teleoperate. Whether or not you buy the humanoid thesis, that’s the most interesting bet in robotics right now: that generative video is the pretraining corpus physical AI never had.

Fan is also, functionally, robotics’ most effective explainer — his “Physical Turing Test” framing (come home to a spotless apartment and a candlelit dinner, and you can’t tell who did it) has become the field’s shorthand goal. In 2026 he’s sharpened into something more contrarian: language-first VLA models are “head-heavy” and fumble the verbs of physics; teleoperation is a dead end and now makes up under 0.1% of NVIDIA’s training mix; the future is video-first world-action models predicting future pixels and joint torques together. He predicts a passed Physical Turing Test in 2–3 years and a completed robotics “technology tree” by 2040. Read him with the discount you’d apply to anyone whose employer sells the compute — but read him. Unlike most people making these claims, he ships open weights, open datasets, and open code.

Key Articles & Papers

Voyager: An Open-Ended Embodied Agent with Large Language Models 2023 — The first LLM-powered lifelong learning agent — code-writing, self-verification, and a persistent skill library. The blueprint most agent frameworks are still copying. MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge 2022 — NeurIPS Outstanding Paper. Grounded agent learning in YouTube videos and wikis instead of hand-designed rewards — the springboard for everything after. Eureka: Human-Level Reward Design via Coding Large Language Models 2023 — GPT-4 writes reward functions, GPU simulation grades them at scale, and the loop beats human experts on 83% of tasks. Including robot-hand pen spinning. GR00T N1: An Open Foundation Model for Generalist Humanoid Robots 2025 — A 2B-parameter dual-system VLA — slow reasoning VLM plus real-time diffusion action head — released open-weight. The reference architecture for humanoid policies. DreamGen: Unlocking Generalization in Robot Learning through Neural Trajectories 2025 — Video world models as a robot data factory: 22 new behaviors learned from a single pick-and-place demonstration. VIMA: General Robot Manipulation with Multimodal Prompts 2022 — Interleaved text-and-image prompting for manipulation — an early argument that robot tasks are a prompting problem, not a policy-architecture problem. The Physical Turing Test: Jim Fan on NVIDIA's Roadmap for Embodied AI 2025 — The essay version of his signature framing, plus the simulation-scaling argument behind it. The best single-sitting summary of his worldview. GR00T N1.5: An Improved Open Foundation Model for Generalist Humanoid Robots 2025 — The synthetic-data payoff — a model iteration built in 36 hours on DreamGen output rather than months of teleoperation. NVIDIA Isaac GR00T (open-source repo) 2025 — Weights, fine-tuning scripts, and datasets. The fastest way to actually run a humanoid foundation model rather than read about one. A Mine-Blowing Breakthrough: AI Agent Voyager Autonomously Plays Minecraft 2023 — NVIDIA's accessible write-up of Voyager — the right starting point before the paper.

Videos

YouTube video
YouTube video
YouTube video
YouTube video
YouTube video
YouTube video

Controversies

There is no personal scandal here, but there is a real and ongoing argument about credibility in humanoid robotics, and Fan sits on both sides of it.

The messenger problem. Fan is the most quoted voice predicting near-term physical AGI while working for the company that sells the training hardware, the simulator, and now the robot foundation models. Timelines like “Physical Turing Test in 2–3 years” and “technology tree complete by 2040” are unfalsifiable in the short run and commercially useful in the short run — a combination worth noting. Skeptics point out that GR00T demos are shown on partner hardware in curated settings, and that impressive lab manipulation has historically not survived contact with unstructured homes.

Demo integrity — where he’s the critic. To his credit, Fan has been among the loudest people calling out the field’s evaluation hygiene. He has argued that many viral robot videos are the single best take out of hundreds, that the field has no shared reproducible benchmark, and that “in 2026 we must do better and stop treating reproducibility and scientific discipline as second-class citizens.” He has separately vouched publicly for specific third-party demos being real rather than CGI — useful, but also a reminder that the field currently relies on personal say-so instead of standardized evaluation.

Verdict for developers: treat his timelines as advocacy and his artifacts as evidence. The papers, weights, and datasets are open and reproducible; the predictions attached to them are not yet.

Spotify Podcasts

Robotics' End Game with Nvidia's Jim Fan
Robotics' End Game with Nvidia's Jim Fan
Talking Books
2026
Robotics' End Game: Nvidia's Jim Fan
Robotics' End Game: Nvidia's Jim Fan
AI Ascent
2026
Radical Talks, Masterclass Edition: Jim Fan on The Future of Robotics and Embodied AI
Radical Talks, Masterclass Edition: Jim Fan on The Future of Robotics and Embodied AI
Radical Talks
2026
Robotics Will Be Solved by 2040.” — Jim Fan, NVIDIA
Robotics Will Be Solved by 2040.” — Jim Fan, NVIDIA
DROIDS Newsletter
2025
The Physical Turing Test: Jim Fan on Nvidia's Roadmap for Embodied AI
The Physical Turing Test: Jim Fan on Nvidia's Roadmap for Embodied AI
AI Ascent
2025
Jim Fan on Nvidia’s Embodied AI Lab and Jensen Huang’s Prediction that All Robots will be Autonomous
Jim Fan on Nvidia’s Embodied AI Lab and Jensen Huang’s Prediction that All Robots will be Autonomous
Training Data
2024
Episode 10: Dr. Jim Fan, NVIDIA AI
Episode 10: Dr. Jim Fan, NVIDIA AI
Thursday Nights in AI
2023
NVIDIA’s Jim Fan Delves Into Large Language Models and Their Industry Impact - Ep. 204
NVIDIA’s Jim Fan Delves Into Large Language Models and Their Industry Impact - Ep. 204
NVIDIA AI Podcast
2023
01: Applications of Generative AI with NVIDIA's Jim Fan
01: Applications of Generative AI with NVIDIA's Jim Fan
Replit AI Podcast
2023
Jim Fan, NVIDIA: Foundation models for embodied agents, scaling data, and why prompt engineering will become irrelevant
Jim Fan, NVIDIA: Foundation models for embodied agents, scaling data, and why prompt engineering will become irrelevant
Generally Intelligent
2023

YouTube

YouTube video
2026
YouTube video
2026
YouTube video
2024

Related People

pioneer Jensen Huang
© 2026 PrometheusRoot