Nvidia Chief Software Architect, Groq founder and LPU creator
Jonathan Ross
Profile
Jonathan Ross is the rare chip architect whose fingerprints are on two of the most important pieces of silicon in modern AI. At Google, as a side “20% time” project, he initiated and helped design the Tensor Processing Unit (TPU) — the custom accelerator that went on to power a majority of Google’s internal compute and, famously, gave DeepMind’s AlphaGo the inference muscle to beat Lee Sedol in 2016. Watching AlphaGo’s strength climb as its inference got faster taught Ross the lesson that has defined his career ever since: in the deployment phase of AI, speed isn’t a luxury, it’s a capability. A self-taught engineer who never finished a traditional degree path, he left Google in 2016 to bet everything on that idea.
That bet became Groq, where Ross created the Language Processing Unit (LPU) — an architecture built from the opposite direction to a GPU. Where Nvidia’s GPUs are massively parallel, memory-hungry, and brilliant at training, the LPU is a deterministic, software-scheduled, on-chip-SRAM design purpose-built for one thing: generating tokens fast, with sub-millisecond, predictable latency. For developers, the practical payoff arrived through GroqCloud, where open models like Llama ran at hundreds of tokens per second — fast enough that real-time voice agents, low-latency tool use, and interactive reasoning stopped feeling like demos and started feeling like products. Ross’s framing is worth internalizing if you build with LLMs: training is bulk hauling, inference is last-mile delivery, and most of the value users actually feel lives in that last mile.
The company grew aggressively — a $2.8B valuation in 2024, followed by a headline-grabbing $1.5 billion commitment from Saudi Arabia and an Aramco Digital partnership that stood up one of the region’s largest inference clusters in Dammam. Then came the twist: in December 2025, Nvidia — the very incumbent Groq positioned itself against — struck a roughly $20 billion non-exclusive licensing deal for Groq’s inference technology, folding Ross and Groq president Sunny Madra into the company. Ross is now Nvidia’s Chief Software Architect, tasked with integrating LPU-style inference alongside Nvidia’s Vera Rubin GPU generation.
For someone learning AI today, Ross matters because he is the clearest embodiment of a structural shift the field is living through: the center of gravity is moving from training to inference. The training race made Jensen Huang the most important man in hardware; the inference race is where cost, latency, and user experience get decided — and it’s where Ross has spent a decade arguing that the GPU is not the only answer. That Nvidia ultimately paid to bring his answer in-house, rather than merely out-compete it, is the most eloquent endorsement of the LPU thesis anyone could have written.
Key Articles & Papers
Sometimes You Don't Want A GPU: Groq Cofounder Explains Whirlwind Deal With Nvidia Nvidia's 'Aqui-Hire' of Groq Marks Its Entrance Into Non-GPU AI Inference Saudi Arabia Announces $1.5 Billion Expansion to Fuel AI-powered Economy with Groq The Future of AI Compute: A Conversation With Jonathan Ross With Groq, Jonathan Ross Is Taking AI Inference to New Speeds Groq Sends Elon's 'Grok' a Cease & Desist (a Funny One)Videos
Controversies
Nothing scandalous attaches to Ross personally, but two points of friction are worth noting fairly.
The Groq vs. “Grok” naming clash: when Elon Musk’s xAI launched a chatbot called Grok in 2023, Ross’s Groq — which held the earlier trademark — sent a public cease-and-desist, delivered with tongue-in-cheek humor (he suggested Musk name it “Slartibartfast” instead). It’s less a scandal than a reminder that the two names remain a persistent source of public confusion; the serious deepfake lawsuits surrounding Musk’s Grok have nothing to do with Ross or Groq.
The inference benchmark hype: Groq’s marketing leaned hard on eye-catching tokens-per-second figures, and skeptics have long argued that LPU comparisons flatter the architecture by choosing favorable models and batch conditions, and that its SRAM-heavy design requires many chips to hold large models — raising real questions about cost-at-scale versus GPUs. The Nvidia deal arguably settles the technical debate in Ross’s favor, but the economics of inference silicon remain genuinely contested territory.
Spotify Podcasts
YouTube