Co-Founder and CEO of Essential AI, transformer pioneer
Ashish Vaswani
Profile
Ashish Vaswani is the name at the top of Attention Is All You Need, the 2017 Google Brain paper that killed the recurrent neural network and replaced it with the Transformer. If you are learning AI today, essentially everything you touch descends from that eight-page architecture description: GPT, Claude, Gemini, Llama, Stable Diffusion, Whisper, AlphaFold’s Evoformer, the ViT in your image pipeline. It is one of the rare cases where a single paper is not just influential but load-bearing — over 100,000 citations, and the actual blueprint people still implement line for line. Worth knowing, though, that the paper’s author list carries a footnote declaring equal contribution and randomized ordering; Vaswani is “first author” partly by draw. He has been consistent about that publicly. The idea belonged to a room of eight people, several of whom — Noam Shazeer, Aidan Gomez, Niki Parmar, Jakob Uszkoreit — went on to build companies of their own.
The path there was unglamorous and worth noting for anyone who thinks breakthroughs come from prodigies. Vaswani did his B.Tech at BIT Mesra in 2002, a PhD at USC under David Chiang finishing in 2014 — statistical machine translation, the pre-deep-learning world — then USC’s Information Sciences Institute, then Google Brain from 2016 to 2021. He was a mid-career researcher grinding on translation quality when the Transformer landed. After Google he co-founded Adept AI as chief scientist, working on agents that drive software, then left in early 2023 with Parmar. The stated split was about thesis: Adept had committed to an enterprise-agent product surface, and Vaswani wanted to go back to foundational research.
That research bet became Essential AI, founded in San Francisco in 2023, which raised $56.5M in a Series A led by March Capital in December 2023 with Google, Nvidia, AMD, and Thrive Capital all participating — an unusual coalition of chip rivals betting on the same team. Essential’s output is genuinely useful to developers, not just impressive-sounding. Essential-Web v1.0 is 24 trillion tokens of Common Crawl with every one of its 23.6 billion documents labeled across a twelve-category taxonomy, so you can build a domain-specific pretraining corpus with SQL-style filters instead of training your own classifier fleet. It hit #1 on Hugging Face. Then in December 2025 came Rnj-1 (“range-one,” after Ramanujan), an 8B open-weight model on a Gemma-3-style architecture that trades blows with far larger models on coding and math benchmarks — the kind of thing you can actually run on one consumer GPU.
The bet underneath all of it is contrarian, and Vaswani has said so out loud: he thinks the pure-scaling orthodoxy at the closed frontier labs is crowding out the architectural search that produced the Transformer in the first place, and that Big Tech is blinding itself to the next breakthrough. Which makes the ending, so far, complicated. In June 2026 Nvidia quietly acqui-hired Vaswani and a slice of the Essential AI team; reporting says he is now working on Nvidia’s open-source Nemotron models, that fundraising had gotten hard, and that pulling him out of AMD’s orbit was part of Jensen Huang’s motivation. The man who warned about big tech absorbing the research frontier is now inside the biggest company in it — though at least on the open-weights side of it. For developers, the practical read is this: the datasets and the model are open and still on Hugging Face, and the person who wrote the architecture you use every day now shapes the most widely-distributed open model family that ships with the hardware.
Key Articles & Papers
Attention Is All You Need Image Transformer Self-Attention with Relative Position Representations Tensor2Tensor for Neural Machine Translation Stand-Alone Self-Attention in Vision Models Scaling Local Self-Attention for Parameter Efficient Visual Backbones Rethinking Reflection in Pre-Training Practical Efficiency of Muon for Pretraining Essential-Web v1.0: 24T Tokens of Organized Web Data Announcing Rnj-1: Building Instruments of IntelligenceVideos
Controversies
Who actually invented the Transformer. Credit for the 2017 paper has been contested for years. The author list is randomized with an equal-contribution footnote, so “Vaswani et al.” overstates a single person’s role — Uszkoreit is often credited with the original self-attention idea, Shazeer with much of the implementation that made it work. Separately, Jürgen Schmidhuber has argued that his 1990s fast-weight programmers anticipate linear attention. Vaswani himself has generally deflected credit to the group, which is more than can be said for some of the participants in these disputes.
Joining Nvidia after criticizing Big Tech. In a September 2025 Bloomberg profile, Vaswani positioned Essential AI as a counterweight to closed frontier labs and warned that transformer-scaling orthodoxy at big companies was stifling innovation. Nine months later he joined Nvidia via acqui-hire, with reporting citing fundraising difficulty and Nvidia’s interest in pulling him away from AMD, an early Essential AI investor. The charitable reading — that Nemotron is genuinely open-weight and reaches more developers than a struggling startup could — is defensible, but it’s a real reversal and the fate of Essential AI’s independent research program is unclear. Note also that neither company confirmed the move on the record; the reporting rests on sources and LinkedIn changes.
Spotify Podcasts
YouTube