PrometheusRoot
Blog Links Prometheans 100+ AI Books AI Companies Why are you here?
← Prometheans 100+
×
Ashish Vaswani
builder
ResearcherFounder
X / Twitter Wikipedia
transformerattentiongoogle-brainessential-ai

Related

builder Jeff Dean
← Prometheans 100+ Ashish Vaswani

Co-Founder and CEO of Essential AI, transformer pioneer

Ashish Vaswani

Co-Founder and CEO — Essential AI Researcher — Google Brain Co-Founder — Adept AI
Listen — profile
0:00 / 4:13

Profile

Ashish Vaswani is the name at the top of Attention Is All You Need, the 2017 Google Brain paper that killed the recurrent neural network and replaced it with the Transformer. If you are learning AI today, essentially everything you touch descends from that eight-page architecture description: GPT, Claude, Gemini, Llama, Stable Diffusion, Whisper, AlphaFold’s Evoformer, the ViT in your image pipeline. It is one of the rare cases where a single paper is not just influential but load-bearing — over 100,000 citations, and the actual blueprint people still implement line for line. Worth knowing, though, that the paper’s author list carries a footnote declaring equal contribution and randomized ordering; Vaswani is “first author” partly by draw. He has been consistent about that publicly. The idea belonged to a room of eight people, several of whom — Noam Shazeer, Aidan Gomez, Niki Parmar, Jakob Uszkoreit — went on to build companies of their own.

The path there was unglamorous and worth noting for anyone who thinks breakthroughs come from prodigies. Vaswani did his B.Tech at BIT Mesra in 2002, a PhD at USC under David Chiang finishing in 2014 — statistical machine translation, the pre-deep-learning world — then USC’s Information Sciences Institute, then Google Brain from 2016 to 2021. He was a mid-career researcher grinding on translation quality when the Transformer landed. After Google he co-founded Adept AI as chief scientist, working on agents that drive software, then left in early 2023 with Parmar. The stated split was about thesis: Adept had committed to an enterprise-agent product surface, and Vaswani wanted to go back to foundational research.

That research bet became Essential AI, founded in San Francisco in 2023, which raised $56.5M in a Series A led by March Capital in December 2023 with Google, Nvidia, AMD, and Thrive Capital all participating — an unusual coalition of chip rivals betting on the same team. Essential’s output is genuinely useful to developers, not just impressive-sounding. Essential-Web v1.0 is 24 trillion tokens of Common Crawl with every one of its 23.6 billion documents labeled across a twelve-category taxonomy, so you can build a domain-specific pretraining corpus with SQL-style filters instead of training your own classifier fleet. It hit #1 on Hugging Face. Then in December 2025 came Rnj-1 (“range-one,” after Ramanujan), an 8B open-weight model on a Gemma-3-style architecture that trades blows with far larger models on coding and math benchmarks — the kind of thing you can actually run on one consumer GPU.

The bet underneath all of it is contrarian, and Vaswani has said so out loud: he thinks the pure-scaling orthodoxy at the closed frontier labs is crowding out the architectural search that produced the Transformer in the first place, and that Big Tech is blinding itself to the next breakthrough. Which makes the ending, so far, complicated. In June 2026 Nvidia quietly acqui-hired Vaswani and a slice of the Essential AI team; reporting says he is now working on Nvidia’s open-source Nemotron models, that fundraising had gotten hard, and that pulling him out of AMD’s orbit was part of Jensen Huang’s motivation. The man who warned about big tech absorbing the research frontier is now inside the biggest company in it — though at least on the open-weights side of it. For developers, the practical read is this: the datasets and the model are open and still on Hugging Face, and the person who wrote the architecture you use every day now shapes the most widely-distributed open model family that ships with the hardware.

Key Articles & Papers

Attention Is All You Need 2017 — The Transformer. Self-attention replaces recurrence, training parallelizes, and every modern LLM follows. Still the single most valuable paper a developer can read end to end. Image Transformer 2018 — First serious attempt to apply self-attention to image generation, four years before ViT-style architectures became the default in vision. Self-Attention with Relative Position Representations 2018 — Replaced absolute sinusoidal position encodings with relative ones — the ancestor of RoPE and ALiBi, and the reason long-context models work at all. Tensor2Tensor for Neural Machine Translation 2018 — The library that made the Transformer reproducible. A lesson in how shipping usable code, not just a paper, is what spreads an architecture. Stand-Alone Self-Attention in Vision Models 2019 — Showed convolutions could be replaced outright by attention in vision backbones — a key stepping stone toward Vision Transformers. Scaling Local Self-Attention for Parameter Efficient Visual Backbones 2021 — HaloNets. Vaswani first-authored this one: blocked local attention that scales without quadratic cost, an early answer to the context-length problem. Rethinking Reflection in Pre-Training 2025 — Argues self-correction emerges during pretraining, not just RL post-training — a direct challenge to the reasoning-comes-from-RLHF consensus. Practical Efficiency of Muon for Pretraining 2025 — Hard empirical work on whether the Muon optimizer actually beats AdamW at scale. Useful if you train models rather than just fine-tune them. Essential-Web v1.0: 24T Tokens of Organized Web Data 2025 — A fully labeled 24T-token corpus you can filter with SQL to build math, code, STEM, or medical pretraining sets. Genuinely reusable infrastructure. Announcing Rnj-1: Building Instruments of Intelligence 2025 — Essential AI's open-weight 8B model, competitive with much larger models on coding and math. The thesis that architecture beats brute scale, made concrete.

Videos

YouTube video
YouTube video
YouTube video
YouTube video
YouTube video

Controversies

Who actually invented the Transformer. Credit for the 2017 paper has been contested for years. The author list is randomized with an equal-contribution footnote, so “Vaswani et al.” overstates a single person’s role — Uszkoreit is often credited with the original self-attention idea, Shazeer with much of the implementation that made it work. Separately, Jürgen Schmidhuber has argued that his 1990s fast-weight programmers anticipate linear attention. Vaswani himself has generally deflected credit to the group, which is more than can be said for some of the participants in these disputes.

Joining Nvidia after criticizing Big Tech. In a September 2025 Bloomberg profile, Vaswani positioned Essential AI as a counterweight to closed frontier labs and warned that transformer-scaling orthodoxy at big companies was stifling innovation. Nine months later he joined Nvidia via acqui-hire, with reporting citing fundraising difficulty and Nvidia’s interest in pulling him away from AMD, an early Essential AI investor. The charitable reading — that Nemotron is genuinely open-weight and reaches more developers than a struggling startup could — is defensible, but it’s a real reversal and the fate of Essential AI’s independent research program is unclear. Note also that neither company confirmed the move on the record; the reporting rests on sources and LinkedIn changes.

Spotify Podcasts

36 - Attention Is All You Need, with Ashish Vaswani and Jakob Uszkoreit
36 - Attention Is All You Need, with Ashish Vaswani and Jakob Uszkoreit
NLP Highlights
2017

YouTube

YouTube video
2026
YouTube video
2025
YouTube video
2024
YouTube video
2019

Related People

builder Jeff Dean
© 2026 PrometheusRoot