Ndea co-founder, AGI researcher, Keras creator
Francois Chollet
Profile
François Chollet is the rare AI figure who has shaped the field twice, in opposite directions. First as a builder: in March 2015, while working on computer vision problems, he released Keras — a Python API that made neural networks feel like Lego bricks instead of tensor algebra homework. It worked. Keras became the on-ramp for an entire generation of developers, got absorbed into TensorFlow as its official high-level API, and in its Keras 3 rewrite became framework-agnostic, running on top of JAX, TensorFlow, or PyTorch. If you learned deep learning between 2016 and 2022, you almost certainly typed model.add(Dense(...)) at some point. That was him. He also wrote Xception, the depthwise-separable-convolution architecture that quietly became a workhorse in production vision stacks.
Then, as a critic. Chollet’s second act is the argument that the thing he helped popularize is not, by itself, the road to general intelligence. His 2019 paper On the Measure of Intelligence reframed intelligence not as skill but as skill-acquisition efficiency — how fast a system picks up something genuinely new, relative to its priors and experience. A chess engine that crushes grandmasters isn’t intelligent; it’s a very good chess-shaped memory. To operationalize this he built ARC, the Abstraction and Reasoning Corpus: small colored-grid puzzles that any child solves in seconds and that GPT-3 scored a flat 0% on. In 2024 he and Zapier co-founder Mike Knoop turned it into the ARC Prize, seeding it with $1M of their own money. It has since become the benchmark frontier labs are most afraid to publish badly on.
In November 2024, after more than nine years at Google as a Senior Staff Engineer, Chollet left. In January 2025 he and Knoop launched Ndea, an AGI research lab that has since raised roughly $43M from Y Combinator (as a W2026 company), Coatue, Menlo Ventures, Factorial Capital and others. The thesis is specific and contrarian: intelligence comes from fusing deep learning’s pattern intuition with discrete program synthesis — searching the space of programs, guided by learned priors — rather than from another order of magnitude of transformer parameters. Ndea has roughly fifteen people, no product, and multi-year pure-research runway. That is either admirable discipline or an expensive way to be right too late, and Chollet would tell you he’s fine with the bet.
What makes him worth following if you’re learning AI today is that he is intellectually honest in public. When OpenAI’s o3 scored 75.7% on ARC-AGI’s semi-private set (and 87.5% in a high-compute configuration) in December 2024, Chollet — the loudest skeptic of LLM scaling — called it a genuine breakthrough, then explained precisely why it didn’t contradict him: o3 was doing deep-learning-guided program search at test time, not bigger pretraining. He also shortened his own AGI timeline publicly in 2025 rather than defending an old position. Meanwhile ARC-AGI-2 has proven far more resistant than v1 — ARC Prize’s own 2025 analysis put the top verified commercial model at 37.6% and the winning Kaggle entry at 24% — and ARC-AGI-3, the interactive agentic benchmark launched in March 2026, sits at 100% for humans and under 1% for frontier AI. For a developer, that gap is the most useful number in the field: it tells you exactly where today’s models stop being intelligent and start being retrieval.
Books
Key Articles & Papers
On the Measure of Intelligence Xception: Deep Learning with Depthwise Separable Convolutions OpenAI o3 Breakthrough High Score on ARC-AGI-Pub ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence ARC Prize 2025: Technical Report ARC Prize 2024: Technical Report The limitations of deep learning The future of deep learning The implausibility of intelligence explosionVideos
Controversies
Is ARC-AGI really an AGI test? Chollet and Knoop have been accused of overselling ARC as a measure of general intelligence when the definition of AGI is itself contested — some OpenAI staff have argued that by a “better than most humans at most tasks” standard, AGI already exists. Critics on both sides have also charged goalpost-moving: when v1 fell, v2 arrived; when v2 started cracking, v3 arrived. Chollet’s counter is that this is the point — his stated finish line is that AGI has arrived when building tasks that are easy for humans and hard for AI becomes impossible. Reasonable people disagree about whether that’s rigorous or unfalsifiable.
The o3 verification question. The December 2024 o3 result was run on a semi-private evaluation set in collaboration with OpenAI rather than by a fully independent third party, and o3 was tuned on ARC’s public training set. Chollet disclosed both facts, but contemporaneous reporting flagged the benchmark’s design flaws and the absence of independent replication. ARC-AGI-2’s evaluation protocol was tightened partly in response.
The scaling wars. Chollet spent years as the field’s most prominent “LLMs will not scale to AGI” voice, a position that put him in frequent public conflict with scaling advocates and made him an occasional ally of Gary Marcus. When reasoning models began clearing ARC-AGI-1 and he shortened his AGI timeline, some read it as a quiet concession that scaling won. His own framing — that test-time search and program synthesis, not pretraining scale, produced the jump — is defensible, but it’s fair to note he has been more right about the mechanism than about the pace.
Spotify Podcasts
YouTube