Meta VP Research in Generative AI; CMU professor
Russ Salakhutdinov
Profile
Russ Salakhutdinov — Ruslan to the paperwork, “Russ” to everyone else — is one of those researchers whose name you’ve probably read a hundred times without registering it, because it sits in the author list of half the foundational deep-learning papers you were taught from. He did his PhD at the University of Toronto under Geoffrey Hinton, and that lineage matters: his 2006 Science paper with Hinton on training deep autoencoders was one of the sparks that pulled neural networks out of their winter and made “deep learning” a serious phrase again. If you want to understand why representation learning became the whole ballgame, his early work on Boltzmann machines and autoencoders is where the intuitions were forged.
What makes Salakhutdinov worth studying for a developer is his range. He isn’t a one-idea researcher. He co-authored the dropout paper (with Hinton, Ilya Sutskever, and others) that every framework now ships as a one-line regularizer. He was on Show, Attend and Tell, an early and hugely influential demonstration of visual attention for image captioning, alongside Yoshua Bengio’s Montreal group. And when the field pivoted to language, his name shows up again on Transformer-XL and XLNet — architectures that pushed long-context modeling and autoregressive pretraining forward right as the LLM era was taking off. That’s a career spanning RBMs to Transformers, which is unusual and instructive.
He also does the thing academics rarely pull off well: he crosses into industry at scale without disappearing from research. From 2016 to 2020 he was Director of AI Research at Apple — a notoriously secretive shop — while keeping his UPMC Professorship in the Machine Learning Department at Carnegie Mellon. Since June 2024 he’s been Vice President of Research in Generative AI at Meta, where his mandate is multimodal LLMs and AI agents: systems that reason across text, images, and video and then go do things. That puts him at the center of the direction most people think the next few years of applied AI will take.
For someone learning AI today, the useful lesson from Salakhutdinov isn’t a single result — it’s the through-line. He keeps working on the same deep question (how do models discover structure in data on their own?) across every architectural fashion cycle. His lectures are genuinely good pedagogy, unusually clear about the statistical machinery underneath, and worth watching even now that the specific models have moved on. He works alongside Meta’s chief AI scientist Yann LeCun under Mark Zuckerberg’s enormous GenAI bet, which means what his team ships will touch a lot of the tooling developers actually use.
Key Articles & Papers
Reducing the Dimensionality of Data with Neural Networks Restricted Boltzmann Machines for Collaborative Filtering Deep Boltzmann Machines Dropout: A Simple Way to Prevent Neural Networks from Overfitting Show, Attend and Tell: Neural Image Caption Generation with Visual Attention Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context XLNet: Generalized Autoregressive Pretraining for Language Understanding
Videos
Spotify Podcasts
YouTube