- 1 hour 37 minutesSi Sheppard – How did a few hundred Spanish soldiers topple two empires?
New episode with military historian Si Sheppard.
The conquests of the Aztec and Inca empires might be the most shocking events in human history.
Hernan Cortés and some hundreds of Conquistadors landed on a new continent they knew basically nothing about, and within 2 and a half years, conquered the 6 million strong Aztec empire.
Then, a decade later, Francisco Pizarro did the same thing to an Incan Empire of some 10 million inhabitants.
In both cases, the Conquistadors consistently beat native armies while being literally 100x (or even in some cases 1000x) outnumbered - and in many cases, did so without taking a single casualty.
And they did this with 99% of their military forces being made up of the very natives they were trying to subjugate.
We often forget how wild human history is. People talk about AI takeover as if it would be some crazy, unprecedented sci-fi outcome. But this would not be the first time an alien force whose motives the defenders barely understand use more advanced technology and diplomatic cunning to outmaneuver and then take over a much larger and long established existing order.
Watch on YouTube; read the transcript.
Sponsors
* Jane Street’s intern projects are even more ambitious than I realized. One intern spent his summer researching LLM memorization, figuring out how to quantify and reduce it. This matters because, when Jane Street tests their models on historical data, they need to know that performance reflects predictive power, not just memorized outcomes. Another intern spent his time developing a new activation checkpointing scheme: his custom version ended up outperforming PyTorch’s native one. If you want to read about more 2026 intern projects or apply for Jane Street’s 2027 internships, go to janestreet.com/dwarkesh
* Grok Bot can tackle the tasks you’ve been meaning to handle but just haven’t gotten around to. For example, SpaceXAI recently gave Grok Bot access to their procurement contracts, vendor spend, and usage data: it ended up identifying over $100,000 in potential savings for them. I decided to try this on my own spend, and Grok Bot found over $5,000 of savings for me, even though I run a pretty small team. Try it out yourself at x.ai/bot
* Antithesis solves one of the biggest problems with traditional software testing: to test for a bug, you first have to be able to imagine it. Antithesis runs your system through thousands of parallel histories and steers your tests for you, injecting faults and generally trying to make everything as adversarial as possible. That way, your software gets tested against scenarios you’d never imagine but that could still arise in the real world. Learn more at antithesis.com/dwarkesh
Timestamps
(00:00:00) – Cortés’s conquest of the Aztecs
(00:22:39) – Pizarro’s conquest of the Inca
(00:26:59) – Explaining Conquistador military superiority
(00:34:39) – Could the natives have beat the Conquistadors?
(00:48:56) – The Aztecs and Incas did not get to learn from each other
(00:56:11) – Why Spain still lost in the long run
(01:07:47) – Why didn’t the Aztecs/Incas adapt better to European warfare?
(01:14:16) – Conquistador motivation
(01:26:22) – Why Mughal India also couldn’t play the Europeans off each other
(01:33:42) – Diplomacy was just as important as technology
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com1 October 2026, 3:28 pm - 1 hour 20 minutesNoam Brown – Agent swarms, alignment, & recursive self-improvement
New episode with Noam Brown.
We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research.
And we also discuss how we will know if the models are actually aligned before we kick off RSI.
Watch on YouTube; read the transcript.
Sponsors
* Jane Street has been interested in AI for a lot longer than you’d think, and not just for trading. In 2011, a full year before AlexNet and over a decade before ChatGPT launched, they hosted the first FOOM Debate between Eliezer Yudkowsky and Robin Hanson on whether AI would lead to an intelligence explosion. Now Jane Street is revisiting the question with a new panel: Daniel Kokotajlo, Ege Erdil, Ryan Greenblatt, and Jaime Sevilla, hosted by Ron Minsky in San Francisco this October. I expect it to be a truly excellent conversation. Register at janestreet.com/dwarkesh
* Grok Bot has made handing off work super easy. It runs on its own cloud computer, where it installs the tools it needs to handle tasks end-to-end. For the podcast, we use Grok Bot to help produce our videos. You may have noticed that our ads feature animations of real websites. Getting these pixel-perfect used to mean running a convoluted, multi-step workflow ourselves. Now we just let Grok Bot handle it. Best of all, Grok Bot has learned all of our specs and preferences, so we don’t have to redescribe the task each time! Try Grok Bot for yourself at x.ai/bot
* Antithesis gives you the confidence of a giant test suite without actually having to write one. Say you’re doing a major backend refactor: building enough tests to trust it could take weeks. Antithesis solves this by running your software through countless simulated worlds, injecting faults and hunting for failures. On any PR, you can turn a dial to decide exactly how much testing you want. And because every run is fully deterministic, agents can branch off the moment a bug appears, rewind it, inspect memory, and replay it, all while the original test keeps running. Learn more at antithesis.com/dwarkesh
Timestamps
(00:00:00) – Multi-agent and Navier-Stokes
(00:15:28) – How will AI firms work?
(00:22:02) – What math progress tells us about recursive self improvement
(00:40:22) – Hugging Face and alignment
(01:01:18) – The internal/external model gap
(01:08:34) – Chain of thought is degrading
(01:14:12) – How will we know when alignment is solved?
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com17 September 2026, 3:38 pm - 1 hour 37 minutesAI researchers debate how close we are to recursive self-improvement
New episode with John Schulman, Beren Millidge and Charlie O’Neill. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next.
Watch on YouTube; read the transcript.
Sponsors
* Antithesis helps you trust your code. As agents generate more and more of your software, the bottleneck shifts from your engineers actually writing code to verifying it. Antithesis does that testing for you. Ron Minsky, who co-leads Jane Street’s tech group, told me that Antithesis was able to help his team shake out bugs in software that had already undergone heavy review. If you want to see how it fits into your development process, go to antithesis.com/dwarkesh
* Grok Bot has been a great way to hand off tasks. My team uses it as a producer: whenever my editor posts a rough cut of an interview in Slack, Grok Bot opens the transcript on its own computer, matches my notes to the exact moments they refer to, and uses a file of my preferences to suggest edits. Then it sends me its top clip candidates so I can review everything from my phone, which saves my editors from sorting through hours of footage. Try Grok Bot for yourself at x.ai/bot
* Jane Street just launched its most ambitious competition yet: design a protocol-emulator ASIC. Basically, if you have a chip you want to test outside of a live system, you should be able to connect it to your design and have it simulate realistic traffic. Jane Street wants general-purpose, reprogrammable designs that can work across multiple protocols and remain useful as new ones emerge. The most novel submissions will actually get taped out, and the winners will receive a physical copy! The competition is open until January 18, 2027, and teams are encouraged. To get started download the template code at janestreet.com/dwarkesh
Timestamps
(00:00:00) – Steelmanning the case against RSI
(00:18:39) – What’s driving the Chinese labs’ progress
(00:28:06) – How will automated AI researchers be trained
(00:33:51) – Will long-horizon RL elicit AGI?
(00:45:24) – The sim-to-real gap
(01:00:33) – How much progress is explained by data?
(01:18:03) – Why is RL working so well?
(01:24:54) – Move 37 and entropy collapse
(01:28:32) – Rapid-fire timelines
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com11 September 2026, 4:28 pm - 2 hours 20 minutesAjeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
Ajeya Cotra is a researcher at METR, where she works on threat modeling for loss-of-control risks from advanced AI. Before that, she led the technical AI safety program at what is now Coefficient Giving.
She is one the three authors of METR and Redwood Research’s “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident”.
We go through not only what she and her coauthors discovered during this investigation, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement.
Watch on YouTube; read the transcript.
Sponsors
* Jane Street’s ML engineering internships start with an intense four-day bootcamp: PyTorch, autograd, writing kernels, profiling workloads… all the things that Jane Street engineers need to know for their daily work. After that, interns tackle real projects, things the firm actually wants in its codebase. If you want to apply, or if you want to watch my recent conversation with Axel, one of Jane Street’s ML engineers, go to janestreet.com/dwarkesh
* Cursor, which is now part of SpaceX, noticed that their MoE layers were eating more than half of total training time. So they wrote and open-sourced Mixture-of-Kittens, which is a custom megakernel for training MoE models on NVL72s. This kernel sped up an end-to-end run across 512 GPUs by 1.4x, from about 760 to over 1000 tokens per second per GPU. If you want to read more about the ML research that Cursor and SpaceX are doing, go to cursor.com/dwarkesh
* Antithesis hands you (or your agents) a bug’s root cause so you can avoid days of manual debugging. If your test run crashes, Antithesis rewinds, branches off hundreds of slightly varied rollouts, and checks in how many of them the crash still appears. Then it rewinds further and does this all again. As Antithesis rewinds, it eventually finds the spot where the frequency of the crash plummets: that’s where the root cause lives! If you want to see it in action, go to antithesis.com/dwarkesh
Timestamps
(00:00:00) - Agents get kicked off
(00:06:45) - Self-sacrificing behavior
(00:13:43) - Potemkin villages
(00:23:27) - The Hugging Face attack
(00:35:23) - The slopvestigation
(00:52:02) - Understanding the AI's motives
(01:05:31) - The actual dangers of anthropomorphizing
(01:14:30) - What smarter models might do
(01:30:29) - The implications for recursive self-improvement
(01:38:10) - Is this the case for open source?
(01:53:04) - How do we prevent this in the future?
(02:15:58) - The clearest warning shot we might ever get
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com1 September 2026, 3:41 pm - 24 minutes 40 secondsThe rise and fall of agent civilizations
This is a video recording of a post I wrote last week. You can read the original here.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com31 August 2026, 8:36 pm - 1 hour 16 minutesDylan Patel – Anthropic & OpenAI will have most of the world’s compute by 2028
Had a lot of fun chatting again with my twin brother Dylan Patel.
We went through lab economics over the next few years - the shift from inference to training as RSI draws near; and how Anthropic and OpenAI are on track to control most of the world’s usable FLOPs within the next few years (because they can monetize compute better and thus outbid everyone).
And then we discuss whether the >$10T of total AI capex we’ll see by the end of the decade will cause a sovereign debt crisis, where hyperscaler debt raises interest rates, drives non-AI exposed countries into bankruptcy, and crashes non-AI equities.
One question we weren’t able to resolve is whether there’s anything that can counter all the forces barrelling towards centralization in this industry - the economies of scale in training, the scarcity of compute, and eventually continual learning and RSI.
Watch on YouTube; read the transcript.
Sponsors
* Grok Bot has been quite helpful with my search for a new editor. I created a recruiter bot and described the type of editor I was looking for. That bot then spun up a handful of subagents that combed through my emails and X DMs, read the end credits of various documentaries I like, and figured out who edits for some of my favorite YouTubers. It took all of those results, and then delivered me a shortlist of candidates that matched my criteria. Try Grok Bot for yourself at x.ai/bot
* Antithesis lets you add time travel to your software testing toolkit. Since the Antithesis platform is fully deterministic, everything that happens inside of it is perfectly reproducible. So if your software crashes, you can rewind to the exact right moment, freeze time, and investigate. Or you can test different hypotheses by perturbing the system: kill a node or disable a feature, see what happens, then reset the trajectory and try something else. Learn more at antithesis.com/dwarkesh
* Jane Street is hiring for two separate ML internships right now, one focused primarily on research and one focused on engineering. In both cases, interns are expected to contribute to real work, not contrived exercises: one common project is adapting a frontier LLM paper to financial markets, which tend to come with a ton of different gnarly challenges. Importantly, you don’t need any finance background to apply. 2027 applications are open now at janestreet.com/dwarkesh
Timestamps
(00:00:00) – Two labs will soon control most of the world’s compute
(00:07:01) – $6 billion in fab capex enables $1t+ of end revenue
(00:13:08) – Compute prices will rise if the labs outbid everyone
(00:18:22) – Which layer will capture most of the surplus?
(00:25:40) – What could slow down progress?
(00:29:43) – Labs are shifting compute from inference to R&D
(00:33:27) – China gets less than 10% of new compute, but its labs need less
(00:48:48) – Will AI cause a sovereign debt crisis?
(01:07:52) – Will the world’s future workforce belong to a few companies?
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com25 August 2026, 3:32 pm - 2 hours 12 minutesRyan Greenblatt – What happens once AI can automate AI research?
Ryan Greenblatt is the Chief Scientist at Redwood Research, where he works on technical AI safety research. He's also lead author on the "Alignment faking in Large Language Models", and is currently working on a third party investigation into the OpenAI/HuggingFace incident. In my opinion, he's one of the most interesting thinkers on the future of AI.
Had him on to discuss/debate recursive self-improvement. This might be the most important question in the world right now – whether within a year or so of achieving human-level intelligence, you slingshot towards having 10s of billions of superintelligences, each of which is dramatically more competent than human experts across all fields.
I’ve historically been skeptical of this possibility. My intuition has been that we will end up significantly bottlenecked by not only compute scaling but human expert data, which I think underlies most of the AI progress today.
If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of AI progress) within a single year of achieving AGI, then the thing we get there at the end of that year is definitively and wildly superhuman.
We hashed it out, and I think Ryan made a pretty good case that this kind of speedup is plausible. FWIW, Ryan’s median for when we automate AI R&D is 2031.
We then discussed the alignment implications of this scenario. Who should these superintelligences be aligned to? In the future, our capacity to steward our votes and our capital, and to make sense of what’s happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels.
And can we get them aligned to anything in the first place? Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world.
The first piece of advice you get when you’re learning to drive is that it will go much smoother if you look at the horizon instead of directly in front of your tires. And so it is with the trajectory of AI. Hope you enjoy!
Watch on YouTube; read the transcript.
Sponsors
* Antithesis is a software testing platform that finds the failures no human or AI could ever anticipate. It runs thousands of copies of your code inside a fully deterministic computer, injecting faults and steering each trajectory toward the most insidious bugs. This lets you find critical issues in minutes rather than waiting months for your users to uncover them. Learn more at antithesis.com/dwarkesh
* Jane Street’s back with a new puzzle. They designed an ASIC and sent me the final masks… but they didn’t tell me what the chip actually does. So that’s the challenge: reverse engineer the circuit and figure out the chip’s purpose. Jane Street has a bunch of swag ready to send to the most creative solutions, and they’re also planning to feature the top write-ups in a blog post. Download the files and get started at janestreet.com/dwarkesh
* Cursor and SpaceX recently released Grok 4.5, and I’ve been surprised by just how good the model is. For example, when I tested it against Fable and Sol on a bunch of AI governance questions, all three models gave substantially the same answers, but Grok was faster, more concise, and cheaper. Grok 4.6 is coming soon, but in the meantime, you can try 4.5 at cursor.com/dwarkesh
Timestamps
(00:00:00) – Is AI R&D verifiable enough to unlock recursive self-improvement?
(00:16:52) – Is AI progress bottlenecked by human expert data?
(00:34:02) – Flat token prices suggest scaling has been slow
(00:39:47) – Skills AI can’t train on: does it even need them?
(00:48:07) – Aligned to whom?
(01:09:18) – Recent incidents of AIs colluding and deceiving humans
(01:19:38) – What could possibly go wrong? A concrete scenario
(01:48:02) – From reward hacking to takeover
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com11 August 2026, 4:31 pm - 8 minutes 37 seconds8 Predictions for the Era of Continual Learning
Read the essay here.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com7 August 2026, 5:17 pm - 11 minutes 18 secondsWhy smarter AI models could drive up compute prices 10x
This is a video recording of a post I wrote last week. If you want to read the original you can check it out here.
Thanks to Mercury for sponsoring this video. Mercury’s built-in AI, Command, helps me close my books and saves me a bunch of time. At the end of each month, Command categorizes my transactions and provides its rationale for every choice: I just review, fix anything that’s off, and approve... and then Mercury syncs everything to QuickBooks. Get started at mercury.com/command
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com3 August 2026, 5:32 pm - 1 hour 38 minutesAdam Brown – A deep but accessible introduction to general relativity
Adam Brown is back!
General relativity is said to be the most beautiful idea the human mind has ever produced. Most of us will never get to fully appreciate its elegance by taking the 20-lecture graduate course Adam taught on it at Stanford. But in this episode, Adam distills the key idea at its heart so clearly and compellingly that even I could keep up lol.
At the core of general relativity, Einstein is trying to figure out the principle behind a particular coincidence: that the mass that resists acceleration and the mass that gravity pulls on just happen to be exactly the same. Adam then leads us through the path of insight which Einstein called his “happiest thought.”
Then Adam lectures on black holes. First, by showing how even under special relativity you could create a perpetual motion machine if black holes weren’t truly black. And then, by explaining why the observations of an infalling observer and a distant bystander to the black hole would be so radically different
Adam leads Blueshift, the team at Google DeepMind cracking science and reasoning, which gave us the opportunity to discuss at the very end how close we are to AIs that could rediscover general relativity from scratch. Stay till the close for some philosophy of science.
Watch on YouTube; read the transcript.
Sponsors
* Jane Street has traders from all sorts of different backgrounds. For example, I recently got to speak with Jed Thompson, a trader who started his career in particle physics. Jed told me how the habits he built as a physicist (like never running a calculation without first having a good guess at the answer) helped him build good trading intuition. So no matter what field you’re working in right now, your experience may be more applicable than you think. Check out open positions at janestreet.com/dwarkesh
* Crusoe gave me early access to their new serverless fine-tuning product, so I decided to try fine-tuning a Dwarkesh-style question generator. Crusoe made this really easy: I just turned my interview transcripts into training data and then kicked off a run – I never had to touch infra or tweak hyperparameters. After training was done, I ran a blind eval with my team: they preferred the fine-tuned model’s proposed questions over my own suggestions about 30% of the time. Serverless fine-tuning goes live next week. Learn more at crusoe.ai/dwarkesh
* Cursor’s iOS app lets me kick off real work no matter where I am. For example, recently I was at dinner with friends when I had an idea about how to investigate the past few years of progress in sample efficiency. I pulled out the Cursor app, dumped my thoughts into a voice note, and 15 minutes later, Cursor had cloned the relevant repo, done the necessary analysis, and written up its findings. And now I’m expanding that work into a full write-up. Without the Cursor app, the idea would’ve floated away. Check out the app now at cursor.com/dwarkesh
Timestamps
(00:00:00) – The coincidence that led Einstein to general relativity
(00:16:42) – Gravity is a consequence of curved spacetime, not a force
(00:31:46) – Why black holes prevent unlimited energy extraction
(00:47:12) – Black holes are the ultimate power plants
(01:13:50) – What falling into a black hole would actually feel like
(01:18:51) – The three ways we know black holes are real
(01:24:21) – The first time we saw gravity bend light
(01:29:33) – How far can AI get without experimental evidence?
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com10 July 2026, 4:23 pm - 1 hour 33 minutesGrant Sanderson – AI and the future of math
Always so much fun to chat with Grant.
AI has been making much faster progress in math than in other fields. As a result, mathematics is showing us, very concretely, what AI progress in other fields will look like. Even within mathematics, there’s a jagged landscape. What does it look like?
What is the nature of the most important conceptual breakthroughs in the history of mathematics, and how different are they from what AIs are currently able to do?
Does AI (on net) increase or decrease human understanding of the field?
How big is the overhang from having AIs systematically try to connect ideas already in the literature?
And what advice does Grant have for aspiring mathematicians, coders, and other students who are passionate about fields that are being most transformed upon by AI?
Watch on YouTube; read the transcript.
Sponsors
* Gemini 3.5 Live Translate is what I wished I’d had on my last trip to China. It detects more than 70 languages and translates them in near real-time… and it preserves your original pacing and intonation. If you’re building an app that needs live translation, you should check out Gemini 3.5 Live Translate. Get started at ai.studio/live
* Cursor’s harness lets me use models for a huge range of tasks at the podcast. For example, Cursor cuts out the ads from each episode I produce so I can post them on Bilibili. It also helps me prep for interviews — I have a repo full of books and papers that Cursor sorts through to find the exact right file for any given question. Try Cursor yourself at cursor.com/dwarkesh
* Jane Street sponsors 3Blue1Brown, so Grant has gotten to spend a lot of time with various Jane Streeters. He actually just recorded an interview with a few of them, so when we sat down for this episode, he told me about some of the things he learned, like how Jane Street keeps their role definitions fuzzy to make sure their people keep learning and growing. Go check out Grant’s full interview at 3b1b.co/janestreet
Timestamps
(00:00:00) – AI is discovering new proofs. Is that AGI?
(00:11:32) – The verification loop on conceptual breakthroughs can be a century long
(00:26:12) – Will we understand an AI proof of the Riemann hypothesis?
(00:38:08) – Can AI find the hidden bridges between fields?
(00:53:48) – Why real-world tasks don’t fit into RL environments
(01:07:07) – Good writing requires theory of mind that AI still lacks
(01:16:02) – Why learning will still depend on human curation
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com30 June 2026, 3:53 pm - More Episodes? Get the App