- 10 minutes 15 secondsPoolside’s model factory is shipping: Laguna S 2.1 and the Desktop Assistant
Poolside trained Laguna S 2.1 on 4,096 NVIDIA H200s and published the open weights sixty days later. The 118-billion-parameter mixture-of-experts model runs multi-hour autonomous coding sessions, handles up to a million tokens of context, and fits quantized on a single desktop-class machine.
Nathan Benaich of Air Street Capital walks through the Model Factory that made that schedule possible, the system co-founders Jason Warner, GitHub’s former CTO, and Eiso Kant built to take a research idea to a production run in under a week. The episode covers the 409,000 training environments behind the model, the 50-minute session in which it built a browser rendering engine with no human intervention and no ability to see its own output, why the licenses attached to open weights differ so sharply between Poolside, Tencent and Moonshot, and what it means for a company to run a capable coding model inside its own security boundary rather than renting one through an API.
Read the full piece: press.airstreet.com Poolside’s launch post: poolside.ai/blog/introducing-laguna-s-2-1 Evaluation trajectories: trajectories.poolside.ai RAAIS 2025 - Eiso Kant, Inside poolside’s path to AGI: press.airstreet.com/p/eiso-kant-poolside-ai-raais-2025
From Air Street Press. Subscribe at press.airstreet.com, read the State of AI Report at stateof.ai, and leave a rating - it helps the show.
29 July 2026, 2:07 pm - 7 minutes 18 secondsBlack Forest Labs FLUX 3: from video generation to robot control
Description:
Black Forest Labs has launched FLUX 3, a multimodal foundation model that learns jointly from images, video and audio within a single architecture, and mimic robotics has released FLUX-mimic, a video-action model built on that backbone and being tested on real assembly work in Audi's Production Lab. Nathan Benaich of Air Street Capital, an investor in Black Forest Labs, reads his Air Street Press essay on why a model trained to predict how scenes evolve turns out to be a usable robot controller. Covers the FLUX 3 preference results against Runway Gen-4.5, Grok Imagine Video, Kling v3 Pro, Seedance 2.0 and Gemini Omni Flash; how a lightweight action decoder reads the video prediction path without ever generating video; the frozen-backbone ablation against π0.5; the 101-millisecond system reaction time on a single RTX 5090; and what Air Street's own robotics deal flow says about where the constraint really sits.Chapters (estimated at ~150 wpm, slide proportionally against final audio):
- 0:00 A robot arm in Audi's Production Lab
- 0:40 What Black Forest Labs and mimic released
- 1:20 Early access and open weights
- 1:50 The preference-test results
- 2:40 What a model must represent to predict video
- 3:30 Reading actions off the video path
- 4:20 The frozen-backbone ablation
- 5:00 101 milliseconds, and the Audi tasks
- 5:50 What we see in robotics deal flow
- 6:30 The road to physical intelligence
Links: FLUX 3 · FLUX-mimic (BFL) · FLUX-mimic (mimic) · Odyssey Series B · BFL Series B ·
27 July 2026, 1:19 pm - 6 minutes 35 secondsThe case for intelligence beyond language
Raia Hadsell, VP of Research at Google DeepMind, makes the case that intelligence is more than language: the same recipe that learns the patterns of text can learn the patterns of any complex system. She walks through DiffusionGemma and text diffusion, the Genie world models, and DeepMind's robotics stack, where world models now generate training data you cannot tell from the real thing. Recorded at RAAIS 2026.
Timestamps
0:00 Introduction (Nathan Benaich)
0:35 From philosophy to DeepMind: the frontiers of intelligence
2:37 The twenty-year lesson: one recipe for complex systems
4:34 DiffusionGemma and the Gemma 4 open models
5:32 How text diffusion works
7:34 Speed, self-correction, and the sudoku test
10:51 World models: better agents need better worlds
12:56 Genie 1 to Genie 3
15:01 Genie 3 demos: typing a world into being
18:20 World models for education
19:38 Grounding Genie in Street View
20:41 Robotics: a brain and a spine
23:14 Gemini Robotics-ER 1.6 and Boston Dynamics' Spot
24:42 The vision-language-action model
25:45 The data bottleneck and closing the loop
27:24 Beyond language: the domains still to crack
21 July 2026, 1:20 pm - 6 minutes 52 secondsEurope’s air defense gap
At RAAIS 2026, Nathan Benaich sits down with Hadrien Canter, co-founder and CEO of Alta Ares, the air-defense company building AI-guided interceptors to shoot down cheap drones and cruise missiles. They get into why Europe has lost air superiority for the first time in modern history, what the data loop looks like when you run it on a freezing front line instead of a laptop, why "quantity is the quality" in the industrialization race, and how the talent pool is shifting toward European defense. Recorded live at RAAIS 2026 in London.
Chapters
00:00 — Introducing Hadrien Canter and Alta Ares
00:47 — 2022 in Ukraine, and how Europe lost air superiority
03:23 — The Series A and the Airbus partnership
05:13 — The data loop, edge AI, and the three phases of a mission
08:22 — What no simulation can reproduce
10:09 — Hiring for the mission; defense as the precondition for peace
12:36 — Open research questions and the human in the loop
14:22 — How the adversary uses AI: evasive Shaheds, drone mesh, China
17:27 — Two interceptors, and why quantity is the quality
20:50 — Iron Dome math, budgets, and peace-time vs war-time
23:41 — Q&A: keeping pace with a fast-changing front
26:31 — Q&A: talent and the shift toward European defense
29:30 — Q&A: re-arming without permanent war
32:56 — Freedom doesn't come for free
16 July 2026, 1:20 pm - 8 minutes 19 secondsFrom driving the world to dreaming it
World models are the bet that AI should learn the world by watching it and acting in it, not just by reading about it. At RAAIS 2026, Odyssey co-founder and CTO Jeff Hawke makes the case: a world model is a neural simulator - an interactive stream of pixels that runs in real time, models physics, and answers back.
He walks through Odyssey's four research fronts - streaming interactive pixels (Odyssey-2), joint audio and video (Starchild-1), shared multiplayer worlds (Agora-1, demoed live as a fully generated game of GoldenEye), and PROWL, which sends a reinforcement-learning agent to find and fix a world model's own failures - and argues the field is at its GPT-2 moment: promising, but pre-ChatGPT, with the GPT-3-style commercial unlock still ahead.
Recorded at the 10th Research and Applied AI Summit (RAAIS), London, June 2026.
Timestamps
00:00 Intro: Nathan on Odyssey and world models
01:05 Jeff Hawke: from self-driving to world models
01:40 The bet — a missing form of intelligence
02:40 Why world models suddenly matter (the late-2025 flip)
03:16 What a world model actually is (and isn't)
04:45 The neural simulator
06:34 Two principles: end-to-end learning and generality
07:19 The "GPT-3 of world models" and four research themes
08:46 Odyssey-2: streaming, interactive pixels
10:33 Starchild-1: generating audio and video together
13:03 Agora-1: multiplayer world models
13:57 Live demo: the room plays GoldenEye
16:20 PROWL: improving the model by breaking it
18:39 Where Odyssey goes next
19:55 Still the GPT-2 era
21:30 Q&A: physics limits, safety, compute cost, merging with LLMs
14 July 2026, 1:20 pm - 7 minutes 15 secondsTurning compute into intelligence
Ted Moskovitz leads the Science of Scaling team at Anthropic, the group that works out how to turn compute into smarter models. In this RAAIS 2026 fireside with Air Street Capital's Nathan Benaich, he argues that frontier scaling has become an empirical science - a discipline for cutting uncertainty before spending the compute, not just buying more of it.
They get into the honest measure of AI acceleration (it's the counterfactual, not the benchmark), why a bigger model can be cheaper than splitting a task across small ones, whether a model can have research taste, and why safety and capability turn out to be the same axis. Plus the highest-leverage AI work to do in 2026, and why Anthropic's London office no longer feels like a satellite.
Recorded live at RAAIS 2026 in London.
Timestamp:
00:00 - Meet Ted Moskovitz and the Science of Scaling team
00:45 - What "the science of scaling" actually means
01:18 - Why scaling is a science, not an art
02:55 - Big labs vs the new "neo labs"
04:47 - How a research finding reaches the product
06:44 - What neuroscience carries over to AI (and what doesn't)
09:12 - "When AI builds itself" and the real measure of acceleration
10:33 - Trust, bypass mode, and the latest model jumps
11:42 - One big model vs many small ones
13:13 - Can a model have research taste?
15:39 - How safety research makes products better
17:36 - Emergent misalignment and the alignment race
19:14 - The highest-leverage AI work in 2026
20:21 - Inside Anthropic's London office
21:34 - Audience Q&A
9 July 2026, 1:19 pm - 7 minutes 54 secondsAccelerating science and medicine with collaborative agents
Vivek Natarajan, Research Lead for AI, science and medicine at Google DeepMind, on porting the self-play and search recipe behind AlphaGo into scientific and clinical reasoning. He walks through the AI co-scientist, which generates and debates hypotheses (one matched a decade of lab work in two days), and AMIE, a diagnostic dialogue system trained in simulation. Recorded at RAAIS 2026.
Chapters:
0:00 Welcome and introducing the AI co-scientist
1:41 Origins: Med-PaLM and the leap to hypothesis generation
5:10 System 1 versus System 2 thinking
6:34 Borrowing from AlphaGo: self-play and search
8:02 Generate, debate, evolve, and tournaments
11:47 Testing in real labs: Imperial College and antimicrobial resistance
13:29 Ten years in two days: Penadés reacts
15:44 More breakthroughs: leukemia, liver fibrosis and vorinostat
18:44 Plant immunity and protein design
20:09 Democratizing medicine: from Med-PaLM benchmarks
21:28 AMIE and the value of experience
23:12 Diagnosis, empathy and augmenting doctors
25:19 Real patients: the Beth Israel feasibility study
27:21 The co-clinician and the new triad of care
28:31 Audience Q&A7 July 2026, 1:19 pm - 7 minutes 20 secondsBeyond hill climbing: the path to superhuman scientific discovery
At RAAIS 2026, Google DeepMind's Roberta Raileanu lays out a recipe for superhuman scientific discovery: AI systems that make groundbreaking discoveries across domains faster than people can. She walks through three ingredients - reinforcement learning to discover solutions where progress can be measured, open-ended divergent search to find new problems rather than climb known ones, and meta-learning to speed up discovery on problems no one has posed yet. The through-line: we can search for anything we can measure, but we still cannot measure what makes a discovery good. The bottleneck isn't the search. It's the signal.
Chapters:
00:00 - Introduction
00:51 - Defining superhuman scientific discovery
01:40 - The state of play: real progress, real plateau
06:58 - Ingredient one: discovery as reinforcement learning (Move 37, MLGym)
12:51 - Ingredient two: open-ended search and why greatness cannot be planned
18:30 - Rainbow Teaming: quality-diversity in practice
21:07 - Ingredient three: meta-learning the process of discovery (DiscoBench)
25:06 - The recipe, and the missing signal
2 July 2026, 1:18 pm - 7 minutes 44 secondsCompute scarcity is an engineering problem
Angelos Perivolaropoulos, a research engineer at ElevenLabs, on turning GPU scarcity into an inference-engineering problem: how to serve far more users on the same hardware, from batching to frontier architecture changes. Recorded at RAAIS 2026.
00:00 Introduction: ElevenLabs and the GPU squeeze
00:38 The question: how to scale when you can't add capacity
01:11 About Angelos: Scribe, speech-to-text and text-to-speech
01:56 GPU scarcity meets exponential demand
02:44 What a token actually costs: compute vs memory bandwidth
03:38 Prefill, decode and the KV cache
05:53 Batching and continuous batching (1 → 15 users/GPU)
08:37 FP8 quantization and quantize-aware training (→ 20)
11:29 Speculative decoding and multi-token prediction (→ 28)
15:13 Compressing the KV cache: TurboQuant and distillation (→ 70)
17:27 Frontier architectures: MLA, linear attention, state-space (→ 140)
20:39 Trade-offs: nothing is free
22:03 Q&A: papers vs production, token subsidies, TTS evals
30 June 2026, 1:18 pm - 14 minutes 35 secondsState of AI: Compute Index 2026
The fifth State of AI Compute Index, in collaboration with Zeta Alpha. After a soft 2025, open AI research citations rebounded in 2026 - and NVIDIA still appears in ~91% of them. But the bigger story has moved off the page: Hopper is now the live installed base, Blackwell is mostly still pipeline, and frontier labs have started buying compute by the gigawatt. Nathan walks through what changed, what didn't, and why "GPU count" is becoming the wrong question.
Read the full piece and explore the live charts: https://www.stateof.ai/compute
Chapters
- (00:00) What's new in v5 - citations, infrastructure, and gigawatts
- (01:25) The breather was short: 2025 was a pause, not a rollover
- (03:15) NVIDIA at ~91%, and the challengers - AMD, Huawei, Apple, TPU
- (05:05) Inside NVIDIA: the handover from A100 to Hopper to Blackwell
- (06:55) Startup silicon fragments - Groq, Cerebras, and the NVIDIA deal
- (08:20) Hopper is the installed base: 460k deployed GPUs
- (09:50) Blackwell is mostly pipeline: 80% still announced
- (11:00) The demand side, measured in gigawatts
- (12:15) Looking ahead, and why a GPU order isn't a cluster
Links:
- Full index and charts: https://www.stateof.ai/compute
- State of AI Report: https://www.stateof.ai
- Air Street Press: https://press.airstreet.com
If you found this useful, rate State of AI with Nathan Benaich five stars and share it with someone building in AI infrastructure - it genuinely helps.
29 June 2026, 8:44 pm - 4 minutes 48 secondsSTARK: Europe's next defense prime
Air Street Capital backs Stark, the German multi-domain defense company, in its €500M led by Founders Fund and Sequoia.
In this episode, we discuss why cheap, software-defined unmanned systems in the air and at sea are the decisive lesson of Ukraine, and why we think co-founder and CEO Uwe Horstmann - a Project A GP and Bundeswehr reservist - is building the German neoprime Europe needs. Round led by Sequoia and Founders Fund, with the NATO Innovation Fund, Project A, and Air Street.
Links: stark-defence.com · full post at press.airstreet.com · YouTube version
26 June 2026, 1:32 pm - More Episodes? Get the App