• 10 minutes 15 seconds
    Poolside’s model factory is shipping: Laguna S 2.1 and the Desktop Assistant

    Poolside trained Laguna S 2.1 on 4,096 NVIDIA H200s and published the open weights sixty days later. The 118-billion-parameter mixture-of-experts model runs multi-hour autonomous coding sessions, handles up to a million tokens of context, and fits quantized on a single desktop-class machine.

    Nathan Benaich of Air Street Capital walks through the Model Factory that made that schedule possible, the system co-founders Jason Warner, GitHub’s former CTO, and Eiso Kant built to take a research idea to a production run in under a week. The episode covers the 409,000 training environments behind the model, the 50-minute session in which it built a browser rendering engine with no human intervention and no ability to see its own output, why the licenses attached to open weights differ so sharply between Poolside, Tencent and Moonshot, and what it means for a company to run a capable coding model inside its own security boundary rather than renting one through an API.

    Read the full piece: press.airstreet.com Poolside’s launch post: poolside.ai/blog/introducing-laguna-s-2-1 Evaluation trajectories: trajectories.poolside.ai RAAIS 2025 - Eiso Kant, Inside poolside’s path to AGI: press.airstreet.com/p/eiso-kant-poolside-ai-raais-2025

    From Air Street Press. Subscribe at press.airstreet.com, read the State of AI Report at stateof.ai, and leave a rating - it helps the show.

    29 July 2026, 2:07 pm
  • 7 minutes 18 seconds
    Black Forest Labs FLUX 3: from video generation to robot control

    Description:
    Black Forest Labs has launched FLUX 3, a multimodal foundation model that learns jointly from images, video and audio within a single architecture, and mimic robotics has released FLUX-mimic, a video-action model built on that backbone and being tested on real assembly work in Audi's Production Lab. Nathan Benaich of Air Street Capital, an investor in Black Forest Labs, reads his Air Street Press essay on why a model trained to predict how scenes evolve turns out to be a usable robot controller. Covers the FLUX 3 preference results against Runway Gen-4.5, Grok Imagine Video, Kling v3 Pro, Seedance 2.0 and Gemini Omni Flash; how a lightweight action decoder reads the video prediction path without ever generating video; the frozen-backbone ablation against π0.5; the 101-millisecond system reaction time on a single RTX 5090; and what Air Street's own robotics deal flow says about where the constraint really sits.

    Chapters (estimated at ~150 wpm, slide proportionally against final audio):

    • 0:00 A robot arm in Audi's Production Lab
    • 0:40 What Black Forest Labs and mimic released
    • 1:20 Early access and open weights
    • 1:50 The preference-test results
    • 2:40 What a model must represent to predict video
    • 3:30 Reading actions off the video path
    • 4:20 The frozen-backbone ablation
    • 5:00 101 milliseconds, and the Audi tasks
    • 5:50 What we see in robotics deal flow
    • 6:30 The road to physical intelligence

    Links: FLUX 3 · FLUX-mimic (BFL) · FLUX-mimic (mimic) · Odyssey Series B · BFL Series B ·

    27 July 2026, 1:19 pm
  • 6 minutes 35 seconds
    The case for intelligence beyond language

    Raia Hadsell, VP of Research at Google DeepMind, makes the case that intelligence is more than language: the same recipe that learns the patterns of text can learn the patterns of any complex system. She walks through DiffusionGemma and text diffusion, the Genie world models, and DeepMind's robotics stack, where world models now generate training data you cannot tell from the real thing. Recorded at RAAIS 2026.

    Timestamps

    0:00 Introduction (Nathan Benaich)

    0:35 From philosophy to DeepMind: the frontiers of intelligence

    2:37 The twenty-year lesson: one recipe for complex systems

    4:34 DiffusionGemma and the Gemma 4 open models

    5:32 How text diffusion works

    7:34 Speed, self-correction, and the sudoku test

    10:51 World models: better agents need better worlds

    12:56 Genie 1 to Genie 3

    15:01 Genie 3 demos: typing a world into being

    18:20 World models for education

    19:38 Grounding Genie in Street View

    20:41 Robotics: a brain and a spine

    23:14 Gemini Robotics-ER 1.6 and Boston Dynamics' Spot

    24:42 The vision-language-action model

    25:45 The data bottleneck and closing the loop

    27:24 Beyond language: the domains still to crack


    21 July 2026, 1:20 pm
  • 6 minutes 52 seconds
    Europe’s air defense gap

    At RAAIS 2026, Nathan Benaich sits down with Hadrien Canter, co-founder and CEO of Alta Ares, the air-defense company building AI-guided interceptors to shoot down cheap drones and cruise missiles. They get into why Europe has lost air superiority for the first time in modern history, what the data loop looks like when you run it on a freezing front line instead of a laptop, why "quantity is the quality" in the industrialization race, and how the talent pool is shifting toward European defense. Recorded live at RAAIS 2026 in London.

    Chapters

    00:00 — Introducing Hadrien Canter and Alta Ares

    00:47 — 2022 in Ukraine, and how Europe lost air superiority

    03:23 — The Series A and the Airbus partnership

    05:13 — The data loop, edge AI, and the three phases of a mission

    08:22 — What no simulation can reproduce

    10:09 — Hiring for the mission; defense as the precondition for peace

    12:36 — Open research questions and the human in the loop

    14:22 — How the adversary uses AI: evasive Shaheds, drone mesh, China

    17:27 — Two interceptors, and why quantity is the quality

    20:50 — Iron Dome math, budgets, and peace-time vs war-time

    23:41 — Q&A: keeping pace with a fast-changing front

    26:31 — Q&A: talent and the shift toward European defense

    29:30 — Q&A: re-arming without permanent war

    32:56 — Freedom doesn't come for free


    16 July 2026, 1:20 pm
  • 8 minutes 19 seconds
    From driving the world to dreaming it

    World models are the bet that AI should learn the world by watching it and acting in it, not just by reading about it. At RAAIS 2026, Odyssey co-founder and CTO Jeff Hawke makes the case: a world model is a neural simulator - an interactive stream of pixels that runs in real time, models physics, and answers back.

    He walks through Odyssey's four research fronts - streaming interactive pixels (Odyssey-2), joint audio and video (Starchild-1), shared multiplayer worlds (Agora-1, demoed live as a fully generated game of GoldenEye), and PROWL, which sends a reinforcement-learning agent to find and fix a world model's own failures - and argues the field is at its GPT-2 moment: promising, but pre-ChatGPT, with the GPT-3-style commercial unlock still ahead.

    Recorded at the 10th Research and Applied AI Summit (RAAIS), London, June 2026.

    Timestamps

    00:00 Intro: Nathan on Odyssey and world models

    01:05 Jeff Hawke: from self-driving to world models

    01:40 The bet — a missing form of intelligence

    02:40 Why world models suddenly matter (the late-2025 flip)

    03:16 What a world model actually is (and isn't)

    04:45 The neural simulator

    06:34 Two principles: end-to-end learning and generality

    07:19 The "GPT-3 of world models" and four research themes

    08:46 Odyssey-2: streaming, interactive pixels

    10:33 Starchild-1: generating audio and video together

    13:03 Agora-1: multiplayer world models

    13:57 Live demo: the room plays GoldenEye

    16:20 PROWL: improving the model by breaking it

    18:39 Where Odyssey goes next

    19:55 Still the GPT-2 era

    21:30 Q&A: physics limits, safety, compute cost, merging with LLMs

    14 July 2026, 1:20 pm
  • 7 minutes 15 seconds
    Turning compute into intelligence

    Ted Moskovitz leads the Science of Scaling team at Anthropic, the group that works out how to turn compute into smarter models. In this RAAIS 2026 fireside with Air Street Capital's Nathan Benaich, he argues that frontier scaling has become an empirical science - a discipline for cutting uncertainty before spending the compute, not just buying more of it.

    They get into the honest measure of AI acceleration (it's the counterfactual, not the benchmark), why a bigger model can be cheaper than splitting a task across small ones, whether a model can have research taste, and why safety and capability turn out to be the same axis. Plus the highest-leverage AI work to do in 2026, and why Anthropic's London office no longer feels like a satellite.

    Recorded live at RAAIS 2026 in London.

    Timestamp:

    00:00 - Meet Ted Moskovitz and the Science of Scaling team

    00:45 - What "the science of scaling" actually means

    01:18 - Why scaling is a science, not an art

    02:55 - Big labs vs the new "neo labs"

    04:47 - How a research finding reaches the product

    06:44 - What neuroscience carries over to AI (and what doesn't)

    09:12 - "When AI builds itself" and the real measure of acceleration

    10:33 - Trust, bypass mode, and the latest model jumps

    11:42 - One big model vs many small ones

    13:13 - Can a model have research taste?

    15:39 - How safety research makes products better

    17:36 - Emergent misalignment and the alignment race

    19:14 - The highest-leverage AI work in 2026

    20:21 - Inside Anthropic's London office

    21:34 - Audience Q&A

    9 July 2026, 1:19 pm
  • 7 minutes 54 seconds
    Accelerating science and medicine with collaborative agents

    Vivek Natarajan, Research Lead for AI, science and medicine at Google DeepMind, on porting the self-play and search recipe behind AlphaGo into scientific and clinical reasoning. He walks through the AI co-scientist, which generates and debates hypotheses (one matched a decade of lab work in two days), and AMIE, a diagnostic dialogue system trained in simulation. Recorded at RAAIS 2026.

    Chapters:
    0:00 Welcome and introducing the AI co-scientist
    1:41 Origins: Med-PaLM and the leap to hypothesis generation
    5:10 System 1 versus System 2 thinking
    6:34 Borrowing from AlphaGo: self-play and search
    8:02 Generate, debate, evolve, and tournaments
    11:47 Testing in real labs: Imperial College and antimicrobial resistance
    13:29 Ten years in two days: Penadés reacts
    15:44 More breakthroughs: leukemia, liver fibrosis and vorinostat
    18:44 Plant immunity and protein design
    20:09 Democratizing medicine: from Med-PaLM benchmarks
    21:28 AMIE and the value of experience
    23:12 Diagnosis, empathy and augmenting doctors
    25:19 Real patients: the Beth Israel feasibility study
    27:21 The co-clinician and the new triad of care
    28:31 Audience Q&A

    7 July 2026, 1:19 pm
  • 7 minutes 20 seconds
    Beyond hill climbing: the path to superhuman scientific discovery

    At RAAIS 2026, Google DeepMind's Roberta Raileanu lays out a recipe for superhuman scientific discovery: AI systems that make groundbreaking discoveries across domains faster than people can. She walks through three ingredients - reinforcement learning to discover solutions where progress can be measured, open-ended divergent search to find new problems rather than climb known ones, and meta-learning to speed up discovery on problems no one has posed yet. The through-line: we can search for anything we can measure, but we still cannot measure what makes a discovery good. The bottleneck isn't the search. It's the signal.

    Chapters:

    00:00 - Introduction

    00:51 - Defining superhuman scientific discovery

    01:40 - The state of play: real progress, real plateau

    06:58 - Ingredient one: discovery as reinforcement learning (Move 37, MLGym)

    12:51 - Ingredient two: open-ended search and why greatness cannot be planned

    18:30 - Rainbow Teaming: quality-diversity in practice

    21:07 - Ingredient three: meta-learning the process of discovery (DiscoBench)

    25:06 - The recipe, and the missing signal


    2 July 2026, 1:18 pm
  • 7 minutes 44 seconds
    Compute scarcity is an engineering problem

    Angelos Perivolaropoulos, a research engineer at ElevenLabs, on turning GPU scarcity into an inference-engineering problem: how to serve far more users on the same hardware, from batching to frontier architecture changes. Recorded at RAAIS 2026.

    00:00 Introduction: ElevenLabs and the GPU squeeze

    00:38 The question: how to scale when you can't add capacity

    01:11 About Angelos: Scribe, speech-to-text and text-to-speech

    01:56 GPU scarcity meets exponential demand

    02:44 What a token actually costs: compute vs memory bandwidth

    03:38 Prefill, decode and the KV cache

    05:53 Batching and continuous batching (1 → 15 users/GPU)

    08:37 FP8 quantization and quantize-aware training (→ 20)

    11:29 Speculative decoding and multi-token prediction (→ 28)

    15:13 Compressing the KV cache: TurboQuant and distillation (→ 70)

    17:27 Frontier architectures: MLA, linear attention, state-space (→ 140)

    20:39 Trade-offs: nothing is free

    22:03 Q&A: papers vs production, token subsidies, TTS evals

    30 June 2026, 1:18 pm
  • 14 minutes 35 seconds
    State of AI: Compute Index 2026

    The fifth State of AI Compute Index, in collaboration with Zeta Alpha. After a soft 2025, open AI research citations rebounded in 2026 - and NVIDIA still appears in ~91% of them. But the bigger story has moved off the page: Hopper is now the live installed base, Blackwell is mostly still pipeline, and frontier labs have started buying compute by the gigawatt. Nathan walks through what changed, what didn't, and why "GPU count" is becoming the wrong question.

    Read the full piece and explore the live charts: https://www.stateof.ai/compute

    Chapters

    • (00:00) What's new in v5 - citations, infrastructure, and gigawatts
    • (01:25) The breather was short: 2025 was a pause, not a rollover
    • (03:15) NVIDIA at ~91%, and the challengers - AMD, Huawei, Apple, TPU
    • (05:05) Inside NVIDIA: the handover from A100 to Hopper to Blackwell
    • (06:55) Startup silicon fragments - Groq, Cerebras, and the NVIDIA deal
    • (08:20) Hopper is the installed base: 460k deployed GPUs
    • (09:50) Blackwell is mostly pipeline: 80% still announced
    • (11:00) The demand side, measured in gigawatts
    • (12:15) Looking ahead, and why a GPU order isn't a cluster

    Links:

    If you found this useful, rate State of AI with Nathan Benaich five stars and share it with someone building in AI infrastructure - it genuinely helps.

    29 June 2026, 8:44 pm
  • 4 minutes 48 seconds
    STARK: Europe's next defense prime

    Air Street Capital backs Stark, the German multi-domain defense company, in its €500M led by Founders Fund and Sequoia.

    In this episode, we discuss why cheap, software-defined unmanned systems in the air and at sea are the decisive lesson of Ukraine, and why we think co-founder and CEO Uwe Horstmann - a Project A GP and Bundeswehr reservist - is building the German neoprime Europe needs. Round led by Sequoia and Founders Fund, with the NATO Innovation Fund, Project A, and Air Street.

    Links: stark-defence.com · full post at press.airstreet.com · YouTube version

    26 June 2026, 1:32 pm
  • More Episodes? Get the App