• 10 minutes
    AI@70 EPISODE TWO: THE WATERSHED

    THE NEXT BILLION SECONDS: AI@70

    EP 2: THE WATERSHED

    ChatGPT opened the door to artificial intelligence. Almost immediately, researchers built 'agents' using it - but three years of hard yards passed before agents were 'good enough' to be useful. In this episode we trace the path from 'bad' to 'good enough' agents - and what it unlocked as we 'crossed the watershed'.

    No one really thought that on the last day of November in 2022, the world would change completely. Not even the folks closest to it. Like Sam Altman, CEO of OpenAI, who said, "We always knew we'd hit a tipping point," 

    "But we didn't know what the moment would be."

    His firm modestly promoted a 'research preview' of their new AI chatbot.ChatGPT.

    G'day, I'm Mark Pesce and this billion seconds are already unfolding as the most significant of this century.

    In this miniseries, celebrating the 70th anniversary of artificial intelligence, we're looking at where we’ve come from, how we got here - and where we seem to be going.

    Because we’re travelling at the speed of thought.

    ChatGPT promised artificial intelligence on tap for everyone. But it took another three years for that promise to be realised. 

    This episode traces that story, from something kids used to do their homework for them - into a tool reshaping our economy.

    That's on this episode of This Billion Seconds.

    I reckon anyone who used a chatbot in 2023 or 2024 deserves a gold star.

    Two gold stars if they used one to get work done.

    ChatGPT was interesting.

    But it wasn't very accurate.

    The earliest chatbots, trained to be eager to please, would sometimes make up their answers. 

    With no basis in fact. We call these 'hallucinations', but that's not what they are.

    It's that we'd caught the AI out. Asked a question it didn't know how to answer. So, it did its best with whatever it could pull together.

    It didn't mean to lie. It didn't even know it was lying. It was just generating a response. In a chatbot conversation that isn't a big deal. You could fact check - against another chatbot, against the web - even, if you can imagine it, ask another human being.

    Those sorts of hallucinations could be caught - if you went to some effort.

    But then again - why ask a chatbot if you want to make an effort?

    Folks took chatbots at their word.

    And some them paid a price.

    Hallucinations are an annoyance for people.

    But they're show-stoppers for autonomous agents.

    Oh yes, the nearly forty year old dream of Apple CEO John Sculley of autonomous agents running around and doing all of our work for us - that dream flickered back into life alongside ChatGPT.

    Researchers reckoned that with 'good enough' AI, they'd be able to build agents.

    To read your email, Keep your calendar. Book a reservation. That sort of thing.

    Just three months after ChatGPT landed, the first of those autonomous agents popped up.

    A piece of software known as AutoGPT. AutoGPT turns an AI like ChatGPT into an agent.

    It does that by providing the three things an agent needs. Memory - so that the agent can remember what it's supposed to be doing, and keep notes on its progress as it does it.

    Tools - So that it can do things like read and write files, respond to emails, or add items to the calendar, And goal logic - this is the thing that turns an AI into a single-minded goal-oriented piece of software.

    Basically, it's the same quality as the Terminator: the agent has one goal and it will stop at nothing to achieve it. AutoGPT gave ChatGPT memory, tools and goal logic - everything need to turn it into an autonomous agent.

    And it should have been absolutely amazing. Except for one thing.

    Here's how an agent works: you give it a goal, and it 'decomposes' that goal into a series of steps, then breaks those steps down into discrete actions.

    The agent then methodically works its way through each of the actions.

    At the end of every action, an agent 'reflects' - it checks the results of that action.  Did it work? Does the agent get to go on to the next action, or does it need to do this action again?

    And this is where the problems arise. Because if ChatGPT happens to hallucinate in the midst of this process, the whole thing very quickly falls over.

    Maybe an agent thinks an action worked, when it didn't. Or thinks it didn't, even though it did.

    Now there's always a change that - for any given action - there will be a hallucination that will make the agent fall over.

    You can deal with that if there are just a few actions. Chances are it will get through all of the actions before problems arise.

    But for anything even modestly complex, there are many actions. Tens to hundreds. Maybe even thousands.

    So the probability for just one hallucination - that's all it takes to make the agent fall over - grows higher and higher as the list of actions grows longer and longer.

    Now here's the thing: people measured this.

    A group known as Model Evaluation and Threat Research or METR, they've kept a running record of how long agents can run before they fall over.

    And back in early 2023, they'd make it no more than about 4 minutes.

    An agent can't do a lot in 4 minutes. So although we could make agents using ChatGPT, we couldn't make them work well.

    For that, we'd need better AI. Funny thing about that. All of us can take some credit for making AI better.

    You see, most everyone who's using ChatGPT or Claude or Gemini or any of the others, is making those models better with every question we put the them. Those questions and the answers to them get fed back in, to train the next generation of models.

    It's why AI is so much better now than it was three years ago - and why I say anyone who used AI day-to-day in 2023 or 2024 deserves a gold star. Those were hard yards because the AI just wasn't very good. How do we know AI is better to day than it was two or three years ago? Ah, here's where we come back to those folks at METR.

    They've been tracking what they call the 'task horizon' of agents. How long they can perform actions before they fall over. And they noticed that with every subsequent generation of AI, that task horizon got longer. Dependably. It soon became clear that the task horizon doubled an average in seven months. 4 minutes becomes eight minutes toward the end of 2023, then sixteen minutes in mid 2024 thirty-two minutes in early 2025, and then sixty-four minutes. Just over an hour In November 2025. Three years after ChatGPT launched. Now that moment in time - just under a year ago - saw the launch of three brand new models: Google Gemini 3, Anthropic Claude Opus 4.5, and OpenAI GPT-5.2.

    Each of these models could be used to create agents with task horizons greater than an hour.

    An hour is "long enough" that you can assign an agent a reasonably complex task, let it go off and do the work, with confidence that the work will be done correctly.

    That's kind of a magic length of time. It takes agents out of the theoretical - where they'd been for nearly 40 years - and makes them very practical. That's the reason I call this moment "The Watershed". It's the moment when AI gets "good enough" to do real work. I'm far from the only person to notice this. Anyone using AI tools to write software noticed a big shift, as they stopped fighting with their tools. Because the tools had gotten 'smart enough' to handle the task.

    This is the moment where we first hear the term 'vibe coding' - tell the agent what kind of software you want to create, and the agent will go off and build it for you. And now that we're on the other side of the watershed, it's all downhill. Gaining speed. Because the task horizon didn't stop growing at an hour. We're more than seven months beyond November 2025, and task horizons have doubled again. Two hours.

    By early next year, four hours. And by the end of next year - eight hours. An entire work day. Tell the agent what to do - for the whole of the day - and let it do its thing. We've come a long, long way since November 2023. But we're through the hardest bits. You're using the worst AI you'll ever use.

    And boy, will it get better from here. In our next episode, we'll look at what autonomous agents mean for business - and whether business is actually prepared to pay for them. That's on the next episode of This Billion seconds.

    THIS BILLION SECONDS was written and recorded by Mark Pesce. Produced with assistance from Myrtle and Pine. If you like this show, please share it with a friend. And make sure to follow or subscribe to get all of the episodes in this series.This is Mark Pesce, thanking you for listening.

    See omnystudio.com/listener for privacy information.

    21 August 2026, 4:39 am
  • 11 minutes 17 seconds
    THIS BILLION SECONDS: AI @ 70 Ep 1: AI TAKES THE PENSION

    THIS BILLION SECONDS

    AI @ 70 -- EPISODE ONE: AI TAKES THE PENSION 

    On the 17th August 1956, artificial intelligence was born. The founders of the field reckoned they'd have thinking machines equivalent to humans in five year, perhaps ten at the most. What looked easy at the outset turned into a seven-decade struggle - across two "AI winters", when the discipline nearly faded away - before we got to ChatGPT.

     

    TRANSCRIPT: 

    MARK PESCE, HOST: On the 17th August 1956, artificial intelligence was born.

    The final day of a six week workshop that ran through a hot lazy summer at Dartmouth University. 

    A workshop on “artificial intelligence”. 

    Words that had never been put together before - and would never be separate again. 

    G'day, I'm Mark Pesce and this billion seconds are already unfolding as the most significant of this century.

    In this miniseries, celebrating the 70th anniversary of artificial intelligence, we're looking at where we’ve come from, how we got here - and where we seem to be going.

    Because we’re travelling at the speed of thought.

    The future looked bright at the beginning of artificial intelligence.

    This episode looks at how what seemed like a very straightforward goal - making machines that could think - turned into a 70 year struggle. 

    That’s on this episode of THIS billion seconds.

    Claude Shannon.

    Marvin Minsky.

    John Conway.

    These aren’t household names, but in the halls of computer science and in the history of artificial intelligence they are renowned. 

    Three of the founders of the field.

    They were among around 60 different individuals who passed through Dartmouth in the summer of 1956.

    The workshop was purposely constructed very casually. 

    You could come in and stay for the whole thing - or just drop by check it out and move on.

    People were encouraged to give talks. 

    They were encouraged to share their ideas. 

    They were encouraged to brainstorm.

    On the very last day, the 17th of August, they had a set of presentations covering some of the things they’d learned from one another over those six weeks.

    This is the first moment when the shape of artificial intelligence - as we think of it today - becomes visible.

    The workshop participants returned to their respective institutions filled good ideas, high hopes - and completely unrealistic expectations.

    They reckoned they’d have machines that could think in five years. 

    10 years, tops.

    It didn’t work out that way.

    The most obvious obstacle was that no one had a workable definition of intelligence. 

    Not human intelligence, and certainly not artificial intelligence.

    If you don’t know how to describe intelligence how can you build that into a machine?

    So the search for artificial intelligence became a quest to understand intelligence more broadly.

    That’s a hard problem. 

    70 years later it’s not clear that we have a much better idea of what intelligence is. 

    But that didn’t stop these first pioneers.

    Instead, they borrowed from what they all agreed was a marker of high intelligence, deciding that if the computer could reproduce that capability, the computer must be intelligent.

    It's all so logical. And all so very wrong.

    Almost all of them were excellent chess players. 

    All of them agreed that chess was an excellent example of intelligence. 

    So why not teach the computer to play chess?

    It’s an interesting idea. But it was quickly disproven.  Yes, you could teach a computer to play chess. Turns out, that wasn't even terrifically hard.

    A computer can think so much faster than a person it can simply work its way through every chess move available to it, selecting the best one.

    It has the advantage of speed. But not intelligence.

    By 1977 you could buy a little computer that could play chess - and beat the pants off of almost anyone except a master player. 

    Those computers weren’t intelligent.

    Another approach taken early on worked from the idea that it might be possible to teach a computer enough facts about the world that the computer would be able to make intelligent decisions about how to operate in real world.

    This is called the 'top-down' approach to artificial intelligence. 

    On the surface it sounds like a good idea. In practice it’s basically unworkable. 

    Because there is so much knowledge that is embedded in this world, it is almost impossible to teach it to a machine. 

    There have been multiple attempts. 

    The last of them only drew to a close in 2017.

    None of them worked.

    Artificial intelligence was an incredibly active field of research in the 1960s and the 1970s as that first generation of pioneers worked its way through the hard problems.

    Only to learn that the hard problems didn’t get any easier.

    The people paying for all of this work got disappointed in a lack of progress, withdrew their funding, and forward progress in AI ground to a halt.

    This is known as the first “AI Winter“.

    Then came along the second generation of researchers, led by a brilliant Australian by the name of Rodney Brooks.

    Brooks took a look at the world. 

    Creatures as simple as insects could make very good decisions about the world. 

    With very little intelligence. 

    How did they do that?

    Brooks decided to design computers and robots that took their cues from those insects. 

    Suddenly robots could climb staircases. 

    They could walk upright. 

    All sorts of things that have been very hard to do before with the top-down approach - and which we don’t think of as intelligence but actually require a great deal of smarts - they all worked worked with Brooks' 'bottom-up' approach to artificial ingtelligence.

    This caused an enormous burst of enthusiasm. 

    If it was possible to use simple techniques to solve complex problems, maybe, just maybe, we could build artificial intelligence from the bottom up rather than the top down.

    They gave it a red-hot go. 

    And it's not that the approach failed.

    But there are limits. 

    Insect intelligence is incredibly useful.

    It's given us all sorts of crazy robots that can balance themselves and perform dexterous tasks.

    But those body smarts aren't the smarts we think of as smarts - even though they are legitimately smart.

    The kinds of things researchers wanted to do with artificial intelligence... 

    [ John Sculley on Agents.wav  ] 

    That's John Sculley, the man who famously fired Steve Jobs when running Apple Computer. 

    Sculley wanted to leave his own mark on computing, so he pointed to an amazing future of autonomous agents.

    But there was one problem.

    Autonomous agents need artificial intelligence.

    And in 1987 - when Sculley introduced agents to the world - there wasn't any artificial intelligence they could use.

    Insects can't do research.

    And that was the great disappointment of the 2nd age of AI.

    Not that it didn't work.It did.

    But it didn't do hold the answer to the great promise of artificial intelligence: machines that think like people do.

    And so - despite the great robots and the robot vacuum cleaners that popped up in living rooms - interest in AI waned again.

    The second AI winter.

    That winter lasted for more than a decade. From the late 1990s through around 2010.

    And we come to what we think of as the 'modern' age of AI.

    Two great minds played complementary roles.

    Fei-Fei Li and Ilya Sutskever.

    Fei-Fei Li realised that you were going to need a lot of data to train artificial intelligence systems.

    So she put together the IMAGENET project.

    14 million images.

    Each of them carefully labeled.

    A massive project that took years.

    When that was done, then Ilya Sustskever built software known as AlexNet on top of it.

    IMAGENET provided raw information, AlexNet provided the intelligence that allowed a computer program to recognise the difference bertween a dog, a boat and a cloud.

    To do that took huge amounts of information - and massive, massive amounts of computer power.

    But it worked. And because it worked, it kicked off the modern age of AI.

    Five years after AlexNet, eight researchers at Google invented the 'Transformer'.

    The foundation for all modern language models like ChatGPT, Claude - oh, yes, named after Claude Shannon, one of the founders of the field - and Google's Gemini.

    But it took another five years for for the Transformer to prove its worth.

    Because, like IMAGENET before it, it had to be trained on billions and billions and billions of words.

    And who was doing that?  Ilya Sutskever - who had moved on from AlexNet to work at a not-for-profit research company.

    OpenAI.

    Everything he'd learned from AlexNet went into the tool OpenAI was building on top of the Transformer.

    On the 30th of November, 2022 - more than 66 years after the Dartmouth workshop on artificial intelligence, OpenAI launched their tool.

    ChatGPT.

    And - finally - AI had arrived.

    Sort of.

    In our next episode we'll take a look at the last three years of horrible AI, bad AI, mid AI, and finally, 'good enough' AI.

    And how that changed everything.

    That's on the next episode of THIS BILLION SECONDS.

    THIS BILLION SECONDS was written and recorded by Mark Pesce.

    Produced with assistance from Ampel Myrtle and Pine.

    If you like this show, please share it with a friend.

    And make sure to follow or subscribe to get all of the episodes in this series.

    This is Mark Pesce, thanking you for listening.

    See omnystudio.com/listener for privacy information.

    16 August 2026, 6:25 pm
  • 15 minutes 28 seconds
    Pokemon Go used to train killer drones (RNZ Nine to Noon rebroadcast)

    On RNZ's "Nine to Noon" Mark gave host Kathryn Ryan an update on the US export controls on the "Fable 5" AI model - and how that's causing them trouble with their allies, who had already integrated it into workflows. Then they'll look at what happens when a security researcher goes rogue - threatening the world's fourth most valuable company. The episode includes the unpleasant story of how millions of Pokemon Go players unwittingly collected geospatial data now licensed to be used in military drones

    See omnystudio.com/listener for privacy information.

    18 June 2026, 4:10 am
  • 12 minutes 8 seconds
    Unsafe AI? Anthropic's Fable 5 - an AI that may be too smart for its own good (RNZ "Nine To Noon" rebroadcast)

    On Friday 12 June, AI powerhouse Anthropic cut off all access to their latest-and-greatest "Fable 5" AI model. Based on the incredibly powerful Mythos - which has the capacity to hack into almost any computer - Fable had safeguards to block that sort of use - or so Anthropic believed. A hack to 'jailbreak' Fable 5 went viral just a day after it went live - and very quickly, the White House got alarmed. Mark Pesce shares the whole tale with RNZ "Nine To Noon" host Kathryn Ryan.



    See omnystudio.com/listener for privacy information.

    15 June 2026, 2:06 am
  • 17 minutes 56 seconds
    RNZ "Nine To Noon" - Vulnageddon, RAMageddon and ChatGPT Goblins?

    Host Kathryn Ryan in conversation with Mark, profiling the concerning capability of Anthropic's Mythos AI model to hack into the most protected systems - now it seems other AI models are just as powerful. RAMageddon could be on the way - with a major shortage of RAM forcing prices higher and why does ChatGPT have a goblin problem?

    Thanks to RNZ - Nine To Noon

    The Next Billion Seconds with Mark Pesce is produced by Ampel and Myrtle and Pine  

    Listen on Spotify, Apple 

    Sign up for 'The Practical Futurist' newsletter here.

    https://nextbillionseconds.com 

    See omnystudio.com/listener for privacy information.

    14 May 2026, 12:23 am
  • 16 minutes 41 seconds
    RNZ rebroadcast- Does Anthropic's Mythos threaten us all?

    On "Nine to Noon" host Kathryn Ryan and Mark discuss Anthropic's latest 'Mythos' AI model,  considered so dangerous to cybersecurity the company has set up a cross-company partnership to work on patching all of the hacks it found. Will other companies be as responsible if similar, powerful models of AI are created? Ronan Farrow's New Yorker piece on Sam Altman paints him in a diabolical light. And the new White House app has some really concerning security features...

    Thanks to RNZ - Nine To Noon

    The Next Billion Seconds with Mark Pesce is produced by Ampel and Myrtle and Pine  

    Listen on Spotify, Apple 

    Sign up for 'The Practical Futurist' newsletter here.

    https://nextbillionseconds.com 

    See omnystudio.com/listener for privacy information.

    9 April 2026, 1:30 am
  • 18 minutes 19 seconds
    RNZ Nine to Noon Replay: Anthropic vs. the Department of War

    Leading AI firm Anthropic drew two red lines on their work with the newly renamed US 'Department of War': No mass surveillance of Americans, and no autonomous weapons. US Secretary of Defense Pete Hesgeth threatened to declare Anthropic a 'supply chain risk' (basically, an enemy of the state) unless they erased those red lines. Anthropic held fast - and that's changed the politics of AI. Plus: Let the app know you're not dead!

    See omnystudio.com/listener for privacy information.

    5 March 2026, 1:54 am
  • 18 minutes 32 seconds
    RNZ Nine To Noon Replay: French social media ban, Microsoft encryption sharing

    RNZ replay 30 January 2026

    In conversation with RNZ "Nine to Noon" host Kathryn Ryan, we look at the recent massive layoffs at Amazon, France's social media laws for under 15s, Microsoft giving the keys to your data kingdom to the FBI, and what happens when an organisation tasked with maintaining the safety of the airways turns to a computer to 'flood the zone' with regulation?

    Thanks to RNZ - Nine To Noon

    The Next Billion Seconds with Mark Pesce is produced by Ampel and Myrtle and Pine  

    Listen on Spotify, Apple 

    Sign up for 'The Practical Futurist' newsletter here.

    https://nextbillionseconds.com 

    See omnystudio.com/listener for privacy information.

    29 January 2026, 3:10 am
  • 22 minutes 45 seconds
    THE AGETECH AEON #2 - CAN TECH CARE FOR US?

    Abby Bloom, Catherine Ball and I (along with Sally Dominguez) are heading back to the Consumer Electronics Show in Las Vegas to see the latest innovations in 'agetech' - technologies that help restore ability and increase agency as we age. Is agetech "the" solution to the crisis in care, or is it, instead, kit that helps support us as we transition toward a culture where care is a central and valued activity?

    The second of two SXSW Sydney 2025 live recordings. 

    https://cadense.com/ - grippy sneakers

    Hapta

    The Skip with Joy

    See omnystudio.com/listener for privacy information.

    9 December 2025, 5:45 pm
  • 20 minutes 50 seconds
    AGETECH AEON - Ep 1

    Abby Bloom wants us to face an uncomfortable truth: we need a lot more space and time for care in our culture. A rapidly aging population - living twenty to thirty years longer than the generations preceding - means we're going to have to rethink and redesign our lives, our work and our whole civilisation so that we don't disappear beneath a burden of care that no one can bear alone. How do we make that transition? Abby's new book "The Cost of Not Caring", offers some suggestions. It's the shadow looming over all our bright futures - and it's front and centre on THE NEXT BILLION SECONDS.

    The first of two SXSW Sydney 2025 live recordings. 

    See omnystudio.com/listener for privacy information.

    2 December 2025, 6:15 pm
  • 18 minutes 34 seconds
    ALWAYS IMPERFECT - COME TOGETHER FIGHT NOW

    Could it be that the smartphone and social media brought us together only to make us completely intolerant? Research seems to be pointing that way - so what can we do? Also, fake receipts, fake voices and fake apologies - the more AI we use, the more we get faked out. Finally, 'mind captioning' - could bring a voice to those who can not use their own! Recorded live on Radio New Zealand's "Nine To Noon" show on Thursday 20 November 2025 - with big thanks to host Kathryn Ryan!

    The Next Billion Seconds with Mark Pesce is produced by Ampel and Myrtle and Pine  

    Listen on Spotify, Apple 

    Sign up for 'The Practical Futurist' newsletter here.

    https://nextbillionseconds.com 

    See omnystudio.com/listener for privacy information.

    24 November 2025, 8:39 am
  • More Episodes? Get the App