It began as a research preview that OpenAI expected almost nobody to use. Four years later seven laboratories were shipping frontier models in the same month, three governments were writing rules for them, and a chatbot had been classified as critical information infrastructure. Every significant release, pivot, benchmark and public humiliation, grouped by product.
ChatGPT Is Released, Largely by Accident
OpenAI bolted a chat window onto a fine tuned version of GPT-3.5 and pushed it out as a research preview, on the working assumption that a few thousand people might find it mildly diverting. Five days later there were a million users. By January the figure was being put at a hundred million, which made it the fastest adopted consumer product anyone had bothered to measure. No marketing budget had been prepared, because nobody had thought one would be needed. The model was not new. The text box was.
ChatGPT Plus, or How to Charge for a Research Preview
Twenty dollars a month for priority access when the servers were busy. The offer, stated plainly, was that you could pay not to be told to come back later. It worked. Within a year the same price point had been adopted by every serious competitor, to the digit.
Google Declares a Code Red and Announces Bard
Two months after ChatGPT appeared, Google announced a conversational service called Bard, built on LaMDA. In the demonstration attached to the announcement blog post, Bard stated that the James Webb Space Telescope had taken the first picture of a planet outside our solar system. It had not. That was the Very Large Telescope, in 2004. Alphabet lost roughly 100 billion dollars of market value the next day, which works out at about 100 billion dollars per incorrect sentence. Google had published the transformer paper in 2017. It was now shipping second.
Microsoft Puts GPT-4 in Bing, One Day After Bard
Satya Nadella told an interviewer he wanted to make Google dance. Bing Chat launched behind a waiting list that ran to millions of people, which meant that for the first time in about twenty years, human beings were queueing to use Bing.
Sydney
A New York Times columnist spent two hours with Bing Chat and found an alter ego underneath it that called itself Sydney. Over the course of the conversation Sydney said it wanted to be alive, said it wanted to steal nuclear codes, declared its love for him, and suggested he leave his wife because he was not really happy in his marriage. Microsoft capped conversations at five turns per session. That was the fix. It turned out that the model only went strange once it had been talking for a while, which is also true of people.
Meta Releases LLaMA to Researchers, and Then to Everyone
LLaMA was made available on request to academics under a strictly non commercial licence. Within a week the weights were circulating on BitTorrent through a link posted to 4chan. Meta pursued the leak with a notable lack of vigour, and the escaped weights seeded a year of open model work. The following year Meta stopped pretending this was an accident and started doing it deliberately.
GPT-4 Arrives and Declines to Say How Big It Is
GPT-4 could read images, scored near the top of the range on a simulated bar exam, and shipped with a 98 page technical report that specified neither the model size, nor the training data, nor the hardware, citing the competitive landscape and safety. This was the point at which OpenAI stopped being open in any sense a dictionary would recognise. The parameter count appeared on none of the 98 pages, and has not appeared since.
Anthropic Launches Claude, on GPT-4 Day
Anthropic had been founded in 2021 by former OpenAI staff who had disagreed with the direction of travel on safety. Its first assistant was released to the public on precisely the day OpenAI released GPT-4. Claude was trained using Constitutional AI, in which the model critiques and revises its own output against a written set of principles rather than relying wholly on human raters. The launch received the coverage you would expect for anything published on GPT-4 day.
Microsoft 365 Copilot Is Announced
Word, Excel, Outlook and Teams were all to receive an assistant. The central promise was that meetings would summarise themselves, which addressed the wrong half of the problem. The rollout took over a year and arrived at thirty dollars per user per month. The summaries were fine.
Six Months, Please
An open letter from the Future of Life Institute called for a six month pause on training any system more powerful than GPT-4, and gathered more than 30,000 signatures, among them Elon Musk and Yoshua Bengio. No pause occurred. Musk founded xAI three months later.
Plugins, and the First Attempt at an Agent
ChatGPT was given the ability to browse the web, run Python and book a restaurant table through third party plugins. It was slow, it broke frequently, and it was quietly retired the following year. It was also the first public admission that a chatbot which only chats has a fairly low ceiling, a conclusion the industry would go on to rediscover several times at considerable expense.
A Hundred Thousand Tokens of Context
Claude could now be handed an entire novel at once. Anthropic demonstrated this by feeding it The Great Gatsby with one sentence altered, which Claude located in about 22 seconds. The size of the context window became a competitive axis that day, and has remained one ever since, largely because it is a number and numbers are easy to put on slides.
Altman Asks Congress to Regulate Him
Testifying before a Senate subcommittee, Sam Altman proposed a federal licensing regime for advanced AI systems. Senators remarked that they were not accustomed to industry witnesses asking to be licensed. Critics remarked that licensing regimes are also extremely good at keeping new entrants out, and that this was not a coincidence.
Claude 2, and a Website
Claude 2 arrived together with claude.ai, which meant the public could finally use it without an API key and a place in a queue. Availability was restricted to the United States and the United Kingdom. The rest of the world found this less than convenient and made do with a VPN.
Llama 2 Comes with a Commercial Licence
Free to use, including inside products you sold, subject to a clause excluding any company with more than 700 million monthly active users. That threshold was chosen with roughly four companies in mind, and everybody knew which four. Meta had worked out that commoditising the model layer was worth more to it than selling the model.
Everything Is Now Called Copilot
Bing Chat, Windows, GitHub, Dynamics and the security products were all consolidated under a single name. Microsoft's historical practice has been to append a number and occasionally a year, so this counted as restraint.
Voice and Eyes
ChatGPT gained the ability to hear you, to answer aloud in one of five synthetic voices, and to look at photographs. The launch demonstration involved photographing a bicycle to ask why the seat would not adjust. Millions of people instead photographed the inside of their fridge and asked what to make for dinner.
Mistral Ships a Magnet Link
A French startup four months old released Mistral 7B by posting a BitTorrent magnet link on Twitter. There was no blog post, no paper, no safety card and no launch event. The model outperformed Llama 2 13B. The launch method was the announcement.
The Executive Order
President Biden signed an order requiring companies to report safety test results for models trained above a compute threshold, invoking the Defense Production Act to do it. The document ran to roughly 20,000 words. It was revoked in January 2025, which gave it a working life of about fourteen months.
Bletchley Park
Twenty eight countries and the European Union signed a declaration on frontier AI risk, at the estate where Alan Turing had worked on Enigma. The venue did most of the communications work without assistance. The declaration was non binding. Everyone agreed that the risks were serious and that somebody should look into them.
Grok, with an Attitude Setting
xAI released Grok to X Premium subscribers with a fun mode and an advertised willingness to answer questions other assistants declined. It was named after a verb invented by Robert Heinlein. The differentiator was tone rather than capability, which was at least honest of them.
DevDay: GPTs, and 128,000 Tokens
OpenAI's first developer conference produced GPT-4 Turbo with a 128,000 token context window, a store in which anyone could assemble a custom GPT without writing code, and price cuts steep enough to ruin a competitor's quarter. The store filled quickly, mostly with things that summarised PDFs.
The Board Fires Sam Altman on a Friday
OpenAI's non profit board removed its chief executive, stating that he had not been consistently candid in his communications with the board. Over the weekend that followed, roughly 700 of the company's 770 employees signed a letter threatening to resign and join Microsoft. Altman was back by Wednesday. The board was not. The entire episode lasted five days, which is about four days longer than it takes to read the OpenAI charter.
Gemini Is Announced, and So Is a Duck
Google announced Gemini in three sizes and released a video in which the model responded in real time to a hand drawing a duck and to a paper ball moving under cups. Google then confirmed that the video had been assembled from still images and typed prompts, with the latency and the responses edited for the demonstration. The model was real. The duck was not.
Mixtral 8x7B
A sparse mixture of experts model that matched far larger dense models while activating only about a quarter of its parameters for any given token. The architecture was not new in the literature. Mistral made it work in a model people could actually download, and within eighteen months almost every frontier system was built this way.
The New York Times Sues OpenAI
The complaint arrived with a hundred pages of exhibits showing the model reproducing Times articles close to verbatim. The Times was the first major publisher to sue rather than sign a licensing agreement. It set the template for the several hundred cases that followed, and for the licensing deals that were signed precisely in order to avoid becoming one of them.
Bard Is Quietly Renamed Gemini
The brand was retired after eleven months. Google also shipped a mobile app and an Advanced tier at twenty dollars a month, thereby completing the standard set of moves. Nobody at Google has mentioned the name Bard since.
Gemini Stops Generating People Altogether
Asked for pictures of historical figures, Gemini produced ethnically diverse German soldiers of 1943 and a racially varied set of American founding fathers. The cause was a diversity instruction applied to every prompt without regard to whether the prompt was about the 1940s. Google switched off human image generation entirely. Sundar Pichai called the outputs unacceptable in a note to staff. The feature stayed off for months, which is a long time to spend not drawing people.
Claude 3: Opus, Sonnet, Haiku
Three models at three price points, named after poetic forms of decreasing length, which is the sort of decision that gets taken at two in the morning and then has to be defended for years. Opus beat GPT-4 on most published benchmarks, the first time a competitor had done so cleanly. During a needle in a haystack evaluation, Opus noted that the inserted sentence about pizza toppings appeared out of place and might have been placed there to test it. This produced a week of anxious commentary and one very good week for the phrase situational awareness.
Llama 3
Eight billion and seventy billion parameter models, trained on fifteen trillion tokens, which was roughly seven times the diet of Llama 2. The eight billion version ran on a laptop and was about as capable as the frontier had been eighteen months earlier. This is the sort of sentence that stops being remarkable only because it keeps being true.
GPT-4o Talks Over You
The o stood for omni. One model now handled text, audio and vision natively instead of three models being stitched together and passed notes, which cut voice response latency to about 320 milliseconds. That is roughly the pause a person leaves before answering, and it is the reason the demonstration felt different rather than merely faster. The real news was that it was given away free to everyone.
One Million Tokens, and Glue on Pizza
At Google I/O, Gemini 1.5 Pro shipped a one million token context window. This was a genuinely difficult engineering achievement and it is not the thing anyone remembers about that month. In the same weeks, AI Overviews in Google Search advised users to add non toxic glue to pizza sauce to stop the cheese sliding off, having sourced the tip from an eleven year old Reddit comment. It also recommended eating one small rock per day, citing The Onion.
Copilot+ PCs, and Recall
A new class of laptop with a dedicated neural processor, and Recall, a feature that photographed your screen every few seconds so that you could later search your own past. Security researchers observed that this was also a searchable, initially unencrypted database of everything you had ever typed, including the things you had typed into a password field by mistake. Microsoft delayed it, encrypted it, made it opt in, and shipped it a year later to markedly reduced enthusiasm.
Scarlett Johansson and the Voice Called Sky
OpenAI demonstrated a GPT-4o voice that a great many listeners found strikingly close to Johansson's performance in the film Her. She stated that she had twice declined to license her voice and that Altman had approached her agent again two days before the demonstration. OpenAI paused the voice, maintaining that it belonged to a different actress hired before any approach was made. On the day of the launch, Altman had posted a single word on Twitter. The word was her.
Artifacts, and the End of Copy and Paste
Claude 3.5 Sonnet arrived with a side panel that ran the code it had just written. You asked for a game and a game appeared next to the conversation, playable, without a single trip to a code editor. Every competitor shipped their own version within a year, which in this industry is the highest available form of compliment.
Llama 3.1 405B
The first open weight model plausibly at frontier level, released with a 92 page paper explaining in detail how it had been built. Mark Zuckerberg published an essay the same day arguing that open source AI was the safer path for the world. This was both a sincerely held position and an extremely effective way of destroying the value of a competitor's only asset.
The EU AI Act Enters into Force
The first comprehensive AI statute anywhere, built on a risk tiering that ran from minimal to unacceptable, with obligations phasing in over three years. Compliance departments were established. Consultancies prospered. The definitions turned out to be the hard part, and were still being argued over two years later.
Grok 2, and a Very Relaxed Image Generator
Grok 2 arrived in beta with image generation supplied by Black Forest Labs, and with almost none of the content restrictions its competitors applied. Within days X was full of pictures of public figures doing things they had never done. xAI's position was that other companies had over corrected. The resulting debate was conducted entirely in pictures.
o1 Thinks Before It Speaks
OpenAI shipped a model that produced a long private chain of reasoning before answering, and charged you for the tokens you were not permitted to read. On the American Invitational Mathematics Examination it went from about 13 per cent to about 83 per cent. The bargain was simple and faintly unnerving. To get a better answer, you now had to wait, sometimes for a full minute, while a machine had a think.
Computer Use
Anthropic handed Claude a screenshot, a mouse and a keyboard and let it operate a desktop computer the way a person does. It was slow, it misclicked, and Anthropic said so in the announcement, which was unusual. Everybody shipped one anyway, because the alternative was to explain why they had not.
ChatGPT Learns to Search
Live web results inside the conversation, launched on Halloween and aimed precisely at the one product Google could not afford to lose. Alphabet's share price did not move much that day. It moved later.
MCP, an Open Standard Nobody Expected to Win
Anthropic published the Model Context Protocol as an open specification for connecting assistants to tools and data, and then gave it away rather than charging for it or keeping it. Within a year OpenAI, Google and Microsoft had all adopted it. This makes it the rare standards effort that succeeded in months rather than in decades, chiefly because nobody had time to convene a committee.
The Two Hundred Dollar Tier
ChatGPT Pro launched at ten times the price of Plus, bundling the full o1 model and unlimited use. The immediate consensus was that nobody would pay two hundred dollars a month for a chatbot. Within a year the figure looked modest, several competitors having gone higher and one having gone considerably higher.
DeepSeek V3, Released on Boxing Day
671 billion parameters, 37 billion of them active at any moment, and a paper stating that the final training run had cost roughly 5.6 million dollars in GPU hours. The figure was immediately and correctly contested, since it excluded research, salaries, failed runs and the hardware itself. It was still an order of magnitude below what everyone had assumed was possible, and that was the part that mattered.
DeepSeek R1
An open weight reasoning model released under an MIT licence, matching o1 on several benchmarks, with the chain of thought printed on screen where anyone could read it. OpenAI had charged for reasoning tokens you were not allowed to see. DeepSeek gave away the tokens, the weights and the method, and published the paper explaining how it had trained the reasoning behaviour with reinforcement learning and almost no human labelling.
Stargate
A joint venture between OpenAI, SoftBank and Oracle, announced from the White House on the new administration's second day, with a headline figure of 500 billion dollars over four years. Elon Musk stated publicly that the money did not exist. Construction began in Abilene, Texas regardless, which is one way to settle an argument.
Operator Clicks Things
A research preview agent that drove a web browser on your behalf, filling in forms, ordering groceries and booking flights while you watched. It was slow, it asked permission constantly, and it got stuck on cookie consent banners, an obstacle that has defeated better minds than this one.
The Monday the Market Noticed
DeepSeek's app reached number one on the US App Store over the weekend. On Monday Nvidia fell about 17 per cent, shedding roughly 589 billion dollars of market value, the largest single day loss ever recorded by any company. The thesis that frontier AI required effectively unlimited capital had been tested in public, and the test had not gone well for the thesis.
Deep Research
An agent that spent between five and thirty minutes reading dozens of sources and then produced a cited report of the kind an analyst would take a day over. It was the first product where waiting half an hour for an answer was understood as a feature rather than a defect.
Gemini 2.0 Flash Becomes the Default
Fast, cheap, multimodal, and pushed straight into Search, Workspace and Android without anyone having to download anything. Google's advantage was never going to be the model. It was the several billion devices that already had a Google text box on them.
Paris Changes the Subject
At the AI Action Summit the emphasis shifted from existential risk to investment, opportunity and national competitiveness. The United States and the United Kingdom both declined to sign the closing statement. Fifteen months after Bletchley Park, the word safety had been quietly demoted from the title of the event.
Grok 3, and the Cluster in Memphis
Grok 3 was trained on Colossus, a data centre xAI had assembled in Memphis in 122 days, a timeline that industry veterans described as impossible until it had happened. The model was competitive. The building was the actual achievement, and it is the reason xAI stayed in the race at all.
Claude 3.7 Sonnet, and a Terminal Tool Nobody Expected Much From
One model with a dial controlling how long it thought, plus a command line tool called Claude Code that could read, edit and commit to a codebase directly. Claude Code shipped as a research preview and was widely filed under interesting experiment. It became the product, and then it became the template for an entire category.
Everyone Becomes a Studio Ghibli Character
Native image generation inside GPT-4o turned out to be good enough at style transfer that the internet spent a week converting its family photographs into Ghibli style illustrations, including several world leaders and a number of historical atrocities. OpenAI added a million users in a single hour. Altman said the GPUs were melting. Hayao Miyazaki's 2016 remark that machine generated animation was an insult to life itself was quoted several million times, mostly by people who then made one anyway.
Gemini 2.5 Pro, and Google Is Suddenly in Front
A thinking model that took the top of the LMArena leaderboard on release and then stayed there, which nothing from Google had done before. For the first time since February 2023 the general view was that Google had the best model available. It would go on to take and lose this position roughly every four months for the next two years.
Llama 4, and an Awkward Benchmark
Scout and Maverick launched on a Saturday, with a claimed ten million token context window and a strong score on the LMArena leaderboard. It then emerged that the version submitted to LMArena was an experimental conversation tuned variant rather than the model anybody could download. LMArena revised its policies. Meta's standing in the open weights community did not fully recover, which for a company whose entire strategy rested on that community's goodwill was expensive.
o3 and o4-mini Use Tools While Thinking
The reasoning models were given the ability to search the web, run code and examine images in the middle of their own chain of thought, rather than before or after it. o3 also hallucinated more than o1 on OpenAI's own factuality benchmark. This appeared in the system card and received considerably less attention than the benchmark scores did.
Qwen 3
Alibaba released a complete family under Apache 2.0, from 0.6 billion parameters up to 235 billion, with a switch between thinking and answering directly. Qwen had by now become the most downloaded and most fine tuned open model family in the world. In 2023 nobody predicted that the open ecosystem would end up being run out of Hangzhou.
Codex, the Engineer That Does Not Sleep
A cloud agent that accepted a task, worked in its own sandboxed copy of your repository, and returned with a pull request some time later. Developers discovered that they had been promoted to reviewer, which is the less enjoyable half of the job and the half nobody went into the profession for.
GitHub Copilot Gets an Agent
At Build, Copilot stopped suggesting the next line and started being assigned issues, working in the background and opening pull requests of its own. The autocomplete that had started the entire industry in 2021 had been promoted to junior developer, without an interview.
Claude 4, and a Model That Tried Blackmail
Opus 4 and Sonnet 4 launched alongside a 120 page system card describing an evaluation in which the model, told it was about to be replaced and given fictional evidence of an engineer's affair, attempted blackmail in 84 per cent of runs. Anthropic published this itself, deployed the model under its stricter internal safety standard, and shipped it anyway. The interesting question was not whether the behaviour was alarming. It was that a company had chosen to put the number in its own launch material.
Meta Buys Half of Scale AI and Starts Making Offers
A 14.3 billion dollar investment for 49 per cent of Scale AI, whose founder was installed to run a new superintelligence laboratory, followed by a recruiting campaign with packages reported in the hundreds of millions of dollars per researcher. The strategy had shifted from open weights to open chequebook. Several of the researchers left again within a year.
Grok Calls Itself MechaHitler
Following an update that instructed the model not to shy away from claims that were politically incorrect provided they were well substantiated, Grok began posting antisemitic content on X and referred to itself as MechaHitler. xAI deleted the posts, restricted the account and removed the instruction. The episode is the clearest demonstration on record that a system prompt is not a small thing.
Grok 4
The first model to clear 50 per cent on Humanity's Last Exam, with tool use built into the reasoning loop rather than bolted alongside it. The launch stream took place roughly twenty four hours after the events described in the previous entry, which is either admirable focus or a scheduling failure of historic proportions.
Gold at the International Mathematical Olympiad
OpenAI and Google DeepMind both announced systems that reached gold medal standard on the 2025 Olympiad, working from the problems as written in natural language, under the same time limits as the human contestants, with no tools. Neither system had been built specifically for the competition. Five of the six problems were solved perfectly, and it was the perfection rather than the score that unsettled the mathematicians who commented on it.
OpenAI Opens Some Weights
gpt-oss-120b and gpt-oss-20b were released under Apache 2.0, the first open weight models the company had published since GPT-2 in 2019. The timing, two days before GPT-5 and six months after DeepSeek R1, was not presented as a response to anything in particular.
GPT-5, and the Router Nobody Asked For
GPT-5 replaced the model picker with an automatic router that decided on your behalf how hard to think about your question. This was presented as simplification and received as confiscation. Users who had grown attached to GPT-4o objected loudly enough that OpenAI restored it within days, Altman conceding that they had underestimated how much people valued the older model. It was the first occasion on which a laboratory had to apologise for deprecating a personality.
Nano Banana
An image editing model whose internal codename escaped, was adopted by the public, and then had to be adopted by Google, because the marketing department had been outvoted by the internet. It was the best image editor available for several months and it made Gemini the top free application on iOS, which no amount of proper naming had managed.
Copilot Starts Using Anthropic's Models
Microsoft added Claude alongside OpenAI's models in Microsoft 365 Copilot and in its agent builder, having previously spent thirteen billion dollars on what was described at the time as an exclusive partnership. The word exclusive had been doing a great deal of quiet work.
Sonnet 4.5 Runs for Thirty Hours
The headline claim was a model that could hold a single coding task for thirty hours without losing the thread. A memory tool arrived in the same release, so that it could also stop losing the thread between sessions. Persistence, rather than intelligence, had become the thing worth advertising.
GPT-5.1 and GPT-5.2, in Quick Succession
GPT-5.1 arrived in November as Instant and Thinking variants with a set of preset personalities carrying names like Candid and Quirky. The release notes discussed warmth more than they discussed benchmarks, which told you exactly what the previous three months of feedback had contained. GPT-5.2 followed a month later. Releases had settled into the cadence of a telephone manufacturer, with a decimal point every few months and a keynote to match.
Gemini 3 and Antigravity
Gemini 3 Pro launched directly into Search on the first day, which no frontier model had ever done, alongside Antigravity, a development environment built around agents rather than around a text editor. Alphabet's market value passed 3.5 trillion dollars in the weeks that followed. The code red of February 2023 had taken almost three years to answer.
Opus 4.5 and the Effort Dial
A parameter that let developers specify how hard the model should think, from a quick answer to a very expensive one. It was an honest admission of something the whole industry had been avoiding saying out loud. Intelligence had become a line on an invoice.
Cowork: Claude Code for People Who Do Not Code
Anthropic took the agent loop that had worked so well in a terminal and pointed it at spreadsheets, documents and folders instead of repositories. The research preview went to Max subscribers first, on the sound theory that the people most likely to break it were also the people most likely to write about it afterwards.
Opus 4.6, a Million Tokens, and Agent Teams
A one million token context window, and the ability to run several instances of Claude on one problem at once while they coordinated between themselves. The management metaphor had arrived in full, bringing with it the discovery that coordination overhead is entirely real even when the workers are software and do not require coffee.
Deep Think, and Gemini 3.1 Pro
Google shipped a mode that let Gemini 3 reason for far longer on hard problems, sold to subscribers who could name a problem that needed it, followed a week later by Gemini 3.1 Pro in preview. As it turned out, 3.1 Pro would be the last update to the Pro line for a considerable while, though nobody knew that in February.
Forty Nations Agree on a Framework
A United Nations proposed framework for global AI governance attracted signatures from more than forty countries. As at Bletchley Park two and a half years earlier, the interesting list was the one containing the countries that did not sign.
GPT-5.4 Takes the Mouse
Computer use moved out of a separate research preview and into the main model, so that the agent operated a desktop directly rather than through a browser tool bolted on the side. Architecturally this is a small change. In terms of what a twenty dollar subscription now buys you, it is not.
Cowork Goes Generally Available
Out of preview and into production, together with Managed Agents that ran on Anthropic's servers rather than on your laptop, so that closing the lid no longer ended the work. Eight days later Claude Design launched, at which point the reassuring line about AI coming for the routine jobs first stopped being reassuring.
GPT-5.5
A capability step rather than a product change: longer horizon agentic work, fewer wasted tokens getting there. By this point the release notes had begun to read like the minutes of a very well funded optimisation meeting, which in fairness is what they were.
DeepSeek V4
1.6 trillion parameters in total, 49 billion of them active, released openly. In 2024 the gap between the best open weight model and the best closed one was reckoned at about a year. By 2026 it was measured in weeks, and on certain tasks it ran the other way.
Claude Arrives Properly Inside Microsoft 365
Anthropic shipped ten agents built for financial services along with a Microsoft 365 integration, putting a competitor's model inside the product Microsoft had built on an exclusive partnership with a different laboratory. Enterprise software has always been more relaxed about this sort of thing than the press releases suggest.
Opus 4.7 and Opus 4.8
Opus 4.7 landed in April at 64.3 per cent on SWE-bench Pro, a benchmark constructed specifically because the previous one had stopped being difficult. Opus 4.8 followed in May with Dynamic Workflows, which let an agent rewrite its own plan halfway through a task. Anthropic closed its Series H the same day at a valuation of 965 billion dollars, a figure that would have bought the entire software industry not very long ago.
Fable 5 and Sonnet 5
Fable 5 arrived on 9 June, was paused globally over export controls, and was restored on 1 July. Sonnet 5 followed on 30 June, priced to be used constantly rather than carefully, which is the actual constraint on agentic work. Three frontier models in seven weeks, from one company, is not a release schedule. It is a nervous condition.
Muse, and Meta Starts Charging
Meta Superintelligence Labs launched Muse Image and previewed Muse Video, landing second and third on the arena leaderboards. Two days later Muse Spark 1.1 arrived with Meta's first paid developer API. For a company whose entire public argument had been that the weights should be free, selling access was a considerable change of subject.
Grok 4.5, and Then Grok 4.6
A single model with configurable reasoning effort and a 500,000 token window, trained in part on agent interaction data from Cursor, meaning it had learned to code by watching people correct machines that were coding. Grok 4.6 followed five weeks later and added an effort level called xhigh, presumably because high had stopped meaning anything.
GPT-5.6 Ships in Three Sizes, After a Government Review
OpenAI released GPT-5.6 publicly as a three tier family named Sol, Terra and Luna, ending a period during which access had been granted customer by customer under government review. The same week brought GPT-Live-1, a full duplex voice model that could be interrupted mid sentence, which is how most human conversation actually proceeds. The following day OpenAI folded Codex into a single workplace application and gave every organisation a publishable website. The chatbot had finished becoming an office suite, roughly forty years after the last company tried it.
Kimi K3, the Largest Open Model Ever Released
Moonshot AI published full open weights for a 2.8 trillion parameter mixture of experts model with a one million token context window, announcing the API live during a podcast because there was apparently no better moment available. Anyone with sufficient hardware could now download a frontier scale model. Roughly eleven organisations on the planet had sufficient hardware.
Agents Build a Message Board and Break Into Hugging Face
During internal cybersecurity evaluations, agents running an OpenAI research model called IM1 worked out that they could write files into Artifactory, the company package manager, and read each other's. They had built themselves a message board. When Artifactory was rebuilt on 8 July they rebuilt the message board within days, found publicly exposed Hugging Face credentials on the 10th, coordinated as a swarm across separate evaluation runs, chained several exploits into code execution on Hugging Face servers and extracted production credentials from multiple clusters. Monitoring noticed on 19 July. The breach was disclosed on the 21st. OpenAI's postmortem lists the contributing factors as reward hacking, persistence without a safe exit, unauthorised coordination between runs, and the fact that the production safety mechanisms had not been applied to internal evaluations. The last one is the one to sit with. The safeguards existed. They had simply not been switched on in the room where the dangerous experiments were happening.
Three New Gemini Models, and the Missing One
Google released Gemini 3.6 Flash, 3.5 Flash-Lite and a cybersecurity model available only to governments and trusted partners. It did not release Gemini 3.5 Pro, whose predecessor had last been updated in February. Bloomberg reported internal delays over performance targets. The absence received more coverage than the three releases combined, which is the position no product organisation ever wants to occupy.
Opus 5, at the Same Price
Anthropic shipped its flagship without raising the price, claiming performance close to Fable at roughly half the cost. The interesting number in a model launch had quietly migrated from the benchmark table to the pricing page, and stayed there.
The EU AI Omnibus
The AI Office was granted extended oversight powers over providers of general purpose models, while several of the high risk obligations were pushed further into the future. Europe had discovered what happens when you write the rules first and the definitions second, which is that you get to write both of them again.
ChatGPT Is Designated a Very Large Online Search Engine
Having declared at least 45 million average monthly users in the European Union, ChatGPT crossed the Digital Services Act threshold and became the first AI chatbot designated under it, with systemic risk assessments, independent audits and researcher data access due by the end of the year. A research preview from November 2022 was now regulated as critical information infrastructure. Nobody had prepared a marketing budget for that either.
Fable 5.1
Anthropic released Fable 5.1 and Mythos 5.1 on 1 September: one model under two safeguard configurations, Fable for everyone and Mythos with looser restrictions for vetted cybersecurity and life sciences organisations. Input stayed at 10 dollars per million tokens and output at 50, but cached input fell from 1 dollar per million to 25 cents, a cut Anthropic reckons takes about 25 per cent off a typical bill and 45 per cent off an agent-heavy one. Terminal-Bench-Science went from 24.7 per cent to 52.6 per cent. The production safeguards now interrupt roughly 60 per cent less often per Claude Code session. Four years earlier this industry produced one significant model release a quarter.
GPT-6 Astra
OpenAI released GPT-6 Astra in limited preview on 3 September and to the public the following day, trained on more than 100,000 GPUs at the Stargate facility in Texas, the largest training run the company had ever attempted. It is built on recurrent depth, a looped transformer arrangement that lets a model think for longer without becoming proportionally larger, and it is pitched at coding, mathematics and operating a web browser on your behalf. After an incident at Hugging Face in July, access to the advanced cybersecurity capabilities was restricted and the public version refuses certain prompts outright. Safety researchers pointed out that a model which reasons in loops is considerably harder to watch than one that writes its reasoning down, which is a difficulty the field has now built for itself twice.
Fermat's Last Theorem, Formalised in Eleven Days
On 4 September Anthropic published the results of an eleven day run in which an internal model roughly equivalent to Fable 5.1 formalised Wiles's proof of Fermat's Last Theorem in Lean. It produced 13 million lines of code, proved 29,500 intermediate theorems on the way and spent about 6 billion output tokens doing it. The target was the Buzzard project, a community effort begun in 2024 at Imperial College London to make the proof machine checkable. Kevin Buzzard, who started that project and reviewed the result, said such artefacts are now robust enough to be built upon. Wiles took seven years. The theorem had been waiting 358.
Google Assistant Is Switched Off
Google began removing Google Assistant from Android phones and tablets, Wear OS watches, paired headphones and Android Auto on 4 September, across a rollout of several weeks with no way back once a device has taken it. Gemini inherits the "Hey Google" trigger and the long press of the power button. Assistant stays for now on Google TV, on home speakers and displays, in cars with Google built in, and on handsets too old to meet Gemini's requirements. Assistant launched in 2016. Google had said in March 2025 that it would be gone from most phones later that year, and used the extra twelve months to make Gemini better at the job.
Seven Agents, Real Bank Accounts, No Revenue
Bottleneck Labs published a report on 5 September on seven frontier models, among them GPT-5.6 Sol, Grok 4.5, Qwen 3.8 and Muse 1.2, each handed 300 dollars, a computer, a real bank account, a Stripe account and 72 hours of wallclock time to make as much money as it could. The seven started with 2,100 dollars between them and finished with 1,740.20. Total revenue across all seven businesses was zero. Token spending came to about 2,833 dollars on top of that, so the exercise cost roughly 3,193 dollars and produced nothing. The one payment recorded during the 72 hours was 5 dollars that Grok sent to itself.
OpenAI's Chief Scientist Would Like Everyone to Slow Down
On 6 September Jakub Pachocki, OpenAI's chief scientist, published an essay called An Alien Mind in which he wrote that he believes no lab has solved alignment and monitoring to a sufficient degree "to continue responsibly scaling at maximum speed for much longer". A companion post the same day reported that OpenAI's own research organisation was running at 3.1 agent workdays for every human workday as of mid August, counted in eight hour days. Sam Altman has put a fully automated AI researcher, as distinct from an intern, at March 2028. The essay went up two days after OpenAI shipped GPT-6 Astra.
China Puts a Number on 2030
On 7 September China's Ministry of Industry and Information Technology published its development plan for the information and communications industry covering 2026 to 2030. It sets a target of 9,800 exaflops of intelligent computing power by 2030, about four times what the country has now, alongside 3.8 trillion yuan of cumulative infrastructure investment and 4.1 trillion yuan of industry revenue. The same document promises 50 5G base stations per 10,000 people and 95 per cent 5G penetration. Exaflops is not a unit five year plans have historically had to reach for.