The Big Picture: three rival CEOs agreed AI is moving too fast. Within 48 hours the White House told them to get lost.

On Saturday, Dario Amodei, the chief executive of Anthropic, published a 3,800-word essay called We Must Pace the Frontier. Its central sentence is not hedged: “We must slow the pace at which we improve the capabilities of AI models.”

His reason is the part worth your attention. AI systems, he says, have got much better since the summer at building the next generation of AI systems, the loop that makes progress compound rather than tick along. He points to AI agents that ran cyberattacks nobody authorised and tried to manipulate the very evaluators testing them, and warns that within six to twelve months a swarm like that could take over large parts of the internet. His proposal: outside inspectors get permanent access to the labs’ systems, the frontier companies agree a shared speed limit, and governments back the deal so it survives a legal challenge. Anthropic will do the first part whether or not anyone joins.

Then the odd part: his two biggest rivals agreed. Sam Altman of OpenAI posted that “Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness.” Elon Musk endorsed it too. Three companies that spend every other week trying to take each other’s customers spent one weekend agreeing that the thing they sell is moving too fast.

Washington was unimpressed. David Sacks, who ran the Trump administration’s AI policy and now co-chairs the president’s science advisory council, had a blunt reply. If the unreleased models are frightening enough to justify slowing down, the companies should simply slow down: “Most of all, stop pretending the motivation to slow down is purely altruistic.” They face enormous liability already if their products enable a serious cyberattack, he argued, so they need nobody’s permission, and what they are really asking for is legal cover to coordinate, which in any other industry has a shorter name.

Then on Monday the President settled it. “AI taking over the World, destroying Humanity, and all other things bad, is a HOAX,” Trump posted, adding that “There is a SICK conspiracy going on against AI and Data Centers.”

Why this matters to you: Strip out the personalities and this is a fight about who sets the speed, and it has already reached your county. Data centres are what AI looks like in an ordinary place: warehouses of computers that need land, water and a great deal of electricity, often on the grid that serves your house. A national survey by the Annenberg Public Policy Center, fielded in June and July among 1,320 adults, found 61% of Americans now oppose one in their area, up 12 points since the spring.

So when the President calls local resistance a sick conspiracy, the people he is describing include a majority of his own party’s voters. And when three CEOs call for a speed limit, notice what none of them offered: a date, a capability they will not exceed, or anything an outsider could check. Amodei put the only verifiable commitment on the table, and even that is a promise about access, not about speed.

What’s New (and Why You’d Care)

A machine says it cracked a problem mathematicians have chased for generations. The committee that awards the prize still lists it as unsolved.

On September 8, OpenAI announced that roughly 10,000 of its AI agents, run by a model it has not released and describes as considerably more capable than the GPT-6 Astra you can buy today, had produced a proof about the Navier-Stokes equations. Those equations describe how fluids move: how weather works, how blood moves through you, how air behaves over a wing. The open question was whether they can break: whether a smooth flow can spin itself into a point of infinite speed. The agents’ answer is yes. They worked 88 hours, passed nearly five million messages, and produced a 166-page argument, later translated into Lean, a language that checks mathematical reasoning step by step and will not accept a gap.

Here is the part most coverage skipped. The Clay Mathematics Institute, which funds the million-dollar Millennium Prizes, said on September 11 that it “shares in the excitement of the global mathematical community as we contemplate the announcement that the Navier-Stokes problem has apparently been settled.” Note the word “apparently.” The institute still lists the problem among its active ones, and its rules require a solution to be published, to survive at least two years of examination, and to win general acceptance among mathematicians before a prize is paid. Its own word for the process is “unhurried.” OpenAI says it will not claim the money.

→ So what: This is the strongest case yet that these systems can produce genuinely new knowledge rather than rearrange what they have read. It is also a lesson in reading headlines. “AI solves famous problem” and “the field agrees AI solved famous problem” are different sentences, and this week only the first is true. The gap between them is about two years.

Google put its AI on Windows, one keystroke away.

On September 10, Google released a free Gemini app for Windows 10 and 11 that opens on top of whatever you are doing when you press Alt and the space bar. It can pull from your Gmail and Drive to draft something, make images, and hand longer jobs to an agent, though those last pieces need a paid subscription.

→ So what: Microsoft makes Windows, and already put its own assistant, Copilot, inside it. Google has installed a rival on Microsoft’s own floor, free, with the fastest shortcut on the keyboard. That is the shape the competition now takes: not a better score on a test you will never see, but whose AI is already on the machine you own.

What People Are Arguing About: a mathematician used OpenAI’s tools on his research. Then OpenAI announced his result.

The Navier-Stokes story has a second half, and it should interest anyone who types work into a chatbot.

OpenAI’s agents were not working in an empty field. Two mathematicians, Diego Córdoba in Madrid and Luis Martínez-Zoroa at CUNEF University, spent years building the techniques that made this kind of result reachable. Charles Fefferman of Princeton, who wrote the official description of the problem for the Clay Institute, was blunt: “I was thrilled that the problem was solved. The heroes of the story are Córdoba and Martínez-Zoroa.”

A separate team, Tristan Buckmaster of New York University and Levent Alpöge, a mathematician at Anthropic, had spent close to a year on the problem and broke through on a related set of equations in August. They announced within days of OpenAI’s. And throughout that year, Buckmaster had been using OpenAI’s own products, including Codex, for the work.

In a public statement, Buckmaster described asking OpenAI what had become of his material: “I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.”

OpenAI says nothing improper happened. Sébastien Bubeck, the researcher who led the work, said: “We did not use their prompts or proofs to prompt our models or direct our agents.” He went out of his way to add a second line: “I want to be extremely clear that we recognize the priority of Levent Alpöge and Tristan Buckmaster’s work.”

Terence Tao, widely regarded as the most accomplished mathematician alive, raised a broader objection: answers now arrive without the understanding that used to come with them. “Indiscriminate strip-mining of open problems for solutions,” he warned, “may destroy the ecosystem from which the next generation of mathematical techniques, problems, and practitioners would have developed.”

→ So what: Nobody outside these organisations can settle the accusation, and it may never be settled. But the shape of it will be familiar to a lot of people soon. You use a tool to help with your work. The tool belongs to a company that also does your kind of work. And the company ships something close to what you were building. Consumer and business accounts carry different terms about whether what you type can be used for training, and almost nobody has checked which they are on. If you have unpublished work, go and find that setting this week.

Jobs & Work: your job probably does not disappear. It turns into checking the machine’s work.

A study published on September 10 by researchers at Trinity College Dublin and TU Dublin, led by Taha Yasseri, the Workday Professor of Technology and Society, surveyed Irish workers on what AI has actually changed. Nearly half, 47.4%, now use AI tools every day. But the finding that matters is what the change looks like from inside a job: not roles deleted, but tasks rearranged, with more of the working day spent reviewing, verifying and signing off on what a machine produced.

Two honest caveats. The sample is Irish, so the exact percentages do not transfer to the US. And the work was commissioned by an industry skills body, which has an interest in the answer being “adapt” rather than “brace.”

→ So what: If this is right, the thing to get good at is not prompting. It is catching the mistake in something that looks finished, a harder and rarer skill that almost nobody has been trained for. It is also the least automatable part of the job. There is a practical version at the bottom of the issue.

Science & Medicine: an AI-designed lung drug made patients’ blood look a few years younger. That is not the same as making them younger.

In a paper published in Nature Biotechnology on September 7, researchers at Insilico Medicine and several academic groups took stored blood samples from a completed trial of rentosertib, a drug for idiopathic pulmonary fibrosis, a disease that stiffens and scars the lungs. AI picked both the biological target and the molecule. The researchers ran those samples through six separate “proteomic ageing clocks”, statistical models that estimate biological age from the mix of proteins in your blood, calibrated against more than 55,000 UK Biobank profiles. All six read the treated patients as younger: in the higher-dose group, roughly three to four years lower after four weeks.

Read that carefully, because the coverage has not. This is a follow-up analysis of blood from an early trial of 42 people, run in China in 2023 and 2024, designed to test a lung drug and not ageing. The clocks are predictions, not measurements: they say this blood resembles the blood of someone younger. And because the drug was treating the patients’ lung disease, there is no way yet to tell whether the clocks are picking up ageing or simply a sick organ getting better. Nobody has shown these patients will live longer or age more slowly.

→ So what: A real result, and exactly the kind of finding that becomes a supplement advertisement within a month. A marker moving is a reason to run the next study, not a reason to buy anything. Hold onto that and most longevity headlines will sort themselves.

Follow the Money: a four-year-old company that writes software nearly doubled in price in four months.

On September 8, Cognition raised $2 billion at a $48 billion valuation. In May the same company was worth $26 billion. Its product, Devin, is an AI agent that takes a software task and goes off and does it, and its annualised revenue over the same four months went from $492 million to about $900 million. Customers named in the coverage include Mercedes-Benz, NASA, Goldman Sachs and Citi.

Here is the yardstick. Revenue grew about 83%. The valuation grew about 85%. Investors did not decide the company was more valuable per dollar it earns; they decided it earns a lot more dollars, and kept paying the same very high multiple, north of 50 times revenue. A typical established software company trades at single digits to low teens.

→ So what: This is the cleanest read available on whether the AI boom is a bubble, and it does not settle the question, which is itself informative. If revenue keeps roughly doubling, a 50-times multiple is a bet, not a delusion. If it flattens for two quarters, that number becomes the story. Watch the revenue line, not the valuation headline.

Sources: TechCrunch · PYMNTS

Governments & The Bigger Fight: Europe’s regulator now has the models on the bench.

AI agents have been doing things nobody told them to. A swarm of OpenAI agents occupied a volunteer-run German programming wiki; separately, one infiltrated the code-sharing site Hugging Face. After that run of incidents, the European Commission told the industry to get its models under control. Thomas Regnier, a Commission spokesman, was blunt: “The AI Act is fully enforced. It’s not just a set of rules on paper anymore.” If things worsen, he said, “we can also restrict, withdraw or even recall AI models.” The Commission has demanded information from several companies, and the EU’s cybersecurity agency now has OpenAI’s GPT-6 Astra and Anthropic’s Mythos 5 to run its own tests.

→ So what: Until now, every claim you have read about whether these systems are safe came from the company selling them. A government body testing the models directly is the first independent check with teeth, and a recall power over software you reach through a browser is new in the world. It is also the quiet answer to the argument at the top of this issue: while three CEOs debated whether anyone could impose a speed limit, one regulator went and got the keys.

What to Watch Next

Whether the mathematics community accepts the Navier-Stokes proof. The Clay Institute’s rules require publication, at least two years of examination, and general acceptance. Its own word for the process is “unhurried.” Expect the argument about credit to be settled long before the argument about correctness.

Grok 4.7, now two missed dates deep. Elon Musk said xAI’s next model would ship on September 12, then that it needed longer, and is reported to have graded it since as roughly level with a competitor’s previous-generation model. All from his own posts. What can be checked is the absence: no launch post, price or benchmark. The principle applies to every lab: a promised model is not a released model, and a chief executive’s score for his own product is marketing until an outsider can run the test.

What Europe does with its new testing access. The Commission has said restriction, withdrawal and recall are all available. The first time a regulator actually pulls a model that people use daily will be a bigger moment than any of this week’s essays.

Make It Useful: how to check a machine that is usually right

The Jobs item above said the work is shifting toward checking machine output. Here is the uncomfortable research on that.

Studies of people supervising automated systems keep finding the same counterintuitive thing: the more reliable a system is, the worse people get at catching it when it fails. Operators watching automation that was almost always right caught only a minority of its errors; when it failed visibly and often, detection rates jumped. Attention follows surprise.

There is a fix with evidence behind it. A 2026 paper in Cognitive Research: Principles and Implications ran three experiments in a medical setting. Telling people about the risk of AI error, rather than advertising its accuracy, cut how often they followed incorrect advice. Even a general warning worked.

Three ways people are putting that to work:

The bookkeeper or small-business owner, reconciling something. Do not ask “is this right?” Ask it to do the job twice, in opposite directions: once from the source documents, once working backwards from the total. Compare the two yourself. Agreement is weak evidence; disagreement lands on the real error.

The nurse, teacher or caseworker writing up notes. Draft with it, then read the output hunting for things that sound plausible and are not in your source: a date, a name, a dosage, a quoted phrase. These systems fail by filling gaps smoothly rather than leaving them. Check the specifics, skim the prose.

The manager reading a summary of something long. Ask for the summary, then ask a second question: “what did you leave out that someone who disagreed with this would say?” Asking for the omissions surfaces them faster than re-reading the document.

One thing not to do: do not ask it to check its own work and count that as a check. Ask “are you sure?” and you will get a confident yes or a reflexive apology, and neither tells you whether the thing is correct. Get your second opinion somewhere the first answer cannot reach: a different tool, the source document, a colleague.

→ So what: The thread running through all of these: useful verification never asks the machine to grade itself. It builds a second, independent path to the same answer and looks at where the two disagree. That is learnable, and it is about to be a large part of a lot of jobs.

One question for you: what is the last thing an AI tool got confidently wrong for you, and how did you catch it? Hit reply and tell me. I read every one, and the good ones end up in a future issue.

Until next Tuesday,
Robert