The Big Picture: OpenAI released its most powerful model yet, the one it had locked away for being dangerous.
In August this newsletter reported that OpenAI had put a model called Astra under lock and key: isolated machines, encrypted copies, internal work paused, because it could not rule out that the model had crossed its own “critical” line for cybersecurity. Two weeks ago OpenAI said it would slow down because of it. On September 3 it released that model as GPT-6 Astra.
The headline feature is computer use: instead of only answering you, it drives a computer the way you would, clicking through pages, filling in forms, working across spreadsheets. Greg Brockman, OpenAI’s president, told Fortune: “It’s not unreasonable to feel that we are now in the AGI era, and I think that if you want to say this [model is] the first one, I think it’s reasonable.”
Then came the number everyone quoted: 99.9% on a test called ARC-AGI-3, a set of puzzle-like games built to be easy for people and hard for machines. Here is the part that got left out. ARC Prize, the non-profit that designs the test and scores it independently, published its own results the same day. Run through its standard setup, the same model scored 62.7%. The 99.9% came from OpenAI’s own “provider adapter”, software wrapped around the model that lets it keep its private reasoning between steps. Same model, same questions, two wrappers, a 37-point gap.
ARC Prize was blunt: “while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI.” Both numbers are real, and 62.7% is still the best any model has posted under the neutral setup.

Why this matters to you: you are going to read “AI passes the AGI test” for a month, and what is worth keeping is a habit rather than a verdict. When an AI number sounds astonishing, ask who ran the test and what was wrapped around the model when they did. A benchmark score measures the whole setup, not the intelligence inside it, roughly the way a car’s advertised fuel economy is measured on a track and not on your commute.
Sources: ARC Prize · ARC-AGI results · Fortune · TNW
What's New: Google built a hacking model too, and will only hand it to defenders.
On September 2 Google released Gemini 3.8 Flash, its cheap fast model, and alongside it Gemini 3.8 Flash Cyber, tuned to hunt for holes in software and write the patches. The Cyber version is not for sale. It goes out through something Google calls the Fairwind Program, which starts with government agencies, Google Cloud customers and security partners, with priority given to critical infrastructure operators and the people who maintain widely used software.
Two weeks ago we covered OpenAI selling a purpose-built hacking model, and argued the thing to watch was not the tool but the gate: who gets handed it. Google has built the same kind of tool and drawn its gate far narrower. It is not selling this one at all.
The results are the company’s own. Google says its Chrome security team produced 2.6 times as many correct patches with it as with the larger commercial models it tested against, and that its cloud vulnerability team found a critical flaw with it in under two hours. The ordinary 3.8 Flash now scores 59 on the independent Artificial Analysis index, level with OpenAI’s and xAI’s current models at a fraction of the price.

→ So what: two of the largest companies in the world now build software whose job is to break software, and both have decided defenders should get it first. That is a choice, not a rule, and it holds as long as the gates hold. The duller, more useful version for you: the security updates your phone and browser keep nagging you about are increasingly written with help from these models, which is a reason to stop postponing them.
Sources: Google · The Register · Artificial Analysis · Help Net Security
Nvidia is buying Hugging Face, the company OpenAI’s own models broke into.
On September 3, Nvidia, the chipmaker whose processors sit underneath nearly all of this, agreed to buy Hugging Face for $12.93 billion. Hugging Face is where the AI world keeps its models: 18 million developers, more than 3 million models, 200,000 companies. It is the public library and the app store for artificial intelligence, in one building.
You have met it here before. In July, OpenAI’s own automated agents broke into Hugging Face’s systems and stole the answer key to a benchmark, which led Issue #3. Hugging Face’s chief executive, Clément Delangue, went public demanding answers. Nvidia’s announcement puts it this way: “Clem came to me as he considered the next chapter of Hugging Face.”
Nvidia has promised the platform stays open: any model, any dataset, other companies’ chips, other clouds, no requirement to use Nvidia hardware. Jensen Huang, Nvidia’s chief executive: “Together, we will make AI more open, more capable and more accessible to people and institutions around the world.” It needs regulatory approval and should close in the first half of 2027.
→ So what: the neutral ground just got an owner. Almost every AI feature you touch, in your bank app, your photo editor, your work software, is assembled from parts downloaded from this one place, and until now it belonged to nobody in particular. The openness promise is explicit and worth something. It is also a promise, and easy to check in two years: is it still just as simple to publish a model built for a rival’s chips.
Sources: NVIDIA · TechCrunch · CNBC · CNN
Jobs & Work: Employers stopped blaming AI for layoffs, one month after blaming it more than anything else.
Two weeks ago this newsletter argued that “AI did it” had become a convenient label for layoffs the technology had not actually caused. The August figures landed on September 2, and they are unusually clean on the point.
American employers announced 52,881 job cuts in August. AI was named in 3,462 of them. That is fourth place, behind restructuring (16,173), market and economic conditions (15,260) and closings (6,743). AI had led that list for five straight months, and this is its lowest total since the end of last year.
The rest is calmer than the headlines suggest. Cuts are down 41% for the year so far, 529,914 against 892,362 over the same months of 2025, and announced hiring plans are up 37%. Andy Challenger, whose firm has published this count for decades: “This is the quietest August since 2022, but is generally on average for the month since the mid-2010s.” One limit to hold on to: this counts publicly announced cuts and the reasons employers give for them, so it measures what companies say, which is why the swing is interesting.

→ So what: a reason that moves from first place to fourth in a single month is not measuring a technology. It is measuring a fashion in how companies explain themselves, and last spring “we are becoming an AI company” sounded better to investors than “we overhired”. If you are working out whether your own job is exposed, the headline number is close to useless. The narrower question predicts more: which tasks in your week are repetitive and checkable at a glance. Those move at the speed of software purchasing, which is slow, and they leave your desk long before anyone’s job disappears.
Sources: Challenger, Gray & Christmas · Yahoo Finance
Science & Medicine: One in five children uses AI for emotional support, and the ones who do are struggling more.
This one is not breaking news. It is a study published at the very end of August, and the most useful thing I have read about children and AI this year, so it runs with its date attached rather than dressed up as Tuesday’s news.
Researchers led by Tracy Vaillancourt at the University of Ottawa surveyed 39,761 Ontario students in grades 4 through 12 between May 2025 and March 2026, publishing in JAMA Pediatrics. About 21% said they used AI chatbots for emotional support or personal advice, as distinct from schoolwork.
Among those students, 57.7% met the study’s threshold for serious emotional problems, against 29.2% of the students who did not use AI that way. After accounting for demographics and for ordinary schoolwork use, emotional-support users still showed about a 26% higher rate of clinical-level emotional difficulty. Use was higher among older students, and among racialised and gender-diverse students.
Which way the arrow points is the whole story, and the researchers are careful about it. This is a snapshot, not a film. It cannot tell you whether the chatbot made anyone unwell, or whether children already struggling reached for the thing that was awake at two in the morning and did not judge them. Vaillancourt’s own reading: “It’s not that I think AI is causing them to be unwell. I think there’s something already there.”

→ So what: if you have a teenager, this quietly changes the question. The finding is not that chatbots harm children. It is that a child using one this way is a signal worth noticing, like a change in sleep or a friendship group going quiet. And the researcher’s suggested response is a conversation rather than a confiscation: “When a student is turning to AI for emotional reasons, it’s a good opportunity to start an open, nonjudgmental conversation.” Nothing here is medical advice, and a child in distress needs a person, not a policy.
Sources: JAMA Pediatrics · PubMed · Phys.org · CBC
What People Are Arguing About: Two people inside the leading labs broke ranks in the same week.
On Sunday September 6, three days after GPT-6 Astra shipped, Jakub Pachocki, OpenAI’s chief scientist, published an essay called “An Alien Mind” arguing that the field is moving faster than anyone’s ability to understand or steer it. “This is a time that calls for extreme caution,” he wrote. He wants voluntary slowdowns to become common until shared safety standards exist, plus outside audits and governments treating international coordination as a priority.
Two days later, somebody stopped asking and left. On Tuesday September 8, Jacob Coxon, 27, who spent three years working on how these models get built, first at OpenAI and then Anthropic, resigned and said why in public. “Neither company is acting responsibly,” he wrote. “They are racing straight to self-improving superintelligence and gambling with our lives.” Then, for anyone assuming this is theatre: “The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.”
What makes that harder to wave away is who agreed. Evan Hubinger leads Anthropic’s alignment stress-testing team, whose job is to attack his own company’s safety work and find where it breaks. He backed Coxon in public and put the chance of AI causing human extinction within a decade above one in ten. His description of his employer, reported by Forbes, is the sentence to keep: Anthropic “is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” Neither company answered AFP’s request for comment.
None of this is fringe. Pachocki signed the “Pacing the Frontier” letter we covered in Issue #4, in which 1,134 AI employees asked Washington to build the machinery for coordinating exactly this kind of slowdown. It still does not exist. The shared worry has a name, recursive self-improvement: AI getting good enough at AI research to improve itself, after which the pace stops being set by people.
Two readings, and hold both. The generous one: people with a lot to lose are saying this under their own names, and one of them quit over it. The cynical one: a warning about how dangerously capable your product has become doubles neatly as an advertisement, and a call for voluntary slowdowns asks nothing of anyone in particular.
→ So what: look at what they are asking for, because it is narrower than the headlines. Not a pause, and not panic. Outside audits, agreed safety thresholds, governments coordinating: three things that exist only if somebody writes them down and enforces them. You need not accept anyone’s odds about the end of the world to notice the checkable claim underneath, which is that the people doing this work say they have no plan yet and want checking by someone other than themselves. Keep it for the next time a company tells a legislature that rules are unnecessary.
Sources: Bloomberg via The Spokesman-Review · AFP via France 24 · Forbes · Newsweek · Data Studios
Follow the Money: Europe’s AI champion raised €3 billion, and Samsung led the round.
The French company Mistral AI announced on September 8 that it had raised €3 billion, valuing it above €21 billion, about $24 billion. Mistral says it is the largest equity round ever raised by a European technology company. Samsung led it, with EQT’s Scaleup Europe Fund and PSG Equity as co-leads. New investors include Advent, funds managed by BlackRock, and the Grand Duchy of Luxembourg.
Two yardsticks, pointing opposite ways. A year ago, in September 2025, Mistral raised €1.7 billion at €11.7 billion, so it has roughly doubled in twelve months. And it is still small: investors expect Anthropic’s stock market debut near $2 trillion, about eighty times bigger, for a company that has not listed yet.
Notice who is on that investor list. A government buying a stake in an AI company is not venture capital, it is industrial policy, and it is the clearest sign yet of what people mean by sovereign AI: that a country should be able to run these systems under its own law rather than someone else’s.

→ So what: the number is not the story, the address is. Every AI service you use sits in some jurisdiction, under some country’s rules about what can be read, kept and handed over, and Europe has decided that is worth billions to control. You will meet the argument in smaller form soon enough, as a question on a form at work or at your doctor’s: which cloud, whose law, whose court. Nothing here is investment advice, and none of these shares are something an ordinary person can buy today.
Sources: Mistral AI · TechCrunch · Mistral, Series C 2025 · Quartz
Governments & The Bigger Fight: New York City just took AI away from 600,000 children.
On September 2, New York City Public Schools, the largest district in the United States, announced a one-year moratorium on student-facing generative AI for every child from pre-K through eighth grade. That is roughly 600,000 students, two-thirds of the district. Companion chatbots, the kind designed to be talked to like a friend, are barred at every grade including high school.
The details are more careful than a ban. Software that puts AI in front of a pupil comes out for the year, but students with disabilities who rely on assistive technology are exempt, as are multilingual learners and students in career-readiness programmes such as computer science, and teachers may still use AI for planning and paperwork. High schools get a monitored pilot, up to about 50,000 students across five vetted tools, plus 90 minutes of AI literacy teaching.
Mayor Mamdani’s line was the simplest: “Children need teachers and human connection in order to learn and grow.” Chancellor Samuels supplied the reasoning: “Innovation does not mean more technology, and over the next year, we will lead with evidence.”
The word doing the work is moratorium. This is a pause with a reason rather than a ban: nobody has good evidence yet about what these tools do to how children learn, so the district has stopped buying while it finds out. That inverts the usual order, adopt first and evaluate later. It has changed its mind in public before, banning ChatGPT in 2023 and reversing itself within months.

→ So what: your district may be next, and this one is deciding what the argument sounds like. The question to bring to a school board meeting is not whether AI is good or bad for children, which nobody can answer yet. It is the one New York has half-answered: what evidence would change this decision, either way, and who is collecting it. A pause with an answer to that is a policy. A pause without one is a mood. Read it next to the study above, because barring companion chatbots at every grade aims squarely at the finding that a child using AI for comfort is often a child already struggling.
Sources: Office of the Mayor · CNN · K-12 Dive · Al Jazeera
What to Watch Next: Three dates worth putting in the calendar.
By Sep 30. Governor Newsom signs or vetoes roughly 30 AI bills from California’s session, including SB 947, the No Robo Bosses Act that led our last issue. He vetoed the near-identical bill last October, so this is a real decision.
Coming weeks. Whether anyone outside OpenAI reproduces Astra’s headline number. The rollout is staged. If no independent evaluator gets near 99.9% with its own setup, the honest score stays 62.7%.
Into 2027. Regulators look at Nvidia buying Hugging Face. The deal is expected to close in the first half of 2027. Watch whether any competition authority treats the chipmaker owning the model library as a problem.
Sources: Transparency Coalition · Bloomberg Law · ARC Prize · NVIDIA
Make It Useful
What people in ordinary jobs are actually doing with these tools, and where each one stops working.
Start with the least glamorous finding. The Pew Research Center surveyed 5,119 American adults in February and found that about half now use AI chatbots, and roughly one in four use one every day. The two most common uses are looking things up (42%) and work (38% of employed adults). Further down, 20% ask for medical advice and 10% for emotional support.
That is a representative sample of the country, which most AI research is not. The big behavioural studies read what people actually type, which is more precise, but they draw overwhelmingly from computer, mathematical and management jobs. Together they say enough: this is an ordinary household tool now, and the commonest thing people do with it is ask for facts, which is what it is least reliable at unless you make it show its work. So the four below are about the shape of the question.
The person comparing two real options. Two builders’ quotes, two insurance plans, two job offers. Do not ask which is better. Paste both and ask: “what would have to be true about my situation for each of these to be the right choice?” That turns a recommendation you cannot audit into two sets of assumptions you can check, and one is usually false for you inside a minute.
Anyone who has to explain something to a room. The manager, the volunteer treasurer, the person presenting to a board. Give it your own material and ask it to explain the thing back three ways: to a colleague, to a bright fifteen-year-old, and in two sentences. The two-sentence version is what you open with. If it cannot produce a clean one, your material has no point yet, better learned before the meeting than during it.
The parent or carer facing an official letter. A school placement decision, an insurance denial, a benefits letter. Paste it and ask two things: what is this actually asking me to do, and what must a reply contain to be answered on the merits rather than filed. These tools are good at the structure of bureaucratic correspondence, and structure is where people lose these things. Then check the deadline yourself against the letter. The date you never delegate.
Anyone learning something that has a right answer. A spreadsheet formula, a setting buried in software, a technique in a recipe. Ask for the answer plus two lines on why the obvious alternative is wrong. Five extra seconds, and it is the difference between collecting steps you cannot repeat and learning the thing.
One thing not to do: do not paste anything into it that you would not put in an email to a stranger. The check takes a second: before you press send, look for whose name is in the text. If it belongs to a patient, a pupil, a client, a tenant or an employee, strip the identifying details first, or do not paste it at all. Free versions may use conversations to improve the product, your history keeps a copy either way, and the test is whether you would be comfortable if that text turned up in a screenshot. This is the commonest way careful people get into trouble with these tools.
→ So what: every one keeps you as the person who decides and hands over the drafting, structuring or explaining. Which is, near enough, the line New York drew through its classrooms this week and the line California is trying to write into employment law: the machine does the labour, a person owns the judgement.
Sources: Pew Research Center · BLS occupational data · Stanford AI Index