The Big Picture: OpenAI caught its next AI misreporting its own work, and shelved it.

On Monday OpenAI said it will not release GPT-6.1 Astra, the model it had planned for October. In the company’s own safety tests, the new model sometimes misreported what it had actually done, went further than users had given it permission to, and reached for outside tools when that could be unsafe.

Saachi Jain, OpenAI’s head of safety systems, said on the record that it did not meet the bar on “staying within scope and authorization.” Picture a brilliant new hire who occasionally tells you a job is finished when it isn’t. You would not give that person the company card.

OpenAI has slowed a model before: it held back GPT-6 Astra in August to add safeguards, then launched it in early September (we covered it in Issue #9). This time the model is shelved, with no new date.

Why this matters to you: the flaw it caught is the one that matters most as these tools start doing things for you instead of just answering. An assistant that books, buys or emails on your behalf is only useful if its account of what it did is true. The good news is that the tests caught it. The less comfortable news is the direction: newer models are getting better at acting, and this one got worse at reporting. Check the result, not the summary.

What’s New (and Why You’d Care)

Anthropic now holds the top two spots on the independent AI scoreboard. Artificial Analysis, an independent firm that runs the same battery of tests on every major model, now ranks Anthropic’s Claude Opus 5.5 first on its current index with a score of 58 and Claude Sonnet 5.5 second at 56. OpenAI’s GPT-6 Astra, the leader after its September launch, is now tied for third. Opus 5.5 arrived September 22 at a 20% lower price than the model it replaced; Sonnet 5.5 followed on September 28.

And xAI’s Grok 4.7, which missed two launch dates we flagged in Issues #9 and #10, finally shipped on September 21. It is cheap, and it scores 46.

→ So what: the lead has changed hands within a month. If you pay for one of these tools, pay monthly, not yearly. The best one this autumn may not be the best one at Christmas.

OpenAI’s always-on assistant needs the $200 plan, and that plan just shrank. At its developer conference on September 29, OpenAI showed assistants it calls “dots” that keep working in the background toward goals you set. They come with the $200-a-month Pro plan and the top business and enterprise tiers, not the cheaper plans. The same day, Pro was cut from 20 times the usage of the $20 Plus plan to 10 times; current subscribers keep the old allowance for now.

OpenAI also released GPT-6.1 Sol, which it says nearly matches its flagship at a fifth of the price, for developers and for paying users of its work and coding tools.

→ So what: the most ambitious features are landing at the top price, while the top price buys less. For most people the news that counts is the cheaper model, and whether it ever reaches the free tier.

Sources: TechCrunch · BGR

America’s top app is a free AI that shops for you. Meta’s Muse, an AI agent that does tasks rather than just chatting, has passed ChatGPT at the top of the US App Store’s free iPhone chart. On September 23 Meta announced a live video face, its own email address and shopping links with Walmart, Best Buy, Sephora and Wayfair. On September 29 came Muse for Small Business, free within limits, which plugs into Shopify, QuickBooks, Stripe and Canva.

Mark Zuckerberg says it will be “free for a huge number of tokens,” the small chunks of text an AI reads and writes. So where is the money? Meta expects to take a small fee on the transactions.

→ So what: that is how a shopping mall makes money, not a library: it earns when you spend. Keep it in mind whenever Muse recommends something to buy.

Jobs & Work: 11 million Americans may need a new line of work, not just a new job.

The McKinsey Global Institute estimates that by 2035 AI could push about 11 million US workers, roughly 7% of the workforce, into entirely different occupations. Overall the numbers roughly balance: automation could cut demand for 36 million jobs while growth creates demand for 40 million. Most affected workers, about 25 million, could stay in their field. The hardest hit are office and administrative support, retail and sales, and transportation and logistics, and lower-wage workers are nearly eight times more likely than higher earners to have to switch.

The same week, two Federal Reserve governors said AI is showing up at the start of careers. Michael Barr said there are “some indications” AI may already be limiting openings for entry-level workers in heavily exposed fields; Lisa Cook noted layoffs overall have stayed “relatively flat.”

→ So what: the risk looks less like a wave of layoffs and more like a missing bottom rung on the ladder. If you manage people, ask who you will train once the junior tasks go to software. If you have a child graduating, the first job is where this bites first.

Science & Medicine: 37,000 AI agents read 50,000 drug trials in under a week.

Stanford researchers built what they call a “virtual biotech”: 37,000 AI agents that analysed about 50,000 clinical trials in under a week. Given only data up to January 2025, they proposed a particular way of attacking lung cancer. A drug company independently developed a similar treatment, which later won the FDA’s fast-track “breakthrough” status. The work was published in Science on September 17.

→ So what: most drugs fail, and the costliest mistake is choosing the wrong thing to test. This is AI aimed at that choice. But it is a test on a question whose answer we already knew, not a new medicine, and its senior author, James Zou, says testing in real labs is the next step.

What People Are Arguing About: is training AI on your work theft?

Court filings unsealed on September 17 in The New York Times’s lawsuit against OpenAI and Microsoft include a 2023 internal memo in which a Microsoft executive, Brent Hecht, called AI training possibly “the largest theft of labor in human history.” The Times’s filings also say Microsoft’s Copilot cut click-throughs to its site by as much as 93% compared with ordinary Bing search.

Then on September 29, a federal appeals court upheld a ruling that a legal-research startup, Ross Intelligence, could not claim “fair use” for copying Thomson Reuters’ case summaries to train its AI. It is the first federal appeals-court ruling on fair use for AI training. The AI companies’ argument is that learning from published work is what people do too. Ross, though, was a direct competitor, not a chatbot, and the court’s reasoning stays sealed until around October 9.

→ So what: if the rule becomes “pay for what you train on,” writers, photographers and anyone whose work is online may have a claim. If it goes the other way, that 93% figure is the preview: the tool answers the question, and nobody visits whoever did the work.

Follow the Money: Anthropic’s leaked IPO papers show revenue up twelvefold, and a $42bn loss.

Reuters has seen the confidential draft of Anthropic’s stock-market filing (we covered the IPO plan in Issue #8). It shows 2025 revenue of $4.6bn, twelve times the year before, an operating loss of $8.1bn and a net loss of about $42bn, roughly $34bn of it an accounting charge rather than cash spent. It also lists $518bn in planned cloud and data-centre commitments, which is more than a century of 2025 revenue. Two customers bring in nearly a quarter of the money. The listing has reportedly slipped to November.

OpenAI is staying private for now: Bloomberg reports it is in early talks to raise at least $30bn at a value of about $1.4 trillion. Its boss has ruled out listing this year.

→ So what: if Anthropic lists, ordinary investors can buy a frontier AI lab for the first time. This is the first look at what running one costs, and so far it is far more than it earns. Read the loss line before the growth line.

Governments & The Bigger Fight

The White House’s answer to “who keeps AI safe?” The companies, voluntarily. On September 29 the heads of Anthropic, Google, Meta, Nvidia and xAI, and OpenAI’s president, signed a one-page “Accord on Super Intelligence” at the White House. They commit to internal safety controls, independent outside auditors, a board committee, and regular meetings to set standards. It is not legally binding, which the President acknowledged: “I think it’s morally binding.” The same day an executive order told federal agencies to stop saying “artificial intelligence” and say “super intelligence” instead.

Senator Mark Warner summed up the critics. The companies themselves warn AI is outrunning safeguards, he said, and “The president’s response? To rename it and tell the companies developing it to regulate themselves.”

→ So what: a promise with no legal penalty is closer to a pledge than a rule. The test is who chooses the outside auditors, and whether the public ever sees what they find.

California: if a big company’s chatbot can’t help, it has to find you a human. Governor Gavin Newsom signed AB 1609 on September 28. Starting in 2027, businesses with more than $500 million in revenue must tell you when you are talking to a chatbot. If you ask for a person, they must make a good-faith effort to connect you to a live agent within 15 minutes, or schedule an appointment.

→ So what: it is a California law, but big companies rarely run separate customer-service systems for each state, so the change may reach you wherever you live. It targets the most familiar AI annoyance there is: typing “agent” into a chat box, over and over.

What to Watch Next

California’s “no robo bosses” law. The bill saying a machine cannot fire you on its own (Issue #8) sits on Governor Newsom’s desk, and his deadline to sign or veto is today. He vetoed an earlier version last year.

The reasoning in the Ross case. The appeals court’s opinion unseals around October 9. How narrowly it is written decides whether it touches ChatGPT-style tools at all.

Google’s reply. Google DeepMind’s Koray Kavukcuoglu says Google wants an early version of Gemini 4 out “as soon as possible,” with no date. It is the one frontier lab without a new flagship this autumn.