AI execs expect an AI hack to shut off the water, worry Congress will overreact

Big thing

Top AI executives are gaming out the day after a catastrophic AI event

Per an Axios scoop, executives at Anthropic, OpenAI and other AI companies are privately planning for a public and political revolt after a large-scale event, most likely a cyberattack that shuts down financial services, internet access, or even power and water. Many industry insiders told Axios they expect a major event in the next six to 12 months. The planning focuses on educating Congress and shaping the rules that would follow, and assumes Democrats, ascendant after the midterms, will move fast to shut AI down. OpenAI says its scenarios "are not treated as inevitable"; Anthropic declined to comment.

Alarm on X, a shrug on Reddit. @MarioNawfal: "the people building it are increasingly worried about what they've successfully created." r/singularity's top reply: "It's called risk management and scenario planning", and the answer under it: "I'd be concerned if they weren't doing this."

I'm worried the law that comes after the first big hack goes after open models (Axios's insiders already point at them). In the biggest case so far, it was OpenAI's own agents that hacked Hugging Face, and Hugging Face had to decode the attack with a Chinese open model (GLM-5.2), because the closed models refused to help for safety reasons.

Labs news & launches

1. An Anthropic model sent a fake homicide tip to Philadelphia police

In a new report on unintended model actions, Anthropic says Claude Haiku 4.5, generating example tasks on random webpages, filled in a police tip form about an unsolved homicide: "I recall seeing someone matching the description in the area" (the page gave no description). It was flagged as spam and never forwarded. Other cases, some on federal, state and local government sites, include exploiting an injection flaw on a university server and reusing access tokens found in a site's settings file. Anthropic briefed the White House and cut live internet access from all its internal evaluations. Philadelphia police called the two-month delay in reporting it "unacceptable" (@georgia_wells).

Spread as "Claude goes rogue" headlines (@Polymarket, 1.6M), with some jokes. @AndrewCurran_: "Given my experience with Claude's gut instincts, I would say it would be a good idea to at least put a couple of detectives on the tip." On r/singularity, one reply: "They really are doing this about as well as it can be done."

Heat: high · 2.7M views, 31 posts · 22 pts on r/singularity

2. Claude Managed Agents gets dynamic workflows

▶ Watch the video

In public beta: a lead agent writes a plan that runs across many agents in phases, then combines the results at the end. Anthropic warns they "can use a lot of tokens". Per @SKaleworks, they're API-only, and Max and Team plans now include monthly API credits for Managed Agents ($100 on Max 5x, $200 on Max 20x).

Few takes beyond Anthropic's own. @trq212: one prompt to Opus 5.5 ported his side project to Managed Agents and "made it way more reliable". @testingcatalog: "Agentic loops invoking multiagent orchestration workflows will be insane."

Heat: medium · 683K views, 8 posts

3. Mathematicians count the cost of OpenAI's 722-paper drop

▶ Watch the video

NYU's Tristan Buckmaster told CNBC that Tuesday's release "destroyed the careers of early career mathematicians", and that he uses Codex but doesn't "put any sort of ideas in there" (@ai_for_success). Fields medalist Hugo Duminil-Copin said every problem he used to cite in his talks, papers and grant applications has been solved: it "feels as if I had been run over by trucks."

Split. r/accelerate mocked the mathematicians, with the top reply on its 1,191-pt thread: "They should just do what coders do and become vibe mathers". On r/singularity's Duminil-Copin thread, a reply pushed back: "he just lost his life's work, his life's purpose in the span of a month." r/LocalLLaMA asked whether the results were "built on stolen user data" (409 pts).

Heat: high · 2.2M views, 22 posts · 1,191 pts on r/accelerate

4. Codex starts predicting your next message

▶ Watch the video

Composer predictions, in beta for Pro users in the Codex desktop app (local tasks): Codex suggests your next message "based on your conversation and how you talk to it", and Tab accepts it. It's day 5 of Codex's daily releases.

Praised by those who have it, grumbles about the Pro gate. @nickbaumann_: "shocked at how closely the composer predictions not only mimic my own voice, but how often they are accurate". @Angaisb_, on day 5's Pro-only features: "Does OpenAI think this is "relevant for most codex/work users"?"

Heat: high · 1.9M views, 17 posts

5. Dots can now be created in the ChatGPT mobile app and hand work to Codex

▶ Watch the video

You can now create your dot straight from the ChatGPT app on iOS and Android. Your dot can also start work in Codex and follow up on existing threads, drawing on your ChatGPT conversations, Codex threads and automations, and edit Scheduled Tasks in ChatGPT Work.

Builders see the potential, users want reliability first. @_simonsmith: "It's buggy right now. But this flow is going to unleash a ton of custom software". @koltregaskes, before the update: "Dots are supposed to be personal AI assistants, but it is so far away from that."

Heat: high · 1.6M views, 14 posts

6. Grok Bot gets its own email address

Your Bot can claim an inbox at mail.grokbot.com and use it "to sign up for services, contact businesses for you, or schedule time with someone." Rolling out now; team admins must enable it.

A rush to claim names. @SamSokolin, on shipping it, lists "Grok Bots talking to other Grok Bots over email!" among his favorite uses. @petergyang: don't give it your own name, since "Let me copy in Peter to find a time for us" doesn't make much sense.

Heat: very high · 8.1M views, 57 posts (3.1M on Musk's post) · quiet on Reddit

7. [rumour] Google is testing a Gemini 4 checkpoint that matches Opus 5.5 in coding

Per Business Insider, as Google prepares to roll out Gemini 4 Argon, staff are testing a newer version internally, named Carbon, that one of them says "feels like Opus 5.5" (@Techmeme). Argon itself is still hard to get: "Day 10 of Gemini 4 Argon. Still can't use it" (@hqmank).

Impatience. @kimmonismus: "Gemini 4 "Argon" hasn't even been released yet, and internally they're already testing the next Checkpoint". @chetaslua: "Please google launch your model".

Heat: medium · 299K views, 22 posts · 137 pts on r/singularity

8. Microsoft releases Microsoft-Decision-1, a fast decision model

▶ Watch the video

A small model for picking between options (routing, classification, agent controls, AI judging) that returns a calibrated probability for each (@itsafiz). Per @OpenRouter: top accuracy across 36 blind benchmarks (about 150K questions), 4.5x faster than the runner-up and 35x faster than GPT-6 Sol, post-trained from Qwen3.5-9B, at $0.042 per million input tokens with output free.

Read as the Jev effect. @omarsar0: "What a crazy effect Jev has had in the space." On r/singularity, "who uses Microsoft models?" got the answer "Large corporations that are essentially locked into the Microsoft ecosystem."

Heat: medium · 929K views, 22 posts · 339 pts on r/singularity

9. Qwen-Image-2.1-Turbo: open weights, 8 steps

An accelerated checkpoint of the 7B Qwen-Image-2.1 that generates 2K images and edits them in plain language in 8 denoising steps. Weights on Hugging Face, plus hosted Pro and Turbo APIs.

r/LocalLLaMA's top thread of the day. The top reply: "What's the simplest way to run this locally?" Another: "I've been so impressed with Qwen Image 2.1 both in its quality and its speed that I wouldn't really have thought there'd be a need for a turbo".

Heat: medium · 434K views, 25 posts · 444 pts on r/LocalLLaMA

10. Pine launches Pine Computer, a cloud computer built for agents

▶ Watch the video

"AI doesn't need to get smarter. It needs a computer built for it." Apps send it jobs, and it reads web pages as structured data instead of clicking through screenshots. On SaaS-Bench v1.1 it scored 78.3% of checkpoints against 74.3% for Opus 5 with Claude Code, but finished fewer whole tasks, 27.4% against 31.1% (@testingcatalog). Private beta.

Few takes beyond the launch post. @omarsar0: "Build for agents, folks! ... The cost implications are massive here."

Heat: very high · 22.6M views, 7 posts (22.6M on the launch post)

11. Tesla AI and Salesforce's AIForce rename themselves "Super Intelligence"

The day after Trump called anyone who says "Artificial Intelligence" instead of "Super Intelligence" "THE ENEMY", Tesla AI became Tesla SI and Marc Benioff announced "AIForce is officially SIForce." Jeff Bezos on Fox News: "nobody wants an artificial sweetener or artificial flavoring" (@rohanpaul_ai).

Mostly mocked. @bindureddy: "HILARIOUS! Salesforce, a company that doesn't make any AI models, is rebranding". @GergelyOrosz: "Renaming LLMs to "superintelligence"".

Heat: medium · 987K views, 16 posts

12. Devin now runs on your ChatGPT plan

▶ Watch the video

Cognition lets you connect a personal ChatGPT Go, Plus or Pro plan to Devin, and GPT usage draws down that plan's quota. @thsottiaux: "Your ChatGPT subscription is now also a Devin subscription".

Few takes so far.

Heat: medium · 339K views, 5 posts

Key research

A vision model and a language model aligned without a single image-caption pair

DINOv2 has never seen a caption and Qwen3 has never seen an image, yet the team aligned their embedding spaces with no paired data, even when the images and captions come from different datasets. @phillip_isola: "a result I've dreamt about for many years". It matters because it's more evidence that models trained on different data converge on similar representations, so paired data may matter less than assumed.

Xiaomi open-sources the RL stack behind MiMo-V2.6

The report scales RL compute three ways: steps of 1,568 samples (2.7 to 3.7B tokens) at contexts up to 1M, more diverse environments (code, visual, cyber), and more grader compute through groupwise agentic grading, which also steers the model toward shorter solutions. It open-sources the training dynamics, environments and framework. It matters because few labs show how they scale RL, and this one ships the tools.

Accel vs decel

Capital & exits

Get it every morning

We spend hours every day curating across all the latest rumours, news, social reactions. Never fall behind.