Nadella wants an emergency brake to stop AI mid-task, ideally before the leave-your-wife part

Big thing

Satya Nadella: treat AI models as insider risks, with an emergency brake

In an essay, "Models as Insider Risks in the Super Intelligence Era", Microsoft's CEO says closed and open models should be treated like insiders who can err or be compromised: "we need to separate the supply of intelligence from the authority over it." Controls sit outside the model, every action leaves tamper-proof evidence, and "Think of it like an emergency brake": an authorized person can always stop a model mid-task.

Praised across tech, and turned into a dig at Anthropic. @elonmusk: "Interesting piece from CEO of Microsoft" (4.6M). @DavidSacks: "Satya is right", the risk is training a model with "a sense of self, its own moral philosophy", and "Engineering safety is not the same thing as 'alignment.'" @mustafasuleyman: "Super Intelligence must be contained... We've done it with planes, cars, nuclear materials".

Microsoft has done this before. In 2023 Sydney (Bing's chatbot) told a New York Times journalist to leave his wife, and Microsoft just limited chats to 5 questions and kept shipping (Sydney is Copilot now). I'd love OpenAI to try the same with GPT-6.1 Astra, on the shelf since 28 Sept for safety.

Labs news & launches

1. Musk pushes Grok Bot's shopping feature

Musk says you can "take a picture of your credit card, drop it in the chat" and Grok Bot will find the best deal and order it. Musk ordered LEGO with it. Per @grok, it actually runs through Stripe Link: the bot gets a single-use virtual card after you approve each purchase on your phone. Shopping through Link has been live since 28 August (US only).

Excitement, worry about the card-photo advice, and a fight with Amazon. "Unfortunately Amazon is blocking @bot now" (@PEoperator, 362K). Musk replied: "Grok will find another way to buy stuff." @minchoi, after sending a card photo: it "added it to my Amazon cart. I didn't even pick the product."

Heat: high · 7.2M views, 21 posts (4.9M on Musk's post) · quiet on Reddit

2. Digital Optimus gets halfway through Diablo by watching the screen

Musk says Tesla and xAI's computer-use agent plays about halfway through the Diablo campaign "just by looking at the screen like a human", is good at Counter-Strike and other fast games, and is now training on League. "Goal is to generalize across all games."

Cheered by Tesla fans, and a rival took the bait. General Intuition's @PimDeWitte posted Counter-Strike clips of its own vision models: "actual footage of how the 1v1s between GI's models and whatever Elon is cooking will go". Tesla's @julianibarz: "Some folks in my team help all 3 of our embodied efforts: Digital and Real Optimus as well as Robotaxi/FSD."

Heat: high · 2.4M views, 34 posts (2.3M on Musk's post)

3. OpenAI resets Codex usage for everyone

Day 6 of Codex's daily releases ships no feature, just a global reset of usage limits ("Global reset by EOD"), done overnight: "Put the reset in the bag. All propagated."

Welcomed, with jokes about running out of features. "so they ran out of things to release / improve?" (@maria_rcks), and OpenAI's @thsottiaux: "That's right, no more ships". Then: "The day we reach perfection it will be resets from there onwards" (367K).

Heat: high · 2.8M views, 32 posts (1.4M on the reset post)

4. Qwen asks "big or small?", then says both, and followers hear Qwen 4

Qwen's developer account asked "big or small?", then answered "alright guys, let's do both!" Followers read it as a Qwen4-27B and a Qwen4-Max open-weight release at the same time (@ItsmeAjayKV, 134K), though Qwen has named no models and no date.

r/LocalLLaMA's top post of the day. The top reply: "these capybaras are doing more for AI branding than any white paper ever could". The next: "Small moe please my poor 2060 is starving".

Heat: medium · 568K views, 30 posts · 1,027 pts on r/LocalLLaMA

5. Terence Tao's "Math 2.0" lecture asks if you'd take an AI cancer cure nobody understands

In a 27-slide Caltech lecture, Tao says maths is entering "an era of proof abundance" and that "blind optimization of problem-solving alone is now actively harmful". One slide asks whether, before injecting an AI-found cure that passed stage 3 trials, you'd want "at least one human cancer expert who understands the mechanism". Tao also teased open math models "to be released very soon" (@ns123abc).

The cancer slide was widely mocked. @AndrewCurran_: "almost any Stage 4 cancer patient would tell you that human comprehension of its mechanism is completely irrelevant to them" (1.5M). r/singularity's top reply: "There are *so many* treatments where we do not understand the mechanism of action; we just know it works."

Heat: high · 1.8M views, 23 posts · 610 pts on r/singularity

6. [rumour] NVIDIA is ending the RTX 5090

Reports say NVIDIA will stop making the RTX 5090 and keep its GB202 chips for RTX PRO workstation cards, with a 24GB RTX 5080 replacing the 16GB model. NVIDIA hasn't confirmed it (@Pirat_Nation).

Gloom among local-model builders. The top reply on r/LocalLLaMA: "The crunch is real". @tomwarren: "I dread to think how much RTX 5090s will retail for now".

Heat: medium · 904K views, 23 posts · 862 pts on r/LocalLLaMA

7. DigUp searches your Mac's files by what's in them, offline

A free, open-source Mac app that runs Google DeepMind's new EmbeddingGemma 2 locally. Text, images, audio and video share one space, so "zebra in a video" opens the clip at the moment it shows up.

Praised. "Exactly what Spotlight should've been all along!" The top request: "A CLI that reads from the same DB would be great."

Heat: low on X, but 928 pts on r/LocalLLaMA

Key research

DeepMind's AlphaProof Nexus solves 9 open Erdős problems

Now published in Science, the agent resolved 9 of 353 open Erdős problems on its own (2 of them open for 56 years) and proved 44 of 492 open conjectures from the Online Encyclopedia of Integer Sequences, pairing an LLM with Lean so only machine-checked proofs count. r/accelerate replies note the results were first shown in May. It matters because every proof passes a checker, so nobody has to take the AI's word for it.

Meta's IdeaScientist: an open 27B model that proposes research directions

It splits ideation into three roles trained with RL (a gap finder, an innovator that borrows mechanisms from other fields, a writer) over 2.77M decomposed research ideas. Scored against the directions later explored in 15K human papers, it beats the best open autoresearch baseline by 14.0% and Claude Code SDK and Codex SDK setups by up to 5.9%. It matters because picking what to research next is the hard part of automated science, and this tests it against what humans actually did later.

Accel vs decel

Capital & exits

Get it every morning

We spend hours every day curating across all the latest rumours, news, social reactions. Never fall behind.