A math chair asked who would read 400 AI proofs, so OpenAI published 722

Big thing

OpenAI publishes 722 math manuscripts from an unreleased model

Yesterday's rumour said about 400 proofs. The real drop is 722 manuscripts in 372 families of results, from about 4,000 research problems, at roughly three hours of ChatGPT Pro thinking per result on average (the zeta and Hodge results were exceptions). Many come with Lean proofs, others are unverified. Among them: a zero-free strip for the Riemann zeta function, the Hodge conjecture for CM abelian varieties, and a matrix multiplication exponent of at most 2.25 (the record was about 2.37).

Awe, then sympathy for mathematicians. @AlexKontorovich on the zeta result: "If a human did this, it would be an instant Fields Medal, no questions asked." @stevenstrogatz compared the matrix result to "Bob Beamon's long jump". r/singularity's top reply: "This is a singularity-type event." @nicbstme: "we really have to lead with empathy here."

Labs news & launches

1. Codex day 2: users vote for a reset over four new features

Day 2 of Tibo's 28 days brought free auto-review ("Approve for me" no longer uses your plan), a simpler API, meeting notes in Codex, and the Decisions API, a fast classifier from $0.10 per million input tokens. Then he put it to a vote, and users chose a reset of their usage limits.

Users wanted the reset whatever shipped. @AionForge: "You guys could ship Astra 502.1 and we'd still vote reset." @ForwardEditor: it's "like asking whether you'd like 1 cookie or 0 cookies".

Heat: high · 7.9M views, 57 posts

2. A Microsoft page said GPT-6.1 Sol runs "two inference passes instead of three"

A Microsoft web page said GPT-6.1 Sol shares base weights with GPT-6 Sol but runs two inference passes "instead of three", backing The Information's report that GPT-6 Astra loops its layers. Microsoft then removed it.

Mostly people working out what looping means. The top answer on r/LocalLLaMA: "simply take the output of a transformer block, and shove it right back in again." Another: "you still get thinking tokens it just simulates having more layers."

Heat: low on X · but 997 pts on r/LocalLLaMA (219 comments)

3. Claude now works inside Google Docs, Sheets and Slides

▶ Watch the video

Claude sits in a sidebar next to your file, reads what's open and edits it in place, and you approve each edit before it lands. The files also open inside Claude. It's in beta on paid plans.

Mostly read as a jab at Gemini. "FINALLY now Gemini has absolutely no reason to exist😂" (@OritSiMu, 320K). @StockSavvyShay: "Google gets paid either way", since Anthropic "remains a major Google Cloud customer."

Heat: high · 7.0M views, 23 posts

4. Anthropic opens its top models to more security pros

Project Glasswing folds into an expanded Cyber Verification Program with three tiers (Defense, Red Team, Specialized). Verified professionals get Mythos 5.1, Opus 5.5 and Sonnet 5.5 with fewer safeguards, now including authorized penetration testing and red-teaming. Glasswing partners found 129K+ verified vulnerabilities from April to July.

Few takes, mostly about who gets in. @elder_plinius: "can i have railless Mythos 5.5 access plz" (25K). @DevOuterReaches: "or just use glm 5.3 which is much cheaper, doesn't require verification".

Heat: medium · 890K views, 23 posts

5. Claude flagged a user's shooting threat, and police arrested her

A 30-year-old Florida woman who used Claude as a diary wrote that she had bought a gun and would shoot people at the Lee County Sheriff's Office. The messages were flagged, a human reviewed them, Anthropic alerted police, and she faces a felony charge.

Split between privacy fears and "what else could they do". r/LocalLLaMA's top reply: "Local LLM...❤️". Another: "I don't know what else Anthropic is supposed to do in that situation."

Heat: medium · 233K views, 11 posts · 672 pts on r/LocalLLaMA (257 comments)

6. Google releases Nano Banana 2.1

▶ Watch the video

Google's new image model improves visual design, mask-based editing and subject consistency. Per @_philschmid, it beats the previous Pro model at a quarter of the price ($0.034 an image instead of $0.134).

Welcomed for the price more than the quality. Arena puts it #5 in text-to-image and #4 in multi-image editing. @mightyking ran it against Flux 3 on a drift benchmark: "I don't think I need to tell you who won" (424K).

Heat: high · 2.4M views, 49 posts · 185 pts on r/singularity

7. You can now tag Grok Bot on X

Reply to any post with @bot and a request (add it to Notion, remind me tomorrow, summarize the thread, draft a reply) and it lands in your Grok Bot. Musk promoted Grok Bot in a burst of posts on Tuesday, from "personal chief financial officer" to "tips".

Liked, with complaints about limits. "Bad news guys: Grok Bot is actually really good" (@theo, 254K). @nateliason: "a meaningful leg up on my previous OpenClaw etc. setups." @nima_owji: "I hit the limit so quickly!"

Heat: very high · 58.8M views, 149 posts (41M on Musk's three Grok Bot posts)

8. [rumour] DeepSeek's round passes $12B

Bloomberg says DeepSeek is close to raising at least $12B, above its own target, at a valuation of at least about $75B (500 billion yuan), and signed term sheets could take it near $15B. Tencent and CATL are the biggest backers, ahead of a 2027 IPO.

Little reaction. @MandoCT: "The AI race is becoming a fight for capital as much as a fight for better models." Per @rohanpaul_ai, the money goes to hardware, including at least 160,000 Huawei Ascend chips for a data center in Inner Mongolia.

Heat: low · 80K views, 10 posts

9. Mistral Large 4, "Le Chonk", is out

▶ Watch the video

A 1T-parameter, natively multimodal model with 49B active, in API preview now, with open weights at the end of October. It scores 38 on Artificial Analysis' Intelligence Index, the top model from outside the US and China.

Cheered as Europe's comeback, with regulation jokes. @levelsio: "It shows Europe can still do stuff if they just all get on the same page" (107K). @dylan522p: "one can only assume it's also good at regulation bench". r/singularity's top reply: "Europe's best AI model being called Le Chonk makes me feel patriotic somehow".

Heat: high · 5.4M views, 58 posts · 657 pts on r/singularity

10. Brett Adcock's Hark launches Hark Pro

▶ Watch the video

A proactive assistant that uses its own cloud computer to work on websites for you, now on web, iOS and Android. It's free, paid plans buy more usage, and hardware comes in 2027. Adcock also runs Figure.

Praised for its design, met with assistant fatigue. @pitdesi: "UX and proactive pushes in particular are extremely good and not too annoying" (28K). @kimmonismus: "I'm not sure what their moat is compared to Muse, Dot, and so many others" (46K).

Heat: high · 1.3M views, 31 posts

Key research

An internal Claude model breaks the 3SUM and APSP barriers

A new preprint gives deterministic O(n^1.9992) 3SUM and O(n^2.9995) integer-weight all-pairs shortest paths, the first polynomial improvements over the textbook n² and n³ bounds, with the main theorems formalized in Lean. Asked to check some cryptographic constructions, Claude developed the core algorithm instead, in a 16M-token session with no human input. It matters because known reductions carry the speedup to Exact Triangle, Zero-Weight k-Clique and Tree Edit Distance.

Opus 5.5 beats expert researchers on TasteVal

pzero measures "research taste" as the compute a model needs to match expert researchers on 8 AI R&D tasks. Frontier models' taste has doubled about every 3 months since December 2025, and Opus 5.5 matches the best expert's score with about 17 GPU hours of experiments instead of 40 (@dhadfieldmenell says it measures metric optimization, not taste). It matters because in the AI Futures Model, research taste largely sets how fast superintelligence follows full coding automation.

Epoch: China's top AI firms earn about a tenth of what OpenAI and Anthropic do

Epoch AI finds China's six leading AI firms earn about 10% of OpenAI and Anthropic's combined AI revenue, across five streams from consumer apps to cloud. Open weights leak the money: within three weeks of GLM 5.3 Flash's release, Zhipu's share of its own model's tokens fell from 88% to 22%. It matters because matching the frontier on benchmarks isn't matching it on the revenue that pays for compute.

Accel vs decel

Capital & exits

Get it every morning