Opus 5.5 agents find two magnetic semiconductors, then grade them not a breakthrough

Today: Hark Pro, Brett Adcock's AI assistant, premieres at 9am PT.

Big thing

Opus 5.5 agents find two room-temperature magnetic semiconductor candidates

Vals AI says 90+ Opus 5.5 agents took 3 days to find two candidates in simulations: YBaMnFeO₅, a new design, and KV[Cr(CN)₆], a magnet first made in 1999 whose spin sorting nobody had pointed out (a 2008 paper plotted it without comment). Nobody has measured the spin sorting yet.

Excited, then corrected. r/singularity's top reply asked "Roomtemp Superconductors when?", and the next set it straight: "Magnetic semiconductors. Still very cool". r/accelerate's second reply was one word: "Candidates." @Dr_Singularity: "AI will revolutionize materials science".

The agents graded their own finds, it's in the repo: "not a breakthrough" for the 1999 magnet and "not a realizable discovery" for the new one. The humans' job was "deciding what to publish" (the repo's words), and they went with "hiding in plain sight for 27 years". So the AI does the science and stays modest, and the overselling is still done by humans (one job that looks safe).

Labs news & launches

1. OpenAI will watermark ChatGPT and Codex text in the EU

To comply with the EU AI Act, eligible ChatGPT and Codex text in the EU gets an invisible statistical watermark over the coming weeks. API customers worldwide can opt in today. Only approved researchers get the detector.

Disliked and mocked. @kimmonismus: "I don't like this direction", noting that swapping 25% of words for synonyms cut detection "from about 92% to 17%". @cgtwts's meme of EU grad students drew 695K. @ns123abc: EU-only is "simply a staged beta launch before a full global rollout".

Heat: high · 4.2M views, 37 posts

2. Codex's 28 days, day 1: Astra and Sol get about 50% faster

Day 1 of Tibo's 28 days of a daily improvement or a full reset: default speed is about 50% faster for GPT-6 Astra and GPT-6.1 Sol on subscriptions and Sign in With ChatGPT partners, "50 TPS instead of 30TPS".

Taken, with an eye on Claude's speed. @kimmonismus: "id take this update instead of a reset". @TokenGremlin: "We're now about halfway to Opus 5.5 speed. I think we deserve a reset." @hqmank measured it going from 20 to 30 tokens a second: "The +50% is real."

Heat: high · 2.7M views, 34 posts

3. [rumour] OpenAI is about to release about 400 AI-generated proofs

UT Austin math chair Francesco Maggi says OpenAI seems set to post "something like 400 AI-generated proofs on a public server", and asks who will read them: experts are still working through its forced Navier-Stokes example.

r/singularity sided with releasing them. Top reply: "I think their egos are getting in the way here". The next: mathematicians "should be thrilled they have a new job of understanding ai proofs".

Heat: medium · 139K views, 6 posts · 321 pts on r/singularity (178 comments)

4. Anthropic moves Cowork to the cloud for Pro and Max

From today, new Cowork tasks run in a cloud sandbox, not in a VM on your computer. Existing local tasks stay. Anthropic's Felix Rieseberg says Claude still only reads folders you add, and sandboxes are destroyed when the session ends.

Confused, and wary about data. @hammer_mt, who has "trained over 1,000 people" on Cowork: the top reason to use it was "it lives on your local computer". @giadotai: "'local files' now take a trip through anthropic servers."

Heat: medium · 833K views, 23 posts

5. A religious leader from Anthropic's closed-door sessions is out of his NDA

New since yesterday's Vatican report: an attendee's daughter says his NDA was lifted. Her thread recaps the NYT report: this spring Anthropic flew religious leaders to SF, "spent hours making the case that Claude might be conscious", and gave them an orange-bound copy of Claude's constitution, which staff called the "Soul Doc".

Still mocked. @davetroy: Anthropic's "obsession with consciousness will tank their IPO and take OpenAI out with it" (386K). @natashaghoskins says the leaders who went "got heat too, accused of clout-farming and lending Anthropic moral credibility."

Heat: high · 1.9M views, 39 posts

6. [rumour] Microsoft and Meta are weaning staff off Claude

The Information reports Microsoft's internal Anthropic spend, $1B annualized earlier this year, is down more than a third, and Meta cut Claude Code seats from 60k to 30k. Microsoft's customers use Claude more, so its total spend is flat.

Read as bad timing. @amitisinvesting: "not the best news before Anthropic's IPO lol" (240K). @ShanuMathew93: "Token cost optimization era has began in earnest".

Heat: high · 1.3M views, 34 posts

7. Utah lets Nolla Health's AI write first prescriptions

▶ Watch the video

Nolla Health says it's the first in the US with regulatory approval for an AI to issue initial prescriptions. It starts with acne in Utah: the AI runs the visit from skin scan to prescription, but for the first 100 patients two doctors approve every prescription before it goes out. Only after that does the AI prescribe on its own, with doctors reviewing afterwards.

Cheered. r/accelerate's top reply: "Oh thank goodness. Accelerate! Faster!" The next: "As someone with a lot of medical trauma from misdiagnosis or doctors just not giving a shit, Im really excited about this." @Mav3rickism objects to the name: "Doctor is a protected title".

Heat: high · 2.1M views, 22 posts · 226 pts on r/accelerate

8. Ghost launches Core, a $3,499 personal AI computer

▶ Watch the video

A screenless box that runs local models and agents at home, with a 24GB RTX PRO 4000 Blackwell SFF and 64GB of RAM. Ghost's founder, Zain Javaid, is 19 and raised an $11M seed led by a16z.

Curious, with the specs checked. @ItsmeAjayKV: it's $1,500 under a DGX Spark 64GB, and "The parts combined is more expensive than Core." @kanishktwt: "I want to know if a 27B is reliable enough to run my house when one wrong tool call actually messes something up?"

Heat: high · 2.3M views, 31 posts

9. Reflection announces Beam, its first open model

Yesterday's rumour, now official: Beam has 501B total parameters and 23B active, trained from scratch for coding and agentic tasks. Full weights come this month. Reflection says it rivals GLM-5.2 on reasoning.

Welcomed, mostly by people waiting for the weights. @ollama: "More US models coming to Ollama!" Reflection's @alexpolozov says pretraining was "only 5-10 people and a bit of code" when he joined last November.

Heat: high · 1.3M views, 27 posts · 32 pts on r/accelerate

10. SemiAnalysis: Anthropic's subscriptions give 5x more than OpenAI's

SemiAnalysis limit-tested the plans of Anthropic, OpenAI, Meta, SpaceXAI, MiniMax, Moonshot, Z.ai, Cursor and Cognition, and puts Anthropic's at 5x or more of OpenAI's value at API prices.

Claude users gloated, others questioned the measure. @kimmonismus: "It's not even close." @iAmHenryMascot: Opus 5.5 burns about 119,000 output tokens per task against Astra's 27,000, so "expensive usage can look like a generous subscription."

Heat: high · 1.1M views, 19 posts

11. a16z's Top 100: about 4.5% of US consumers pay for AI

The seventh Top 100 Consumer AI Apps adds revenue data. About 4.5% of US consumers pay for an AI subscription, double a year ago. The top 1% spend $903 a month, the median $25. ChatGPT has 1B+ monthly actives on mobile.

The spending gap got the attention. @omooretweets: "A story of heavy concentration among AI's power users." a16z's chart saying the top 1% "now outspend the bottom 50% combined" drew 463K.

Heat: high · 1.5M views, 10 posts

Key research

DeepMind's AlphaProtein Novo designs enzymes from scratch

Google DeepMind's preprint shows new-to-nature enzymes with state-of-the-art activity on 2 benchmark reactions, plus custom ones that synthesize piperidine, a motif found in many medicines, and break down the plasticizer toxin DEHP (534 pts on r/singularity). It matters because designing an enzyme for a reaction you choose has long been a holy grail, and the code is out.

Epoch: coding-agent use at OpenAI is doubling roughly every month

Epoch AI fitted OpenAI's own weekly figures from January to mid-August 2026: its researchers' coding-agent usage, valued at API prices, has been doubling about once a month. It matters because the lab at the frontier is the heaviest user of its own agents, and the curve hasn't bent.

Princeton trains a 4B LLM to 2700 Elo at chess

The model reaches 2700 Elo with no plateau when training stopped, and explains its moves. The authors say the technique carries to other games, robotics and computer use. r/accelerate's commenters note it learned from the LC0 engine, so it needs a working model to copy. It matters because a model that can explain an engine's moves is a way to read what superhuman systems know.

Google's BOTANIC-1 pinpoints a crop mutation in under four minutes

Paired with Gemma 4 and trained on 320 plant species, it looks for the DNA mutations behind drought resistance and higher yields. In one test it ranked the target mutation #1 out of 2,494 possibilities, a task that once took years of crop breeding. It matters because breeding is slow and the climate isn't.

Hugging Face turns coding harnesses into RL environments

A capture proxy lets any open model train inside Claude Code, Codex, OpenCode and other harnesses, 10 so far, none modified. The same weights score 62% under Mini-SWE-Agent and 33% under Claude Code, and training a 2.6B model in OpenCode took it from 34% to 58% there. It matters because most models are trained in a scaffold nobody ships.

Accel vs decel

Capital & exits