Monday, September 7, 2026HotTea archive editionVerified 2:39 AM PDT

The lead

Claude wrote 13 million Lean lines to check Fermat's Last Theorem

Anthropic says Claude agents formalized an existing proof in 11 days. The code is huge, but two computer checkers accepted it.

Listen to this edition

Prefer audio? The briefing has chapters and a full transcript.

The briefing

The rest of the morning

4 more stories

02

Astra scored 99.9 percent with OpenAI's hidden state and 62.7 percent without it

HotTea covered GPT-6 Astra's release. Since then, outside evaluators have tested the model under different conditions. ARC Prize measured 62.7 percent on ARC-AGI-3 with its standard test setup. That setup lets the model carry visible notes between steps. A provider adapter preserved OpenAI's opaque reasoning state and used compaction. That run reached 99.9 percent. ARC Prize called both results state of the art. Artificial Analysis found a different split. Astra scored 61 on its general Intelligence Index, the same as GPT-5.6 Sol. Its Coding Agent Index rose to 67, up from 65 for Sol.

The 37.2-point ARC gap does not point to one cause by itself. The runs used different context systems, and their best scores came at different reasoning levels. The benchmark also uses bounded game environments, not open-ended work. Artificial Analysis says Astra's token price is 2.5 times Sol's. Astra used fewer tokens, but still cost 75 percent more per task on the general index.

OpenAI ↗ARC Prize ↗Artificial Analysis ↗
Read article →
03

Sanders and Casar put up to 20 years in prison inside an AI pause plan

U.S. Senator Bernie Sanders and Representative Greg Casar announced forthcoming federal legislation that would ban artificial superintelligence. Their summary would also pause advanced AI development until a new cabinet-level regulator sets safety rules and reviews models. The agency could supervise the removal of dangerous capabilities and the destruction of prohibited systems. People who evade the restrictions could face up to 20 years in prison. Companies could lose the right to operate under what the sponsors call a corporate death penalty.

The lawmakers released a policy summary, not full bill text with a bill number and committee referral. Their definition covers systems that match or exceed human performance across many tasks. It also covers systems that could undermine governments or defeat shutdown commands. Those tests leave major questions about measurement, enforcement and ordinary model research. The sponsors are advocates for the proposal, while Fox News framed it through industry and national-security opposition.

Office of Senator Bernie Sanders ↗Office of Senator Bernie Sanders ↗Fox News ↗
Read article →
04

Workers who used AI without clear gains feared job loss most

A Boston Fed research team compared two national surveys. Each surveyed about 1,300 U.S. household heads, in December 2024 and December 2025. The share worried that AI would cost them their own job rose from 5 percent to just over 10 percent. Sixty percent expected AI-related layoffs or fewer workers in their industry. The most worried group used AI for some tasks but reported little or no productivity gain. The strongest gains came from 6 percent of respondents. They felt more secure and had a 14 percent chance of saying AI made them more likely to ask for a raise.

The study measures worker perceptions, not observed layoffs, output or pay. It compares two survey waves taken before publication and cannot show that AI caused the changes. The authors also state that their views do not represent the Federal Reserve Bank of Boston or the Federal Reserve System.

Federal Reserve Bank of Boston ↗
Read article →
05

Texas's 474-gigawatt data-center queue started a hunt for ghost demand

After Texas ordered an audit of data-center grid requests, Reuters reviewed large-load queues across the central United States. Requests exceeded 700 gigawatts, about as much electricity as every U.S. home uses. Texas accounted for more than 474 gigawatts, and ten large utilities elsewhere reported about 270 gigawatts. The numbers are not firm demand. Exelon cut its high-probability data-center estimate by about 40 percent after stricter screening. AEP Ohio's pipeline fell by more than half after state rules added fees and other requirements.

Utilities do not count requests the same way. Some include early inquiries while others require contracts, deposits or permits. Adding the queues together can overstate likely electricity use. Removing weak projects can also hide the remaining problem. Reuters found that confirmed demand still exceeds available generation and grid capacity in several regions.

Reuters ↗Office of the Texas Governor ↗
Read article →

Analysis

Machine checking changes the cost of proof, but people still have to judge the result

Claude's Fermat project shows that a model can produce more formal detail than people could review line by line. That makes machine checks more valuable, and it makes human judgment harder to avoid.

1

Scale can close gaps that people leave implicit

Lean requires every logical step. Claude could keep generating definitions and intermediate theorems until the final statement built without unfinished proof markers.

2

A green build checks a statement and its dependencies

The kernel verifies the encoded chain. It does not decide whether the statement captures the intended mathematics or whether the route is useful to a human reader.

3

Review moves toward selection and structure

Researchers will spend less time filling routine formal gaps. They will spend more time choosing claims, auditing assumptions and explaining why a checked result matters.

The watchlist

Signals that could change the read

Lean contributors reproduce the full build, inspect the proof path and report whether the public dependency chain holds.
Provider-neutral tests publish cost per successful task with the same tools, context limits and stopping rules.
Sanders and Casar file full text with a bill number, measurable thresholds, cosponsors and a committee referral.
Texas reports how much of the 474-gigawatt request queue survives ownership, deposit, permit, power and water checks.
Across the desks Show the proof

Labs, lawmakers and power planners made large claims. The useful evidence came from code, test conditions, bill language and project deposits.

Fermat formalization13M Lean linesAnthropic's public repository says the proof built 60,475 modules and a second Lean kernel accepted 1,052,234 declarations.
Astra test split99.9% versus 62.7%ARC Prize measured 99.9 percent when a provider adapter preserved OpenAI's opaque reasoning state. Its standard test setup measured 62.7 percent.
AI ban penaltyUp to 20 yearsSanders and Casar's summary proposes prison terms of up to 20 years for people who evade the planned restrictions.
Texas grid queue474 GWTexas says about 90 percent of its large new power requests come from data centers, but many requests are not firm projects.

Editorial direction, not a financial index. Each signal is tied to this edition’s reporting.

Edition validated · 5 stories · 11 unique sources

About HotTea & our sources →