Sections 00:00 What we're covering today 00:24 1. Claude wrote 13 million Lean lines to check Fermat's Last Theorem 02:20 2. Astra scored 99.9 percent with OpenAI's hidden state and 62.7 percent without it 03:40 3. Sanders and Casar put up to 20 years in prison inside an AI pause plan 04:58 4. Workers who used AI without clear gains feared job loss most 06:10 5. Texas's 474-gigawatt data-center queue started a hunt for ghost demand 07:23 Visit Hot Tea Disclosure Narration uses an AI-generated voice. Transcript Welcome to Hot Tea for Monday, September 7, 2026. Claude wrote 13 million Lean lines to check Fermat's Last Theorem. Anthropic says its agents turned an existing proof into code that two Lean checkers accepted. We also have Astra's split test scores, a proposed federal A I ban, worker anxiety, and data center power requests. Claude wrote 13 million Lean lines to check Fermat's Last Theorem. Anthropic published what it calls the first complete computer-checked formalization of Fermat's Last Theorem. Claude agents spent 11 days translating an existing mathematical proof into Lean code. The project produced 13 million lines and proved 30,300 intermediate theorems. Anthropic says 29,500 of those theorems were used in the final proof. It also says the run consumed about six billion output tokens. The agents did not discover a new proof. They followed a simplified version of the work behind Andrew Wiles's proof. Dozens of agents used a collaboration system called Prove2Me. A human researcher gave occasional high-level directions. Anthropic says early attempts lost track of the project. That failed work still accounts for about 7 percent of the non-boilerplate lines. The code is public. Its default build checks that the final theorem depends on Lean's three standard axioms. It also checks that the proof contains no unfinished proof markers. The repository says Lean built 60,475 modules. It also says nanoda, a second Lean kernel written in Rust, accepted an export containing 1,052,234 declarations with no errors. There is still a caveat. Anthropic produced the result and owns the public repository. The speed, token count and autonomy figures remain company claims. The formalization checks an existing proof. It does not replace Wiles's mathematical discovery or give a short human explanation. Public code makes outside review possible, but publication is not the same as completed independent review. Independent teams can now clone the repository, rebuild it from scratch and inspect the proof path. The first test is whether Lean and formal-math contributors reproduce Anthropic's checks and accept the dependency chain. The next test is cost and reuse. Watch for signs that the same system can formalize another major theorem with fewer lines, fewer tokens and less human steering. That would show whether this was a one-off feat of scale or a repeatable research tool. Astra scored 99.9 percent with OpenAI's hidden state and 62.7 percent without it. HotTea covered GPT-6 Astra's release. Since then, outside evaluators have tested the model under different conditions. ARC Prize measured 62.7 percent on ARC-AGI-3 with its standard test setup. That setup lets the model carry visible notes between steps. A provider adapter preserved OpenAI's opaque reasoning state and used compaction. That run reached 99.9 percent. ARC Prize called both results state of the art. Artificial Analysis found a different split. Astra scored 61 on its general Intelligence Index, the same as GPT-5.6 Sol. Its Coding Agent Index rose to 67, up from 65 for Sol. The 37.2-point ARC gap does not explain itself. The runs used different context systems. Their best scores also came at different reasoning levels. The benchmark uses bounded game environments, not open-ended work. Artificial Analysis says Astra's token price is 2.5 times Sol's. Astra used fewer tokens, but still cost 75 percent more per task on the general index. Teams should compare Astra and Sol on full jobs with the same tools, context rules and stopping conditions. Cost per successful task matters more than cost per token. More provider-neutral tests should show which gains come from the model and which depend on OpenAI's context system. Sanders and Casar put up to 20 years in prison inside an A I pause plan. U.S. Senator Bernie Sanders and Representative Greg Casar announced forthcoming federal legislation that would ban artificial superintelligence. Their summary would also pause advanced A I development until a new cabinet-level regulator sets safety rules and reviews models. The agency could supervise the removal of dangerous capabilities and the destruction of prohibited systems. People who evade the restrictions could face up to 20 years in prison. Companies could lose the right to operate under what the sponsors call a corporate death penalty. The lawmakers released a policy summary, not full bill text with a bill number and committee referral. Their definition covers systems that match or exceed human performance across many tasks. It also covers systems that could undermine governments or defeat shutdown commands. Those tests leave major questions about measurement, enforcement and ordinary model research. The sponsors are advocates for the proposal. Fox News framed it through industry and national-security opposition. The proposal becomes legislation only when the sponsors file full text. The useful details are the definitions of advanced A I and superintelligence, the powers given to the new agency, cosponsors and committee action. Those details will show whether this is a real regulatory plan or a position statement. Workers who used A I without clear gains feared job loss most. A Boston Fed research team compared two national surveys. Each surveyed about 1,300 U.S. household heads, one in December 2024 and one in December 2025. The share worried that A I would cost them their own job rose from 5 percent to just over 10 percent. Sixty percent expected A I related layoffs or fewer workers in their industry. The most worried group used A I for some tasks but reported little or no productivity gain. The strongest gains came from 6 percent of respondents. They felt more secure. They also had a 14 percent chance of saying A I made them more likely to ask for a raise. This is a survey about worker perceptions, not observed layoffs, output or pay. It compares two waves taken before publication, so it cannot show that A I caused the changes. The authors also state that their views do not represent the Federal Reserve Bank of Boston or the Federal Reserve System. New survey waves need a match against hiring, layoff and wage records. The key split is whether task substitution produces measurable gains for workers. If the middle group grows while pay and hiring weaken, the anxiety will have more economic evidence behind it. Texas's 474-gigawatt data center queue started a hunt for ghost demand. After Texas ordered an audit of data center grid requests, Reuters reviewed large-load queues across the central United States. Requests exceeded 700 gigawatts, about as much electricity as every U S. home uses. Texas accounted for more than 474 gigawatts, and ten large utilities elsewhere reported about 270 gigawatts. The numbers are not firm demand. Exelon cut its high-probability data center estimate by about 40 percent after stricter screening. AEP Ohio's pipeline fell by more than half after state rules added fees and other requirements. The pressure: Utilities do not count requests the same way. Some include early inquiries while others require contracts, deposits or permits. Adding the queues together can overstate likely electricity use. Removing weak projects can also hide the remaining problem. Reuters found that confirmed demand still exceeds available generation and grid capacity in several regions. What to watch: Texas's audit should show how many projects name an owner, power source, water plan and real customer. Deposits, study fees and permit applications are stronger evidence than a place in a queue. Watch how much requested load survives those tests and who pays when utilities build for projects that do not arrive. That is the signal before the noise. This briefing was produced from Hot Tea's verified daily edition. For the complete briefing and every source link, visit Hot Tea dot A I.