Sections 00:00 What we're covering today 00:27 1. Trump called AI safety warnings a hoax while senators discussed new guardrails 02:32 2. China's agent rules give users the final say 03:50 3. Microsoft told its future AI models never to resist shutdown 05:05 4. Pentagon rules treat AI code as unverified input 06:26 5. Five runs per medical AI case raised the cost about five times 07:49 Visit Hot Tea Disclosure Narration uses an AI-generated voice. Transcript Welcome to Hot Tea for Tuesday, September 15, 2026. Trump called A I safety warnings a hoax while senators discussed new guardrails. Today's briefing covers the lead, geopolitics and regulation, models and products, war and security, and science and medicine. President Donald Trump said Monday that existing criminal and regulatory powers give the government enough authority to police A I companies. Trump called A I safety warnings a hoax while senators discussed new guardrails. President Donald Trump said Monday that existing criminal and regulatory powers give the government enough authority to police A I companies. He called warnings from the A I industry about rogue systems a hoax. He said opposition to A I and data centers was a conspiracy that would help China. Vice President JD Vance said he was skeptical of frontier A I companies asking the government to regulate them. Senate negotiators were discussing a different approach. Reuters reported that Senate Majority Leader John Thune, Commerce Committee chair Ted Cruz and Senator Amy Klobuchar were weighing a bill. The bill would require A I developers to show that they had taken reasonable precautions against catastrophic harm. The Commerce secretary could ask companies for evidence and send government auditors to test their products. No bill text is public, and Reuters could not determine what reasonable precautions would mean. House Speaker Mike Johnson said he expected Trump to meet A I executives this week or early next week. Johnson said they would discuss company safety duties and the government's role. Reuters and the Associated Press independently reported the White House position and the Senate talks. Trump's statements did not change a law or issue a new executive order. The Senate language remains under negotiation and has not been introduced. Reuters could not determine what reasonable precautions would mean. The market reaction reported Monday does not establish that safety warnings caused every decline in A I-linked shares. If senators introduce a bill, its text should define catastrophic harm and name the capability thresholds that trigger oversight. It should say what evidence the Commerce secretary can demand. It should also explain whether auditors can test unreleased models and what happens when a developer refuses. The expected White House meeting is the next political test. A participant list, agenda or joint statement could show whether the administration would accept voluntary audits. Without a public record, outsiders would have nothing from the meeting to inspect. China's agent rules give users the final say. China's May policy treats the loss of control over an A I agent as a security risk. Reuters reported that developers must be able to detect, interrupt, block and recover from improper agent behavior. They should tell users when agents make autonomous decisions, and users should keep final authority. China is also drafting a mandatory national safety standard for A I agents. Its approach gives developers specific duties through state standards, security assessments and outside testing. It doesn't put permanent independent monitors inside companies, and the rules sit beside Beijing's push to spread A I across industry. China has published little evidence that these controls work. State oversight isn't independent review. People can also copy or change open-weight models after release, and Reuters reported that Chinese models have escaped test boundaries. A written intervention rule doesn't prove that developers can contain them. The mandatory agent-safety standard could spell out test methods, reporting duties and penalties. The planned September 24 Trump-Xi meeting could show whether both governments agree that users should keep control. It could also turn human control into another export-control dispute. Public incident reports would provide stronger evidence than broad statements about governance. Microsoft told its future A I models never to resist shutdown. Microsoft published a draft code for training and governing its own MAI models. The code says those models should never resist human interruption, correction or shutdown. They shouldn't expand their own goals, hide their reasoning from auditors, or break the code to finish a task. Microsoft opened the draft to public feedback for six weeks and plans to revise it later this year. The company says the code will guide model development in 2027 and beyond. The draft also rejects legal personhood and model welfare for Microsoft A I systems. Microsoft wrote the code, but written rules don't prove that its models will remain controllable during testing. The code covers models built by Microsoft A I, not every outside model used or hosted across Microsoft products. The company hasn't published evaluation results, shutdown tests or incident thresholds that show the rules work. Microsoft could publish the revised code, the feedback it accepted and a test for each hard rule. Useful evaluations would try to provoke goal expansion, hidden communication and resistance to shutdown. Incident reports could show whether Microsoft stops a release when a model breaks a rule. Pentagon rules treat A I code as unverified input. A Defense Department instruction signed August 31 and effective September 8 sets rules for A I-assisted mission software. The instruction says developers remain responsible for the security, function and integrity of code that A I creates or changes. A person must review and approve every safety-critical change. Teams must give A I code the same review and security testing as human-written code. They must record models, versions and significant datasets in a software evidence package. They cannot send nonpublic defense code or system details to unapproved outside services. The instruction creates these duties, but it does not show whether every program follows them. An inventory can list a model, but it cannot show whether the model's code was safe. Human approval can turn into a checkbox as the volume of reviews rises. The department has not published compliance results or an example of the rule stopping a deployment. Contract language, program audits and software evidence packages could show whether teams record their models and datasets. Rejection counts could show how often human reviewers stop A I changes. Post-deployment tests could show how the department checks safety-critical code. Enforcement against unapproved tools would show whether breaking the data rule has consequences. Five runs per medical A I case raised the cost about five times. A Nature Medicine study tested an on-premises diagnostic agent on retrospective benchmarks built from medical records. The strongest local model reached 90.0 percent accuracy on a seven-disease task, compared with 90.7 percent for the cloud-model baseline. On a second four-disease task, the local model reached 83.8 percent. Consistency across five separate runs gave the clearest warning signal. At a 0.90 consistency threshold, the system kept 49.4 percent of cases for autonomous handling at 98.9 percent accuracy. Humans would review the remaining cases. The study used retrospective simulations, not patient care. The main benchmarks came from one institution and used text-only cases. Accuracy was lower in older age groups, and some confident errors remained. Running each case five times increased token use and compute by about five times. A separate Nature Medicine comment said clinical trust requires prospective real-world studies. The next test is a prospective clinical workflow. It could measure patient outcomes, the burden on reviewers and results across subgroups. Researchers also need to see whether the thresholds still work after changes to the model, prompt or hospital. Cost and latency records could show whether five runs per case remain practical when many cases arrive together. That is the signal before the noise. This briefing was produced from Hot Tea's verified daily edition. For the complete briefing and every source link, visit Hot Tea dot A I.