Tuesday, August 25, 2026HotTea verified storyVerified 9:08 AM PDT
← Back to the Tuesday, August 25, 2026 edition

OpenAI published numbers for its own inference chip

OpenAI said Jalapeno is its first custom inference chip. The company said the chip ran 1.5 to 1.9 times more AI work for each watt of power. OpenAI also claimed 1.7 to 3.6 times lower latency than comparison systems across three public models.

Verified 9:08 AM PDT · 2 original sources

The evidence

What the reporting establishes

What happened

OpenAI published the first measured results for Jalapeno on August 25. The company said it built the chip for inference, especially interactive agent workloads. OpenAI said it plans to start using Jalapeno inside its compute infrastructure by the end of 2026.

Why it matters

Inference turns each AI answer into a cost per request. If OpenAI's numbers hold outside its release, Jalapeno gives the company a way to cut latency and power costs. It also gives OpenAI another way to depend less on outside accelerators.

The caveat

OpenAI supplied these benchmark claims. The company compared Jalapeno with commercial systems on public workloads. Those results do not prove that Jalapeno changes production economics.

What to watch

Watch for independent InferenceX results and real deployment volume. Watch customer-facing latency too. The bigger proof is whether Jalapeno handles enough production work to change OpenAI's cloud spending or Nvidia spending.

Audit the story

Original sources

Company claims remain company claims. Follow the reporting and judge the evidence directly.

  1. OpenAIJalapeno's first results show industry-leading speed and efficiency in AI inference
  2. OpenAIThe full stack behind abundant intelligence

Continue the morning

Five stories. One sourced briefing.

Read the full editionListen to the daily audio →