OpenAI published numbers for its own inference chip
OpenAI said Jalapeno is its first custom inference chip. The company said the chip ran 1.5 to 1.9 times more AI work for each watt of power. OpenAI also claimed 1.7 to 3.6 times lower latency than comparison systems across three public models.
Verified 9:08 AM PDT · 2 original sources
The evidence
What the reporting establishes
What happened
OpenAI published the first measured results for Jalapeno on August 25. The company said it built the chip for inference, especially interactive agent workloads. OpenAI said it plans to start using Jalapeno inside its compute infrastructure by the end of 2026.
Why it matters
Inference turns each AI answer into a cost per request. If OpenAI's numbers hold outside its release, Jalapeno gives the company a way to cut latency and power costs. It also gives OpenAI another way to depend less on outside accelerators.
The caveat
OpenAI supplied these benchmark claims. The company compared Jalapeno with commercial systems on public workloads. Those results do not prove that Jalapeno changes production economics.
What to watch
Watch for independent InferenceX results and real deployment volume. Watch customer-facing latency too. The bigger proof is whether Jalapeno handles enough production work to change OpenAI's cloud spending or Nvidia spending.
Audit the story
Original sources
Company claims remain company claims. Follow the reporting and judge the evidence directly.
Continue the morning