Custom inference silicon can improve agent economics and model-owner bargaining power if bounded benchmarks survive production-scale cost, uptime and volume tests.
OpenAI's Jalapeño moved custom inference from roadmap to bounded performance proof
Analysis by Frank Locascio and TheBRRR Research
What happened
OpenAI published first measured results for its 700-watt Jalapeño inference ASIC. Across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, it reports 1.5–1.9 times more peak AI work per watt, 1.7–3.6 times lower end-to-end latency and 2.1–4.1 times higher performance on highly interactive workloads versus compared commercial systems. SemiAnalysis observed InferenceX runs in OpenAI's lab but did not run the full suite or AgentX.
Why it earned coverage
Working first-party silicon produced public multi-model results after the prior cutoff; this is more material than the June tapeout announcement.
Investment transmission
Vertical co-design of models, kernels, memory, chip and network can reduce data movement and sequential decode latency. For multi-step agents, latency compounds across steps, so serving architecture can change successful-task economics more than headline token price.
Affected exposures
Next observable receipt
AgentX; Rubin comparison; absolute rack/facility power; cost per successful agent task; production volume; uptime/yield; share of OpenAI tokens served; customer price or gross-margin evidence.
What would invalidate it
AgentX or independent tests erase the lead, Rubin is superior at comparable dates, yield/reliability or software limits delay volume, or fleet TCO fails to improve.