
newsletter.semianalysis.com
August 25, 2026
24 min read
61/100
Summary
OpenAI has disclosed Jalapeño, a custom AI inference accelerator developed with Broadcom and presented at Hot Chips. The company began designing the chip in mid-2024 and taped out its CoWoS package design in November 2025. Engineering samples use the A0 stepping, while a B0 revision in fabrication is projected by OpenAI to improve performance per watt by about 25%. Production is scheduled to ramp gradually during 2027. SemiAnalysis said it observed OpenAI engineers run parts of its InferenceX benchmark in OpenAI’s lab, but said the reported results were supplied by OpenAI and that it did not run the complete benchmark suite or AgentX’s longer-context, multi-turn tests. SemiAnalysis reported that Jalapeño exceeded Nvidia Blackwell and, in output-token throughput per megawatt, Nvidia Vera Rubin’s published multi-token-prediction results while Jalapeño used single-token prediction. The comparison remains limited by differing models, software maturity, and benchmark configurations. Jalapeño uses HBM4 memory with 15.4 TB/s of package bandwidth, a 700 W TDP, and a TSMC N3P compute die. Each rack contains 128 accelerators, and a scale-up network can link 16 racks, or 2,048 chips. OpenAI designed the chip for a unified inference pool rather than separate prefill and decode pools, and uses its Gluon programming language and Codex-assisted kernel development.
Key Takeaways
What the discussion said
The thread treated the claimed custom OpenAI inference chip less as a settled victory over Nvidia than as a revealing sign of where AI infrastructure is headed. Several commenters were impressed by the prospect of purpose-built silicon driving token costs down, especially for stable, heavily used models where tailoring hardware to a known workload can save enormous amounts at hyperscale. They also saw a crowded field of inference ASICs as healthy pressure on Nvidia’s margins and a possible route to cheaper AI access. But the headline comparison drew substantial skepticism. Readers questioned whether the underlying analysis is independent enough to support grand claims, and whether token-per-joule comparisons against humans measure anything meaningful when a brain handles perception, movement, memory, and biological upkeep alongside speech. The deeper split concerned whether custom hardware creates a durable OpenAI moat. Optimists argue that even a less economical chip can pay off through workload-specific efficiency and supplier leverage. Skeptics counter that semiconductor lead times freeze design choices while model architectures move quickly, so an accelerator can arrive optimized for yesterday’s model. Some also reject the comforting idea that cheaper inference automatically democratizes AI: scale, energy, water, and memory supply could leave the lowest-cost frontier access concentrated in a few giant operators.
Where opinion split
The central dispute is whether a custom inference chip gives OpenAI a lasting advantage over Nvidia and competing AI hardware. Supporters say hyperscale deployment makes model-specific efficiency savings and negotiating leverage valuable even before OpenAI matches Nvidia’s manufacturing scale. Critics say two-year silicon cycles cannot track rapidly changing frontier models, making flexible GPUs and specialized chip vendors the safer long-term bet.
Community Sentiment
Positives
Concerns

OpenAI and Broadcom unveil LLM-optimized inference chip
Jun 24, 2026

OpenAI unveils its first custom chip, built by Broadcom
Jun 24, 2026

Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution
Jul 29, 2026

Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
May 29, 2026

A 10 year old Xeon is all you need
Jun 1, 2026