Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#llms#claude#ai-ethics#code-generation#ai-safety#openai#anthropic#discussion

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

© 2026 Themata.AI • All Rights Reserved

Archive

|

Topics

|

Privacy

|

Cookies

|

Contact
openaillmsinference-chipsai-hardware

OpenAI Jalapeño: Better than Nvidia Blackwell

OpenAI Jalapeño: Better Than Nvidia Blackwell

newsletter.semianalysis.com

August 25, 2026

24 min read

🔥🔥🔥🔥🔥

61/100

Summary

OpenAI has disclosed Jalapeño, a custom AI inference accelerator developed with Broadcom and presented at Hot Chips. The company began designing the chip in mid-2024 and taped out its CoWoS package design in November 2025. Engineering samples use the A0 stepping, while a B0 revision in fabrication is projected by OpenAI to improve performance per watt by about 25%. Production is scheduled to ramp gradually during 2027. SemiAnalysis said it observed OpenAI engineers run parts of its InferenceX benchmark in OpenAI’s lab, but said the reported results were supplied by OpenAI and that it did not run the complete benchmark suite or AgentX’s longer-context, multi-turn tests. SemiAnalysis reported that Jalapeño exceeded Nvidia Blackwell and, in output-token throughput per megawatt, Nvidia Vera Rubin’s published multi-token-prediction results while Jalapeño used single-token prediction. The comparison remains limited by differing models, software maturity, and benchmark configurations. Jalapeño uses HBM4 memory with 15.4 TB/s of package bandwidth, a 700 W TDP, and a TSMC N3P compute die. Each rack contains 128 accelerators, and a scale-up network can link 16 racks, or 2,048 chips. OpenAI designed the chip for a unified inference pool rather than separate prefill and decode pools, and uses its Gluon programming language and Codex-assisted kernel development.

Key Takeaways

  • OpenAI developed the Jalapeño inference chip with Broadcom, beginning design work in mid-2024 and completing its CoWoS package tapeout in November 2025.
  • SemiAnalysis reported that Jalapeño delivered higher measured or published token-throughput-per-megawatt results than compared Nvidia Blackwell and Vera Rubin configurations, while noting that the tests were incomplete and not directly equivalent.
  • Jalapeño uses HBM4 memory, provides 15.4 TB/s of memory bandwidth per package, and has a stated 700 W TDP.
  • A Jalapeño rack contains 128 accelerators, and OpenAI’s scale-up network is designed to connect up to 2,048 accelerators across 16 racks.
  • OpenAI uses a unified hardware pool for prefill, decoding, and speculative-decoding workloads rather than permanently separating chips into prefill and decode pools.

What the discussion said

The thread treated the claimed custom OpenAI inference chip less as a settled victory over Nvidia than as a revealing sign of where AI infrastructure is headed. Several commenters were impressed by the prospect of purpose-built silicon driving token costs down, especially for stable, heavily used models where tailoring hardware to a known workload can save enormous amounts at hyperscale. They also saw a crowded field of inference ASICs as healthy pressure on Nvidia’s margins and a possible route to cheaper AI access. But the headline comparison drew substantial skepticism. Readers questioned whether the underlying analysis is independent enough to support grand claims, and whether token-per-joule comparisons against humans measure anything meaningful when a brain handles perception, movement, memory, and biological upkeep alongside speech. The deeper split concerned whether custom hardware creates a durable OpenAI moat. Optimists argue that even a less economical chip can pay off through workload-specific efficiency and supplier leverage. Skeptics counter that semiconductor lead times freeze design choices while model architectures move quickly, so an accelerator can arrive optimized for yesterday’s model. Some also reject the comforting idea that cheaper inference automatically democratizes AI: scale, energy, water, and memory supply could leave the lowest-cost frontier access concentrated in a few giant operators.

Where opinion split

The central dispute is whether a custom inference chip gives OpenAI a lasting advantage over Nvidia and competing AI hardware. Supporters say hyperscale deployment makes model-specific efficiency savings and negotiating leverage valuable even before OpenAI matches Nvidia’s manufacturing scale. Critics say two-year silicon cycles cannot track rapidly changing frontier models, making flexible GPUs and specialized chip vendors the safer long-term bet.

Read original article

Community Sentiment

Mixed

Positives

  • Purpose-built inference silicon could slash the cost of serving mature, high-volume models, turning a massive recurring GPU bill into a workload-specific optimization problem.
  • Competition from OpenAI, Cerebras, Groq, and other ASIC efforts could finally squeeze Nvidia’s unusually rich margins and push AI compute prices downward.
  • Lower-precision formats such as FP4 show that AI workloads can trade numerical excess for far more useful inference throughput, a direction some readers view as inevitable.
  • Cheaper tokens could unlock demand from sectors currently priced out of large-scale model use, expanding practical AI beyond today’s highest-value applications.

Concerns

  • The claimed superiority over Blackwell rests on industry analysis that readers find difficult to trust, given perceived incentives and a media culture prone to hype.
  • Comparing model token energy to human speech is misleading because the model only emits text while a human brain simultaneously runs perception, embodiment, memory, and survival.
  • Hard-wiring an AI model or narrowly targeted architecture risks shipping after the frontier has moved on; silicon schedules are measured in years, not model-release cycles.
  • Falling per-token costs may not democratize frontier AI at all: energy, cooling water, memory, and deployment scale could consolidate the advantage of the largest labs.
  • Efficiency gains can trigger Jevons-style demand growth, so cheaper inference may increase total compute, environmental load, and data-center buildout rather than reduce them.

Related Articles

OpenAI and Broadcom unveil LLM-optimized inference chip

OpenAI and Broadcom unveil LLM-optimized inference chip

Jun 24, 2026

OpenAI unveils its first custom chip, built by Broadcom | TechCrunch

OpenAI unveils its first custom chip, built by Broadcom

Jun 24, 2026

aistack - How many devs can you fit on a GPU?

Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

Jul 29, 2026

Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)

Real-time LLM Inference on Standard GPUs: 3k tokens/s per request

May 29, 2026

A 10 year old Xeon is all you need - point.free

A 10 year old Xeon is all you need

Jun 1, 2026