US Edition
Your source for latest news
TechnologyAI Hardware

OpenAI's first custom chip claims to outperform Nvidia's Blackwell on inference

Independently verified benchmarks show OpenAI's Jalapeño chip, built with Broadcom, beating Nvidia's current AI processors on speed and efficiency — a rare feat for a first-generation chip.

PT
By PressTemps Technology DeskPublished Today, 01:35 ET · 7 min read
OpenAI's first custom chip claims to outperform Nvidia's Blackwell on inference
OpenAI CEO Sam Altman speaking at TED in April 2025. Photo: Steve Jurvetson / Wikimedia Commons, CC BY 2.0
What to know
OpenAI and Broadcom unveiled independently verified benchmarks for Jalapeño, OpenAI's first custom AI chip, at the Hot Chips conference on August 25
The chip delivered 1.5 to 1.9 times more AI throughput per watt and up to 3.6 times lower latency than comparable Nvidia Blackwell systems
Semiconductor research firm SemiAnalysis verified the benchmark runs in person, using its own independent testing suite
OpenAI plans a small-scale deployment by the end of 2026 and a larger rollout in 2027, while continuing to rely heavily on Nvidia chips in the near term

OpenAI published the first independently verified benchmark results this week for Jalapeño, its first custom-built AI chip, and the numbers suggest the company has produced a serious rival to Nvidia's dominant processors on at least one important measure: how efficiently a trained model can be run to generate answers. The results, unveiled at the Hot Chips semiconductor conference outside San Francisco on August 25, show Jalapeño beating comparable Nvidia Blackwell systems on inference speed and power efficiency in its first generation — a feat semiconductor analysts say new chips rarely pull off.

The chip, formally called an "Intelligence Processor," was co-designed with Broadcom, manufactured on Taiwan Semiconductor Manufacturing Company's N3P process, and integrated into full server racks with help from Celestica, a systems-assembly contractor. It is the first tangible product of a partnership OpenAI and Broadcom announced in October 2025 to build roughly 10 gigawatts of custom AI accelerators — enough generating capacity, in industry terms, to power several large cities — part of a broader shift among the biggest AI developers and cloud providers toward designing their own silicon rather than relying solely on Nvidia's GPUs. Google has built multiple generations of its Tensor Processing Units, Amazon has its Trainium chips, Microsoft has its Maia accelerator, and Meta has its own MTIA line; OpenAI, as primarily a software and model company rather than a cloud infrastructure provider, is the most recent and in some ways the most surprising entrant into that club.

The numbers

According to results OpenAI and Broadcom's investor-relations release, Jalapeño delivers 1.5 to 1.9 times more usable AI throughput per watt than comparison Blackwell systems at peak load, with 1.7 to 3.6 times lower end-to-end latency. On the kind of rapid back-and-forth exchanges that power chatbot conversations, it ran 2.1 to 4.1 times faster in testing. On one narrower test using the Kimi-K2.5 model, an independent analysis found it beat the next-best chip by nearly ninefold.

The engineering specifications help explain why: Jalapeño uses newer HBM4 memory with 15.4 terabytes per second of bandwidth, runs at roughly 700 watts in its current test-silicon form, and supports 32 lanes of 800-gigabit networking that allow racks of up to 2,048 chips to work together as a single system. Chip design took about 16 months from the Broadcom partnership's announcement to a working prototype — unusually fast for custom AI silicon, which typically takes years to go from blueprint to benchmark.

Why the benchmark carries weight

The credibility of the results rests heavily on independent verification. SemiAnalysis, a semiconductor research firm with a reputation for skepticism toward vendor-supplied chip claims, sent its own engineers into OpenAI's labs to observe the benchmark runs firsthand, using its own public "InferenceX" testing suite rather than one supplied by OpenAI. OpenAI supplied the hardware and underlying workloads; SemiAnalysis verified that the runs actually produced the numbers being claimed. That distinction matters in an industry where chip makers regularly publish favorable comparisons using workloads and configurations tuned to flatter their own products; an outside firm physically watching the tests run is a higher bar than most announcements clear. The results were disclosed publicly at Hot Chips, an annual academic and industry conference where chip makers including Nvidia, AMD, Intel and Google's TPU team have historically presented new architectures for peer scrutiny.

"Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin. This is huge news," said Dylan Patel, founder of SemiAnalysis, in the firm's published analysis of the results.

Analysts were careful to note limits on the comparison. Jalapeño's newer HBM4 memory makes a head-to-head match against Nvidia's currently shipping Blackwell chips somewhat uneven; a fairer rival, several noted, is Nvidia's Rubin platform, which also uses HBM4 memory but has not yet shipped in volume. The tested models — DeepSeek R1, Kimi-K2.5 and GPT-OSS — are not the largest frontier models in use today, and OpenAI's disclosed results cover only one workload pattern, short exchanges of roughly 8,000 input and 1,000 output tokens, rather than the long-context, multi-turn tasks increasingly common in AI agents. And only engineering samples of the chip currently exist; nothing is yet in mass production.

What it means for Nvidia

Richard Ho, OpenAI's hardware lead, described the results as "a very, very significant performance advance," saying Jalapeño can serve more AI work per unit of power while maintaining low latency, though he acknowledged Nvidia could still close the gap before Jalapeño reaches full deployment in 2027. Nvidia currently supplies the overwhelming majority of chips used to train and run large AI models — a position built substantially on its Blackwell architecture, which the company has marketed as the fastest general-purpose AI computing platform available at scale — and that dominance has made Nvidia one of the most valuable companies in the world. Nvidia continues to be one of OpenAI's largest suppliers even as OpenAI builds its own alternative; the two relationships, competitor and customer, are not mutually exclusive in an industry where every major AI lab is trying to secure computing capacity from as many sources as possible.

The stakes go beyond bragging rights at a chip conference. Inference — the process of actually running a trained model to answer a user's question — has become the larger and faster-growing cost center for AI companies as products like ChatGPT scale to hundreds of millions of users, overtaking the one-time cost of training a model in many cloud providers' budgets. A chip that meaningfully cuts the cost of serving those queries, even by the 50 to 90 percent efficiency gains OpenAI is claiming, could materially change the economics of running a large consumer AI product, which is why cloud providers have been racing to build their own silicon rather than pay Nvidia's margins indefinitely.

Alexander Harrowell, a senior principal analyst at the research firm Omdia, called the development "the biggest competitive threat to Nvidia," noting that roughly half of all capital spending on AI infrastructure now comes from hyperscale cloud providers that either already have a custom chip program or could reasonably start one. Not everyone is convinced the threat is immediate. CNBC's Jim Cramer offered a skeptical counterpoint on air, saying, "Every day I read about some chip that is superior to Nvidia. And every year I see no real competitors." A person identified by outlets covering the story as a Nvidia representative defended the company's graphics processors, saying they "will remain important given their broad programmability, performance, software ecosystem, and ability to handle a wide range of workloads" — a characterization consistent with Nvidia's long-standing pitch that its chips, unlike narrower custom silicon, can flexibly handle both training and inference workloads across many customers.

Broadcom's role as co-designer and manufacturing partner has also drawn attention on Wall Street, where the company is increasingly viewed as a credible alternative supplier of AI infrastructure alongside Nvidia and TSMC, rather than simply a networking-chip maker, as detailed in its own investor disclosure of the partnership. OpenAI's own economics are part of the story too: building custom silicon is a long-term bet on reducing the company's dependence on a single supplier for the enormous computing capacity its models require, even as its near-term training workloads still run overwhelmingly on Nvidia hardware. Coverage from TechCrunch and Tom's Hardware both noted that OpenAI has been careful to frame Jalapeño as a complement to, not a wholesale replacement for, its continued purchases of Nvidia GPUs, at least for the next several years.

What happens next

OpenAI has said it plans a small-scale internal deployment of Jalapeño by the end of 2026, with a larger rollout targeted for 2027, and the company says work on second- and third-generation versions of the chip is already underway. Nvidia's next earnings call, expected in the coming weeks, will be watched closely for any direct response to the custom-silicon challenge, as will the eventual shipping benchmarks for Nvidia's Rubin platform — the comparison chip several analysts say will offer the fairest test of whether OpenAI's early lead is durable or simply a product of comparing new memory technology against an older generation.

More on this story

All Technology