OpenAI published its first benchmark results Tuesday for Jalapeño, its custom inference chip, claiming 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower latency than Nvidia's Blackwell systems — timed one day ahead of Nvidia's own earnings report. [1][2] Nvidia has been approached for comment and has not yet responded, according to CNBC's coverage. [1]
OpenAI's own technical writeup, published to its research blog, frames the results around a public benchmark called InferenceX, developed by research firm SemiAnalysis, which measures the full process of serving an AI request rather than raw chip throughput alone. Testing across three public models — GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T — Jalapeño delivered a better combination of throughput, power efficiency and latency than comparison systems across the tested operating range, according to OpenAI. On the largest model tested, Kimi K2.5, Jalapeño delivered roughly 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system; OpenAI said its internal testing on frontier OpenAI models, not disclosed publicly, showed the advantage widening further. [2] Jalapeño is rated at 700 watts per package but ran at or below 550 watts sustained on the tested workloads, according to OpenAI's own published figures. [2]
CNBC's coverage treated the benchmark as evidence of a structural shift already underway across the industry. Technology analyst Adrien Sanchez of Yole Group told CNBC that a "hyperscaler-designed chip can now match or beat Nvidia's Blackwell-class GPUs on inference efficiency," while acknowledging Nvidia still owns the "vast majority" of AI compute and retains deep ecosystem lock-in through its CUDA software platform. Sanchez called Jalapeño "a threat to Nvidia's inference margins, which is the field growing the most at the moment." Omdia senior analyst Alexander Harrowell was similarly direct, calling Jalapeño "an impressive achievement, most of all in terms of efficiency," and projecting custom ASIC chips like it will exceed GPUs in shipped volume by 2028, though GPU revenue will remain larger for longer given Nvidia's higher per-chip pricing. [1]
That's where X's chip-focused commentary, tracking the same benchmark release in real time, flagged a caveat CNBC's own reporting eventually noted but did not lead with. Chip analyst Meg McNulty posted a detailed breakdown the day of OpenAI's release: "Jalapeño is rated at 700 W per package and stayed at or below 550 W... OpenAI's chip beats Nvidia. Jalapeño uses HBM4," pointing to the memory-generation mismatch as the load-bearing detail beneath the headline comparison. Nvidia's currently shipping Blackwell platform uses an older memory generation than Jalapeño's HBM4; Nvidia's own next-generation Rubin platform, which uses HBM4 as well, is the genuinely comparable chip — and CNBC's article does, several paragraphs down, quote SemiAnalysis's own blog post acknowledging the comparison was "somewhat incomplete and unfair" for exactly that reason. [1]
SemiAnalysis, which said it visited OpenAI's labs directly to conduct the benchmarking, titled its own blog post on the results "OpenAI Jalapeño: Better Than Nvidia Blackwell" — a framing CNBC's reporting adopted largely at face value in its headline and lead, even while later citing SemiAnalysis's own caveat. "Jalapeño is really competing against chips like Rubin that also use HBM4," the SemiAnalysis analysts wrote, noting that Vera Rubin systems are already shipping to customers while Jalapeño remains in the engineering-sample stage. [1] That distinction matters directly for how investors should read the timing: OpenAI published its results one day before Nvidia's own earnings, in which Nvidia disclosed that its Vera Rubin platform is already "ramping into full production" with partners including CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure and Nebius — meaning the fairer generational comparison is between two chips at very different deployment stages, not two chips launching simultaneously.
The chip is being developed jointly with Broadcom and is expected to begin deployment within OpenAI's own compute infrastructure by the end of the year, with a second generation already deep in development and a third taking shape. [2] OpenAI said the chip's design process itself leaned heavily on AI tools: earlier OpenAI model generations helped design and bring the chip up from initial concept to tapeout in nine months, while newer models are now accelerating how engineers optimize and program it, including AI-generated implementations of selected attention and mixture-of-experts blocks that ran 1.5 to 1.8 times faster than existing human-written code for those specific components. [2] TrendForce analyst Fion Chiu told CNBC the chip could meaningfully reduce OpenAI's reliance on Nvidia for inference workloads specifically, even as she and other analysts expect Nvidia's GPUs to remain essential for large-scale training given their broader programmability and mature software ecosystem. [1]
Nvidia is not the only target of OpenAI's build-your-own-silicon strategy, and CNBC's coverage situates Jalapeño within a broader wave: Google's Tensor Processing Units, Meta's Broadcom-designed chip program, and Amazon's Trainium line, which Anthropic committed more than $100 billion toward earlier this year, all represent hyperscalers and frontier labs moving at least part of their inference workloads off general-purpose Nvidia GPUs. [1] OpenAI, historically one of Nvidia's single largest GPU customers, said in its own writeup that it "will continue to widely deploy accelerators from NVIDIA and other partners for both training and inference workloads" — language that reads less as a break from Nvidia than as a hedge, timed for maximum visibility the day before Nvidia had to answer for its own numbers.
-- THEO KAPLAN, San Francisco