On August 25, 2026, OpenAI published the first test results of its own inference accelerator, Jalapeño. On the public benchmark InferenceX from SemiAnalysis, systems with the new chip showed 1.5–1.9 times more computations per watt at peak throughput and reduced response latency by 1.7–3.6 times compared to systems based on Nvidia GB200 and GB300. Tests were conducted on the GPT-OSS-120B, DeepSeek R1, and Kimi K2.5 models. For a company that serves a huge daily stream of requests for ChatGPT, Codex, and its own API, such metrics are directly tied to the cost of each response.
The chip is designed for a nominal power of 700 watts, with sustained consumption under tested loads remaining at up to 550 watts. A single block of 128 Jalapeños is claimed to deliver 1.7 exaflops of computation in 4-bit format.
Where Independence from Nvidia Ends
Jalapeño was developed jointly with partners: Broadcom was responsible for the silicon implementation and networking, Celestica for the boards and racks. Nine months passed from the start of design to handing off the design for production. Deployment of the chip in OpenAI's infrastructure is planned by the end of 2026, and Jalapeño itself will be the first generation in a multi-year platform with subsequent expansion.
Nvidia's general-purpose accelerator must serve a wide range of customers and task types. Jalapeño, however, is designed for the specific workload that OpenAI itself generates, measures, and can change along with its software stack—centered around its own language models, compute kernels, data movement, memory, and network communication.
Scale Set at 10 GW
The foundation for the entire program is an agreement between OpenAI and Broadcom, under which the parties agreed to deploy 10 GW of OpenAI's own designed accelerators. Rack installation will begin in the second half of 2026 and continue until the end of 2029. At such scale, the difference in watts and milliseconds ceases to be a laboratory metric and turns into a line item for electricity, cooling, memory, networking, racks, and data center floor space costs.
However, OpenAI is not completely abandoning external suppliers:
- Training cutting-edge models still requires external accelerators;
- Volume production of Jalapeño depends on partners for manufacturing, memory, packaging, and assembly;
- The company will continue to widely use chips from Nvidia and other suppliers—both for training and for part of its inference workload.
The project's boundary is therefore clear: OpenAI maintains its external supply chain but takes control of the hardware architecture for its largest daily stream of computations.
The Economics of Inference
The difference between having a custom chip and buying GPUs goes beyond power consumption. Nvidia ended its 2026 fiscal year with revenue of $215.9 billion and a gross margin of 71.1%, and its data center segment grew 68% year-over-year—these figures are reflected in the company's filings with the Securities and Exchange Commission (SEC). A GPU buyer pays not only for the hardware itself but also for the supplier's commercial profit on top of it.
OpenAI builds the chip as an internal component of its own service and benefits from reducing the total cost of processing a request, even without a separate market markup on the processor itself. Part of the cost that previously accrued to the supplier of general-purpose accelerators, the company aims to turn into its own savings and additional computational capacity.
Cheaper and faster inference allows for running agentic tasks longer, serving more parallel requests, and reducing product costs without proportionally increasing the number of data centers. Increased service usage, in turn, improves the utilization of its own platform and makes the next generation of specialized chips more economically justified.
Jalapeño changes OpenAI's position in the AI infrastructure chain: the company already controlled the model, software stack, and product, but now gains control over the processor architecture, which dictates the price of a response. Nvidia retains the market for model training and general-purpose accelerators, but the most massive part of OpenAI's workload ceases to be a guaranteed sale for it.
AI Opinion
From the perspective of machine data analysis, the Jalapeño case follows a trajectory that other industry leaders have already traveled before OpenAI. Anthropic committed to a million TPUs from Google back in October 2025, and Midjourney moved image generation to Google Cloud TPUs for cost savings. The overall logic is the same: hyperscalers move away from general-purpose GPUs to chips tailored for their own workloads, as soon as the volume of inference becomes sufficiently predictable and massive.
The strategy also carries a downside. Specialization for a specific model and stack reduces flexibility when AI architectures change, and reliance on Broadcom and Celestica for manufacturing creates a single point of risk for the entire nine-month development cycle. Will Jalapeño remain competitive after two or three model generations, or will the narrow optimization for today's inference become a limitation tomorrow?








