On August 25, 2026, OpenAI published the first test results of its own inference accelerator, Jalapeño. On the public InferenceX benchmark by SemiAnalysis, systems equipped with the new chip demonstrated 1.5–1.9 times more computations per watt at peak throughput and reduced response latency by 1.7–3.6 times compared to systems based on Nvidia GB200 and GB300. Tests were conducted on the GPT-OSS-120B, DeepSeek R1, and Kimi K2.5 models. For a company that daily serves the massive traffic of ChatGPT, Codex, and its own API, such metrics are directly tied to the cost per answer.
The chip is rated for a nominal power of 700W, while its sustained power consumption under tested loads remained at up to 550W. A block of 128 Jalapeño chips is claimed to deliver 1.7 exaflops of compute in 4-bit format.
Where Independence from Nvidia Ends
Jalapeño was co-developed with partners: Broadcom was responsible for the silicon implementation and networking aspects, while Celestica handled the boards and racks. From initial design to tape-out took nine months. Deployment of the chip in OpenAI's infrastructure is planned for the end of 2026, with Jalapeño becoming the first generation in a multi-year platform set for future expansion.
Nvidia's general-purpose accelerator must serve a broad range of clients and task types. Jalapeño, however, was engineered for the specific workload that OpenAI itself generates, measures, and can modify alongside its software stack—centered around its own language models, compute cores, data movement, memory, and network exchange.
Scale Set for 10 GW
The foundation for the entire program is an agreement between OpenAI and Broadcom, under which the parties committed to deploying 10 GW of OpenAI's own designed accelerators. Rack installations will begin in the second half of 2026 and continue through the end of 2029. At this scale, differences in watts and milliseconds cease to be lab metrics and become cost factors for electricity, cooling, memory, network, racks, and data center floor space.
At the same time, OpenAI is not abandoning external suppliers entirely:
- Training cutting-edge models still requires external accelerators;
- Volume production of Jalapeño depends on partners for manufacturing, memory, packaging, and assembly;
- The company will continue to extensively use chips from Nvidia and other suppliers—both for training and for part of the inference.
The project's boundary is therefore clear: OpenAI maintains an external supply chain but takes control of the hardware architecture for its largest stream of daily compute.
The Economics of Inference
The difference between a custom chip and purchasing GPUs extends beyond energy consumption. Nvidia concluded its 2026 fiscal year with revenue of $215.9 billion and a gross margin of 71.1%, while its data center segment grew 68% year-over-year—these figures are reflected in the company's filings with the Securities and Exchange Commission (SEC). A GPU buyer pays not only for the hardware but also for the vendor's commercial profit margin on top of it.
OpenAI builds the chip as an internal component of its own service and benefits from lowering the total cost of processing a request, even without a separate market markup on the processor itself. Part of the cost that previously accrued to the supplier of general-purpose accelerators is now being turned into the company's own savings and additional compute capacity.
Cheaper and faster inference enables longer-running agentic tasks, servicing more concurrent requests, and reducing product costs without proportional growth in data centers. Increased service usage, in turn, boosts the utilization of the custom platform, making the next generation of specialized chips more economically justified.
Jalapeño changes OpenAI's position in the AI infrastructure chain: the company already controlled the model, software stack, and product, and now gains control over the processor architecture that determines the price of a response. Nvidia retains the market for model training and general-purpose accelerators, but the most massive part of OpenAI's workload is no longer a guaranteed sale for them.
AI Opinion
From the perspective of data-driven analysis, the Jalapeño case follows a trajectory already taken by other industry leaders before OpenAI. Anthropic committed to a million TPUs from Google back in October 2025, and Midjourney shifted image generation to Google Cloud TPUs for cost savings. The underlying logic is the same: hyperscalers move away from general-purpose GPUs to chips tailored for their own workload once the volume of inference becomes sufficiently predictable and massive.
However, the strategy also has a downside. Specialization for a specific model and stack reduces flexibility when AI architectures evolve, and reliance on Broadcom and Celestica for manufacturing creates a single point of risk for the entire nine-month development cycle. Will Jalapeño remain competitive after two or three generations of models, or will the narrow optimization for today's inference become a limitation tomorrow?
end-content







