OpenAI has revealed the first performance results for Jalapeño, its first custom artificial intelligence inference chip. Developed with Broadcom, the processor is designed to run large language models faster and more efficiently — and its early benchmark results put it directly into competition with some of Nvidia’s most powerful AI hardware.
According to results published by OpenAI on August 25, 2026, Jalapeño delivered significantly more AI processing per watt while also reducing response latency compared with the commercial systems used for comparison.
For users, the implications could eventually be very concrete: faster ChatGPT responses, more responsive AI agents, lower infrastructure costs and potentially cheaper AI services.
But Jalapeño also represents something bigger. OpenAI is no longer simply developing AI models. It increasingly wants to control the entire technology stack needed to run them.
What is OpenAI Jalapeño?
Jalapeño is a custom ASIC — Application-Specific Integrated Circuit — designed specifically for AI inference.
OpenAI officially unveiled the processor with Broadcom in June 2026. Unlike general-purpose GPUs that can handle many types of workloads, Jalapeño was designed from the beginning around the requirements of modern large language models.
The distinction between training and inference is important.
Training is the extremely computationally intensive process used to create or improve an AI model.
Inference happens after the model has been trained. Every time a user asks ChatGPT a question, generates code with Codex or calls an AI model through an API, infrastructure has to perform inference to generate the answer.
At the scale of OpenAI, billions of interactions mean that even relatively small improvements in inference efficiency can potentially translate into enormous savings.
Jalapeño Posts Impressive First Benchmarks
OpenAI tested Jalapeño using InferenceX, a public AI inference benchmark developed by SemiAnalysis.
The company tested several major models:
- GPT-OSS 120B
- DeepSeek R1 670B
- Kimi K2.5 1T
Across those workloads, OpenAI says Jalapeño delivered approximately 1.5 to 1.9 times more AI work per watt at peak throughput than the comparison systems.
It also recorded 1.7 to 3.6 times lower end-to-end latency.
For highly interactive workloads, OpenAI reports performance improvements ranging from approximately 2.1 to 4.1 times.
The comparison systems included configurations using Nvidia’s powerful GB200 and GB300 Blackwell accelerators.
These numbers are especially interesting because AI infrastructure usually faces a compromise between two objectives: processing as many requests as possible and responding to individual users as quickly as possible.
OpenAI claims Jalapeño can improve both simultaneously.
Why Power Efficiency Matters So Much for AI
Performance is only part of the AI infrastructure race.
Electricity has become one of the industry’s biggest constraints.
Large AI data centers consume enormous amounts of power, and companies including OpenAI, Microsoft, Google, Meta, Amazon and Nvidia are investing billions of dollars to expand computing capacity.
This makes performance per watt increasingly important.
OpenAI rates Jalapeño at 700 watts, although the company says sustained power consumption remained at or below 550 watts during the workloads it tested.
If the architecture maintains this efficiency when deployed at large scale, OpenAI could serve more AI requests using the same amount of electrical power.
That potentially means:
- lower inference costs;
- greater AI data-center capacity;
- faster responses;
- more simultaneous users;
- more complex AI agents;
- improved reliability during periods of high demand.
For AI companies operating millions or billions of model requests, these improvements can have major economic consequences.
Could Jalapeño Make ChatGPT Faster?
This is ultimately one of OpenAI’s objectives.
The company says improvements in its underlying infrastructure could translate directly into a faster ChatGPT experience.
AI agents could benefit even more.
Unlike a traditional chatbot response, an agent may execute many consecutive reasoning and tool-use steps to complete a task. Small delays at every step can therefore accumulate into significant waiting time.
Reducing inference latency could make agents feel much more responsive.
OpenAI also specifically mentions Codex as one of the products that could benefit from its infrastructure improvements.
Developers using OpenAI’s API may eventually benefit as well if better hardware efficiency reduces the cost of serving models.
However, OpenAI has not announced any specific API price reduction tied directly to Jalapeño.
Is OpenAI Trying to Replace Nvidia?
Not exactly.
Jalapeño clearly introduces a new competitive dimension for Nvidia, but OpenAI does not currently appear to be planning to replace Nvidia GPUs entirely.
Nvidia remains one of the company’s major computing partners.
Instead, OpenAI is building a more diversified infrastructure strategy where its own specialized processors can operate alongside hardware from external suppliers.
That approach is becoming common among the largest AI companies.
Google has its TPUs, Amazon has Trainium and Inferentia, Microsoft has developed its Maia accelerators, and Meta is investing in its own custom AI silicon.
The reason is straightforward: when a company spends billions of dollars every year on computing infrastructure, designing hardware optimized specifically for its workloads can become economically attractive.
Jalapeño brings OpenAI into that same race.
A Chip Designed With Help From AI
There is another interesting detail behind Jalapeño.
OpenAI says its own AI models helped engineers develop the chip.
The processor reportedly progressed from initial design to manufacturing tape-out in approximately nine months, with AI assisting parts of the engineering and optimization process.
This creates an interesting technological feedback loop.
OpenAI builds better AI models.
Those models help engineers build better AI hardware.
That hardware can then run future AI models more efficiently.
OpenAI increasingly describes this strategy as a full-stack approach, covering models, software, networking, memory, chips, data centers and consumer products.
Jalapeño is therefore more than a single processor. It is the first generation of what OpenAI says will become a multi-generation computing platform.
Deployment Begins in Late 2026
Jalapeño is not about to replace existing AI infrastructure overnight.
OpenAI plans an initial deployment in small volumes toward the end of 2026, with production expected to increase during 2027.
The company is already developing future generations of the architecture.
This also means the current benchmark results should be interpreted carefully.
Jalapeño is being compared with today’s Nvidia GB200 and GB300 systems. By the time OpenAI deploys the chip at much larger scale, Nvidia and other semiconductor companies will also have introduced newer hardware.
The AI chip race is moving extremely quickly.
The Bigger Picture: AI Is Becoming a Hardware Battle
The generative AI revolution initially attracted attention because of increasingly capable models such as GPT, Claude and Gemini.
But the next phase of competition may depend just as much on the infrastructure underneath them.
Training and serving increasingly powerful AI systems requires extraordinary computing resources. Companies capable of improving performance while reducing energy consumption and cost will have an important advantage.
OpenAI’s move into custom silicon shows how strategically important that infrastructure has become.
With Jalapeño, the company is attempting to optimize the entire path from the silicon inside a data center all the way to the answer displayed by ChatGPT.
The first benchmarks are promising, but the real test will come when Jalapeño is deployed at scale.
If OpenAI can reproduce these efficiency gains across millions of users and increasingly complex AI agents, the impact could extend far beyond hardware benchmarks.
The next battle in artificial intelligence may not only be about who builds the smartest model — but also about who can run it fastest, cheapest and most efficiently.
Sources
- OpenAI — annonce officielle et benchmarks
- OpenAI — présentation de Jalapeño / Broadcom
- TechCrunch — analyse indépendante
