
Jalapeño’s benchmark still needs a production-shaped comparison
OpenAI’s chip results are promising. The buying question adds contemporary alternatives, realistic agent workloads and a verified power boundary.
OpenAI’s August 25 Jalapeño announcement reports 1.5–1.9 times more work per watt at peak throughput across three tested public models, alongside lower latency. It also says deployment inside OpenAI’s infrastructure is planned to begin by year-end. Measured silicon and a deployment plan are substantial news. They are different stages of availability.
I would keep the performance result and change the question attached to it: what comparison will still be relevant when the intended workload actually reaches these chips?
That is a stricter test than beating a named predecessor. It includes the alternative available at the deployment date, the serving workload and the boundary around the power measurement.
The benchmark authors supply the useful caveat
SemiAnalysis’s firsthand account says its team verified InferenceX runs in OpenAI’s lab. It also says the figures were supplied by OpenAI, the team did not run the full suite, and it had not seen AgentX results. The authors point to long-context, multi-turn workloads as an important further comparison.
They also question Blackwell as the relevant generational baseline and discuss Rubin comparisons. So the criticism cannot honestly be that nobody mentioned Rubin. The stronger issue is whether the available measurements answer the deployment decision at hand.
A single-turn serving test can establish a real advantage without settling how the same stack handles repeated agent turns, changing context lengths and cache pressure. Neither result makes the other fraudulent. They measure different operating conditions.
Power needs the same care. OpenAI explains its normalization using accelerator power ratings. Before translating that into a facility budget, I would ask which other system costs the chosen comparison includes and apply the same boundary to both sides.
Write the acceptance envelope first
Imagine I am evaluating infrastructure for a hypothetical coding-agent service. Its short questions are fast, but its longer jobs return repeatedly with growing histories. Some requests share cached prefixes; others do not. The commercial requirement is to finish those jobs within a service target at an acceptable total cost.
I would describe that workload before looking at a headline multiplier. Then I would request results across the relevant context lengths and arrival rates, including the tail latency and successful completion rate. A single impressive point on the curve would remain a result, rather than become the entire capacity plan.
The comparison should identify the software versions and optimizations too. Hardware and serving software work together. If one system needs a different configuration to achieve its result, that belongs in the evaluation and the operating-cost estimate.
There is a sensible stopping point. A research announcement cannot provide every customer’s acceptance test, and a buyer should not pretend an unavailable test has already disproved the hardware. The appropriate status is promising evidence with a defined next measurement.
I would separate that status from speculation about supplier negotiations. Custom inference hardware may change purchasing options. A published speed chart does not disclose the price concessions, internal economics or strategic intent of a private negotiation.
The practical response to a strong new chip result is to sharpen the comparison, not to inflate it into a shipped fleet or dismiss it as a bargaining prop.
Approve the infrastructure against the workload and alternatives it will actually face, with the same measurement boundary on both sides.


