AI Infrastructure
8 min read

OpenAI’s Jalapeño Claimed 1.9× AI Work per Watt

OpenAI benchmarked its Jalapeño inference chip, while Google, Cisco, Stability AI and TRACE showed control shifting across the AI stack.

A precision aged-brass dovetail locks cream paper, walnut, smoky glass and oxblood fabric into one load-bearing beam.
AI Infrastructure / 8 min read
AIENGINE

8 min read

Share

The 24 hours ending 26 August 2026 at 09:02 in Tehran moved the centre of AI competition away from a model in isolation. OpenAI published the first measured results for its Jalapeño inference chip. Cisco packaged compute, cooling and networking into one sellable system. Google placed models, licensed financial data, agents and controls into an industry product. Stability AI raised capital from companies that also sit inside its creative market. The Linux Foundation took in a specification intended to make runtime evidence portable across those stacks.

The strongest evidence came from OpenAI’s engineering samples: it reported 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems across three public models. Those are benchmark results, not production economics. Google’s finance product is in preview, Cisco’s new configurations are due in October, Stability did not disclose its valuation or operating results, and TRACE calls itself a developer preview. The day supplied credible milestones, not proof that any complete system has won.

The day in five lines

Jalapeño put power and latency in the same test

OpenAI tested Jalapeño on SemiAnalysis’s InferenceX benchmark using GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. The company compared complete serving systems across operating points rather than quoting a chip’s theoretical arithmetic. On GPT-OSS, it reported roughly 1.9 times higher peak mixed-token throughput per kilowatt than an NVIDIA GB200 comparison, alongside end-to-end latency of 1.03 seconds versus 1.80 seconds. On DeepSeek R1, it reported about 1.7 times higher peak throughput per kilowatt and 1.65 seconds versus 5.99 seconds of latency.

The denominator deserves attention. OpenAI normalized the comparison using published package power ratings: 700 watts for Jalapeño, 1,200 watts for GB200 and 1,400 watts for GB300. It said Jalapeño’s measured sustained power stayed at or below 550 watts on the tested workloads. Its architecture keeps model state, including the key-value cache, local while allocating compute, memory and networking across prefill and decode. That is an engineering explanation for the result, not an independently audited cost statement.

SemiAnalysis said it inspected the chip and verified InferenceX runs in OpenAI’s lab, which is stronger evidence than a vendor slide alone. Its 25 August technical account also records the limits: OpenAI supplied the numbers; SemiAnalysis did not run the full InferenceX suite or its longer-context AgentX suite; and a comparison with Blackwell may be stale by the time Jalapeño deploys because NVIDIA’s HBM4-based Rubin systems are beginning to ship.

Production remains ahead. OpenAI said deployment would begin by year-end, while TechCrunch reported on 25 August that initial volumes would be small and more material deployment would come in 2027. OpenAI disclosed no manufacturing yield, system price, depreciation schedule, total ownership cost or external cloud SKU. Jalapeño is for inference, not frontier-model training, and OpenAI says it will keep using NVIDIA and other partners.

What the result could change operationally

Agent latency compounds. A workflow that makes ten sequential model calls does not experience a slow decode once; it waits repeatedly. Higher throughput at a matched time-between-tokens can also let an operator serve more concurrent work from a fixed power envelope. Those are valuable properties if they survive production software, real prompt distributions, retrieval, tool calls and failures.

Procurement teams should therefore translate the headline into a workload test:

  • Measure prefill, decode, networking and cache behaviour separately, then measure the completed task.
  • Reproduce the real distribution of context length, output length, concurrency and model quality.
  • Compare energy at the rack and facility boundary, not only accelerator package ratings.
  • Include queueing, retries, retrieval and tools in latency and cost per successful outcome.
  • Record availability, support, portability, useful life and supply commitments beside performance.

Yesterday’s Groq and Intel infrastructure brief showed why specialized stages need end-to-end evaluation. The earlier VNET capacity review showed why useful capacity cannot be inferred from accelerator specifications without power, commissioning, utilization and finance. Jalapeño strengthens the case for system-level measurement; it does not remove it.

Cisco sold the integration boundary

Cisco’s announcement made a different full-stack claim. It plans to offer Supermicro liquid- and air-cooled compute systems inside its Secure AI Factory with NVIDIA from October 2026. The package combines NVIDIA Vera Rubin NVL72 or HGX Rubin NVL8 systems, Cisco front- and back-end networking, rack-to-fabric cooling, NVIDIA software, observability and a validation service.

That is commercially meaningful because one supplier and partner network can be accountable for more of the rack. It may reduce qualification and support friction for neocloud and sovereign-cloud operators. The release, however, supplied no configuration price, power envelope, delivery volume, performance result or named completed deployment under the new arrangement. “Validated” should therefore become a testable bill of materials, firmware baseline, acceptance procedure and support contract—not a substitute for workload evidence.

Google moved integration into financial work

Google’s release crossed from infrastructure into a regulated workflow. Published at 08:00 ET on 25 August, Gemini Enterprise for Financial Services is a preview for capital-markets and corporate-banking work. Google says its Financial Research agent exposes methodologies, confidence scores, data snapshots and citations, and connects to licensed sources and internal systems through Model Context Protocol integrations. The launch release names more than 50 skills and 13 connectors.

The customer-side evidence is narrower and useful. Deutsche Bank confirmed on 25 August that it was a design partner and plans to use the agent first with Corporate Bank teams serving German mid-sized companies. It helped shape security, auditability, data-residency and access-control requirements. That is a deployment intention, not a measured productivity or risk outcome.

Google says the product can support KYC research, credit analysis, portfolio monitoring and bond work. Its claims about sub-five-minute analysis and faster trade or issuance ideas have not yet been independently validated. Existing data entitlements also do not authorize a model to approve a customer, change exposure or execute a trade. The workload still needs source-level testing, separation of recommendation from action, human approval and narrow downstream credentials. The agent credential guide explains why a controlled connector must preserve subject, actor, audience and operation identity rather than turning access into authority.

Stability’s investors also sit inside its market

At 11:31 ET on 25 August, Stability AI announced $76 million of new Series B capital. It said total funding under its current leadership had reached $232 million across two equity rounds and convertible notes. Electronic Arts, Sony Music Group, Universal Music Group and Warner Music Group joined AMD Ventures and financial investors in the round. Stability plans to fund creative-production products, applied research and professional services.

The investor list is the operational signal. Several participants are not only sources of capital; they are rights holders, distribution channels, users or hardware ecosystem participants. That may reduce coordination friction as Stability builds tools around licensed content and professional workflows. It does not itself grant training rights, guarantee product adoption or validate output quality. The company disclosed no Series B valuation, ownership percentages, revenue, cash burn or customer economics. TechCrunch’s 25 August report provides context on existing entertainment partnerships, but the financing announcement remains the evidence for the amount.

Teams buying or building on creative models should retain the exact licence, asset provenance and deployment route described in the open-weight licence field guide. Strategic alignment can improve access; it cannot replace documented rights.

TRACE offered a portability counterweight

The same day, the Linux Foundation accepted TRACE—Trust, Runtime Attestation and Compliance Evidence—under vendor-neutral governance. Its current documentation defines a signed record containing a model identifier and weights digest, runtime measurement, policy bundle, data classification, tool-transcript hash and an optional external anchor. The aim is for a relying party to verify evidence from hardware roots rather than trust an audit log written only by the operator.

The limitation is explicit: TRACE v0.2 is a developer preview. It has a schema and conformance suite, but its public documentation says the registry is not yet public and warns readers to review scope limits before production reliance. Attestation can help prove which measured environment produced a record. It does not prove that the model’s answer was true, the policy was wise, the human approval was valid or the external action succeeded. Those controls remain separate.

What changed—and what did not

DevelopmentConfirmed milestoneStill unproven
OpenAI JalapeñoWorking engineering silicon and measured public-model resultsProduction volume, cost and long-context advantage
Cisco packageNamed components and October offering planPrice, delivery, power and workload performance
Google finance agentPreview product and Deutsche Bank design/deployment planGeneral availability and measured regulated outcomes
Stability AI$76 million Series B and named investorsValuation, revenue and commercial return
TRACELF governance, v0.2 schema and test suiteBroad interoperability and production assurance

What to watch next

  • Independent Jalapeño results on longer-context, multi-turn and agent workloads, with quality held constant.
  • Production yield, deployed system count, facility-level power and cost per successful task in 2027.
  • Cisco’s October configurations, priced bill of materials, delivery commitments and acceptance evidence.
  • Gemini Finance’s general-availability terms, evaluation results, incident controls and Deutsche Bank’s live scope.
  • Stability AI’s closing disclosures, product revenue and the contracts governing content access and outputs.
  • TRACE security review, public anchoring infrastructure, cross-vendor conformance and named production adopters.

The day’s common direction is clear: advantage is moving into the joins between chip, network, model, data, workflow and evidence. That can improve performance and reduce integration work, but it also concentrates dependencies. Buyers should reward measured outcomes while preserving portable data, explicit authority, reproducible tests and independently verifiable records. Full-stack control is an engineering strategy; it is not permission to stop checking the stack.

TaggedOpenAI JalapeñoAI InferenceAI ChipsFinancial Services AIAI InfrastructureAI Governance
Work With Us

Interested in implementing this for your business?

We help UK businesses put these ideas into practice. Book a call to discuss your specific situation.