Sources
See it in action
Browse the models and styles behind stories like this one — free account, instant gallery.
Explore the catalogDiscuss this with
Pick a companion and get their take on this story

Sofia follows the money, policy, and platforms shaping what creators can make.
Browse the models and styles behind stories like this one — free account, instant gallery.
Explore the catalogPick a companion and get their take on this story
OpenAI's Jalapeño chip has posted benchmark results that beat current state-of-the-art AI hardware on both inference speed and energy efficiency — a combination that, if it holds at production scale, compresses the cost and latency of every model OpenAI runs.

OpenAI's Jalapeño chip, purpose-built for fast AI inference at scale.
Image: The Verge / The Verge AI
Jalapeño is OpenAI's custom-designed AI chip, built specifically for inference — the step where a trained model actually generates output in response to a prompt. That distinction matters: training chips (the H100s and their successors that dominate AI headlines) are optimized for a different workload. Inference silicon is what determines how fast a response arrives and how much electricity it costs to deliver it.
According to TechCrunch, Jalapeño was tested on SemiAnalysis's InferenceX benchmark and registered more tokens per user and more throughput per kilowatt than the current state-of-the-art. For AI-image and AI-video workflows that hit OpenAI's APIs, those two metrics translate directly: more tokens per user means less queuing under load, and more throughput per kilowatt means lower operating cost per generation.

Jalapeño's time-between-tokens (TBT) metric compared to rival inference hardware.
Image: The Verge / The Verge AI
Most inference hardware forces a tradeoff: squeeze out lower latency (faster first token, snappier responses) and you sacrifice raw throughput, or vice versa. That tension is why serving large models at scale is expensive — operators have to overprovision to keep both metrics acceptable.
According to The Verge, OpenAI hardware VP Richard Ho told reporters Jalapeño sidesteps that tradeoff.
"The best of both worlds with lower latency and higher throughput."
— Richard Ho, OpenAI VP of Hardware
If that claim holds under real production traffic — a meaningful caveat — it means OpenAI can serve more concurrent users without the latency spikes that currently make image-generation APIs feel sluggish at peak hours. For creators running batch jobs through the API or building generation pipelines on top of OpenAI models, that's the practical upside.

Jalapeño performs more work per kilowatt than competing inference hardware, according to OpenAI.
Image: The Verge / The Verge AI
OpenAI is the clear near-term winner: proprietary silicon means it no longer depends entirely on Nvidia's supply and pricing for inference capacity. That leverage matters — when Stability AI and others faced GPU shortages in 2023, inference costs spiked and API availability suffered. Owning the chip stack insulates OpenAI from that pressure.
The downstream question for creators is whether efficiency gains reach API pricing. Historically, infrastructure savings at this layer do eventually compress costs — but the timeline is measured in product cycles, not quarters. OpenAI has not announced pricing changes tied to Jalapeño.
One caveat worth holding: the benchmark data comes from OpenAI's own disclosure. SemiAnalysis's InferenceX is a respected methodology, but the specific run conditions, model sizes, and batch configurations haven't been independently replicated at scale. The numbers are credible enough to take seriously; they're not yet settled fact.
For creators choosing between OpenAI-hosted models and self-hosted alternatives on platforms like Charmloop's model catalog, the Jalapeño results suggest OpenAI's hosted inference may widen its speed advantage over third-party GPU deployments — at least until competitors respond with their own silicon. The next concrete data point will be whether OpenAI's API latency numbers shift measurably once Jalapeño deploys at production scale.