Back to latest

The premise that custom inference silicon is a way to undercut Nvidia but not beat it died …

OpenAI's Jalapeño, the Broadcom-built ASIC the company had kept quiet for two years, beat Blackwell on token throughput per megawatt across nearly every configuration …

Top Story

The premise that custom inference silicon is a way to undercut Nvidia but not beat it died at Hot Chips this week. OpenAI’s Jalapeño, the Broadcom-built ASIC the company had kept quiet for two years, beat Blackwell on token throughput per megawatt across nearly every configuration SemiAnalysis could run — without speculative decoding, without multi-token prediction, and without disaggregating prefill from decode. On single-token-prediction output per MW, it also topped Vera Rubin’s published July results, the figure Nvidia and CoreWeave had just run out as their flagship efficiency claim. OpenAI has engineering samples. Rubin is shipping to customers now. The leaderboard, in other words, was drawn from a rearview mirror.

The numbers deserve scrutiny, and SemiAnalysis flags exactly where. All figures were supplied by OpenAI. The runs were verified in person in the lab, but they covered single-turn 8k-1k workloads on open models — DeepSeek R1 above 700 tokens/sec/user at concurrency 1, GPT-OSS at roughly 1,400, and Kimi K2.5 at nearly 700 — not AgentX, SemiAnalysis’ own multi-turn, long-context suite, which is where cache management, prefix caching, and router behavior under production agentic loads tend to punish a new chip. Evals on GSM8k matched Nvidia silicon. That is a necessary caveat and not a disqualifier: first-generation ASICs are supposed to be non-competitive, and Jalapeño is industry-leading where it has been tested, which in itself is the surprising part.

What makes the result structurally interesting rather than just fast is that OpenAI did not do the thing its own press materials implied. This is not a chip over-specialized to OpenAI’s model family. Jalapeño is a generalized inference chip — it ran the full third-party benchmark suite in the lab, and as a joke OpenAI showed it running Doom, ported over with Codex prompts. That undercuts the easy read that OpenAI just hardwired its own weights into silicon and cheated the comparison. A general chip beating Nvidia on perf/W across the curve is a codesign achievement, not a cheat.

The efficiency framing is the right one, and both companies now agree on it. Nvidia’s Vera talk at Hot Chips showed the same revenue-per-watt graph OpenAI’s designers point to; Jensen said at Computex that if you have a gigawatt, throughput per watt is revenue. Power is the binding constraint — grid interconnection and behind-the-meter generation, not silicon, decide how many tokens a company can sell. Every watt has to turn into tokens, and OpenAI designed for exactly that because it is power-limited, not budget-limited.

On total cost of ownership, the comparison is closer than the headline. Vera Rubin and Jalapeño produce nearly the same output tokens per dollar — and Rubin is doing that with speculative decoding, which delivers a roughly 3-5x cost-per-token reduction. Jalapeño gets there without it. Add speculative decoding to Jalapeño and the TCO gap opens further. Part of that advantage is structural in the dumbest possible way: OpenAI traded Nvidia’s margins for Broadcom’s lower ones. But cost alone is not the explanation, and the counterexamples are the tell. Meta and Microsoft have been at their own AI ASIC programs far longer and have nothing shipping; OpenAI went from hiring a team to tapeout in about sixteen months. That speed — the thing the company claims AI-assisted design made possible — is the actual inflection, because it says the barrier to building competitive silicon is falling, and the firms with the best engineering and the least legacy are the ones clearing it.

The honest close is that this story is not yet about beating Nvidia. Blackwell is last year’s chip; the genuine confrontation is Jalapeño against Vera Rubin, head-to-head on HBM4, and it is being run on uneven terms — OpenAI’s part is immature and unproven at agentic scale, while Rubin’s software is already climbing its own maturity curve in production. The number that decides the next six months is not the one in today’s chart. It is which company’s performance-per-watt keeps rising as both chips mature, and whether OpenAI’s lab-verified single-turn results survive contact with real agentic traffic. That test is still ahead of both of them. OpenAI Jalapeño: Better than Nvidia Blackwell

Also Today

Apple introduces M6 and M5 Ultra for a big leap in performance and AI compute · Source Apple’s biggest local-AI silicon generation arrived as a one-two punch: the 2nm M6 in the Mac mini and the quad-die M5 Ultra in the Mac Studio, the latter a first for M-series via UltraFusion’s 4.4TB/s inter-die link. The numbers that matter are memory and neural compute: M6’s Dual 16-core Neural Engine doubles peak AI compute, while M5 Ultra’s up-to-80-core GPU with per-core Neural Accelerators delivers 4.5x M3 Ultra’s AI compute and 1.2TB/s of bandwidth — enough to run hundred-billion-parameter models and fine-tune them entirely on-device. The direction is unmistakable: Apple is betting the agentic-AI workload happens on the desktop, not the data center, and that a 512GB unified-memory pool is the competitive moat against GPU-server economics. That’s a coherent answer to the Nvidia and OpenAI world — but only for whoever can afford the Studio.

AI is hitting entry-level jobs hardest, Stanford study finds · Source The Stanford update on AI’s employment effects sharpens into a distinctly uncomfortable shape: employment for 22-to-25-year-olds in the most AI-exposed occupations now sits 19 percent below their less-exposed peers, up from 13 percent last year — while overall economy-wide employment shows no gap at all. The researchers tie the effect to “codified” knowledge, the formal, textbook-learnable skills that entry-level jobs depend on, which AI automates directly, versus the tacit, mentor-acquired expertise of senior workers that AI mainly augments. College graduation appears to buffer the effect. What this says is that the headline labor numbers are hiding the damage: total employment stays level while the on-ramp for new entrants quietly closes, a dynamic that punishes the young even as it looks like no story at all in aggregate data.

Nitter project received cease and desist · Source Nitter, the unofficial Twitter/X mirror that let people read tweets without an account or the app, has been archived and is now read-only, and every public instance returns the same “Instance has been rate limited” error — the project’s owner says it received a cease and desist. For seven years Nitter was the canonical workaround for the platform’s login walls and API limits, and its death is less a technical defeat than a legal one, since the rate-limiting that killed instances long predates today. The archive is the quiet part: the maintainer chose to preserve the code and walk away rather than fight, which reads as the rational endgame for any project whose entire premise is circumventing another company’s access controls. Another bridge to an open Twitter is gone, and this one isn’t coming back.

FDA authorizes first wearable device that monitors ketone and blood sugar levels · Source The FDA authorized the Libre Duo 10 Day, the first wearable in the U.S. — and the first anywhere — to continuously monitor both ketone and glucose in a single device, for people 2 and up with diabetes. The clinical significance is the real-time element: until now ketone checks were single point-in-time fingersticks that couldn’t show whether levels were rising, whereas the Duo samples every minute and can alert before ketones reach the threshold where diabetic ketoacidosis becomes a medical emergency. The device cleared via De Novo with six studies spanning 600-plus participants, and it fits the FDA’s Home-as-a-Health-Care-Hub push. Continuous ketone sensing turns a reactive, emergency-driven condition into a monitored, trending one — the same wearable pattern that transformed glucose, now applied to the complication that kills the fastest.

Firefox 157 will include JPEG XL by default on all platforms · Source Firefox 157 will flip JPEG XL decoding on by default across all platforms, using the Rust jxl-rs decoder, after years behind a Nightly-only preference and a Firefox Labs checkbox. Timothy Nikkel’s post shows the work was done properly — multithreaded decode landed in jxl-rs 0.6.0 and benchmarks put it slightly ahead of Safari’s C++ libjxl, with progressive rendering and animation that Safari lacks. What’s telling is the browser split: Safari shipped in 2023, Firefox ships now, and Chrome still keeps it behind a flag with no intent to ship, all three using the same Rust library. That asymmetry is a reminder that a technically superior, freely licensed codec can still lose to web-platform inertia, and that Mozilla has quietly become the most adventurous of the three on image formats.

In Brief

  • OpenAI’s new ChatGPT Admin plugin for Workspace lets workspace admins manage users and permissions directly from ChatGPT rather than a separate console. (Source)
  • OpenAI said it disrupted a Russia-origin covert influence campaign that used its AI to promote a fake Israel-based think tank and a “sovereignty” index praising Russia and attacking the West. (Source)
  • Major frontier-model providers have adopted watermarking of AI outputs to satisfy the monitoring requirements of Article 50 of the EU AI Act, which took effect in the EU on 2026-08-02. (Source)
  • Emacs 31.1 is out with no single headline feature, instead a broad set of incremental improvements and refinements across the editor. (Source)
  • EVE Online announced the beginning of its long-anticipated migration to Python 3, one of the largest Python deployments at scale in the industry. (Source)
  • A guide to evaluating LLMs before production argues that strong benchmark scores don’t predict performance on the messy cases that actually matter in deployment. (Source)
  • Run released an SDK for secure evaluation of AI agents, handling authentication and human approval for the TypeScript programs agents write to coordinate tools. (Source)
  • A blogger analyzed how much of Hacker News traffic and content is now AI-related, examining the aggregator’s role amid the AI boom. (Source)
  • AI safety startup Alice raised $140 million to expand stress-testing of advanced OpenAI and Anthropic models and help companies defend against emerging risks. (Source)
  • Autonomous trucking startup Gatik raised $200 million, its largest round yet, to move from development into scaling commercial driverless operations. (Source)
  • SpaceX plans a $100 billion second Starship base in Louisiana with construction from 2027 and a first launch targeted for 2029, though the state’s economic agency pencils in first operations around 2030. (Source)
  • Viral footage from the World Humanoid Robot Games shows sprinting robots beating Usain Bolt’s 100-meter record — alongside runners bursting into flames mid-race. (Source)

One Line

I’m more worried than I was about a labor market that keeps its overall employment level while quietly closing the on-ramp for people starting their careers.

— Erik Brynjolfsson, Stanford economist, speaking to The Washington Post (via Ars Technica)

Search MacWorks

Enter at least two characters.