This week the interesting action is one layer down from model scaling, in who owns the hardware and the distribution. Where last week was about which dial to turn next — params, data, or post-training — this week two moves pull the compute stack in opposite directions. OpenAI announced Jalapeño, its first inference chip, and the numbers are not a first-generation hedge: SemiAnalysis, invited to benchmark it in OpenAI’s own lab, says it beats Blackwell on performance-per-watt across almost every scenario. A company that was purely an NVIDIA customer 18 months ago now has silicon that’s competitive at the top. And NVIDIA is buying Hugging Face for a reported $13 billion — roughly 80× revenue — folding the default, vendor-neutral hub for open models into the dominant AI hardware company. Put together: NVIDIA’s moat is being chipped at from below in silicon and, apparently in response, extended upward into software and distribution. Also This Week covers the rest — the other Hot Chips announcements, Chinese open models still shipping weekly, AGI talk acquiring a date, the build-your-own-agent trend, and platform risk showing up in real API cutoffs.
After first offering roughly $7 billion back in January, NVIDIA is now reported — per The Information’s scoop and follow-up confirmation — to be acquiring Hugging Face for about $13 billion. That is on the order of 80× Hugging Face’s ~$150 million in annual recurring revenue, a multiple that only makes sense if you’re buying strategic position rather than a cash-flow business. Hugging Face had roughly doubled its customer base over 2026, but the price is really about what the company is: the default, vendor-neutral hub where nearly every open model gets hosted, versioned, benchmarked and pulled from.
That is what makes this different from NVIDIA’s other recent moves in the open ecosystem. Buying a startup’s compute contracts or funding open-model training expands demand for GPUs without touching anyone’s independence. Owning the Hub is owning a chokepoint — the piece of infrastructure that every lab, cloud and hobbyist routes through regardless of whose hardware they run on. Latent Space’s own framing was cheerful (“we love it when the good guys win”), and Hugging Face has been a genuinely good steward of open source. But the structural question stands: the dominant seller of AI hardware would now also run the distribution layer for the models, with the visibility into usage and the standard-setting power that comes with it.
The timing is not incidental. The deal landed in the same week that Z.ai’s GLM-5.3-Flash and a new Qwen Flash model — both strong, both open, both trained on Chinese accelerators — had the post-Hot-Chips conversation focused on whether “Western open AI” has a coherent home. The same AI News issue also carried OpenAI’s postmortem on the Hugging Face security incident from late July, a reminder that the Hub is critical enough infrastructure that its failures are industry events.
If this closes, the open-model ecosystem’s main piece of neutral infrastructure will sit inside NVIDIA. Pair it with Jalapeño — a major customer building competitive silicon — and the picture is NVIDIA’s position being pressured in hardware and, seemingly in response, extended into software and distribution. Worth watching for regulatory friction and for how the labs that depend on the Hub react.
OpenAI used the Hot Chips conference to unveil Jalapeño, the inference chip it has been building with Broadcom since mid-2024. The striking part isn’t that it exists — the Broadcom partnership was announced in June — it’s the performance. SemiAnalysis was invited into OpenAI’s lab to run its InferenceX benchmark on the chip, and their verdict is blunt: Jalapeño beats every NVIDIA, AMD and Google chip they’ve tested on several open-source models. First-generation accelerators are almost never competitive; this one leads. They credit aggressive hardware–software co-design, a roughly 16-month path from team hiring to tape-out, and — notably — AI genuinely being used to speed up the chip design.
The headline number is performance-per-watt: token throughput per all-in-utility megawatt, where Jalapeño “smokes every other chip,” and does it without multi-token prediction while the competing chips are shown in their best MTP configurations. It wins in both low-latency and high-throughput regimes rather than being tuned for one point on the curve — over 700 tokens/sec/user at concurrency 1 on DeepSeek R1 with plain single-token decoding, no speculative decoding, no prefill/decode disaggregation. OpenAI’s own figures against NVIDIA GB200/GB300 systems: 1.5–1.9× more work per watt, 1.7–3.6× lower end-to-end latency, and 2.1–4.1× more performance on highly interactive workloads. A separate detail worth its own headline: GPT-Astra and Codex helped write the low-level kernels, bringing three unplanned open-weight models to high performance in about two months, some kernels running 1.5–1.8× faster than expert-written code.
SemiAnalysis is careful about the caveats. Every number came from OpenAI; they verified the InferenceX runs in person but not the full suite, and didn’t see AgentX results — the long-context, multi-turn benchmark that reflects real production cache behavior, where a chip that shines on single-turn tests can stumble. They also call the Blackwell comparison “incomplete and unfair”: Jalapeño’s real competition is NVIDIA’s HBM4-based Rubin generation, and Vera Rubin systems are shipping to customers now while OpenAI won’t put Jalapeño into its own datacenters until year-end. Packaging and foundry capacity (TSMC, CoWoS) remain a hard limit on how fast OpenAI can scale it.
The concrete claim — verified in person by a skeptical third party — is that a company that was a pure NVIDIA customer 18 months ago now has inference silicon competitive with the best. Even with the deployment lag and the Rubin caveat, that’s the strongest evidence yet that frontier labs won’t stay strictly downstream of NVIDIA on inference economics. It’s also the clearest example so far of AI-assisted chip and kernel design producing a real, shipped result.
Jalapeño dominated the coverage, but the 37th Hot Chips also brought Cerebras’s CS-5, Groq’s 3 LPX, and Apple’s M6 — a snapshot of a custom-silicon race that now has frontier AI labs, wafer-scale specialists, and consumer-chip makers all pushing on inference efficiency at once. The through-line matches the focus stories: the interesting competition has moved from “who trains the best model” toward “who runs it cheapest per watt.”
This week’s agent-and-model roundup led with GLM-5.3-Flash (320B total / 18B active params, 1M context, MIT license, trained entirely on Chinese accelerators, scoring 57 on the Artificial Analysis Intelligence Index at ~$0.09/task), plus a Qwen3.8-Flash, an Hy4 preview, Claude’s new built-in browser, and Terminal-Bench-Science. The cadence — a capable open Flash-tier model roughly every week — is the backdrop to the Hugging Face deal.
A roundup covering reports that OpenAI is targeting an internal “AGI bar” by the end of 2026 — less interesting as a prediction than as a marker of how openly frontier labs are now attaching public timelines to the term, and what that does to fundraising and expectations.
Andrew Ng is repositioning DeepLearning.ai around “AI Engineering,” backed by an analysis of 10,000+ job postings — a mainstream signal that the work is shifting from training models to building reliable systems around them. Two companion pieces on what that work actually involves: why the verification layer is now the bottleneck for AI-generated code, and a walkthrough of speculative decoding, the technique behind roughly 3× faster inference with unchanged outputs.
Ramp built its own in-house coding agent, Inspect, rather than adopting a frontier lab’s — and says the bet is paying off. Lovable’s CTO argues the future of SaaS is apps redesigned as MCP-powered “capabilities” that agents, not just humans, can use. And Meta reportedly considered cutting engineering teams by 60% out of fear of leaner AI-native startups — a look at how AI anxiety is reshaping org design at the top.
Cameron Wolfe’s from-first-principles guide to reinforcement learning for LLMs traces the arc from early alignment work to today’s reasoning- and agent-focused frontier — a useful single reference given how central RL has become to capability gains (see last week’s GLM 5.3 story). The recurring research digests also landed: top papers for Aug 17–23 and Aug 24–30.
SemiAnalysis released AgentX 1.0, an open-source (Apache 2.0) multi-turn agentic-coding inference benchmark at 1M context — the first real open benchmark for the long-context, multi-turn workloads that have overtaken fixed-length single-turn traffic as the dominant production pattern. It’s also, per the Jalapeño piece, the suite OpenAI’s chip hasn’t yet been shown running.
OpenAI cut Cursor’s API access amid the Elon-vs-Altman fallout — a concrete case of an app-layer company’s dependence on a model provider that also ships a competing product. Further down the stack, SemiAnalysis argues most GPU neoclouds have weak multi-tenant security — container escapes, kernel bypass, exposed Grafana — a systemic risk as the biggest labs spread workloads across many new vendors (the piece includes a ClusterMAX 3.0 preview).
A walkthrough of whether the encrypted “private” reasoning traces that Anthropic, OpenAI and Google return to API clients can actually be reconstructed by an attacker — the applied version of the reasoning-trace extraction vulnerability from two weeks ago, where a signed reasoning block is replayed into a weaker same-provider model that’s prompted to transcribe it.
An interview with Anima Anandkumar on why physical systems — weather, fusion reactors — still lack the kind of general-purpose foundation models language has, and her argument for what building them would take. A useful counterweight to a week otherwise dominated by inference economics and M&A.