US Edition
Your source for latest news

Nvidia's New Inference Chip, Born of Its $20 Billion Groq Deal, Enters Full Production

The Groq 3 LPX accelerator, built to speed up how fast AI agents generate answers, has begun shipping with Nebius as its first cloud customer, Nvidia said, eight months after it paid $20 billion to license Groq's chip technology.

PT
By PressTemps Technology DeskPublished August 25, 2026 · 6 min read
Nvidia's New Inference Chip, Born of Its $20 Billion Groq Deal, Enters Full Production
Nvidia's headquarters in Santa Clara, California. The company said its Groq 3 LPX inference chip, developed under a licensing deal with startup Groq, has entered full production. (Photo: Coolcaesar / Wikimedia Commons, CC BY-SA 4.0)
What to know
Nvidia's Groq 3 LPX inference chip entered full production this week, the first product from its roughly $20 billion licensing deal with startup Groq announced in December 2025.
The chip targets AI "agents," hitting 3,400 tokens per second in independent benchmarks and up to 4x the responsiveness of rival platforms on long-context tasks.
Samsung Foundry manufactures the chip on its 4nm process, having lifted yields from about 33% to over 80% during the production ramp.
Nebius is the first cloud customer deploying the hardware in production, with SpaceX also named as a Vera Rubin platform customer.

Nvidia said its Groq 3 LPX inference accelerator has entered full production, the first commercial payoff from the roughly $20 billion licensing-and-hiring agreement it struck with the chip startup Groq eight months ago. The company disclosed the milestone at the Hot Chips 2026 conference at Stanford University on Monday, positioning the new hardware as a companion to its flagship Vera Rubin data-center platform rather than a replacement for it.

The chip is aimed squarely at a specific chokepoint in modern AI systems: the "decode" phase of inference, when a model generates each successive word, or token, of its response after it has already processed a prompt. As AI products shift from single-turn chatbot replies toward autonomous agents that reason through dozens or hundreds of steps to complete a task, that generation speed compounds, and Nvidia is betting that data centers will pay for hardware dedicated to it.

The numbers behind the claim

In benchmarks run by the independent testing firm Artificial Analysis, cited in Nvidia's announcement, the Groq 3 LPX system produced 3,400 output tokens per second running the open-source Gemma 4 31B model with a 100,000-token context window — a scenario meant to mimic a long-running agent that has to keep track of an extended conversation or document. Nvidia says that is roughly four times faster than the nearest competing platform for latency-sensitive workloads, and that for very large, 2-trillion-parameter models operating at long context lengths, the system can deliver up to 35 times more inference throughput per megawatt than its own GB200 NVL72 racks.

The hardware itself is built around 256 individual accelerators packed into a single LPX rack, according to Nvidia's engineering blog, which describes the design as spanning seven distinct chips across five rack configurations, tied together with the company's BlueField-4 data-processing units, Vera CPU racks and Spectrum-6 networking gear. Rather than a single monolithic accelerator, Nvidia is selling the LPX as a modular slice of its broader Vera Rubin NVL72 system that customers can add specifically to speed up token generation while leaving other racks to handle the earlier, context-processing stage of inference.

Manufacturing has been handled by Samsung Foundry on its 4-nanometer process, and reporting from South Korea indicates the foundry has pushed production yields from roughly 33 percent up to more than 80 percent as it ramped toward full-scale output, a detail SamMobile reported citing Korean industry outlets. The improvement matters for Nvidia because early yield problems on a new process node can bottleneck an entire product launch regardless of demand.

How Nvidia got here

The chip traces directly back to a deal Nvidia announced on Dec. 24, 2025, in which it agreed to pay about $20 billion — its largest transaction on record, well above the roughly $7 billion it paid for the Israeli networking company Mellanox in 2019 — to license Groq's inference chip technology and bring much of the startup's team in-house. As CNBC reported at the time, the arrangement was structured as a non-exclusive technology license rather than a conventional acquisition, leaving Groq nominally independent while its founder, Jonathan Ross, and its president, Sunny Madra, moved to Nvidia to help commercialize the design.

That structure let Nvidia absorb a would-be rival's core innovation — a chip architecture, sometimes called an LPU, that keeps large amounts of memory directly on the processing die to avoid the data-transfer delays that slow conventional GPU-based inference — without formally taking Groq off the market as an independent company. Groq itself is now listed among the earliest adopters of the resulting Groq 3 LPX hardware, running it inside its own inference cloud.

Who is buying it, and who else is affected

Nebius Group N.V., the Amsterdam-based AI cloud company, is the first customer to put Groq 3 LPX into production, according to Nebius's own announcement, which says the accelerator is being folded into its Nebius Token Factory inference platform alongside existing Vera Rubin NVL72 capacity. Nebius says developers already using its platform will be able to route workloads to the new hardware through the same application programming interface, without migrating to separate infrastructure. Reporting from SiliconANGLE also named SpaceX as a customer planning to build future AI infrastructure around the broader Vera Rubin platform that the LPX extends.

The launch carries stakes beyond Nvidia's immediate customer list. For Samsung, landing the manufacturing contract is a marquee win for its foundry business, which has trailed Taiwan Semiconductor Manufacturing Co. in advanced-chip market share and has been pursuing large AI customers to close that gap. For rival chipmakers, including AMD and the custom AI silicon programs run inside Google, Amazon and Microsoft, the move signals that Nvidia intends to compete directly in the specialized inference-chip niche that Groq and other startups had carved out, rather than ceding that segment while it focused on general-purpose GPUs. For enterprises building AI agents — software that can independently execute multistep tasks such as research, coding or customer support — faster, cheaper token generation directly affects both how responsive those products feel and how much they cost to run at scale.

Reaction and the road ahead

Nvidia framed the launch as a bet on where AI spending is heading. "Inference is the growth engine of AI," the company's chief executive, Jensen Huang, said in the announcement, adding that the new hardware represents "another giant leap in AI throughput, efficiency and responsiveness."

"Generation is the phase of inference that determines how responsive an AI system actually is," said Danila Shtan, chief technology officer of Nebius, describing why the company chose to add the specialized hardware to its cloud platform.

Nvidia has not disclosed pricing for the Groq 3 LPX racks or specified how many customers beyond Nebius and Groq itself currently have access to the hardware. The company's public materials describe the current release as full production rather than a limited pilot, suggesting a broader rollout to additional cloud providers is expected in the coming months as Samsung's yields continue to climb. Investors and rival chipmakers are likely to watch closely for signs of how much of the fast-growing inference market Nvidia can claim with dedicated hardware, on top of the general-purpose GPU business that has driven most of its revenue to date, and how competitors respond with their own specialized designs.

More on this story

All Technology