Mücahithan Avcıoğlu
24 August 2026•Update: 24 August 2026
Nvidia said Monday that its Groq 3 LPX artificial intelligence racks have entered full production, marking the commercialization of technology acquired through the chipmaker’s record $20 billion purchase of Groq assets.
The systems will be deployed alongside Nvidia’s Vera central processing units and Rubin graphics processing units at cloud infrastructure provider Nebius and are expected to become operational later this year.
Nvidia bought assets from AI chip startup Groq in December in its largest-ever transaction.
The Groq 3 LPX is designed for low-latency inference, the process through which trained AI models generate responses. Faster inference is considered particularly important for AI agents and coding assistants, where delays can affect the user experience.
Each rack packages 256 Groq 3 chips and can generate about 3,400 tokens per second, according to a benchmark cited by Nvidia.
The Groq architecture places 500 megabytes of high-speed static random-access memory directly on each chip to reduce memory-related bottlenecks. The chips are manufactured by Samsung Electronics, while Taiwan Semiconductor Manufacturing Company produces Nvidia’s graphics processors.
Nvidia said the specialized systems are intended to complement rather than replace GPUs. While GPUs can perform both AI model training and inference, Groq chips primarily target the latency-sensitive “decode” phase of running models.
The company is also increasing shipments of its Vera Rubin systems, which entered production earlier this year.
Nvidia CEO Jensen Huang said in March that the company expects cumulative sales from its Blackwell and Vera Rubin platforms to reach $1 trillion through 2027. He also said a quarter of the data center capacity allocated to coding applications would use Groq chips.
Competition in specialized inference hardware has intensified as technology companies seek to make AI services faster and more economical. Nvidia rival Advanced Micro Devices has announced plans to integrate rack-scale systems with chips produced by Cerebras.