Nvidia Begins Full Production of Ultra-Fast Artificial Intelligence Chip

Nvidia Begins Full Production of Ultra-Fast Artificial Intelligence Chip

2026-08-25 companies

Santa Clara, Monday, 24 August 2026.
Nvidia’s new Groq 3 LPX chip has entered full production, achieving a record 3,400 tokens per second to make automated corporate AI systems four times more responsive.

Nvidia Enters Full Production of Groq 3 LPX Inference Chip

Nvidia Corporation (NASDAQ: NVDA) announced on Monday, August 24, 2026, that its Groq 3 LPX interactive AI inference accelerator has entered full production [2][5]. This strategic move marks the commercialization of technology following Nvidia’s $20 billion acquisition of Groq assets in December of the previous year [6]. The new hardware is designed to integrate with the Vera Rubin NVL72 architecture to significantly accelerate enterprise agentic AI workloads [1][4]. By focusing on low-latency inference, the company aims to reduce the time required for complex automated tasks, such as software engineering, from hours to minutes [1][2].

Performance Benchmarks and Technical Specifications

In benchmark testing conducted by Artificial Analysis, the Groq 3 LPX achieved a record speed of 3,400 output tokens per second while running the Gemma 4 31B model with a 100,000-token context [2][3]. This performance metric translates to a throughput of 204000 tokens per minute, demonstrating substantial capacity for high-volume data processing [2]. Nvidia claims this speed offers four times higher interactivity for latency-sensitive agentic AI workloads compared to the nearest alternative platform [5]. The implied baseline speed of the nearest alternative would be approximately 850 tokens per second based on these comparative figures [2][5].

Strategic Partnerships and Cloud Deployment

Groq Inc. is among the primary partners bringing the system to cloud markets, leveraging its expertise in operating Language Processing Units (LPUs) [1]. Nebius, an AI cloud provider, has been identified as the first cloud provider to adopt the Groq 3 LPX for its Nebius Token Factory production inference platform [2][6]. Additionally, Groq announced it will deploy the hardware alongside the Vera Rubin NVL72 platform within its purpose-built AI inference cloud in collaboration with Dell Technologies [1]. Dell Technologies is assisting in bringing the NVIDIA Groq 3 LPX and Vera Rubin NVL72 online at scale to turn performance into deployable infrastructure [1].

Market Context and Industry Event Launch

The announcement was highlighted during the Hot Chips conference occurring this week in Palo Alto, California, where Nvidia is showcasing its extreme codesign approach for AI factory infrastructure [4][5]. Industry partners such as SpaceXAI have also announced plans to utilize NVIDIA Vera CPUs for next-generation agentic AI architecture on Earth and in orbital satellites [4][7]. This rollout underscores an evolving hardware push toward high-speed, low-latency inference performance tailored for autonomous corporate AI agents [1][8]. Jensen Huang, founder and CEO of Nvidia, stated that inference is the growth engine of AI and this platform advances the performance frontier for ultrafast token generation [2][5].

Sources


AI Inference Nvidia Chips