Brands
AWS & NVIDIA to deploy 2 million GPUs as AI infrastructure push accelerates
Expanded partnership spans GPUs, CPUs, networking and robotics to scale AI workloads
NEW DELHI: The AI arms race is getting another power boost. Amazon Web Services (AWS) and NVIDIA are expanding their long-running partnership, with the companies planning to deploy 2 million additional NVIDIA GPUs across AWS’s global infrastructure in 2027 and 2028 as demand for artificial intelligence computing continues to surge.
The expanded collaboration will go beyond GPUs, covering CPUs, networking, AI factories, open models, data processing and robotics. The companies said the aim is to give businesses, AI labs and governments more ways to build and run increasingly demanding AI workloads.
The partnership builds on nearly 16 years of collaboration between AWS and NVIDIA. AWS and NVIDIA previously worked together to launch the world’s first GPU-accelerated cloud instance, and AWS now offers a broad range of NVIDIA GPU-based computing options.
AWS had already announced plans at NVIDIA’s GTC 2026 conference to add more than 1 million NVIDIA GPUs to its infrastructure starting in 2026. The latest commitment effectively takes the planned expansion much further, with another 2 million GPUs scheduled for deployment in 2027 and 2028.
The additional capacity will support workloads including agentic AI, scientific research, enterprise automation and physical AI.
The companies are also working on networking technologies designed to connect large numbers of GPUs more efficiently, an increasingly important requirement as AI models become larger and more computationally intensive.
The expansion comes as AI companies and enterprises move beyond experimentation towards deploying AI systems at scale. Applications ranging from drug discovery and fraud detection to autonomous systems and robotics are placing greater demands on computing infrastructure.
The partnership will also bring NVIDIA’s Vera CPUs to AWS, giving customers another option for workloads that require high-performance CPU computing alongside accelerated infrastructure.
Vera is designed to handle the CPU-heavy tasks supporting agentic AI and reinforcement learning, including code execution, tool use, analytics, data pipelines and orchestration.
The companies said Vera can operate both as a host CPU alongside GPUs and as a standalone CPU for AI factory workloads.
For AWS, the move complements its own custom silicon portfolio, including Trainium. Customers will therefore be able to choose between NVIDIA GPUs, AWS’s Trainium chips or combinations of the two depending on their workloads.
AWS and NVIDIA are also deepening their work on NVLink Fusion, NVIDIA’s high-speed chip interconnect technology.
AWS’s Annapurna Labs is set to work with NVIDIA’s custom high-bandwidth memory technology, known as NVHBM, alongside memory suppliers. The combination is intended to provide faster and more power-efficient memory access for AI systems.
By combining NVHBM with NVLink Fusion, the companies aim to enable Trainium and NVIDIA GPUs to work within a common rack-scale architecture, potentially improving performance and efficiency for large AI workloads.
The partnership will extend into government infrastructure as well. AWS and NVIDIA plan to build AI factories for the US government, including 100,000 NVIDIA GPUs on secure AWS infrastructure.
The systems will support federal and national-security workloads classified at Impact Level 6 and above, among the highest security classifications used by the US government.
The move reflects growing government demand for domestic and secure AI computing infrastructure, particularly for national-security applications.
The companies are also expanding integrations across the wider AWS ecosystem. NVIDIA’s Nemotron family of open models is available through Amazon Bedrock as managed, serverless models and through Amazon SageMaker for customers looking to deploy or fine-tune models on their own infrastructure.
On the data-processing side, GPU acceleration using NVIDIA cuDF on Amazon EMR can deliver up to 3.7 times faster processing and 30 per cent better price-performance for Apache Spark workloads compared with CPU-based configurations, according to AWS.
Amazon OpenSearch Service can also use GPU-accelerated vector indexing, with AWS reporting up to nine times faster indexing at a quarter of the cost.
The partnership is moving into physical AI too. Amazon Robotics is integrating NVIDIA’s Jetson, Omniverse and Isaac technologies for warehouse automation, including simulation, synthetic data generation and real-world testing.
With AI demand continuing to outpace earlier expectations, AWS and NVIDIA are betting that the next phase of the boom will require not just more chips, but an entire infrastructure stack built to keep them working together.




