AWS и NVIDIA plan to deploy two million additional GPUs across AWS infrastructure during 2027 and 2028. The planned deployment will include NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra GPUs for AI training, inference, scientific computing, enterprise automation, and physical AI workloads.
The announcement describes a future infrastructure commitment rather than two million immediately available GPUs. It follows an earlier AWS plan to add more than one million NVIDIA GPUs beginning in 2026.
AWS and NVIDIA’s Expanded Infrastructure Plan
AWS и NVIDIA announced the expanded partnership on August 26, 2026. The two companies said the additional GPUs would be deployed across AWS’s global infrastructure, including systems described as AI factories.
The partnership includes more than GPU capacity. AWS and NVIDIA also plan to:
- Introduce infrastructure based on NVIDIA Vera CPUs.
- Продлить NVIDIA NVLink Fusion with custom high-bandwidth memory.
- интегрировать NVIDIA GPUs with AWS Nitro and Elastic Fabric Adapter networking.
- Продолжайте предоставлять NVIDIA Nemotron models through Amazon Bedrock and SageMaker.
- Accelerate data processing and vector indexing through Amazon EMR and OpenSearch.
- Apply NVIDIA’s physical AI platform to Amazon Robotics workloads.
The companies also plan to build AI infrastructure for the US government. This includes 100,000 GPUs on secure AWS systems intended for federal and national-security workloads at Impact Level 6 and above.
Matt Garman, CEO of AWS, said customers “want the freedom to choose the best tools for their AI workloads”. The joint AWS and NVIDIA объявление presents the expansion as a response to organizations moving AI projects from pilots into production.
What Evidence Shows AI Infrastructure Demand Is Growing?
Амазонка и NVIDIA reported substantial growth in their cloud and data center businesses before announcing the additional GPU развертывание.
Amazon reported that AWS generated $42.2 billion in sales during the second quarter of 2026, an increase of 37% from the same period in 2025. Amazon also said its AWS AI business had exceeded a $25 billion annual revenue run rate and was growing at a triple-digit percentage year over year.
Amazon’s purchases of property and equipment increased by $66.1 billion over the trailing 12 months. The company attributed the increase primarily to investments in artificial intelligence. Amazon’s second-quarter results provide financial context for the infrastructure buildout.
NVIDIA reported $89 billion in Data Center revenue for its second fiscal quarter of 2027. This represented an increase of 117% from the same quarter a year earlier.
NVIDIA also reported that its Vera Rubin platform had entered full production, with partner deployments involving CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius. Jensen Huang, NVIDIA’s founder and CEO, said, “And demand is accelerating”. NVIDIA’s financial results filed with the SEC support that statement with the company’s reported Data Center revenue growth.
Что GPU Infrastructure Does AWS Already Provide?
AWS already provides GPU instances based on several NVIDIA architectures. The newly announced two-million-GPU deployment will expand this portfolio over the next two years.
AWS launched EC2 G7 instances using NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs in June 2026. G7 configurations support up to eight GPUs, 256 GB of combined GPU memory, 192 virtual CPUs, 768 GiB of system memory, 7.6 TB of local NVMe storage, and 700 Gbps of network bandwidth.
AWS says G7 instances deliver up to 4.6 times the AI inference performance and 2.1 times the graphics performance of the previous G6 generation. These are AWS-reported comparisons rather than independent HostScore тесты.
G7 instances initially became available in the US East Ohio and US West Oregon regions through On-Demand, Savings Plans, Spot, and selected Dedicated Instance configurations. AWS’s EC2 G7 announcement provides the current instance specifications and regional availability.
HostScoreВзять
The two-million-GPU commitment confirms that AWS and NVIDIA expect AI compute demand to remain high. However, the announcement does not establish how much the new capacity will cost, which regions will receive each GPU model, or when individual instance types will become available.
Potential Effects on AI Hosting Costs
More capacity can ease hardware availability, but it does not guarantee lower hosting prices. GPU hosting costs include the accelerator, CPU allocation, RAM, storage, networking, data transfer, management, and reserved capacity.
Blackwell Ultra, Rubin, and Rubin Ultra systems also target increasingly demanding AI workloads. These systems may deliver more work per GPU, but that does not automatically reduce the entry cost for smaller projects.
Buyers should wait for instance specifications, regional availability, and pricing before calculating the financial effect of AWS’s expansion.
The Role of Smaller GPU Хостинг-провайдеры
AWS suits organizations that need large GPU clusters and close integration with services such as Bedrock, SageMaker, OpenSearch, S3, and Elastic Kubernetes Service. Not every AI project requires that scale or platform depth.
Specialized infrastructure providers can serve businesses that want a known GPU configuration, direct server access, dedicated hardware, or a simpler billing arrangement. Atlantic.Net, for example, currently provides cloud and dedicated GPU конфигурации с использованием NVIDIA L40S and H100 NVL hardware. Its cloud configurations support hourly and longer-term billing, while dedicated servers provide fixed hardware and root access. Atlantic.Net lists its current GPU configurations here.
Atlantic.Net is not a direct replacement for AWS’s largest distributed clusters. It represents another layer of the GPU hosting market for smaller training, inference, research, and dedicated infrastructure workloads.
Последствия для GPU Рынок хостинга
Businesses should match GPU infrastructure to the workload rather than the provider’s total capacity.
Model training depends on GPU memory, multi-GPU networking, storage throughput, and cluster scalability. Fine-tuning may run efficiently on one high-memory GPU, depending on the model and dataset. Inference buyers should measure response latency, concurrency, model-loading time, utilization, and cost per completed request. Long-running workloads should compare reserved cloud pricing with преданный GPU серверы, while experimental projects benefit from hourly billing and fast provisioning.
CPU performance, system RAM, NVMe storage, network bandwidth, and software configuration also affect GPU utilization. A powerful accelerator cannot reach its expected performance when the surrounding infrastructure cannot supply data quickly enough.