How much does an idle NVIDIA A100 cost in the cloud?
The lowest per-GPU rate for the NVIDIA A100 is on p4d.24xlarge on AWS (us-east-1), at $2.74/GPU-hr on-demand. That whole box left running around the clock costs about $16,029/mo. Scheduling it off outside business hours saves up to $11,259/mo (est.).
Cheapest per GPU
$2.74/GPU-hr
p4d.24xlarge · us-east-1
Whole instance 24/7
$16,029/mo
730-hour month
Saved with a schedule (up to, est.)
$11,259/mo
up to 70% off the 24/7 bill (est.)
Instances carrying the NVIDIA A100
| Instance | Provider | GPUs | vCPU / Memory | $/hr | $/GPU-hr | Spot $/GPU-hr |
|---|---|---|---|---|---|---|
| p4d.24xlarge | AWS | 8× GPU | 96 vCPU · 1152 GB | $21.958/hr us-east-1 | $2.74 | $1.21 |
| a2-highgpu-1g | Google Cloud | 1× GPU | 12 vCPU · 85 GB | $3.673/hr us-central1 | $3.67 | $2.12 |
| a2-highgpu-2g | Google Cloud | 2× GPU | 24 vCPU · 170 GB | $7.347/hr us-central1 | $3.67 | $2.12 |
| a2-highgpu-4g | Google Cloud | 4× GPU | 48 vCPU · 340 GB | $14.694/hr us-central1 | $3.67 | $2.12 |
About the NVIDIA A100
The A100 built the current era of ML infrastructure, and it remains widely available on both clouds — AWS in eight-GPU nodes, Google Cloud in machine types from one to multiple GPUs with the silicon bundled in. For teams whose models fit the 40 GB per GPU that these instance types carry, it is often the pragmatic choice: mature, plentiful, and priced below the current flagships.
Its ubiquity is also its risk: A100 boxes are old enough to have become furniture. A machine provisioned two projects ago, still running because nothing broke, is the classic shape of GPU waste — less dramatic than an idle H100 node, but running twenty-four hours a day, every day, for quarters at a time.
Common questions
Is the A100 still a sensible choice for new work?
For fine-tuning, mid-size training, and serving larger models — frequently yes, especially when flagship capacity is quota-limited. The software stack is the most battle-tested in the fleet. For frontier-scale pretraining it has been superseded, but most teams are not doing frontier-scale pretraining.
One big A100 machine or several smaller GPU boxes?
Consolidating onto one multi-GPU machine simplifies data movement and usually idles less than a scattered fleet of single-GPU boxes, each forgotten by a different owner. Whichever shape you pick, per-GPU-hour cost only tells half the story — count the hours each box actually works.
Other GPU tiers
Stop paying for idle cloud GPU servers
Idlefy runs a free audit of your fleet, then flips dev servers to stopped-by-default: lease one when you need it, and
when the lease ends it shuts itself off — stop, not terminate, persistent disks intact. Timers are set by your team, never
guessed, and Idlefy only manages instances tagged idlefy=enabled — on AWS that boundary is enforced by the IAM policy itself.
Prices are on-demand list price for Linux (AWS) / standard (GCP) instances and exclude storage, network, data transfer, and any negotiated or volume discounts. Actual bills vary by usage and account-level pricing. Spot figures are indicative — AWS values are derived from AWS's published typical savings vs on-demand for the region; GCP values are the listed Spot price. Spot capacity can be reclaimed by the provider. Stopping preserves EBS and persistent-disk volumes; data on local instance-store or local-SSD drives does not survive a stop. Prices as of 2026-08-31.