How much does an idle NVIDIA A100 cost in the cloud?

The lowest per-GPU rate for the NVIDIA A100 is on p4d.24xlarge on AWS (us-east-1), at $2.74/GPU-hr on-demand. That whole box left running around the clock costs about $16,029/mo. Scheduling it off outside business hours saves up to $11,259/mo (est.).

Cheapest per GPU

$2.74/GPU-hr

p4d.24xlarge · us-east-1

Whole instance 24/7

$16,029/mo

730-hour month

Saved with a schedule (up to, est.)

$11,259/mo

up to 70% off the 24/7 bill (est.)

VRAM 40 GB per GPU Previous-gen training workhorse 4 instance types

Instances carrying the NVIDIA A100

Instance Provider GPUs vCPU / Memory $/hr $/GPU-hr Spot $/GPU-hr
p4d.24xlarge AWS 8× GPU 96 vCPU · 1152 GB

$21.958/hr

us-east-1

$2.74 $1.21
a2-highgpu-1g Google Cloud 1× GPU 12 vCPU · 85 GB

$3.673/hr

us-central1

$3.67 $2.12
a2-highgpu-2g Google Cloud 2× GPU 24 vCPU · 170 GB

$7.347/hr

us-central1

$3.67 $2.12
a2-highgpu-4g Google Cloud 4× GPU 48 vCPU · 340 GB

$14.694/hr

us-central1

$3.67 $2.12

About the NVIDIA A100

The A100 built the current era of ML infrastructure, and it remains widely available on both clouds — AWS in eight-GPU nodes, Google Cloud in machine types from one to multiple GPUs with the silicon bundled in. For teams whose models fit the 40 GB per GPU that these instance types carry, it is often the pragmatic choice: mature, plentiful, and priced below the current flagships.

Its ubiquity is also its risk: A100 boxes are old enough to have become furniture. A machine provisioned two projects ago, still running because nothing broke, is the classic shape of GPU waste — less dramatic than an idle H100 node, but running twenty-four hours a day, every day, for quarters at a time.

Common questions

Is the A100 still a sensible choice for new work?

For fine-tuning, mid-size training, and serving larger models — frequently yes, especially when flagship capacity is quota-limited. The software stack is the most battle-tested in the fleet. For frontier-scale pretraining it has been superseded, but most teams are not doing frontier-scale pretraining.

One big A100 machine or several smaller GPU boxes?

Consolidating onto one multi-GPU machine simplifies data movement and usually idles less than a scattered fleet of single-GPU boxes, each forgotten by a different owner. Whichever shape you pick, per-GPU-hour cost only tells half the story — count the hours each box actually works.

Other GPU tiers

Stop paying for idle cloud GPU servers

Idlefy runs a free audit of your fleet, then flips dev servers to stopped-by-default: lease one when you need it, and when the lease ends it shuts itself off — stop, not terminate, persistent disks intact. Timers are set by your team, never guessed, and Idlefy only manages instances tagged idlefy=enabled — on AWS that boundary is enforced by the IAM policy itself.

Run a free idle audit

Prices are on-demand list price for Linux (AWS) / standard (GCP) instances and exclude storage, network, data transfer, and any negotiated or volume discounts. Actual bills vary by usage and account-level pricing. Spot figures are indicative — AWS values are derived from AWS's published typical savings vs on-demand for the region; GCP values are the listed Spot price. Spot capacity can be reclaimed by the provider. Stopping preserves EBS and persistent-disk volumes; data on local instance-store or local-SSD drives does not survive a stop. Prices as of 2026-08-31.