How much does an idle NVIDIA B200 cost in the cloud?

The lowest per-GPU rate for the NVIDIA B200 is on p6-b200.48xlarge on AWS (us-east-2), at $14.24/GPU-hr on-demand. That whole box left running around the clock costs about $83,171/mo. Scheduling it off outside business hours saves up to $58,418/mo (est.).

Cheapest per GPU

$14.24/GPU-hr

p6-b200.48xlarge · us-east-2

Whole instance 24/7

$83,171/mo

730-hour month

Saved with a schedule (up to, est.)

$58,418/mo

up to 70% off the 24/7 bill (est.)

VRAM 180 GB per GPU Frontier Blackwell training 2 instance types

Instances carrying the NVIDIA B200

Instance Provider GPUs vCPU / Memory $/hr $/GPU-hr Spot $/GPU-hr
p6-b200.48xlarge AWS 8× GPU 192 vCPU · 2048 GB

$113.933/hr

us-east-2

$14.24 $2.56
a4-highgpu-8g Google Cloud 8× GPU 224 vCPU · 3968 GB

$128.880/hr

us-central1

$16.11 $4.95

About the NVIDIA B200

The B200 is NVIDIA's Blackwell-generation training flagship — roughly 180 GB of HBM3e per GPU, sold as eight-GPU nodes on both clouds (p6-b200 on AWS, A4 on Google Cloud) — and cloud capacity for it is allocated more than it is simply bought: quotas, waitlists, reservation windows, and regional scarcity shape who gets to run on it at all. That scarcity drives a specific behavior — teams that win capacity hold on to it between runs, because releasing an instance can mean losing the slot. Held capacity bills like used capacity.

If you are shopping for B200 time, the listed on-demand rate is only the start: check which regions actually offer it, whether your quota transfers, and what the interconnect looks like at multi-node scale. Check the restart story too — stopping only saves money if you can start again when you need to, which is what capacity reservations and capacity blocks are for. Once restarts are dependable, the difference between "stopped at 2am when training finished" and "stopped Monday morning" is a serious line item.

Common questions

Should we train on B200 or on more H100s?

Benchmark your own workload — the answer depends on model size, parallelism strategy, and what capacity you can actually obtain. Newer silicon usually wins per unit of throughput, but a queue of readily available H100s can beat a waitlist for Blackwell on wall-clock time, which is what deadlines care about.

How do teams avoid paying for idle B200 time?

By making shutdown the default rather than a chore: checkpoints go to durable storage, and the instance stops itself when the lease or the run ends. Keeping scarce capacity warm is sometimes a deliberate, defensible choice — the point is that it should be a choice someone made, with an expiry attached.

Other GPU tiers

Stop paying for idle cloud GPU servers

Idlefy runs a free audit of your fleet, then flips dev servers to stopped-by-default: lease one when you need it, and when the lease ends it shuts itself off — stop, not terminate, persistent disks intact. Timers are set by your team, never guessed, and Idlefy only manages instances tagged idlefy=enabled — on AWS that boundary is enforced by the IAM policy itself.

Run a free idle audit

Prices are on-demand list price for Linux (AWS) / standard (GCP) instances and exclude storage, network, data transfer, and any negotiated or volume discounts. Actual bills vary by usage and account-level pricing. Spot figures are indicative — AWS values are derived from AWS's published typical savings vs on-demand for the region; GCP values are the listed Spot price. Spot capacity can be reclaimed by the provider. Stopping preserves EBS and persistent-disk volumes; data on local instance-store or local-SSD drives does not survive a stop. Prices as of 2026-08-31.