How much does an idle NVIDIA H200 cost in the cloud?

The lowest per-GPU rate for the NVIDIA H200 is on p5en.48xlarge on AWS (us-east-2), at $7.91/GPU-hr on-demand. That whole box left running around the clock costs about $46,206/mo. Scheduling it off outside business hours saves up to $32,454/mo (est.).

Cheapest per GPU

$7.91/GPU-hr

p5en.48xlarge · us-east-2

Whole instance 24/7

$46,206/mo

730-hour month

Saved with a schedule (up to, est.)

$32,454/mo

up to 70% off the 24/7 bill (est.)

VRAM 141 GB per GPU Large-memory training & long-context serving 2 instance types

Instances carrying the NVIDIA H200

Instance Provider GPUs vCPU / Memory $/hr $/GPU-hr Spot $/GPU-hr
p5en.48xlarge AWS 8× GPU 192 vCPU · 2048 GB

$63.296/hr

us-east-2

$7.91 $1.90
a3-ultragpu-8g Google Cloud 8× GPU 224 vCPU · 2952 GB

$83.492/hr

us-west1

$10.44 $6.03

About the NVIDIA H200

The H200 is the H100's large-memory sibling: the same compute generation with 141 GB of HBM3e per GPU, aimed at models, optimizer states, and context windows that no longer fit in 80 GB. Teams reach for it when sharding gymnastics on H100s start costing more engineering time than the hardware premium — a large fine-tune that barely fits becomes a simple job that just runs.

Cloud availability is narrower than for the H100: both clouds sell it — p5en on AWS, A3 Ultra on Google Cloud — but each in a short list of regions, and some of Google's only through reservations or an account team. That concentration makes H200 instances prime candidates for the "hold it, we'll need it Thursday" pattern. The arithmetic is unforgiving: hardware priced for scarce peak work bills the same rate during the quiet days in between.

Common questions

Is the H200 worth the premium over the H100?

When per-GPU memory is your binding constraint, usually yes — fewer shards, simpler code, faster iterations. When your job fits comfortably in 80 GB, you are paying for headroom you never touch; the money is often better spent on more H100 time or on shortening the idle gaps between runs.

What should we check before committing to a cloud's H200 capacity?

Region coverage first — a short region list means your data, your team's latency, and your failover story are all pinned to those few regions. Then quota mechanics and how quickly you can restart after a stop. The rate matters, but est. total cost tracks running hours more than list price.

Other GPU tiers

Stop paying for idle cloud GPU servers

Idlefy runs a free audit of your fleet, then flips dev servers to stopped-by-default: lease one when you need it, and when the lease ends it shuts itself off — stop, not terminate, persistent disks intact. Timers are set by your team, never guessed, and Idlefy only manages instances tagged idlefy=enabled — on AWS that boundary is enforced by the IAM policy itself.

Run a free idle audit

Prices are on-demand list price for Linux (AWS) / standard (GCP) instances and exclude storage, network, data transfer, and any negotiated or volume discounts. Actual bills vary by usage and account-level pricing. Spot figures are indicative — AWS values are derived from AWS's published typical savings vs on-demand for the region; GCP values are the listed Spot price. Spot capacity can be reclaimed by the provider. Stopping preserves EBS and persistent-disk volumes; data on local instance-store or local-SSD drives does not survive a stop. Prices as of 2026-08-31.