How much does an idle NVIDIA L40S cost in the cloud?

The lowest per-GPU rate for the NVIDIA L40S is on g6e.xlarge on AWS (us-east-2), at $1.86/GPU-hr on-demand. That whole box left running around the clock costs about $1,359/mo. Scheduling it off outside business hours saves up to $954/mo (est.).

Cheapest per GPU

$1.86/GPU-hr

g6e.xlarge · us-east-2

Whole instance 24/7

$1,359/mo

730-hour month

Saved with a schedule (up to, est.)

$954/mo

up to 70% off the 24/7 bill (est.)

VRAM 48 GB per GPU Single-GPU fine-tuning & heavy inference 6 instance types

Instances carrying the NVIDIA L40S

Instance Provider GPUs vCPU / Memory $/hr $/GPU-hr Spot $/GPU-hr
g6e.xlarge AWS 1× GPU 4 vCPU · 32 GB

$1.861/hr

us-east-2

$1.86 $0.67
g6e.2xlarge AWS 1× GPU 8 vCPU · 64 GB

$2.242/hr

us-east-2

$2.24 $0.67
g6e.12xlarge AWS 4× GPU 48 vCPU · 384 GB

$10.493/hr

us-east-2

$2.62 $0.92
g6e.4xlarge AWS 1× GPU 16 vCPU · 128 GB

$3.004/hr

us-east-2

$3.00 $0.84
g6e.8xlarge AWS 1× GPU 32 vCPU · 256 GB

$4.529/hr

us-east-2

$4.53 $1.27
g6e.16xlarge AWS 1× GPU 64 vCPU · 512 GB

$7.577/hr

us-east-2

$7.58 $1.59

About the NVIDIA L40S

The L40S occupies a useful middle: 48 GB of VRAM in a single GPU, enough to fine-tune mid-size models or serve models that overflow an L4, without stepping up to a multi-GPU training node. On AWS it ships in the g6e family across six sizes, from the single-GPU g6e.xlarge through the four-GPU g6e.12xlarge to the largest single-GPU host, g6e.16xlarge, so you can pick a host shape that fits the job instead of over-provisioning around it.

Its workloads — fine-tuning sessions, evaluation batches, heavier inference experiments — are bursty and working-hours-shaped. That session shape is the point: the box is genuinely needed while someone is iterating and genuinely useless the moment they stand up — so a lease that expires with the session removes the idle hours without asking anyone to remember anything.

Common questions

When do I need an L40S instead of an L4?

When the model or batch no longer fits in the L4's 24 GB, or throughput per box starts dictating fleet size. Doubling VRAM to 48 GB covers most mid-size fine-tunes and larger serving footprints. If you find yourself sharding across several L40S boxes, though, price out a proper multi-GPU training node instead.

Is an L40S box worth keeping up around the clock?

Only if it serves production traffic around the clock. For development and fine-tuning use, someone is actively iterating on it a few hours a day — the rest is idle time at premium-silicon rates. A lease that stops the box between sessions typically saves up to the entire off-hours share (est.) with no workflow change.

Other GPU tiers

Stop paying for idle cloud GPU servers

Idlefy runs a free audit of your fleet, then flips dev servers to stopped-by-default: lease one when you need it, and when the lease ends it shuts itself off — stop, not terminate, persistent disks intact. Timers are set by your team, never guessed, and Idlefy only manages instances tagged idlefy=enabled — on AWS that boundary is enforced by the IAM policy itself.

Run a free idle audit

Prices are on-demand list price for Linux (AWS) / standard (GCP) instances and exclude storage, network, data transfer, and any negotiated or volume discounts. Actual bills vary by usage and account-level pricing. Spot figures are indicative — AWS values are derived from AWS's published typical savings vs on-demand for the region; GCP values are the listed Spot price. Spot capacity can be reclaimed by the provider. Stopping preserves EBS and persistent-disk volumes; data on local instance-store or local-SSD drives does not survive a stop. Prices as of 2026-08-31.