Posted on Mar 22, 2026 · Updated Jun 24, 2026 · 14 min read
Machine Learning Cloud Costs 2026: Training, Inference & GPU Pricing
Machine learning in the cloud costs most application teams $500–$5,000 per training run — and then far more once a model is live. The bigger surprise is where the money goes: inference, not training, now eats roughly 80% of AI infrastructure budgets, up from a training-dominated split just two years ago (GPUnex, 2026). This is the complete cost map for an ML workload — GPU compute, training vs inference, storage, data transfer, and managed-platform fees — plus the optimizations that cut 40–60% off the bill.
Looking for something narrower? For a head-to-head of GPU instance prices across providers (including specialist clouds like CoreWeave and Lambda), see GPU cloud pricing compared. For the cost of serving a model, see LLM inference cost per 1M tokens, or the self-host vs OpenAI API break-even. Want the numbers for your exact workload? Run the ML cost calculator.
Disclosure: Some links to cloud providers on this page are affiliate links. If you sign up through them we may earn a commission, at no extra cost to you. It never affects our pricing data or recommendations.
TL;DR — ML Cloud Cost Snapshot (2026)
- AWS p4d.24xlarge (8x A100 40GB): $32.77/hr on-demand, ~$9/hr on spot
- AWS p5.48xlarge (8x H100 80GB): $98.32/hr on-demand, ~$30/hr on spot
- Azure NCads H100 v5 (8x H100): $98.46/hr on-demand
- GCP a3-highgpu-8g (8x H100 80GB): $98.32/hr on-demand
- Cost split shifts with maturity: one fine-tuning run is ~60-70% training, but across a live product inference dominates — ~80% of AI infrastructure budgets industry-wide
- Spot/preemptible GPU instances save 60-90% — viable for all training jobs with checkpointing
- Managed ML platform surcharges: SageMaker adds 30-40%, Vertex AI adds 20-30%, Azure ML adds ~25%
The 2026 backdrop: why ML costs are climbing so fast
- Hyperscaler capex is forecast above $600B in 2026 — a 36% jump, with ~75% tied to AI infrastructure (IEEE ComSoc, 2026)
- GPU now accounts for 18% of spend at AI-forward enterprises, up from just 4% in 2023 (State of FinOps 2026)
- Static GPU deployments run at only 30–40% utilization — idle accelerators are the single biggest source of ML waste
Here's the counterintuitive part. Per-token inference cost has fallen roughly 1,000x in three years ($20 down to about $0.40 per million tokens), yet total inference spending keeps climbing — usage scales faster than unit cost drops (GPUnex, 2026). Cheaper tokens haven't produced smaller bills. That's why optimizing how you serve a model now matters more than how you train it. For the full market picture, see the State of AI Infrastructure Costs 2026.
Table of contents
What drives machine learning cloud costs?
ML workloads have a more complex cost anatomy than standard web infrastructure. Five components account for essentially all of the bill.
1. GPU compute
GPU instances are the dominant cost for any non-trivial ML workload. A single H100 GPU costs roughly $2.00–$2.50/hr to rent on-demand across major clouds. An 8-GPU instance (the minimum for most serious training runs) therefore starts at $16–$20/hr just for the accelerators. Training a mid-sized model for 72 hours on a p5.48xlarge costs approximately $7,079 at on-demand rates.
GPU pricing is highly generation-sensitive. A100-based instances (p4d, NC A100 v4) cost roughly a third of H100-based instances (p5, NCads H100 v5, a3-highgpu) for the same GPU count. For fine-tuning and smaller experiments, A100 instances often provide better cost efficiency per token trained.
2. Training time
Training cost scales linearly with wall-clock time. Doubling batch size can reduce training time — but GPU memory constraints limit this. Multi-node distributed training (using 4–32 GPU instances simultaneously) cuts elapsed time but multiplies the per-hour spend. The total GPU-hours required is the invariant; the question is how fast you want to burn them.
3. Storage
Model checkpoints, training datasets, and intermediate activations can consume terabytes of storage. EBS gp3 in AWS costs $0.08/GB/month. A 5 TB dataset stored for one month costs $400. S3 at $0.023/GB/month is cheaper for infrequently accessed datasets ($115/month for 5 TB), but reading from S3 during training introduces I/O latency that can meaningfully increase total training time on bottlenecked pipelines.
High-performance parallel file systems like AWS FSx for Lustre ($0.145/GB/month for SSD-backed) or GCP Filestore are often used to feed GPUs fast enough to avoid idle time between batches. At $0.145/GB/month, 100 TB of Lustre costs $14,500/month — often more than the GPU compute itself for data-heavy workloads.
4. Data transfer (egress)
Pulling training data into a cloud region from on-premises or another region incurs egress charges. AWS charges $0.09/GB for internet egress out of us-east-1. Loading a 10 TB dataset from external storage costs $921 in egress alone. For cross-region training (data in us-west-2, training in us-east-1), inter-region transfer runs $0.02/GB.
See our guide on cloud egress costs for a full breakdown of transfer pricing across providers.
5. Inference serving
Once trained, models need to serve predictions. A real-time inference endpoint running 24/7 on a single ml.g5.2xlarge in SageMaker costs approximately $1,760/month. That is often more than the training run that produced the model. Inference costs are ongoing; training is episodic. For high-traffic production inference, the cost structure is fundamentally different from training — and demands its own optimization strategy.
How much do GPU instances cost on AWS, Azure & GCP?
All three major clouds use NVIDIA GPUs for ML workloads. The key differentiator is not GPU hardware (largely similar across providers) but pricing, availability, and the ecosystem around each instance family. Use the ML cost calculator to compare specific instance configurations side by side.
A100 GPU instances (40GB and 80GB)
| Instance | Provider | GPUs | GPU RAM | vCPUs | On-Demand/hr |
|---|---|---|---|---|---|
| p4d.24xlarge | AWS | 8x A100 | 320 GB | 96 | $32.77 |
| p4de.24xlarge | AWS | 8x A100 | 640 GB (80GB each) | 96 | $40.97 |
| NC A100 v4 (8-GPU) | Azure | 8x A100 | 320 GB | 96 | $36.19 |
| a2-ultragpu-8g | GCP | 8x A100 | 640 GB (80GB each) | 96 | $40.49 |
| a2-highgpu-8g | GCP | 8x A100 | 320 GB | 96 | $26.27 |
H100 GPU instances (80GB)
| Instance | Provider | GPUs | GPU RAM | vCPUs | On-Demand/hr |
|---|---|---|---|---|---|
| p5.48xlarge | AWS | 8x H100 | 640 GB | 192 | $98.32 |
| NCads H100 v5 (8-GPU) | Azure | 8x H100 | 640 GB | 96 | $98.46 |
| a3-highgpu-8g | GCP | 8x H100 | 640 GB | 208 | $98.32 |
H100 instances cost almost exactly 3x as much as their A100 equivalents across all three clouds — a remarkably consistent ratio. The H100 delivers roughly 3x the training throughput for most transformer workloads (per NVIDIA H100 benchmarks), which means cost-per-token-trained is roughly equivalent between A100 and H100 instances. The practical advantage of H100 is speed: finishing a training run in 24 hours instead of 72 hours, which matters for iteration velocity and spot-interruption risk.
GCP's a2-highgpu-8g stands out as the lowest-cost A100 option at $3.28/GPU-hour, roughly 20% cheaper than the equivalent AWS p4d. For long training runs where you can tolerate GCP's preemption model, this translates to material savings. H100 pricing is effectively identical across all three clouds.
Single-GPU and development instances
Not every ML workload needs 8 GPUs. For development, experimentation, and smaller fine-tuning jobs, single- and dual-GPU instances are often more cost-effective. AWS g5.xlarge (1x A10G, 24GB) costs $1.006/hr. Azure NC4as T4 v3 (1x T4, 16GB) costs $0.526/hr. GCP n1-standard-8 with 1x T4 costs $0.952/hr including GPU.
For iterative experimentation on smaller models, starting with T4-based instances and scaling up to A100/H100 only for final training runs can cut development-phase GPU spend by 60-80%.
Training vs inference: where does the money actually go?
For most teams, inference wins. Across AI-forward organizations, serving models now accounts for roughly 80% of AI infrastructure budgets, versus about 20% for training (GPUnex, 2026). The conventional wisdom that training dominates only holds for foundational pretraining from scratch — a GPT-scale run still costs tens of millions. But if you're fine-tuning an existing model and shipping it, inference usually overtakes training cost within six months of launch.
The crossover point happens when cumulative inference cost exceeds cumulative training cost. For a product serving 10,000 requests/day at 500ms average latency on a GPU instance, that crossover typically occurs within 3–6 months of launch. Plan for it.
Inference optimization: the bigger lever
Inference cost per request depends on three factors: instance type, model size, and batching efficiency. Switching from a dedicated GPU endpoint (ml.g5.xlarge at $0.528/hr) to a shared multi-model endpoint can cut inference cost per request by 60-80% for low-traffic models. Quantizing models from fp16 to int8 cuts memory requirements roughly in half, letting you run on smaller (cheaper) instances or pack more replicas onto the same instance.
For a detailed breakdown of how purchasing model affects inference costs, see the reserved vs spot vs on-demand guide. Real-time inference requires reserved or on-demand instances; batch inference is a strong candidate for spot.
The ML cost calculator lets you model both training and inference costs separately, including multi-region inference deployments.
How do spot instances cut training costs 60-90%?
GPU spot instances are the single highest-impact cost lever for ML training workloads. AWS p4d.24xlarge spot pricing runs approximately $9.83/hr — a 70% discount from the $32.77 on-demand rate (AWS Spot Instance Advisor). A 72-hour training run that costs $2,360 on spot versus $7,079 on demand is the same training run — just with checkpoint-aware fault tolerance built in.
GPU spot markets behave differently from general-purpose compute spot markets. The demand for p4d and p5 instances in us-east-1 is high enough that interruption rates are elevated, particularly during business hours. Teams running serious training workloads should distribute across multiple availability zones and use mixed instance type pools.
| Instance | On-Demand/hr | Typical Spot/hr | Savings | 72-hr Training Savings |
|---|---|---|---|---|
| p4d.24xlarge | $32.77 | ~$9.83 | 70% | $1,653 |
| p5.48xlarge | $98.32 | ~$29.50 | 70% | $4,959 |
| g5.48xlarge (8x A10G) | $16.29 | ~$4.89 | 70% | $820 |
| GCP a2-highgpu-8g (Spot) | $26.27 | ~$7.88 | 70% | $1,326 |
Making training fault-tolerant for spot
Every training framework supports checkpointing. PyTorch Lightning checkpoints every N steps natively. Hugging Face Trainer saves checkpoints at configurable intervals. Checkpointing every 30 minutes means you lose at most 30 minutes of compute on interruption — the cost of one checkpoint save is trivial compared to the spot savings on a multi-day run.
The practical approach for multi-node training: use Elastic Fabric Adapter (EFA) networking on AWS to maintain low-latency inter-node communication, set checkpoint intervals at 15–30 minutes, and configure your job runner (SageMaker, Ray, or custom scripts) to automatically resume from the latest checkpoint on a new spot allocation. This architecture eliminates essentially all the risk of spot interruption for training.
For context on the broader tradeoffs between spot, reserved, and on-demand purchasing across all workload types, see our reserved vs spot vs on-demand comparison, including the section specifically covering GPU spot availability patterns in 2026.
Azure Spot VMs and GCP Spot VMs for ML
Azure Spot VMs for NC-series (GPU) instances provide similar 60-80% discounts but use an eviction model with 30 seconds notice rather than AWS's 2-minute warning. This makes checkpoint intervals even more critical on Azure. GCP Preemptible/Spot VMs for A2 and A3 instances are capped at 24-hour maximum runtime, which is not a concern for most training runs but matters for very long runs that exceed the 24-hour mark without a restart.
How much do SageMaker, Vertex AI & Azure ML cost?
Managed ML platforms add a markup on top of raw instance prices in exchange for managed infrastructure, experiment tracking, model registry, and deployment tooling. The surcharge is significant — and often invisible until you see the bill.
AWS SageMaker
SageMaker charges a surcharge on every instance type it manages. For GPU instances, the SageMaker Training price is approximately 30-40% higher than the underlying EC2 on-demand price. An ml.p4d.24xlarge in SageMaker Training costs $42.074/hr versus $32.77/hr for the equivalent EC2 p4d.24xlarge — a 28% surcharge (AWS SageMaker Pricing).
SageMaker Savings Plans (1-year and 3-year) cover the SageMaker-specific surcharge and can bring the effective rate down to near raw EC2 pricing. SageMaker also supports spot training natively through Managed Spot Training, which automatically handles checkpointing and restarts. For teams using SageMaker, Managed Spot Training + Savings Plans is typically the lowest-cost configuration.
Google Vertex AI
Vertex AI Training charges the underlying GCP Compute Engine rate plus a managed service surcharge of approximately 20-30% for GPU workloads. The Vertex AI Training price for an a2-highgpu-8g is $33.07/hr versus $26.27/hr for the raw GCP instance — a 26% surcharge. Vertex AI offers custom training with preemptible (spot) VMs, cutting the effective rate to approximately $9.92/hr for the same machine (Google Cloud Vertex AI Pricing).
GCP Committed Use Discounts (1-year or 3-year) apply to the underlying Compute Engine resource but not the Vertex AI management surcharge. Teams running continuous training pipelines should compare total cost between Vertex AI with CUDs versus raw GKE clusters with GPU node pools.
Azure Machine Learning
Azure ML computes prices at approximately 25% above raw Azure VM on-demand rates for GPU compute clusters. An NC96ads A100 v4 compute cluster in Azure ML costs $39.91/hr versus $31.93/hr for the raw VM — a 25% surcharge (Azure Machine Learning Pricing). Azure Reserved VM Instances reduce the underlying VM cost by up to 63% on a 3-year term, though the Azure ML layer is not covered by reserved pricing.
| Platform | Raw Instance/hr | Managed Platform/hr | Surcharge | 72-hr Overhead |
|---|---|---|---|---|
| SageMaker (p4d) | $32.77 | $42.07 | +28% | +$670 |
| Vertex AI (a2-high) | $26.27 | $33.07 | +26% | +$490 |
| Azure ML (NC96 A100) | $31.93 | $39.91 | +25% | +$574 |
The managed platform surcharge is worth paying if it replaces significant MLOps engineering effort. For teams without dedicated infrastructure engineers, SageMaker and Vertex AI provide real value. For teams with strong platform engineering, running training jobs directly on EC2, GCE, or Azure VMs via Kubernetes or Ray clusters eliminates the 25-30% surcharge while maintaining full control over the training environment.
5 optimization tips
Tip 1: Use spot instances for all training (with checkpointing)
Implement checkpoint-save every 15-30 minutes and use automatic restart on interruption. On AWS, SageMaker Managed Spot Training handles this natively. On GCP, use Vertex AI custom training with preemptible VMs and the TFX/PyTorch checkpoint callbacks. The operational complexity is low; the savings are 60-70% on every training job. A team running $20,000/month in training jobs can cut that to $6,000–$8,000 with no change to model quality.
Tip 2: Right-size inference endpoints
The default SageMaker endpoint instance type (ml.g4dn.xlarge) is frequently over-provisioned for low-traffic models. Multi-model endpoints (MMEs) let you pack multiple models onto a single endpoint and pay only when a model is actively serving. For models receiving fewer than 100 requests per hour, serverless inference endpoints on SageMaker or Vertex AI eliminate idle GPU costs entirely — you pay per invocation rather than per hour.
Tip 3: Profile GPU utilization before scaling
Under-utilizing a GPU instance is the most common ML cost inefficiency. A training job at 30% GPU utilization is paying for 70% of compute capacity that sits idle. Profile with NVIDIA SMI, PyTorch Profiler, or AWS CloudWatch GPU metrics before adding more instances. Common causes of low utilization: data loading bottlenecks (fix with more DataLoader workers or FSx for Lustre), small batch sizes (increase if memory allows), and single-GPU code on multi-GPU instances (use DataParallel or DeepSpeed).
Tip 4: Apply quantization and model compression for inference
int8 quantization (via bitsandbytes, GPTQ, or TensorRT) cuts model memory in half with minimal accuracy loss for most production use cases. Half the memory means you can run on a smaller GPU instance or serve twice as many concurrent requests on the same instance. For a model that was bottlenecked at ml.g5.xlarge ($0.528/hr), int8 quantization may allow it to run on ml.g4dn.xlarge ($0.736/hr — actually more expensive) or significantly reduce latency-driven scaling, cutting the number of instances behind an endpoint from 3 to 2. Run latency benchmarks after quantization; results vary by architecture.
Tip 5: Delete idle notebooks and kernels immediately
SageMaker Studio notebook kernels and Azure ML compute instances continue running — and billing — until explicitly stopped. A ml.g5.xlarge notebook left open over a weekend costs $25.27 (3 days × 24 hours × $0.528/hr) for zero compute work. Set lifecycle configurations to auto-stop idle kernels after 30–60 minutes in SageMaker. In Azure ML, configure auto-shutdown on compute instances. This is a behavioral cost rather than an architectural one — but across a team of 10 ML engineers, idle notebook waste commonly reaches $500–$2,000/month.
Keep going deeper on AI/ML costs: compare GPU prices across hyperscalers and specialist clouds in our GPU cloud pricing comparison, price out serving a model in LLM inference cost per 1M tokens, run the self-host vs OpenAI API break-even, and see the full strategy in cloud cost management for AI/ML startups. For broader cost context, our reserved vs spot vs on-demand guide covers the commitment-model decision, and the Databricks cost breakdown digs into managed ML platform pricing. For build economics, see what it costs to train an LLM in 2026, the monthly cost of a production RAG system, and a vector database pricing comparison.
Estimate your ML cloud costs before you commit
Use the ML cost calculator to model training cost (GPU type, number of nodes, training hours, spot vs on-demand), inference cost (requests per day, latency target, endpoint type), and storage cost across AWS, Azure, and GCP — side by side, before you write a single line of training code.
Frequently asked questions
How much does it cost to train a large language model on AWS?
It depends on model size, training data volume, and hardware. A 7B parameter model fine-tuned on a proprietary dataset typically requires 10–50 GPU-hours on A100 instances, costing $327–$1,638 at p4d on-demand rates. A 70B parameter model from scratch can require hundreds of thousands of GPU-hours — millions of dollars. For most application teams fine-tuning existing open-source models, total training cost per experiment run is $500–$5,000. Use the ML cost calculator to estimate your specific configuration.
Is SageMaker more expensive than raw EC2 for ML training?
Yes, by approximately 25-40% on GPU instances. SageMaker ml.p4d.24xlarge costs $42.074/hr versus EC2 p4d.24xlarge at $32.77/hr — a 28% surcharge. The surcharge covers managed infrastructure, experiment tracking, model registry, and deployment tooling. Whether it's worth paying depends on your team's MLOps maturity. SageMaker Managed Spot Training partially offsets the surcharge by making spot GPU training far more practical.
Can I use spot instances for ML training without losing work?
Yes, with checkpointing. Checkpoint every 15-30 minutes to durable storage (S3, GCS, Azure Blob). On interruption, the job restarts from the latest checkpoint — losing at most one checkpoint interval of compute. AWS SageMaker Managed Spot Training handles this automatically. PyTorch Lightning, Hugging Face Trainer, and TensorFlow all support checkpoint callbacks natively. The additional checkpoint storage cost (a few GB per checkpoint at $0.023/GB on S3) is negligible compared to 60-70% compute savings.
What is the cheapest cloud for GPU training in 2026?
GCP a2-highgpu-8g (8x A100 40GB) at $26.27/hr is the cheapest on-demand A100 option among the major clouds — roughly 20% below AWS p4d. For H100 workloads, all three clouds charge approximately $98.32–$98.46/hr for 8-GPU instances, with no meaningful price difference. Spot discounts are roughly equivalent across clouds (60-70%). Specialized GPU cloud providers (CoreWeave, Lambda Labs, RunPod) offer H100 at $2.00–$3.00/GPU-hour on-demand, which is 35-50% cheaper than hyperscaler on-demand rates — at the cost of fewer managed services and smaller geographic footprint.
How do I reduce ML inference costs in production?
Four levers in order of impact: (1) Use serverless or multi-model endpoints for low-traffic models — eliminate idle GPU hours entirely. (2) Apply int8 or int4 quantization to reduce memory and enable smaller/cheaper instances. (3) Implement request batching — serving 8 requests per GPU call costs the same as serving 1. (4) For batch inference workloads (not real-time), use spot instances. Teams that apply all four typically cut inference costs by 60-80% versus a naive always-on dedicated GPU endpoint.
Does AWS charge extra for SageMaker model storage?
SageMaker model artifacts are stored in S3, billed at standard S3 rates ($0.023/GB/month in us-east-1). Model registry entries in SageMaker do not carry a per-model storage fee beyond the underlying S3 cost. SageMaker Feature Store charges separately — $0.882/GB/month for the online store (DynamoDB-backed) and $0.023/GB/month for the offline store (S3-backed). Teams with large feature sets should monitor Feature Store costs closely; they can grow significantly with high write throughput.
Estimate your cloud costs — for free
Compare AWS, Azure, and GCP pricing side by side with our free calculator, and dig into the guides to learn how to cut cloud waste. No sign-up required.
Matt Anderson
FinOps Analyst, spendark
Matt is a FinOps practitioner who has spent the better part of a decade helping startups and SMBs right-size compute spend across AWS, Azure, and GCP. At spendark he focuses on AI/ML cost — GPU procurement, training-vs-inference economics, and the optimizations that keep a model affordable in production.