Posted on Oct 5, 2026 · Updated Oct 5, 2026 · 13 min read

Datadog Cost Optimization: 10 Ways to Cut Your Bill (2026)

The short answer: most Datadog overruns don't come from the per-host price everyone negotiates — they come from three line items most teams never watch: custom metrics, indexed log events, and APM/trace span ingestion. Datadog bills on roughly 20 independent SKUs, and host count is the one number that's visible on every dashboard, so it's the one people fixate on when the invoice jumps. It's usually the wrong one. Custom metric cardinality, how much of your log volume gets indexed versus just ingested, and how many spans your tracers actually ship all scale independently of host count and can double a bill in a single deploy.

Every price in this article is an illustrative 2026 list price. Datadog prices each SKU separately, annual-commit and on-demand rates differ meaningfully, and the company changes pricing and packaging often enough that any number you read today is a model to rebuild with your own rate card, not a quote to budget against. For the wider pattern of where cloud spend gets wasted beyond observability tooling specifically, see our cloud cost optimization checklist.

TL;DR — Datadog cost optimization (2026)

  • Host count is rarely the problem. Custom metrics, indexed logs, and APM span ingestion usually drive the overage that shows up as a surprise
  • Custom metric cardinality is the single most expensive mistake — tagging one metric with a high-cardinality field like user_id can multiply its billed timeseries count 10–100x
  • Metrics Without Limits lets you query high-cardinality tags without paying to index every combination — apply it before you ever add a risky tag
  • Logs bill twice: once to ingest, once to index. Most teams index far more than they ever query — exclusion filters routinely cut the indexed share to 15–25% of ingest
  • APM spans have the same ingest/index split as logs. Ingestion-time (head-based) sampling at the tracer cuts the bill before Datadog ever sees the volume
  • Annual commitments save 10–30% over on-demand rates, but bursting past a commit tier costs the full on-demand rate on the overage — and over-committing wastes the unused allowance just as badly
  • For some teams, CloudWatch, self-hosted Grafana + Prometheus, or an OpenTelemetry pipeline into cheaper storage is a legitimately better answer once observability spend outgrows what a single SaaS vendor should cost
Analytics dashboard charts representing observability monitoring spend

Why host count is the number people watch — and the wrong one

Datadog's original pitch was infrastructure monitoring priced per host, so host count is still the metric that shows up on the account overview, gets reported to finance, and gets negotiated hardest at renewal. For a traditional fleet of long-lived VMs, that was a reasonable proxy for cost. It stopped being one once workloads moved to containers, autoscaling groups, Kubernetes nodes, serverless functions, and ephemeral CI runners — none of which map cleanly to "a host" — and once Datadog's own product surface expanded to roughly 20 separately metered SKUs covering infrastructure, APM, logs, RUM, synthetics, security, and more.

Observability and monitoring tooling has become one of the fastest-growing non-compute line items on a cloud bill, and broad industry benchmarking (Flexera's 2026 report) puts overall cloud waste at roughly 29%. A meaningful share of that waste hides specifically in observability spend, because unlike compute or storage, nobody reviews a "custom metrics" line item in a monthly infrastructure review — it doesn't correspond to a resource anyone owns. A host someone can right-size. A spike in custom metric cardinality from a code change three deploys ago is invisible until the invoice lands.

The fix starts with reading the bill by SKU instead of by host count. Once you can see which line items are growing independently of your infrastructure footprint, the rest of this guide is about closing each one down.

Anatomy of a Datadog bill: the SKUs that actually matter

Most of a Datadog invoice comes from a handful of SKUs, not all 20. The table below lists illustrative 2026 list prices (annual commit, Pro-tier equivalents where applicable) for the ones that move real bills — treat every number as a placeholder for your own rate card.

SKUIllustrative list priceWhat drives it up
Infra host~$15/host-monthFleet size — usually the least volatile line
Custom metrics~$5 per 100 metrics/month beyond the included allowanceTag cardinality, not metric count — the biggest surprise on this list
Log ingestion~$0.10/GB ingestedRaw volume shipped, regardless of whether it's ever read
Indexed log events~$1.70–$2.55 per million events/month (15–30 day retention)Percentage of ingested logs flagged for indexing — usually the real driver
APM host~$31/host-monthHosts running traced services
Ingested/indexed spans~$0.10/GB ingested, ~$1.50–$2/million spans retained at full fidelityTrace volume and sampling rate at the tracer
Synthetic tests~$0.0012/API test run, ~$0.0075/browser test runTest frequency and browser-test overuse
RUM sessions~$1.50 per 1,000 sessionsSampling rate — most teams don't need 100% of sessions
Serverless invocations~$5 per million function invocations monitoredInvocation volume on high-frequency, low-value functions
Where a Typical Datadog Bill GoesInfra hosts30%Indexed logs25%APM + spans20%Custom metrics15%RUM6%Synthetics4%

Illustrative composite split for a mid-size account, built to match the worked examples in this article — not survey data. Your own split depends entirely on how aggressively you've tagged metrics and indexed logs.

Notice that the three SKUs behind the thesis of this article — indexed logs, APM/spans, and custom metrics — add up to roughly 60% of the illustrative bill above, well ahead of the infra host line most negotiations target. That's the pattern worth internalizing: hosts are the stable, boring part of the bill. The volatile part is almost always somewhere else.

The cardinality trap in custom metrics

Datadog doesn't bill custom metrics by metric name — it bills by unique timeseries, which means every distinct combination of a metric name and its tag values counts separately. A metric like checkout.latency tagged only by region might produce a handful of timeseries. The same metric tagged by user_id produces one timeseries per unique user who ever triggers it, and that count only grows as your user base grows.

This is the cardinality trap: a single line of code adding one high-cardinality tag to one metric can turn a five-dollar line item into a five-figure one, and it usually ships without anyone noticing until billing flags it weeks later. The worked example later in this article walks through exactly that scenario with real numbers.

Metrics Without Limits is Datadog's answer to this: it lets you ingest a metric with its full tag set for querying ad hoc, while choosing a smaller set of tags that actually get indexed and billed as separate timeseries. Applied correctly, a team can keep user_id-level detail available for a one-off investigation without paying to permanently index every user as a distinct billed series. The catch is that it has to be set up before the cardinality explosion happens, or applied as the fix once it's caught — it doesn't retroactively refund the metrics already billed that month.

Two habits prevent most cardinality incidents: never tag a custom metric with anything that scales with your user base or request volume (user IDs, request IDs, email addresses, raw IP addresses), and run a periodic audit of custom metrics against the dashboards and monitors that actually reference them — metrics nobody queries are pure waste regardless of their cardinality.

Logs: ingest vs. index, and where the fat is

Datadog logs bill twice, and the two charges measure different things. Ingestioncharges for every log line you ship, regardless of what happens to it next. Indexingcharges separately, per event, for the subset you choose to make searchable and keep at a given retention window. Most teams default to indexing everything they ingest, which means they pay full price to make permanently searchable a mountain of health-check noise, verbose debug output, and duplicate log lines nobody will ever query.

Exclusion filters are the lever: they run at the point logs enter the pipeline and decide what gets indexed, without affecting ingestion. A filter that drops health-check 200s, known-noisy debug namespaces, and duplicate retry logs typically cuts the indexed share from close to 100% of ingest down to 15–25%, while ingestion cost (already the cheaper of the two charges) stays flat. For logs you're required to retain for compliance but rarely query, Flex Logs (or an equivalent low-cost archive tier) stores them searchably at a fraction of the indexed-log price, at the cost of slower query latency when you do need them.

Sampling is the complementary move for genuinely high-volume, low-value log streams — access logs on a high-traffic endpoint, for instance — where even exclusion from indexing doesn't make sense because the ingestion volume itself is the problem. Sampling a fixed percentage at the source (in the logging library or sidecar, before Datadog ever sees it) cuts both the ingest and index charges proportionally.

These same ingest/index economics show up on every major logging platform, not just Datadog — our guide to cutting logging costs across vendors covers the pattern in more detail if you're comparing Datadog against Splunk, Elastic, or CloudWatch Logs specifically.

APM and trace spans: sampling vs. retention filters

APM pricing layers the same ingest/index pattern on top of a per-host APM charge. You pay a flat rate per APM-enabled host, then pay again for the volume of trace spans ingested, and again for whichever subset of spans you choose to retain at full fidelity for searching and analytics. The two levers that control this are sampling and retention filters, and they operate at different points in the pipeline.

Head-based (ingestion-time) sampling happens in the tracer, before a trace ever leaves your service. You decide up front — say, keep 10% of traces — and the decision is made without knowing whether that particular trace was interesting (a slow request, an error). It's simple, cuts ingestion volume (and therefore cost) at the source, but risks dropping the rare, high-value traces you most want to see.

Retention filters operate after ingestion: you ingest everything (or a high percentage), then apply rules that keep 100% of error traces and slow requests at full retention while sampling down the high-volume, uninteresting bulk — successful, fast, routine requests. This is closer to what's often called tail-based sampling in spirit: the decision about what to keep accounts for how the trace actually turned out. It costs more at the ingestion stage than head-based sampling, but avoids throwing away the traces most useful for debugging.

The practical approach for most teams is a blend: apply moderate head-based sampling to cut raw ingestion volume on very high-throughput services, then layer retention filters on top so errors and outliers are never sampled away before they can be flagged for full retention.

Ten levers, ranked by savings per hour of work

Roughly ordered from fastest payoff to longest-term structural change. Savings ranges are illustrative percentages of the relevant line item, not of the total bill — apply them to whichever SKU each lever targets.

  1. Log exclusion filters on noisy sources. An hour of work on health-check, debug, and retry-log patterns. Savings: 30–70% of the indexed-logs line.
  2. Metrics Without Limits + drop risky tags. Audit custom metrics for high-cardinality tags (user IDs, request IDs) and either exclude them from indexing or remove them outright. Savings: 20–90% of the custom-metrics line, depending on how bad the sprawl is.
  3. APM ingestion sampling with error/slow-request retention filters. A half-day of tracer config. Savings: 40–80% of span ingestion and indexing cost.
  4. Reduce RUM session sampling rate. Most products don't need 100% of sessions recorded to spot real UX regressions. Savings: directly proportional — sampling to 30% cuts the RUM line roughly 70%.
  5. Delete unused custom metrics and dashboards. Cross-reference metrics against monitors and dashboards that actually query them. Savings: 10–30% of the custom-metrics line, with zero ongoing effort once done.
  6. Move cold, rarely-queried logs to Flex Logs / archive tier. Keeps compliance retention intact at a fraction of the indexed price. Savings: 15–40% of the portion of indexed logs that's compliance-driven rather than operationally queried.
  7. Right-size host count and remove orphaned agents. Decommissioned instances and duplicate container agents keep reporting long after the workload is gone. Savings: 5–15% of the infra-host line, with no ongoing maintenance.
  8. Trim redundant synthetic tests and test frequency. Consolidate overlapping checks and reduce browser-test frequency on low-traffic paths. Savings: 20–50% of the synthetics line.
  9. Right-size annual commitment tiers. Requires a few months of usage data. Savings: 10–30% versus on-demand, but only if the committed tier matches actual usage — see the commitment section below.
  10. Consolidate overlapping observability tools. The longest-term move: if you're running Datadog alongside a second APM or logging tool out of habit or acquisition debt, picking one eliminates duplicate ingestion entirely. Savings: variable, but often the single largest number on this list once untangled.

Two worked examples with full dollar math

Example 1: a 120-host startup with cardinality sprawl. Assumptions: 120 hosts on an infra Pro plan at an illustrative $15/host-month, each generating around 150 custom metrics on average (business and infrastructure metrics combined), with each host's plan including a free allowance of 100 custom metrics.

Line itemBefore tagging incidentAfter adding user_id to 5 metrics
Infra hosts (120 × $15)$1,800/mo$1,800/mo (unchanged)
Custom metric series18,000 total (120 × 150)~93,000 (5 metrics × 15,000 monthly active users replace the original 5)
Included allowance12,000 (120 × 100)12,000 (unchanged)
Billable overage6,000 series → 60 blocks of 10081,000 series → 810 blocks of 100
Custom metrics cost (×$5/block)$300/mo$4,050/mo

A developer debugging per-user latency adds a user_id tag to five checkout-flow metrics. With roughly 15,000 monthly active users, those five metrics stop being five timeseries each and become 15,000 timeseries each — 75,000 new series replacing the original five. Total custom metric count jumps from 18,000 to about 93,000, overage jumps from 6,000 to 81,000 billable series, and the custom-metrics line goes from $300/month to $4,050/month— a $3,750/month increase, roughly 13x, from one tag on five metrics.

The fix: apply Metrics Without Limits to exclude user_id from the indexed tag set (keeping it queryable only for ad hoc investigation, not permanently billed) and remove the tag once the debugging session is done. That restores the custom-metrics line to roughly its original $300–400/month, recovering the full $3,650–3,750/month delta.

Example 2: a log-heavy platform team. Assumptions: a platform team ingests roughly 5 billion log events per month across their services (about 2,500 GB at a typical ~500 bytes/event), and — like most teams starting out — indexes everything by default at 15-day retention.

Line itemIndex everything (default)After exclusion filters (~20% indexed)
Ingestion (2,500 GB × $0.10)$250/mo$250/mo (unchanged — filters act at index time)
Indexed events5,000M (100% of ingest)1,000M (~20% of ingest)
Indexing cost (×$1.70/million)$8,500/mo$1,700/mo
Total logs bill$8,750/mo$1,950/mo

Applying exclusion filters to drop health-check 200s, verbose debug namespaces, and duplicate retry logs from the indexed set — without reducing what's ingested, so nothing is lost for later re-indexing if needed — brings indexed volume down to roughly 1 billion events a month. The logs bill drops from $8,750/month to $1,950/month, a $6,800/month reduction, about 78%. Rebuild both examples with your own host count, metric tagging habits, and log volume — the mechanics transfer directly even when the raw numbers don't.

Annual commitments, overage, and the over-commit trap

Datadog, like most usage-billed SaaS, prices annual commitments 10–30% below on-demand rates in exchange for committing to a minimum spend across a set of SKUs for the year. For a team with stable, well-understood usage, that discount is close to free money. For a team still finding its footing on custom metrics or log volume, it's a bet that can go either way.

Bursting past a committed tier bills the overage at the full on-demand rate, not the discounted committed rate — so a team that commits to a tier sized for average usage and then has a cardinality incident (see Example 1 above) pays on-demand pricing on top of the explosion, compounding the damage. Over-committing is the quieter version of the same mistake: paying up front for a usage tier the team never grows into wastes the unused allowance just as thoroughly as an overage wastes the on-demand premium, just less visibly, because it never shows up as a spike on an invoice.

The practical rule: don't commit to a tier until you have at least two or three months of stable usage data to size it against, build in headroom for growth you can actually forecast rather than growth you hope for, and revisit the commitment at renewal rather than letting it auto-renew at a tier set a year earlier under different usage patterns.

See exactly which SKU is driving your observability spend

spendark pulls your cloud and tooling bills apart by line item so you can tell whether a custom-metrics spike, an indexing change, or a genuine host-count increase is behind a jump — before you spend a renewal negotiation guessing.

When Datadog is the wrong tool for the budget you have

Not every team should optimize a Datadog bill — some should leave, and it's worth saying honestly when that's true rather than only ever selling the next lever.

CloudWatch makes sense if your workload is almost entirely AWS-native, your dashboards are simple, and you don't need cross-cloud or on-prem visibility — you give up Datadog's unified APM/logs/metrics correlation and its much better UI, but for an AWS-only shop with modest needs the savings can be substantial. Our breakdown of CloudWatch's own pricing covers its ingestion and custom-metrics charges, which follow a similar shape to Datadog's.

Self-hosted Grafana + Prometheus (plus Loki for logs) trades Datadog's subscription cost for infrastructure cost and operational ownership — you pay for the servers and the on-call burden of running your own observability stack, but at high enough volume the compute cost of self-hosting undercuts per-GB and per-metric SaaS pricing significantly. The trade is real: you lose Datadog's polish, its integrations, and the time your team would otherwise spend on product work.

An OpenTelemetry pipeline into cheaper storage (routing traces, metrics, and logs through an OTel collector into a lower-cost backend, keeping Datadog or another SaaS only for the subset of data that benefits from its analysis tools) is the middle path: you keep vendor flexibility and avoid being fully locked into one pricing model, at the cost of running and maintaining the collector pipeline yourself.

The honest trigger for considering any of these: when observability spend crosses roughly 10% of total infrastructure spend and keeps growing faster than the business, it's worth building the comparison rather than applying the tenth incremental lever to the SaaS bill. For most teams below that line, the levers above get you most of the way there without the migration cost.

Frequently asked questions

Why is my Datadog bill so much higher than I expected?

It's almost never host count. The usual culprits are a spike in custom metric cardinality (a tag like user_id added to even a few metrics), a jump in indexed log volume (as opposed to ingested volume, which is cheaper), or increased APM span ingestion from a new or more heavily traced service. Read the bill by SKU, not by host count, to find which one moved.

What is Metrics Without Limits and how does it save money?

It lets you ingest a custom metric with its full tag set for ad hoc querying while choosing a smaller set of tags that actually get indexed and billed as separate timeseries. It prevents a high-cardinality tag (user IDs, request IDs) from multiplying your billed metric count, which is the single most common cause of a custom-metrics cost spike.

What's the difference between log ingestion and log indexing costs?

Ingestion charges for every log line shipped to Datadog, at a flat per-GB rate. Indexing charges separately, per event, for whichever subset you make searchable at a chosen retention window — and indexing is the more expensive of the two. Exclusion filters reduce what gets indexed without touching what gets ingested, which is usually where most of the savings sit.

Should I use head-based or tail-based sampling for APM spans?

Head-based sampling decides at the tracer, before a trace is known to be interesting, and cuts ingestion cost at the source but risks dropping rare error traces. Retention filters (Datadog's version of tail-aware sampling) ingest more but keep 100% of errors and slow requests while sampling down routine traffic. Most teams get the best cost-to-visibility ratio from a blend of both.

When should a team move off Datadog to CloudWatch or a self-hosted stack?

When observability spend is a large and fast-growing share of total infrastructure cost — roughly 10% and climbing is a reasonable trigger to investigate — and the workload fits a narrower tool well: CloudWatch for AWS-only shops, self-hosted Grafana + Prometheus + Loki at high enough volume to justify the operational cost, or an OpenTelemetry pipeline for vendor flexibility. Below that threshold, applying the cost levers in this article usually closes most of the gap without a migration.

Estimate your cloud costs — for free

Compare AWS, Azure, and GCP pricing side by side with our free calculator, and dig into the guides to learn how to cut cloud waste. No sign-up required.