Self-hosting an LLM usually wins on cost only past a real, sustained volume threshold — roughly the point where you're saturating GPU capacity most hours of the day, not bursting into it. Below that line, idle GPU-hours, MLOps headcount, and reliability engineering erase the per-token savings. Above it, self-hosting can cut unit economics substantially, but only if utilization stays high.
Quick answer: Self-hosting tends to beat API pricing once you're running sustained inference at meaningfully high, consistent utilization — often cited informally around hundreds of millions of tokens per month for mid-size open models. Below that, or with spiky traffic, hosted APIs are cheaper once you count idle compute and ops labor.
Most build-vs-buy debates about LLM inference get framed as a GPU spec sheet argument: which card, which quantization, which serving framework. That's the wrong first question. The right first question is a total cost of ownership question, and TCO for inference behaves nothing like TCO for a web server. GPUs don't degrade gracefully under low utilization — they just burn money while idle. This piece walks through the real cost stack on both sides, gives you a directional breakeven, and flags the utilization pattern that should end the debate before it starts.
What self-hosting actually costs beyond the GPU price tag
The sticker price of a GPU-hour is the smallest line item in a realistic self-hosting budget. The larger, harder-to-see costs are utilization inefficiency, the MLOps headcount required to keep serving infrastructure healthy, and the reliability engineering needed to hit production SLAs. Teams that only model "GPU-hour times hours needed" consistently underestimate true cost by 2-4x.
Break the stack into four buckets:
- Compute capacity. Reserved or on-demand GPU instances (A100s, H100s, or cheaper inference-optimized silicon), sized for peak load, not average load.
- Utilization drag. GPUs provisioned for peak traffic sit mostly idle during off-peak hours — and you pay for the whole reservation regardless.
- MLOps and platform headcount. Someone has to manage model versioning, quantization, batching, autoscaling, and upgrades as new model checkpoints ship.
- Reliability engineering. Failover, load balancing across regions, monitoring, and incident response for a system that, unlike a stateless API call, holds a live model in GPU memory.
Key term:
utilization rate— the percentage of provisioned GPU-hours actually used for productive inference. This single number swings self-hosting economics more than any pricing negotiation.
The utilization trap
A GPU provisioned for a traffic peak that happens two hours a day is idle 22 hours a day. If your workload is spiky — customer support bursts at 9am, a marketing campaign spike, seasonal demand — you're paying full reservation cost for single-digit utilization. This is the single most common reason self-hosting projects blow their business case in year one.
Contrast that with API providers, who pool demand across thousands of customers and can run their own GPU fleets at far higher aggregate utilization than any single company's workload alone would justify. That pooling is the entire economic argument for the hosted model, and it's why low or spiky utilization almost always favors APIs — you're renting a slice of someone else's high-utilization fleet instead of building your own low-utilization one.
What the API side actually costs beyond the per-token rate
Hosted APIs look expensive per token compared to raw GPU-hour math, but that comparison ignores what you're buying: elasticity, zero ops burden, and instant access to frontier model upgrades. The real API cost stack is the per-token rate plus the engineering discipline required to keep spend predictable at scale.
The per-token price is the visible cost. The costs teams miss are:
- Spend variance. Usage-based pricing means a traffic spike or a poorly bounded agent loop can produce a surprise bill.
- Vendor lock-in risk. Prompts, tool schemas, and evals tuned to one provider's model behavior don't always port cleanly to another.
- Latency and rate-limit ceilings. Shared infrastructure means you're subject to provider-side throttling during their peak demand, not just yours.
- Data residency and compliance constraints. Some regulated workloads simply can't send data to a third-party endpoint regardless of price.
None of these costs are fixed — they're manageable. Techniques like semantic caching of repeated LLM responses and prompt caching mechanics can cut effective per-token spend on the API side substantially, often more cheaply than standing up your own infrastructure to chase the same savings. Model routing — sending easy queries to a cheap, fast default model and reserving frontier models for hard cases — is the other lever; see how LLM model routing between cheap and default models works in practice.
Where API elasticity genuinely wins
APIs shine when demand is unpredictable, when you need to swap in a new model within days of release, or when your team has no bandwidth to run a serving stack. Elasticity has a price, but it's the price of not building a factory to make one product.
The breakeven: when self-hosting actually pays off
Self-hosting starts to make financial sense when three conditions hold simultaneously: sustained high utilization, a stable model choice you're not swapping monthly, and volume large enough to amortize a dedicated MLOps function. Absent any one of these, the API's convenience premium is usually cheaper than it looks.
There's no single universal token number — hardware costs, model size, and negotiated API rates all shift the line — but directionally, teams and infrastructure vendors (including analyses from a16z and Andreessen Horowitz's LLM cost-modeling writeups, and benchmarking work published by MosaicML/Databricks) have converged on a rough pattern: self-hosting mid-size open-weight models starts to beat API pricing once monthly volume reaches the hundreds of millions of tokens range and utilization stays above roughly 50-70% of provisioned capacity.
| Factor | Favors self-hosting | Favors API |
|---|---|---|
| Monthly token volume | Very high, sustained | Low to moderate |
| Traffic pattern | Flat, predictable | Spiky, bursty, seasonal |
| Model stability | Fixed model, infrequent swaps | Wants latest frontier model often |
| MLOps headcount available | Dedicated team exists | No dedicated ML infra team |
| Compliance/data residency | Hard requirement to keep data in-house | No such constraint |
| Latency sensitivity | Needs custom hardware tuning | Standard latency acceptable |
A simple breakeven framing to run internally
- Estimate monthly token volume at 6- and 12-month horizons, not just today.
- Model utilization honestly — use your actual traffic curve, not an averaged daily rate.
- Price the fully-loaded GPU cost: instance cost, plus a fractional MLOps salary, plus reliability engineering time.
- Compare against your actual API spend after applying caching and routing optimizations, not list price.
- Re-run the model at the volume you expect in 12 months, since inference costs and open-model quality both shift fast.
Treat this as a directional exercise, not a spreadsheet you defend to the cent — the inputs (GPU pricing, API rates, open-model quality) are moving targets, and the honest output is a range, not a single number.
Reliability and talent: the hidden variable most TCO models skip
Self-hosting inference means owning uptime for a stateful, memory-resident system — a fundamentally harder reliability problem than fronting a stateless API call, and one that requires ML infrastructure talent that's scarce and expensive relative to general backend engineering.
When an API provider has an outage, it's their incident, their status page, their postmortem. When your self-hosted cluster has an outage, it's your on-call rotation, your postmortem, and often your customers watching a support ticket queue back up. Reliability isn't a line item you can skip — it's a standing commitment.
Google's Site Reliability Engineering practice (the SRE book and its error-budget framework) is the standard reference here: it treats reliability as an explicit budget you spend deliberately, not an afterthought. Apply that lens to self-hosted inference and the honest question becomes: does this team have the standing capacity to run an error budget for a GPU-backed service, or would that capacity be better spent on product work?
The talent cost nobody prices upfront
MLOps engineers who can tune batching, manage quantization tradeoffs, and debug GPU memory fragmentation under load are a narrower talent pool than general infrastructure engineers. Hiring for this function isn't just a salary line — it's a recruiting timeline that can stall a self-hosting decision by quarters.
Anchoring the decision before it becomes a spec argument
Build-vs-buy debates for infrastructure decisions have a predictable failure mode: they turn into an argument about GPU specs, quantization schemes, or which serving framework is fastest, before anyone has agreed on what actually matters most for this decision — cost, speed to market, control, or risk tolerance. That's a sequencing error, not a technical one.
Prodinja's Trade-off Triangle exists for exactly this moment. It's designed to force the conversation to name its dominant constraint — cost, speed, or control — before the debate splinters into competing technical arguments that each optimize for a different, unstated priority. Naming the constraint doesn't replace the TCO math above; it makes sure the team doing that math agrees on what "winning" even means before someone opens a GPU pricing spreadsheet. This kind of framing pairs naturally with a broader trade-off analysis discipline for product decisions, which covers the same anchor-before-you-argue principle across other build-vs-buy calls a PM will face.
Key Takeaways
- GPU sticker price is the smallest cost in self-hosting — utilization drag, MLOps headcount, and reliability engineering dominate real TCO.
- Low or spiky utilization almost always favors APIs, because providers pool demand across customers to reach utilization no single company's workload can match alone.
- Self-hosting breakeven is a volume-and-consistency question, not a per-token price question — directionally, hundreds of millions of tokens per month at sustained high utilization is the rough zone where it starts to pay off.
- Caching and routing can shrink API costs significantly before you ever need to consider owning infrastructure — cheaper to implement than a GPU fleet.
- Reliability for self-hosted inference is a standing commitment, not a one-time setup cost — budget for it like an SRE error budget, not a launch checklist item.
- Name the dominant constraint (cost, speed, or control) before debating hardware specs — it keeps the team arguing about the same thing.
Frequently Asked Questions
Is self-hosting an LLM cheaper than using an API?
Only above a real utilization and volume threshold — for most teams with moderate or spiky traffic, hosted APIs remain cheaper once idle GPU-hours and MLOps labor are counted. Self-hosting pays off mainly at sustained, high-volume, high-utilization workloads.
What volume of usage justifies self-hosting an LLM?
There's no universal number, but directionally, teams tend to see self-hosting pay off once monthly volume reaches the hundreds of millions of tokens range with utilization sustained above roughly 50-70% of provisioned capacity. Below that, API convenience and pooled-demand pricing usually win.
Does self-hosting give better data privacy than an API?
Self-hosting can offer stronger data residency and compliance control since data never leaves your infrastructure, which matters for regulated workloads. That control comes with the tradeoff of owning reliability and security operations for the hosted system yourself.
Can caching or routing reduce API costs enough to avoid self-hosting?
Yes — semantic caching and model routing (sending simple queries to cheaper, faster models) can meaningfully cut effective per-token API spend, often more cheaply than standing up dedicated inference infrastructure. Many teams should exhaust these levers before evaluating self-hosting at all.
What's the biggest hidden cost in self-hosting LLMs?
MLOps headcount and reliability engineering, not GPU pricing. Running a stateful, memory-resident inference service reliably requires scarce ML infrastructure talent and a standing operational commitment comparable to an SRE error-budget practice.