A flat monthly LLM API bill tells you nothing about margin. Real token cost FinOps means tagging every request with feature, tenant, model, and prompt version, then rolling those tags into a dimensional cost model—so you can see which features and customers are profitable before finance asks why AI gross margin dropped ten points.

Quick answer: Attribute every LLM call to four dimensions—feature, tenant, model/version, prompt version—at the point of the request, not after the fact. Without that, you're managing cost with a single number that hides where the money actually goes.

Why a Flat API Bill Is a Governance Failure, Not Just a Reporting Gap

A single line item from OpenAI, Anthropic, or your inference provider tells you total spend, not unit economics. That's the same mistake as reporting total AWS spend without knowing per-customer infrastructure cost—except LLM spend scales with usage in a much more volatile, less predictable way than compute.

Unit economics for an AI feature means cost per request, per user, per outcome—not cost per month. Traditional SaaS cost allocation assumed relatively flat marginal cost per user. Generative AI breaks that assumption: a single power user running long chains-of-thought or large-context RAG queries can cost 50-100x a typical user in the same billing period.

This isn't a hypothetical governance nice-to-have. It's the difference between:

  • Knowing your AI copilot feature has healthy 70% gross margin, versus
  • Discovering three months in that one enterprise tenant's usage pattern has quietly pushed a "included" feature to negative contribution margin.

The LLMOps and AI observability discipline that tracks latency, quality, and drift needs a cost lane too—and most teams build the quality and latency dashboards first, leaving cost as an afterthought until finance escalates it.

Build a Cost-Attribution Dimensional Model Before You Build a Dashboard

A cost-attribution model works like a data warehouse star schema: a fact table of token events (timestamp, tokens in/out, model, latency, cost) joined to dimension tables—feature, tenant, prompt version, and model version. Without this structure at ingestion time, you cannot retroactively attribute historical spend accurately.

The four core dimensions

DimensionWhat it capturesWhy it matters for margin
FeatureWhich product surface triggered the call (e.g., summarize-doc, agent-triage)Lets you compute margin per feature, not just per app
Tenant/customerWhich org or account owns the requestSurfaces which accounts are subsidized vs. profitable
Model + versione.g., claude-sonnet-4-5 vs. a cheaper fallback modelQuantifies the cost of quality/latency tradeoffs
Prompt versionWhich prompt template or chain version ranTies cost regressions to a specific deploy, not a mystery

Each LLM call should emit a structured event carrying all four tags plus token counts (input, output, cached) and provider-reported cost. This is a data modelling problem before it's a FinOps problem: get the entity relationships right—request, feature, tenant, prompt_version, model_version—and the rollups become simple aggregations rather than forensic reconstructions.

Why prompt version matters more than teams expect

Prompt engineering changes are usually tracked in a changelog.md file, if at all. But a prompt-version tag on every cost event lets you answer: "Did the v14 prompt rewrite that fixed hallucinations also double our output tokens?" Without it, cost regressions get attributed to "the model got more expensive" when it was actually a verbose system prompt shipped two sprints ago.

Treat every prompt change as a versioned artifact with its own cost signature, the same way you'd version an API endpoint schema.

From Token Counts to Unit Economics: The Calculation Chain

Unit economics requires converting raw token counts into a per-request cost, then rolling that up to cost-per-feature and cost-per-tenant, before comparing against the revenue or value each generates. Skipping any link in this chain produces numbers that look precise but mean nothing.

The calculation chain, in order:

  1. Raw token cost: (input_tokens × input_rate) + (output_tokens × output_rate), per provider pricing tier.
  2. Effective cost with caching discounts: prompt caching (available on Anthropic's and OpenAI's APIs) can cut input costs 50-90% for repeated context—exclude this and you overstate cost by a lot.
  3. Cost per successful outcome: divide total cost for a feature by completed, accepted outputs—not raw calls—so retries and failed generations count against the denominator honestly.
  4. Cost per active user or tenant: aggregate #3 across a billing period per tenant.
  5. Contribution margin: subtract #4 from the revenue or allocated value that tenant/feature generates.

A worked example: the power-user cohort problem

Say an AI writing-assistant feature serves 10,000 monthly active users at an average cost of $0.40/user/month—a healthy number against a $15/seat price point. But when you slice by usage percentile, a different picture emerges:

User cohort% of users% of token spendAvg cost/user
Light (1-10 requests/mo)70%12%$0.07
Medium (11-50 requests/mo)25%28%$0.45
Power (50+ requests/mo)5%60%$4.80

The blended average hides that 5% of users consume 60% of spend. If those power users are on a flat-rate plan, this cohort is your effective loss leader—invisible until you attribute at the tenant-percentile level. This is precisely the kind of distortion that flat "total API spend / total users" reporting cannot surface, and precisely why per-dimension attribution matters more than aggregate cost tracking.

What to do about it:

  • Usage-based tiering: introduce metered pricing or overage fees once a tenant crosses a token threshold—common practice among API-metered SaaS products.
  • Feature-level throttling or downgrade: route power-user requests to a cheaper model tier (e.g., Haiku-class instead of Opus-class) once a soft cap is hit, if quality tolerance allows.
  • Prompt/context optimization for high-frequency paths: power users often hit the same feature repeatedly—this is where prompt caching and context trimming pay off disproportionately. See the context engineering discipline for managing what actually goes in the prompt for how uncontrolled context growth silently inflates cost alongside quality risk.
  • Cohort-aware pricing conversations: flag the cohort to sales/success before renewal, not after a margin post-mortem.

Wire Cost Into the Same Observability Layer as Quality and Latency

Cost should be the third pillar alongside quality and latency in your AI feature's day-one dashboard, not a separate spreadsheet finance builds quarterly. Treating cost as a first-class observability signal means it degrades gracefully into alerts, not a surprise line item.

Most teams under-invest here because cost feels like a finance problem rather than a product/engineering one. It isn't—cost is a product design output: prompt length, retrieval strategy, model choice, and retry logic are all product decisions with direct cost consequences.

What to actually monitor

  • Cost per feature per day, trended against a target band (not just a raw number).
  • Cost per tenant, top-20 by spend, refreshed daily—this is where the power-user problem surfaces early.
  • Cache hit rate on cacheable prompt segments—a leading indicator of avoidable spend.
  • Cost-per-successful-outcome, not cost-per-call, so a spike in retries or failures shows up as a cost anomaly, not just a quality one.

This pairs directly with the discipline laid out in setting up the three dashboards an AI product needs from day one—cost, quality, and latency should share the same event schema so a spike in one can be correlated against the others without a separate data-pull.

And when a cost anomaly does trace back to a quality regression—say, a model hallucinating longer, more token-hungry responses—the fix belongs in your eval suite, not a one-off patch. Closing the loop from production failures to eval test cases is how a cost-driven incident becomes a permanent regression guard instead of a recurring fire drill.

Turn Cost Data Into Prioritization Inputs, Not Just a Postmortem

Cost attribution only pays off when it changes a decision—specifically, when it feeds the reach, impact, and effort inputs of your feature prioritization framework, so a cost-heavy AI feature is weighed against its actual value rather than judged in a vacuum.

Once you have per-feature token economics—actual cost-per-outcome, not a guess—you can answer the question that usually gets fudged in planning: "Is this AI feature worth what it costs to run at scale?" A feature might have great engagement (high reach) but if its effort input (including ongoing inference cost, not just build cost) is quietly 10x a comparable feature, that changes the prioritization math.

This is exactly the gap Prodinja's RICE and Kano prioritization tooling is built to close: once you have real per-feature token-economics numbers from the attribution model above, you can feed them into the reach and effort inputs and see the cost-adjusted priority score shift—rather than prioritizing AI features on engagement alone and discovering the margin problem after ship. It's not a magic cost calculator; it's a place to put the numbers you've already earned through attribution so they actually influence the roadmap.

Key Takeaways

  • A flat LLM API bill is not a cost model—it hides which features and tenants are actually profitable.
  • Build a dimensional cost model: feature, tenant, model/version, and prompt version, tagged at the point of the API call.
  • Calculate cost per successful outcome, not per raw call, and account for prompt-caching discounts before comparing periods.
  • Power-user cohorts routinely consume 50-60%+ of token spend while representing a small fraction of users—segment by percentile, not average.
  • Wire cost into the same observability pipeline as quality and latency so anomalies correlate instead of living in separate spreadsheets.
  • Use real per-feature unit economics as an input to prioritization frameworks like RICE and Kano, not as a postmortem exercise.

Frequently Asked Questions

How do you attribute LLM costs to specific features?

Tag every request at call-time with a feature identifier, then join token counts and provider-reported cost against that tag in your event pipeline. Retrofitting attribution after the fact is unreliable because provider billing rarely preserves per-request metadata—capture it yourself at the source.

What's the difference between cost per call and cost per outcome?

Cost per call averages spend across every API request, including retries and failed generations, understating the true cost of a working result. Cost per successful outcome divides total spend by only the accepted, completed outputs—the number that actually reflects unit economics finance cares about.

How much does prompt caching actually save on LLM costs?

Directionally, prompt caching can cut input-token costs by roughly 50-90% for repeated context segments, depending on provider and how much of the prompt is cacheable versus dynamic (Anthropic and OpenAI both publish caching discount tiers). The savings compound most on high-frequency, high-context features—exactly where power-user cohorts concentrate spend.

Why do a small percentage of users consume most of the AI cost?

Usage in LLM features follows a heavy-tailed distribution similar to Pareto-style patterns long documented in usage analytics (a legacy of research from analysts like Clayton Christensen–adjacent jobs-to-be-done segmentation work, applied here to consumption rather than adoption): a small cohort with intensive, repeated, or long-context usage drives a disproportionate share of variable cost. Segment by percentile, not average, to see it—see the worked cohort example above for the mechanics.

Should cost attribution live in engineering, product, or finance?

It should be a shared schema that all three consume, owned operationally by whoever owns the LLMOps observability pipeline—usually platform or AI engineering—but reviewed jointly with product and finance monthly. Understanding your customers' actual jobs-to-be-done and usage journey alongside cost data helps distinguish "expensive but core to the job" usage from usage worth re-architecting or re-pricing, and mapping cost spikes against the customer journey shows whether high-cost moments coincide with high-value moments or just inefficient prompting.