FinOps for AI

Grounded Token Economics on Amazon Bedrock

Deterra = Deterministic + Terra (solid ground, earth).

Deterra focuses on one use case only: AI on Amazon Bedrock.
Deterra puts OpenTelemetry together with billed data (AWS Cost and Usage Report) the old-fashioned way: rule-based and deterministic.

Our hot takes:

  • You can't govern what you haven't measured yet.
  • Attributing estimated cost to business metrics is like budgeting your living expenses from your pre-tax income. The starting point is wrong.
  • With LLMs shipping at this speed, measurement is unglamorous grunt work. Other products paint you a rosy picture, then sneak in a disclaimer about “potential discrepancies from your bill.” We start with the bill. We end with the bill. We match the bill. Did I mention the bill?

LLM observability shows what happened.
Deterra shows what it actually cost

Many tools touch the same OpenTelemetry data, which are only estimates. Deterra joins the cost from those Signals with the cost in the Ledger.
One code path to join them all, and in darkness Attribute them.

LLM observability
Traces prompts, agents, tools, evals, and output quality.
Did the AI behave correctly?
Deterra
Reconciles telemetry against the bill and attributes it to teams, features, and users.
Who caused the spend, and does it match the bill?
APM
Connects AI calls to services, latency, errors, and incidents.
Is it healthy?
AI gateways
Routes models, caches calls, enforces limits. Built by engineers, for engineers.
Which model handles this?

Legacy FinOps renovates a platform built for a different cost primitive. Emerging AI FinOps tools go wide across model providers, stop at token count estimates, and reach for control before they've measured anything. The billing platforms focus on the model-direct sell side.

Deterra stubbornly ploughs the other way. From token counts, through logs and metrics, all the way down into the weeds of the billing file. One hyperscaler at a time. Building trust the only way it can be built: one correct number after another.

The signals scatter. The ground moves. We hold the number still.

Your cost signals sit in different places, speak different units, and change every time AWS ships. Chasing them down is unglamorous, sweaty, exhausting work. We built this, so you don't have to.

SCATTERED SIGNALS
RECONCILIATION
ATTRIBUTED SPEND
ONE RECONCILED VIEW
0 attributed
01 / RECONCILE

Signal vs Ledger

One shows spend trend and distribution in real time. One shows the final dollar you pay. The gap in between is where the money hides. It's the first thing you see in Deterra.

02 / ATTRIBUTE

Every dollar has a cause

Each token mapped to the unit that produced it. A user, agent, feature, prompt... metrics your business already runs on. Every token accounted for. Unit economics on solid ground.

03 / OPTIMIZE

Usage and rate down, human feedback in

Tokens made “optimal” a judgment call. A smaller model saves money until the answers get worse. Caching saves money unless the prompt changes before the cache pays back. Deterra puts the cost next to the trade-off, and a human makes the call.

04 / GOVERN

Enforcement, after evidence

Allowlists, ceilings, limits, agent RBAC. Set from what you measured, not from what you guessed.

A token is not a server-hour

Cloud FinOps was built to meter deterministic consumption. A vCPU-hour costs what a vCPU-hour costs. Infrastructure is provisioned by a small number of engineers, under a fixed budget, and it serves the whole business.

Tokens behave nothing like that. Everyone consumes them, everywhere, inside the org and out. And no two consumers cost the same.

Different prompts, different bills

One user summarizes a paragraph. The next pastes a 200-page contract.

Same feature. Same model. Same button. Two orders of magnitude apart.

No engineer made that choice. The user did.

Same prompt, different bills

A system prompt that quietly grew. A max-tokens ceiling nobody revisited. Extended thinking left on for a workload that never needed reasoning.

Extra cost stacks up quietly. You see the number. You don't see why.

A vCPU-hour is a unit of resource. You know how to pay for it, meter it, control it.

A token is a decision to spend, made by any user, at any time. You pay for it. You don't control it.

Finance can close the books on everything except AI

Every enterprise running AI on the cloud is having some version of the same conversation. It usually starts as a question and ends as a shrug ¯\_(ツ)_/¯

CFO:“I don't know where the AI cost went.”
CEO:“We're spending more on AI than ever. Margin hasn't moved.”
VP Engineering:“I shipped three AI features this year. Which one is making money?”
Head of Product:“Which users take up more AI cost? Are they my premium paying customers?”
Controller:“Finance can close the books on everything except AI. That line item is a mystery every single month.”

You can't govern what you haven't measured

We start with visibility and visibility only. Read-only. No proxy, no prompt caching. We connect through a single cross-account IAM role, scoped to the minimum permissions we need.

NOW

Visibility

Reconciled cost across every signal. Full attribution to business metrics. The gap between estimated and billed cost, surfaced, analyzed, proven. And then, we start to optimize.

NEXT

Optimization

Every optimization starts as a conversation. What's possible, what suits your use case, what survives simulation. Then we implement at scale.

LATER

Governance

Agent RBAC, policy enforcement, model allowlists, spend ceilings. One bill, one trail, and an audit history that compounds from day one.

We're deliberate on sequence and we take our time to build trust, because that's the only order this works in.

We are a small team building a sharp tool. We pair deep FinOps experience with solid engineering, and we are collecting hot field notes from the most advanced Amazon Bedrock adopters.

Your first call with us will be short and sweet. You have three types of AI cost: live, logged, and billed. We show all three on one page and make the math make sense.

If you want to try it out, we ask for read-only permissions to your CUR and OTel data. That's it.

Then, we explore further, together.

Logan Stone

Logan Stone

Co-Founder & CEO
  • AWS 5x certified, sales gone tech
  • SME of (AWS+FinOps+Partners)
  • Career TAM and that AWS billing guy
Shidi Zhao

Shidi Zhao

Co-Founder & CTO
  • BSc and MSc in CS, University of Toronto
  • Citibank → Amazon → DoorDash → Deterra
  • Software engineer and that math guy