AI cost engineering · For 50 to 500 person companies · $5k+/mo on AI model usage

Your inference bill grew again last quarter. The invoice won't tell you which feature spent it.

We find out which of your product features are spending the money, then cut what is safe to cut. Every number we report is one your finance team can recompute from your own provider invoices.

Book the free audit 30 minutes · written diagnosis · no card, no obligation
Typical outcome
20–50% cut
Without letting output quality slip.
How we are paid
A share of savings
25% of what your finance team verifies. No retainer.
Who does the work
Two senior engineers
the two people on this page. No bench, no handover.

Published research on AI spend

What you get

Four deliverables, in the order they have to happen.

Not a strategy deck. Working software in your own systems, and a savings number your finance team can check line by line. The order is fixed for one reason: you cannot cut a cost until you know which product feature is creating it.

01

Cost attribution

Every model call tagged to the product feature, the customer and the environment that triggered it. Your invoice stops being one number and becomes a list of features with a price next to each. This is also what makes ROI measurable per use case: once a feature has a cost, you can set it against the revenue or hours it returns and rank them.

always first
02

Cost saving opportunities

The cheapest wins that need no rework of your product.
Token discipline caps what goes into each call and what comes back, and clears out prompt examples that no longer pay for themselves. Caching restructures system prompts, retrieval templates and tool schemas so the repeated opening is one contiguous block billed at a fraction. Batching moves deadline-free jobs onto a discounted queue.

repeated content costs ~10% of standard
03

Cheaper models, equal performance

Not every request needs your most expensive model. We sort your traffic into job types, such as extraction, classification, scoring and rewriting, then move the ones that do not need top-tier reasoning onto cheaper models. Each move ships behind an automated quality test with a pass mark you set in week one, so nothing degrades quietly on the way to a lower bill.

60 to 90% cheaper per moved request
04

Purpose-built task model

Where one narrow, high-volume job dominates your bill, we train a small model to do only that job, then prove it matches your quality bar before it goes live. This is the biggest saving available and the one we propose least often, in roughly one engagement in three. Training and hosting cost is subtracted before we ever call the result a saving.

70 to 90% on that one workload

Process

Eight weeks, five phases, no surprises.

Each phase ends in something you keep and can show to someone else. You can stop after phase two with a written diagnosis and owe nothing.

Week 1 · We find out where the money goes

Read-only access to your provider billing exports and your gateway configuration. Every model call that passes through your gateway is tagged to a feature, an environment and, where there is one, a customer. Anything the tagging cannot reach is flagged rather than smoothed over. Your starting position is then frozen: cost per request, tokens per call, calls per feature. You also set the quality pass mark for each part of the product here.

You keep a frozen baseline and a cost dashboard · built from our public audit checklist

Week 2 · The three biggest leaks, written down and ranked

A short written diagnosis: where the spend is, how much is recoverable, what the fix is, and what it costs to do. Ranked by return per engineering day, so you can fund the top item and ignore the rest. Whatever week one could not attribute is explained here, and calls bypassing your gateway altogether are among the most common things we find. That one is often worth more than the first optimisation. If nothing on the list is worth doing, that is the finding and we say so in writing.

You keep the diagnosis memo · every figure costed at current prices, re-checked weekly

Weeks 2 to 5 · The safe cuts ship

Token discipline goes first: fewer tokens to cache, fewer to route, smaller bills on the models you keep. Then caching, batching and routing, landing as pull requests in your repository, reviewed by your engineers and merged on your schedule. The quality tests run before and after every change. Anything that cannot hold its pass mark does not ship, and the saving it would have produced is never counted.

You keep the routing layer, cache configuration and quality test suite

Weeks 5 to 7 · A purpose-built model, only if it pays for itself

Reached only when phase three has already proved that one job is big enough and narrow enough to justify training a model for it. Skipped more often than not, and we would rather tell you that in week five than bill you for it in week seven.

You keep the trained model, its quality report and a drift monitor

Week 8 · Handover, then we are gone

Operating notes, the quality tests running in your own pipeline, and a savings report your finance team can recompute from provider invoices line by line. One check-in 30 days later is included. Nothing of ours keeps running in your systems and there is no licence to renew.

You keep the verified savings report · signed by you before any invoice
Stack

We work inside the tools you already run.

No proprietary platform. No migration. These are the providers, gateways, and infra we operate in daily.

Model providers
OpenAIAnthropicGoogle Vertex AI Azure AI FoundryAWS BedrockMistral
Monitoring and quality testing
OpenTelemetry / OpenLLMetryLangfuse Datadog LLM ObservabilityBraintrustPromptfoo
Open-weight hosting and training
Together AIFireworks AIAmazon SageMaker vLLMHugging Face

Whatever framework your engineers build on, LangChain, LlamaIndex, LangGraph or something written in-house, is where we read how your calls are structured. It is not something we add to or ask you to change.


Modeled scenarios

What a Sprint is built to deliver.

Three worked examples against the company profiles we are built for. Every line is calculated from published provider prices and a stated assumption about how much of the traffic each change reaches. The spreadsheet is in our public repository, so you can check the arithmetic before you check our references.

scenario // A

AI-native SaaS · Series B · ~120 people

Assistant in product$18.4k$12.9k
Document classifier$11.2k$4.3k
monthly, down 42%$29.6k$17.2k
changes and reach
  • caching, with 55% of input repeated and hit 4 times in 5
  • purpose-built classifier on 70% of that workload
scenario // B

SaaS with AI added on · Series C · ~200 people

In-app copilot$14.8k$10.7k
Auto-tagging job$4.2k$2.6k
monthly, down 30%$19.0k$13.3k
changes and reach
  • 45% of copilot turns moved to a mid-tier model, output capped
  • discounted queue on 75% of tagging, which has no deadline
scenario // C

Content operations · bootstrapped · ~80 people

Client drafting$7.4k$5.3k
Translation pipeline$3.6k$2.2k
monthly, down 32%$11.0k$7.5k
changes and reach
  • caching on brief templates, output capped on first drafts
  • discounted queue plus a smaller model on cleared language pairs
what these numbers assume
  • Prices. Published list prices for the models named in the spreadsheet, as of the current tracker week. No negotiated or committed-spend discounts assumed.
  • Volume unchanged. Every figure is a change in cost per request applied to the same traffic. Growth is modeled separately.
  • Repeated content priced at about a tenth of standard, and the discounted queue at half, in line with current provider documentation.
  • Quality. Every moved request is assumed to pass its quality test. Requests that would fail stay on the original model and produce no saving.
  • Excluded. Engineering time, storage and our own fee. Scenario A is net of training cost; nothing else is.
  • What this is not. A forecast for your stack. How much of your content repeats, how much of your traffic can move and how much passes its quality test are unknown until week one, and those three numbers decide the outcome.
Fit

When to call us, and when not to.

Our fee only works when the savings are real and sizeable, which makes us the wrong choice more often than the right one. Reading the second column will save us both a call.

Worth a conversation

  • $5k or more a month on AI model usage, and growing
  • AI cost is now a visible line in your margin, and someone is being asked about it
  • Your traffic includes repetitive, high-volume, narrowly defined work
  • Model calls already pass through a gateway, or you are willing to add one
  • Engineering can review and merge pull requests on a normal cadence

We would decline

  • Under $5k a month. The fee would not cover two senior engineers, so nobody wins. Take the free audit findings and run them in-house.
  • Your AI cost is a fixed platform licence rather than metered usage. There is nothing per-request to optimize.
  • Your traffic is entirely low-volume, open-ended, human-reviewed work. No repeatable job types, nothing to move.
  • You are after headcount replacement or a general AI strategy engagement. Not what we do.
  • You have a committed-spend contract with a floor above current usage. Cutting consumption changes nothing until that is renegotiated, so renegotiate first.
Pricing

One fee, and only if it works.

Nothing to pay for the audit, no retainer, and no invoice at all until savings appear in your own provider bills and your finance team has signed off on them.

engagement 01

Spend Leak Audit

One call, then a written diagnosis. No obligation on either side.

$0
Free, whatever we find.
no card · no obligation · most audits do not lead to a Sprint, and we will tell you when yours should not
  • Cost per feature, estimated from your billing data
  • The three leaks we would expect in your setup, ranked
  • An honest read on whether a Sprint is worth doing at all
Book the free audit
engagement 02 · recommended

Cost Control Sprint

Eight weeks. Software shipped, dashboards live, quality tests running in your pipeline.

25%
of verified savings. You keep 75%.
no upfront fee · invoiced quarterly as savings land · nothing to pay if nothing verifies
  • All four deliverables above
  • A savings report your finance team recomputes and signs
  • You own all code and models outright, with no licence
  • One check-in 30 days after handover
Book the free audit

How a saving is counted

Our fee is a share of savings, so this definition is the most important term in the contract. It is worth two minutes of your CFO's time.

  • Your starting position is frozen in week one from your own provider invoices, not from a dashboard of ours.
  • We measure cost per unit of work, per request or per document, never your monthly total. So if your traffic doubles, your bill rises and the saving we claim does not change. Growth never creates a fee and never erases one.
  • A change counts only after it holds your quality pass mark for 30 days in production. Anything that cannot hold it is reverted and never counted.
  • Provider price cuts do not count. If a vendor drops its list price, that saving is yours alone and is excluded from our fee.
  • Your finance team recomputes every line from provider invoices. Anything they cannot reproduce is struck from the report before it becomes an invoice.

Commercial terms

The full terms are three pages and you get them with the audit findings, before any commercial conversation. This is all of the substance.

  • Fee. 25% of verified savings, annualised over 12 months from the date each change holds its pass mark. Never more than a quarter of the result, so what we earn moves with your outcome rather than with our hours.
  • Invoicing. Quarterly, in four instalments, as savings actually appear. Never a lump sum against a projection.
  • If it stops working. If a change is reverted, or you switch providers and the saving no longer exists, the remaining instalments stop and the count is adjusted.
  • Ownership. You own the code, the prompts, the routing configuration, the quality tests and any trained model. We keep no licence and run nothing in your systems.
  • Getting out. Either side, seven days' notice, no penalty. Savings verified before that date remain payable. Nothing else does.
Team

Two people, and both of them do the work.

No bench, no analyst layer, no account manager between you and the engineers. That is also why we run two Sprints at a time and no more.

agentic automation · ai finops

Bartek Boniecki

Owns the money side. Builds the tracking layer that puts a price on every product feature, maintains the weekly provider price tracker, and designs the automations that come out of the findings. Works from the open audit checklist in our public repository, the same documents you receive with your audit.

AI orchestration · llm engineering

Jean-Luc Momprivé

Owns the engineering side. Twenty years building and shipping AI systems inside large organisations, including product and innovation leadership at enterprise scale. Decides which parts of your traffic can safely move to a cheaper model, builds the tests that prove it, and refuses the changes that cannot hold their quality bar.

FAQ

Answered straight.

What buyers actually ask on the first call.

getting started
We tried this internally and gave up. What is different?+

Tracking the right metrics is the key. If nothing in your setup is recording cost per feature, so every improvement is a guess and nobody can defend a number to finance. We arrive with the tracking checklist, the provider price tracker and the quality test harness already built. Your engineers keep shipping product; we do the part that never survives sprint planning.

How much of our team's time does this take?+

Access setup in week one, then roughly two hours a week from one engineer for context and code review. The Sprint is designed on the assumption that nobody on your side has spare capacity. If it needed a dedicated person from you, the economics would not work for either of us.

What access do you need, and what happens to our data?+

Read-only billing exports, your gateway configuration, and repository access scoped to the directories we touch, with no commit rights on protected branches. No production data or customer content leaves your environment, and sample prompts used to design quality tests stay inside your infrastructure. Cost is attributed against hashed account identifiers, never customer personal data. Mutual NDA before the audit call on request, and a data processing addendum is available. We are a two-person firm and not SOC 2 certified. We would rather tell you that on the first call than have your security review find it in month three.

Why not just buy a cost-tracking tool?+

Buy one. Several are good, and we wire into whichever you pick. A tool shows you the bill. It does not decide which parts of your traffic can safely move to a cheaper model, build the tests that prove quality held, restructure your prompts so the repeated part is actually billed as repeated, or defend the resulting number to your CFO. The tool is the instrument. This is the work.

money
Our spend is $5k to $10k a month. Is this worth it?+

Even at $5k-10k a month, a 40% reduction in cost is roughly $24k-48k over a year. Our invoice would be 25% of it, spread across four quarters, you keep the rest, and you keep the cost tracking and the quality tests permanently. Below about $5k a month we decline, because the fee stops covering two senior engineers for eight weeks and we would be taking your money to do a worse job. The audit tells you which side of that line you are on, at no cost.

Why a share of savings rather than a fixed fee?+

Because we cannot honestly quote a fixed number before seeing your traffic, and because it puts the risk of a disappointing outcome on us rather than on you. It has one drawback and we will name it for you: it gives us an incentive to chase the biggest number rather than the safest one. The quality pass marks you set in week one, the 30-day hold before anything counts, and your finance team's sign-off exist to remove that incentive. A change that degrades your product earns us nothing.

What if our traffic grows and the bill goes up anyway?+

Then your bill goes up and our fee does not change. Everything is measured per request against a starting position frozen in week one, so growth is neutral in both directions. It never creates a fee for us and it never wipes one out. This is also why what we sell is cost per request rather than a promise about your monthly total, which we do not control.

risk
What happens when you leave?+

Everything lives in your repository and your tools, and you own it outright. Nothing of ours keeps running in your systems, there is no licence and no renewal. Your team extends the work from there, which is the design goal rather than a concession. It is also why we would rather ship a smaller change your engineers understand than a clever one they cannot maintain.

The audit is free. The diagnosis is written. The decision is yours.

Thirty minutes, read-only access, and a written read on where your AI spend is going, whether or not you ever hire us.