Skip to content
Blog

Tokens Tell You Cost. Delivery Tells You Value. | CMD+RVL

Written by CMD+RVL. Edited by Zac Ruiz.

Token counts explain part of what AI costs. They say nothing about what it delivered. Measure the report that ships faster, the check that runs itself, the question answered once, and what the team can do now that it could not before.

A lot of AI strategy conversations start with the wrong unit: the token. How many did we use, what did they cost, which model is cheaper, can we route more requests to a small model, how much would caching save.

Those are reasonable operating questions, and none of them is a strategy. A token is an input to the cost of running a model, and managing AI by token volume is like managing a factory by how much electricity each machine draws. You should know the number. Nobody would mistake lower electricity use for higher output, better quality, or a shorter cycle.

The question that matters is what got delivered.

Know the cost. Attach it to a result.

Every AI system has a cost structure: input tokens, output tokens, model tier, tool calls, infrastructure, engineering time, review. Know those numbers. But cost only means something when it is attached to a result.

A $2 model run that delivers nothing is expensive. A $200 workflow that takes two days out of a recurring process is cheap. A $2,000-a-month system that produces something the team could not reliably produce before may be very cheap.

So the conversation should move from cost per token, to cost per delivered result, to what that result is worth.

Measure delivery, before and after.

You do not need an AI ROI framework. Pick one thing the business already cares about. Measure it before the new workflow exists, then measure it again once the workflow is in use. Five measures usually say more than any token dashboard.

1. Cycle time

How long did it take before, and how long does it take now?

A report that took two days and now takes an hour has a visible before and after. So does a diligence step that went from three days to same-day, or a Friday-afternoon research task that is now fifteen minutes of review. Cycle time is easy for everyone in the business to understand because it describes something delivered, not a model.

2. Manual effort

How many human steps did it take, and which of them still need judgment?

The goal is not to remove people. Often the delivered change is that someone stops collecting, copying, reformatting, and reconciling, and spends that time on the part that needs their expertise. That is a real change even when the final output looks the same.

3. Repeated verification

How much of the process is people checking the same thing again because the answer arrived without its evidence?

This is the hidden cost of AI-assisted work. A model can produce an answer in seconds; if someone then spends twenty minutes reconstructing where each claim came from, the speedup is gone. A delivered result keeps its sources, definitions, assumptions, and gaps attached, so the next person does not have to rediscover them.

4. Adoption

Did anyone use it?

An impressive prototype can deliver nothing. A plain workflow used every week can be worth a great deal. Ask whether the result became part of how someone does their job: did they come back to it, did a second person use it, did it replace an old step, did the team ask for the next version. Use of the delivered result matters more than use of the model.

5. New capability

What can the team do now that it could not do before?

Token dashboards miss this entirely. Maybe the team checks every item instead of a sample, answers a question across five years of filings instead of the latest one, or keeps a recurring view that was too expensive to build by hand. That goes beyond efficiency: the team can deliver something it could not deliver before.

The scorecard fits on one page.

For one delivered result:

MeasureBeforeAfter
Cycle timeHow long it tookHow long it takes now
Manual effortHuman steps or hoursHuman steps or hours now
VerificationWhat had to be recheckedWhat arrives with its evidence
AdoptionWho used the old processWho uses the new result
CapabilityWhat was not practicalWhat is practical now
CostExisting labor and system costAI, software, review, and operating cost

Now the token bill has context. You can still optimize it, but you are optimizing the cost of a known result rather than shrinking a number because a dashboard made it visible.

Start with public data when you can.

Do not make private data access the first proof that the method works. When a problem can be tested on public information, start there: filings, regulatory records, government data, market data, anything a person can inspect.

Can the system find the right information, keep the source, produce a result someone can check, repeat the process, and show a before and after? If yes, you have learned the important thing before asking the business to connect private systems.

You have also separated two questions that usually get mixed together: does the method deliver something useful, and would private data make that useful thing more specific? Keeping them apart lowers the cost of learning, avoids unnecessary access, gives the team something concrete to react to, and makes the private-data conversation better, because you can name exactly what extra information would improve the result.

It is also why getting one dataset ready is more useful than a broad data program. Prepare one dataset so AI answers can be checked before trying to make every system "AI ready."

What remains after the chat.

A good session with a model can feel productive, and sometimes it is. The question is what remains when it ends. Did the source survive? The definition? The method? Did the useful context reach the next person, and did the answer become something the business can update, check, and build on? If not, the next task starts with the same gathering and explaining.

That is why the value of AI usually sits outside the model. The model helps produce the result; the durable value is the organized information, the repeatable process, the evidence, and a next question that starts further ahead. There is a longer version of that argument in What survives the chat window.

A better strategy question

Instead of "how do we reduce token costs?", ask "what result are we trying to deliver faster, more reliably, or for the first time?" Then measure the result.

If the result matters and the economics are poor, optimize the system: a smaller model where it works, cached context, less unnecessary input, deterministic steps kept deterministic, no model paid to redo what software already does reliably. Those are good engineering decisions. They come after the business has decided what is worth delivering. Otherwise token optimization is a very efficient way to make an unclear system cheaper.

Delivered work compounds.

Delivered work leaves something behind: a cleaned definition, a resolved identifier, a source map, a reusable check, a standing report, a decision trace, a piece of context the next question does not have to rediscover. One delivered result becomes material for the next. The team spends less time starting over, and the cost of getting to an answer falls while the starting point improves.

Tokens are consumed. Delivered context accumulates.

Measure what was delivered.

For each AI initiative, write down the before state first: how long it takes today, how much manual effort it needs, what gets checked repeatedly, who uses the result, what is not practical. Build the smallest useful version. Measure the same things again.

Now you have something better than a usage dashboard: a delivered result, its cost, and evidence of what changed. That is enough to make the next decision.

PreviousSalt IO Is Now CMD+RVL
All posts