That gives you a more useful starting point for a buying decision. A subscription allowance tells you something about access. A usage dashboard tells you something about consumption. You still need to know what the organisation got back.
Decide what counts as finished
Take one recurring task and write down its acceptance conditions before measuring it.
For a research brief, that could mean answering three specified questions, linking the supporting sources and flagging anything unresolved. For an enquiry summary, it might mean capturing the request accurately and passing it to the right person for review.
Keep the definition steady while comparing approaches. A rough first draft and a checked, usable brief are different outputs. If the standard changes halfway through the trial, record that and explain what it means for the comparison.
The FinOps Foundation’s unit economics guidance makes the wider principle clear: connect technology spending to a meaningful business result. For this exercise, the useful unit is an accepted piece of work.
- PrepareAgree the task and sources.
- RunRecord usage and attempts.
- ReviewCheck and correct the result.
- Accepted resultMeets the agreed standard.
Retry or correct: Review → Run.
- Usage
- People
- Setup and upkeep
- Unsuccessful work
Make the cost record complete
Usage. Record the provider’s reported consumption, additional charges and relevant tool or hosting costs. Separate subscription fees, included allowance and extra spending. Avoid counting the same cost twice. Where you allocate a shared subscription across tasks, explain the method.
People. Include preparing the input, supervising the run, checking the result and correcting mistakes. Record minutes even if you are not yet putting a monetary value against them. A task that needs repeated intervention has a different operating cost from one that needs a short final check.
Setup and upkeep. Someone has to configure the workflow, maintain its connections and investigate failures. Keep these costs visible alongside routine running costs, with a clear assumption about how many tasks will share them.
Unsuccessful work. Include retries and rejected attempts within the relevant usage and people totals. Dividing total trial cost by accepted results prevents unsuccessful runs disappearing from view.
The FinOps for AI guidance describes the difficulty of forecasting across varied services and distinguishing successful outputs. Treat an early estimate as something to test, with its assumptions attached.
Look at what happens between the request and the answer
A request to “check these websites” leaves several questions open. How many sites? Which pages? How current must the information be? What happens when a page is unavailable?
Inspect a sample run. Look for repeated searches, unnecessary page visits, the same information being retrieved again and attempts that continue after a useful answer is already available. Ask whether an approved integration, export or ordinary rule could handle a predictable part of the work.
Try a change on the same test cases and check the outputs again. A shorter run is worthwhile only if it still meets the agreed standard. Switching models or moving work to your own equipment should face the same test. Include equipment, support, maintenance and energy in that comparison.
A small example
Imagine a fictional team preparing weekly supplier summaries. Its first trial produces twelve summaries, but only eight meet the agreed acceptance criteria by the end of the trial.
The team records all twelve attempts, the reruns and the reviewer’s time. It finds that unclear source requirements cause much of the correction. It then repeats the exercise with a fixed source list and explicit questions, using comparable cases and the same acceptance criteria.
There is no assumed saving here. The second trial might be better or worse. What matters is that the team can explain whether the change reduced effort, improved the result or simply moved work from the AI to a colleague.
Put boundaries around the trial
Before a long or repeated run, agree:
-
The task, owner and acceptance criteria
-
The permitted sources and actions
-
A usage or spending ceiling, plus a time limit
-
How many retries are reasonable before someone reviews the problem
-
Who receives alerts and can stop the work
-
Where results, exceptions and actual costs are recorded
Check whether a provider’s limit stops further usage or only sends an alert. Review any extra-usage setting before relying on it.
Finally, compare accepted results, total effort and quality with your existing process. Time released is useful capacity; call it a cash saving only when expenditure actually falls. Decide how the team will use the capacity and when to review the evidence again.
Start with one task you understand well. A clear result, an honest cost record and a short comparison will tell you more than a large allowance used enthusiastically.