Unico Connect
Choosing an AI model for production using evaluation, cost per task and latency
Back to Blog
AIUpdated September 26, 202613 min read

How to Choose an AI Model for Production

Saurav Jagdale

Saurav Jagdale

Technical Lead, Unico Connect

In this article

Most teams choose a model the same way. Somebody reads a benchmark table, picks whichever name is highest, wires it in, and finds out six weeks later that it is too slow, too expensive, or wrong in a way the benchmark never tested. Then they swap models and repeat the process.

The problem is the question. There is no answer to which model is best, because a model is only best for a given task, at a price and a latency you can live with. A model that wins every public leaderboard can still be the wrong choice for extracting fields from your invoices.

What follows is the method we use to pick models for production systems. It is deliberately boring, takes about a week, and gives you a decision you can defend and repeat. If you want a vendor by vendor verdict instead, read Claude versus GPT versus Gemini. Knowing how to run the choice yourself matters more, because that verdict goes stale while the method keeps working.

Quick Answer

Choose a model in five steps. Build an evaluation set of 50 to 200 real cases from your own traffic, with correct answers written down. Define pass and fail at the level of the whole task, so a partly right answer counts as a fail. Run three or four candidate models across price tiers against that set and record accuracy, cost and latency together. Compute cost per successful task, the only cost number that means anything, because a cheap model that fails a third of the time costs more than its token price suggests. Then set a latency budget at the 95th percentile and eliminate anything that misses it. The model that survives is usually smaller than your biggest candidate and almost never the leaderboard favourite. Unico Connect uses this method to pick models for production systems.

Key Takeaways

  • Use public benchmarks to cut the field to three or four candidates, then put them away. Choosing between the finalists takes your own eval set, since it is the only evidence that transfers to production.
  • Cost per successful task is the unit to compare. Divide the total spend for a run by the number of tasks that passed. A model at one fifth the price that fails twice as often can cost more once failures escalate to a human, and the token price never shows you that.
  • Count retries in the price. If your system retries on failure, the failure rate multiplies your bill and your latency at the same time, so cheap and unreliable compounds on both counts.
  • Measure latency at the 95th percentile, never the average. The average hides the slow tail, and users experience that tail as the product being broken. Reasoning heavy models in particular have a long right tail.
  • Plan to make the choice again. Model versions change under you, prices move and your traffic drifts, so keep the eval set in version control and rerun it on a schedule. The next choice then costs a day instead of another six weeks.

Step One, Build the Eval Set

Everything depends on this and most teams skip it.

Take 50 to 200 real examples from your own traffic. Leave out synthetic examples, the happy path from your demo and any case you wrote while thinking about the prompt. Keep the messy real inputs, including the malformed ones, the ones in the wrong language and the ones a user pasted from a spreadsheet.

For each one, write down what a correct output looks like. This is the tedious part, and it is where the value comes from, because you are encoding your own standard for the task instead of an abstract notion of quality. Where several outputs would be acceptable, record the rule that makes them acceptable instead of a single golden string.

Three properties matter. The first is a holdout. Keep 20 to 30 percent unseen while you iterate on prompts, or you will tune your way to a number that does not survive contact with new traffic. The second is coverage of the cases you currently get wrong, because a model that fixes your known failures is worth more than one that is marginally better on cases you already handle. The third is version control. Keep the set next to the code so it grows as production surfaces new failure modes.

For this decision, a hundred labelled real cases tell you more than any leaderboard, and you can assemble them in a day.

Step Two, Define What Passing Means

Decide before you run anything, because deciding afterwards means grading on a curve you drew around the result you liked.

Grade at the level of the task. An extraction passes only when every required field is correct. A classification passes when the label matches. A support reply has to answer the question, contain no invented fact and follow the tone rules. Partial credit feels fairer and is worse, because production is usually pass or fail. An invoice with four of five fields right still needs a human.

Where the output is open ended and a human would have to judge it, you can use a strong model as the grader, but two conditions apply. Write the grading rubric explicitly rather than asking for a quality score out of ten, and check the grader against human judgement on a sample before you trust it on the rest. An unchecked model grader mostly measures how much the grader likes the candidate.

Keep a small set of hard rules separate from the graded score. Anything that must never happen, such as leaking data that belongs to another customer, inventing a policy or returning malformed JSON, is a gate that stays outside the graded score. A model that trips a gate is eliminated however well it scores.

Model tiers, and the workload each one actually wins

Model tiers, and the workload each one actually wins
TierWins whenCost shapeThe failure mode to plan for
Frontier modelLow volume, high stakes, or genuinely hard reasoningHighest per call, often lowest per successful task at low volumeA long latency tail that misses an interactive budget
Mid tier modelThe default for most production tasks once prompts are tunedBalanced, and usually the best cost per successful taskQuietly good enough, so nobody ever tests whether a smaller tier would do
Small fast modelHigh volume, low stakes, narrow and well specified tasksLowest per call, worst per successful task if accuracy dropsRetries and human escalation erase the saving without showing on the invoice
Small model plus fallbackHigh volume where a validation gate can catch failures reliablyClose to the small model, plus the escalation rate times the large modelA weak gate, which passes bad output through and hides the problem

Which should you choose

High volume, low stakes workSmall model behind a validation gatethe token saving is real at volume and a gate catches what it gets wrong
Low volume, high stakes workBuy the accuracyvolume is too small for the saving to matter and a failure costs more than the run
You have not measured anything yetOne mid tier model everywhererouting complexity should be earned with evidence, and most systems never need it

Tiers here are workload categories rather than named products. No vendor list prices appear in this post by design, since published rates move and differ across cloud platforms for the same model. Substitute your own measured accuracy, cost and latency.

Step Three, Run Candidates Across Price Tiers

Shortlist three or four models that span price tiers instead of clustering at the top. A useful default is one frontier model, one mid tier model and one small fast model. You need that spread because the only way to find the cheapest model that clears your bar is to test some you expect to fail.

Hold everything else constant. Same prompt, same tools, same retrieval, same temperature settings, one variable at a time. Then record three numbers for every run and never look at any of them alone.

Accuracy is the pass rate against your eval set under the rules from step two. Cost is the total spend for the run, taken from measured token usage. Latency is the full distribution of response times, tail included.

One caveat trips people up, which is that prompts are not portable between models. A prompt tuned over months against one model is an unfair test of another, and the honest fix is to give each candidate a short, equal tuning pass before you record its numbers. Otherwise you are measuring your prompt history rather than the model.

Step Four, Compute Cost Per Successful Task

This is the arithmetic that changes decisions.

Cost per successful task equals total run cost divided by the number of tasks that passed. The formula folds the failure rate into the price, which token pricing has no way to show you.

Take an illustrative example. Suppose a large model passes 94 percent of your eval set and a small model passes 71 percent, and the small model costs one fifth as much per call. On the token price the small model looks like an obvious win. Per successful task, the large model costs the run divided by 0.94 and the small model costs one fifth of the run divided by 0.71, so the small model is still cheaper, by nearly four times rather than five.

Now add the part teams forget, which is what happens to the failures. If each failure is retried once, the small model runs 29 percent of its volume twice, raising its real cost and its real latency together. Where failures escalate to a human, price the human minutes, and the small model usually loses outright. Failures that reach the user cost a support ticket and some trust, a cost that is real even though it never appears on the invoice.

A rule falls out of this. For high volume and low stakes work, a smaller model with a failure path is usually the right call. For low volume and high stakes work, buy the accuracy, because the volume is too small for the token saving to matter and the failure is too expensive to absorb. Most teams get this backwards, running an expensive model on bulk classification and a cheap one on the thing that touches a customer.

Step Five, Set the Latency Budget

Work backwards from the interface.

A response the user waits for while staring at a spinner needs to feel immediate. Streaming gives you more room, because the perceived wait is the time to first token rather than the time to completion. A background job can take as long as the queue tolerates. Write the number down before you test, because a budget picked after you see the results will tend to fit them.

Then measure the 95th percentile. An average conceals exactly the cases that make users think the product is broken, and reasoning heavy models have a long right tail. Measure under realistic concurrency too, since a model that is fast for one request at a time can slow down when a hundred arrive together.

If a candidate misses the budget it is out, however good its accuracy, because a correct answer that arrives too late does the user no good.

The Routing Decision

You do not have to pick one model, and past a certain scale you should not.

One model everywhere is the right starting point. It is the simplest to reason about and the cheapest to operate, so stay here until you have measured evidence to leave. Most systems never need to.

Tiered routing sends easy cases to a small model and hard ones to a large model, which works when you can classify difficulty cheaply and reliably. The trap is that the classifier is itself a model that can be wrong, so put its error rate in the budget too.

Fallback on failure runs the small model first and escalates only when the output fails a validation gate. This is the pattern we reach for most often, because the escalation trigger is a real check on a real output rather than a guess about difficulty made in advance.

For how these patterns combine once several models run side by side in production, see multi model production AI.

What This Does Not Decide

Model choice is one variable among several, and usually not the biggest. If your answers are wrong because the model cannot see the right information, no model swap fixes that. Retrieval quality, prompt structure and tool design typically move accuracy further than moving up a tier does, and they are cheaper.

Work in this order. Fix retrieval, fix the prompt, fix the tools, and only then change the model. Teams that reverse it spend a lot of money discovering that the frontier model is also unable to guess a fact that was never in its context. For the adaptation side of that work, read RAG versus fine tuning and fine tuning versus prompt engineering.

The number I ask every team for is cost per successful task, and almost nobody has it. They have a token price from a pricing page and an accuracy figure from a different run, never divided into each other. Once you compute it the argument usually ends in about ten minutes, because the model people were defending on price turns out to be the expensive one.

Saurav Jagdale, Technical Lead, Unico Connect

Keep the Decision Fresh

You will run this selection more than once. Model versions change beneath you, prices move, and your own traffic drifts as users learn what the product can do.

Keep the eval set in version control, rerun it when a vendor ships a new version or when you change the prompt, and keep logging accuracy, cost per successful task and 95th percentile latency in production continuously, including long after the selection is done. Drift shows up in production first, and our guide to AI observability explains how to instrument for it.

Because the eval set already exists, the second selection costs a day, and that saving is the return on building it.

Where These Numbers Come From

We left vendor prices out of this post on purpose. Published rates move, they differ across the Anthropic, Amazon Bedrock, Google Gemini Enterprise Agent Platform (formerly Vertex AI) and Microsoft Azure OpenAI platforms for the same model, and any table we printed would be wrong within a quarter. The 94 percent, 71 percent and one fifth figures in the cost section are illustrative inputs chosen to show the arithmetic. They do not describe any specific model, and the point of that section is that you substitute your own measurements. For current vendor positioning see Claude versus GPT versus Gemini, and read list prices from the vendor pricing page on the day you need them, including OpenAI. The method itself is our own delivery practice across AI systems we run in production.

Frequently Asked Questions

How do I choose which AI model to use in production?

Build an eval set of 50 to 200 real cases from your own traffic with correct answers recorded, define pass and fail at task level, run three or four candidates spanning price tiers, then compare on cost per successful task and 95th percentile latency rather than on token price or benchmark scores. Keep the eval set in version control so the next decision is cheap.

What is cost per successful task?

Total cost of a run divided by the number of tasks that passed your quality bar. It folds the failure rate into the price, which token pricing cannot do. A model at one fifth the price that fails twice as often costs more than one fifth as much per successful task, and if failures are retried or escalated to a human it can be more expensive outright.

Should I use the biggest model available?

Usually not. Buy accuracy where volume is low and the cost of being wrong is high, and use a smaller model with a validation gate and a failure path where volume is high and stakes are low. Many teams do the opposite, spending on bulk internal classification while a cheaper model handles the customer facing path.

How many test cases do I need to choose a model?

Between 50 and 200 real cases is enough to separate candidates for most tasks. Below about 50 the differences fall inside noise. What matters more than count is that the cases come from real traffic, include your known failures and the malformed inputs, and that 20 to 30 percent is held back while you iterate on prompts.

Can I use benchmarks instead of building my own evaluation set?

Use them to narrow a shortlist to three or four candidates, then stop. Benchmarks measure general capability on public tasks, and your decision depends on one specific task, your prompt, your retrieval and your latency budget. They also leak into training data over time, which erodes what a high score means.

Should latency be measured as an average?

No. Measure the 95th percentile, because the average hides the slow tail that users experience as the product being broken. Measure under realistic concurrency as well, since throughput under load differs from single request speed. If you stream the response, measure time to first token separately, because that decides whether the wait feels acceptable.

How often should we revisit the model choice?

Rerun the eval set whenever a vendor ships a new model version, whenever you materially change the prompt or retrieval, and on a fixed schedule regardless, quarterly being a reasonable default. Also monitor accuracy, cost per successful task and latency in production continuously, since traffic drift shows up there before it shows up in a scheduled run.

What should we fix before changing models?

Retrieval, then the prompt, then tool design. If the model cannot see the information needed to answer, no upgrade fixes it, and those three are cheaper to change than a tier jump. Change the model last, once you have confirmed that the remaining errors are capability limits.

Conclusion

Model selection looks like a research question, but in practice it is a measurement question. Assemble a hundred real examples, decide what correct means before you look at any output, run a few candidates across price tiers, and divide cost by successes. The answer usually turns out to be a smaller model than the team expected, wrapped in a validation gate that catches the cases it gets wrong.

The lasting artifact is the eval set. Build it once and every future model decision, prompt change and vendor announcement becomes a day of work instead of a debate.

If you would like us to run this against your own workload, see our AI development services and AI integration, or contact us.

Keep reading

Related Articles

View all