Checklist · 5 minute read
AI Model Selection Checklist: Choosing With Evidence
Choose a model on evidence from your own cases, not on benchmarks. Evaluate quality on representative tasks, model cost at realistic volume, measure latency at the tail, check data handling terms, and confirm the deprecation policy before you build a dependency on a specific version.
Model selection made on benchmarks and impressions ages badly, usually within a quarter. This checklist covers choosing with evidence, drawn from FISTA Solutions' AI enablement delivery work.
What should the decision rest on?
Six dimensions, weighted by your situation.
| Dimension | How to assess it |
|---|---|
| Quality on your tasks | Your own evaluation suite |
| Cost at your volume | Modelled, not sampled |
| Latency | Measured at the tail |
| Data handling | Contract terms, not marketing |
| Deprecation policy | Stated notice period |
| Portability | What switching would cost |
Quality evaluation
This is the only assessment that transfers to your situation. See how to build an agent evaluation harness.
- Evaluation cases drawn from real workload, not invented examples
- Correct answers agreed by people who know the domain
- Edge cases and known failure modes included
- Each candidate run against the identical suite
- Results reviewed by a domain expert, not only scored automatically
- Output format reliability tested, not just content quality
- Instruction following tested with your actual prompt structure
Cost modelling
Sample volume and production volume differ by orders of magnitude, and so do the conclusions.
- Token counts measured on representative requests, input and output
- Volume projected for twelve months, not current usage
- Cost compared across candidates at that projected volume
- Caching discounts factored where applicable
- Batch pricing considered for non-urgent work
- A cheaper model tested on the subset of tasks it could handle
- Cost per successful task computed, including retries
Latency and throughput
Measure what users experience rather than what the provider publishes.
- Time to first token measured for streaming use cases
- Total completion time measured at the ninety-ninth percentile
- Measurements taken at your expected concurrency, not serially
- Rate limits confirmed and compared against projected peak
- Behaviour under rate limiting tested
- Regional availability checked against where your users are
- Latency measured over several days, not a single sample
Data handling and terms
Defaults are frequently not what buyers assume. This is general guidance, not legal advice.
- Whether inputs are used for training, stated in the contract
- Retention period for inputs and outputs
- Processing and storage locations confirmed
- Subprocessors disclosed
- Certifications relevant to your sector verified
- Audit trail availability and export confirmed
- Terms reviewed by legal before commitment, not after
Version stability and deprecation
This determines how much unplanned work the choice generates. See how to run a model migration.
- Version pinning available and used
- Stated notice period before a version is withdrawn
- Historical deprecation behaviour checked, not just the policy
- Notification mechanism confirmed and subscribed to
- Evaluation suite ready to run against a new version quickly
- Budget and capacity allowed for periodic revalidation
- Behaviour differences between versions documented when they occur
Portability
Assess the exit before you build the dependency.
- Provider-specific features identified and their value assessed
- An abstraction layer in place over the provider interface
- Prompts stored independently of the provider
- Evaluation suite runnable against alternative providers
- Fine-tuned artefacts' portability understood
- A second provider tested at least once
- Switching effort estimated in person-days
What are the most common failures?
Choosing on benchmarks. Cost modelled at sample volume. Latency measured serially on a quiet day. Terms assumed rather than read. And building on a version with no pinning and no deprecation notice.
Who should own this?
Engineering runs the evaluation; the business owner accepts the quality and cost trade-off; legal reviews the terms. A selection made entirely within engineering misses the commercial and regulatory questions.
How often should it run?
At selection, and again whenever a significant new model is released or your volume changes materially. An annual re-evaluation is the minimum for any system where cost matters.
What evidence should it produce?
The evaluation report comparing candidates on your cases, the cost model at projected volume, latency measurements, and the reviewed contract terms. That package justifies the decision later.
What if you cannot evaluate properly yet?
Pick the model with the best general reputation and the most flexible terms, build the evaluation suite, and revisit within a quarter.
That is a legitimate position early on. What is not legitimate is treating the provisional choice as permanent and building deep provider-specific dependencies before you can measure anything. See why evaluation is the new moat.
What should you do first?
Assemble twenty real cases from your workload with agreed correct answers. That small suite will tell you more than any benchmark table.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: model choices justified by evaluation on real cases at projected volume, with an abstraction layer that keeps the decision reversible, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read how to run a model migration.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why not use published benchmarks?
Because they measure general capability on standardised tasks. Your workload has its own distribution, terminology, and definition of correct, and model rankings frequently reorder on specific tasks.
02What does evaluating properly require?
A suite of representative cases from your actual workload with agreed correct answers, run against each candidate, with results compared on quality, cost, and latency together.
03Why does tail latency matter?
Because the slowest requests are what users notice and what times out. A model with a good average and a poor tail produces a worse experience than its headline numbers suggest.
04What data terms should be checked?
Whether inputs train the provider's models, retention period, processing location, and which subprocessors are involved. Defaults differ by provider and by plan tier.
05What is deprecation policy?
How long a specific model version remains available and how much notice you get before it is withdrawn. Short windows mean recurring migration work you did not plan for.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.