FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Checklist ┬╖ 4 minute read

AI Quarterly Review Checklist: Keeping Systems Honest

Deployed AI systems drift: scope widens, permissions accumulate, corpora go stale, and quality declines without an obvious cause. A quarterly review covering quality trend, scope against the original document, permissions, cost, and corpus health catches all of it before it compounds.

By FISTA Solutions┬╖ AI-Native Engineering Team┬╖
AI Quarterly Review Checklist: Keeping Systems Honest article cover

Deployed AI systems drift in ways nobody decides on. This quarterly review catches it, drawn from FISTA Solutions' AI agents operational practice.

What does the review cover?

Six areas, compared against the original design rather than last quarter.

AreaWhat drift looks like
ScopeHandling cases it was not designed for
QualityGradual decline with no event
PermissionsCapability added incrementally
CostPer-task cost rising quietly
CorpusStale content, no refresh
UsageNobody using it any more

Scope

Compare against the document, not against memory.

  • Original scope document retrieved and read
  • Actual usage compared against documented cases
  • Cases now handled that were not in scope identified
  • Escalation categories compared against what actually escalates
  • Undocumented additions either approved or removed
  • Scope document updated if the change is intentional
  • Risk assessment refreshed if scope widened materially

Quality

The trend matters more than the level. See how to monitor AI quality in production.

  • Evaluation suite re-run and results compared across quarters
  • Production sample reviewed by a domain expert
  • Correction and escalation rates trended
  • New failure modes from the period added as test cases
  • Evaluation cases reviewed for continued representativeness
  • Any decline investigated to a cause
  • Quality criteria checked against current business expectations

Permissions

Recertify rather than assume. See agent permission review checklist.

  • Current capability list compared against approved scope
  • Additions since last review justified or removed
  • Unused capabilities removed
  • Limits still appropriate for current volume and value
  • Service credentials rotated on schedule
  • Human access reviewed
  • Owner confirmed still in post and still accountable

Cost

Per-task cost drifts upward as prompts and context grow.

  • Cost per task trended across quarters
  • Prompt and context size compared against last review
  • Routing decisions re-evaluated against current models
  • Cache hit rate checked
  • Retry and loop rates checked
  • Cost against budget and forecast reviewed
  • Efficiency opportunities identified and prioritised

Corpus and data

Decay is continuous. See AI knowledge base quality checklist.

  • Document age distribution checked against last quarter
  • Refresh pipeline confirmed running
  • New contradictions or duplicates identified
  • Coverage checked against questions users actually asked
  • Content owners confirmed and still engaged
  • Documents with no owner resolved
  • Index lag behind source verified

Usage and value

Some systems should be retired rather than maintained.

  • Usage trended over several quarters
  • Outcome metrics compared against the original success criteria
  • Systems below a usage threshold flagged
  • Systems with no measured outcome flagged
  • Retirement recommended where appropriate, with a plan
  • Maintenance effort compared against value delivered
  • Overlapping systems identified for consolidation

What are the most common failures?

Comparing against last quarter rather than against the design. Reviewing quality without a trend. Permissions assumed unchanged. Corpus decay unnoticed. And maintaining systems nobody uses because retiring them requires a decision.

Who should own this?

The system's business owner leads; engineering supplies the metrics; security participates in the permission recertification. A review run by engineering alone misses scope and value questions.

How often should it run?

Quarterly for production systems, semi-annually for low-risk internal tools. More frequently during the first two quarters after launch, when drift is fastest.

What evidence should it produce?

Dated review records with the metrics, the comparison against the original scope, permission recertification, and decisions taken including retirements.

What if the review finds serious drift?

Treat it as a change requiring the same approval the original deployment had: scope reassessed, permissions reviewed, risk accepted by the owner.

The temptation is to document the current state and move on, which normalises the drift. If the system is doing something that would not have been approved, the question is whether to approve it now or to narrow it back. See AI risk assessment template.

What should you do first?

Retrieve the scope document for your oldest deployed AI system and compare it against what the system actually does. The gap is usually instructive.

How FISTA Solutions helps

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: quarterly reviews compared against original scope and success criteria rather than against the previous quarter, with retirement treated as a legitimate outcome, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To adapt this checklist to your environment, message FISTA on WhatsApp, or read AI agent launch checklist.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why quarterly rather than annually?

Because drift compounds. A quarter of scope creep is correctable; a year of it means the system is doing something nobody approved, with permissions nobody reviewed.

02What is scope drift?

The system handling cases it was not designed for, because they were close enough and it seemed to work. Each addition is small; the accumulation is a system operating outside its assessment.

03How do you spot quality decline?

By trending the same metric over several quarters. A single measurement tells you the current state; only the trend shows gradual degradation, which is the usual shape.

04What makes a system a retirement candidate?

Low usage, no measured outcome, or a process that has changed such that the system no longer fits. Maintaining these consumes attention that working systems need.

05What should you compare against?

The original scope document, permission list, and success criteria. Comparing against last quarter normalises drift; comparing against the design catches it.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.

Start a project