Playbook ¡ 6 minute read
How to Structure an AI Managed Service That Works
An AI managed service works when the scope names quality as well as availability, model and corpus changes are handled explicitly, escalation reaches someone who can decide, and the exit provisions mean the client could leave. Availability-only arrangements miss the failure that matters.
Managed services for AI have to cover something conventional arrangements do not: a system that is up, responsive, and producing materially worse results. This playbook covers structuring one, drawing on FISTA Solutions' forward deployed engineers work.
When is this worth doing?
When an AI system is in production and the client lacks the capacity or the desire to operate it, and where the capability is not so strategic that it must be owned internally.
It is a poor fit where the client wants to build capability, because a managed service that operates the system well can also prevent the client's team from ever learning it.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Define the scope | Quality, not only availability |
| 2. Agree the evaluation set | The basis for quality reporting |
| 3. Set the decision boundary | Client decides, provider operates |
| 4. Define change handling | Models, corpora, prompts |
| 5. Agree the cost mechanism | Volume changes; say how |
| 6. Write real exit provisions | Artefacts, not just notice |
Step 1 â Define the scope including quality
Operation, monitoring, incident response, evaluation, change handling, and reporting â with quality named explicitly rather than implied.
Availability and latency commitments are straightforward and they do not cover the failure that matters. A system returning responses within target while producing worse answers is failing, and an availability-only arrangement is silent about it.
State what the provider is accountable for and what remains with the client. Ambiguity there surfaces during the first incident.
Step 2 â Agree the evaluation set
Both parties agree the cases quality is measured against, and the provider runs them on a schedule with results reported.
This is what makes a quality commitment meaningful. Without an agreed set, quality becomes a matter of impression, and impressions diverge exactly when the relationship is under strain.
Agree how the set grows too. Production failures should become cases, which means a process for adding them and a shared understanding of what that does to the reported score. See how to run an ai evaluation program.
Step 3 â Set the decision boundary
The client decides what the system does, what it may do unsupervised, and where the human boundary sits. The provider operates it and maintains quality within those decisions.
Writing this down prevents the common drift where a provider gradually makes product decisions because they are closest to the system. That is convenient and it moves accountability somewhere the client cannot see.
It also protects the provider. Operating within stated decisions is defensible; making them implicitly is not.
Step 4 â Define how change is handled
Model version updates, corpus changes, prompt adjustments, and configuration changes: who proposes, who approves, what evaluation runs, and how it is recorded.
Unmanaged change is the main way quality declines under a service arrangement. A provider adopting a new model version without evaluation, or a client's team editing the corpus without telling anyone, produces degradation nobody attributes.
Model deprecation deserves its own clause, because it forces work on a timeline neither party controls. See how to handle ai model deprecation.
Step 5 â Agree the cost mechanism
Volume changes and so does cost. Agree how: pass-through of model costs with a margin, banded pricing, or a fixed fee with volume thresholds.
Fixed fees with no mechanism produce a provider incentivised to minimise cost at the expense of quality, or a renegotiation when volume grows. Both damage the relationship.
Make the cost per task visible in reporting. A client who can see it participates in optimisation rather than suspecting it.
Step 6 â Write exit provisions that mean something
Notice period, transfer of the evaluation set, decision records, documentation, and prompt history, plus a transition period with defined support.
Exit provisions that cover only notice leave a client unable to leave in practice, which produces a relationship maintained by dependency rather than by performance.
Providers who write real exit terms tend to be the ones worth staying with. It is a signal, and it is worth offering rather than conceding. See how to transition ai from vendor to in-house.
What should reporting contain?
Quality against the evaluation set, incidents and their resolution, cost per task, volume, escalation rates, and changes made in the period.
Monthly is usually right. Reporting that arrives quarterly is too slow to act on, and weekly reporting for a stable system produces noise.
Include the things that went wrong. Providers who report only positives are read as incomplete, and the first incident that surfaces independently damages trust disproportionately.
How do you avoid the client losing capability?
By building knowledge transfer into the arrangement rather than treating it as an exit activity.
Joint incident reviews, shared documentation, and periodic sessions where the client's team changes something keep capability alive. Arrangements that keep the client entirely outside the system produce a dependency that neither party intended.
Where a client explicitly wants to remain hands-off, say so in the contract so the expectation is shared.
Who needs to be involved?
An operations lead on the provider side, a named owner on the client side, and someone on each side who can make decisions without escalating.
Arrangements where every decision escalates on both sides are slow in exactly the moments that matter.
How long does it take?
Four to eight weeks to establish, including agreeing the evaluation set and the change process. Arrangements started without those take longer to stabilise.
What are the common failure modes?
Availability-only service levels. No agreed evaluation set. Provider making product decisions. Unmanaged model change. No cost mechanism. And exit provisions that only cover notice.
How do you know it worked?
Quality reported against an agreed set, incidents handled without escalation drama, cost predictable as volume changes, and a client who stays because the service is good rather than because leaving is hard.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Write down what quality means for this system and how it will be measured. Everything else in the arrangement depends on that answer.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: scope covering quality against an agreed evaluation set rather than availability alone, decision boundaries and change handling written down before operation begins, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to run AI operations daily.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What should the scope cover?
Operation, monitoring, incident response, model and corpus change handling, evaluation, and reporting. Availability is table stakes; the failure mode that matters is a system that is up and producing worse results.
02How is quality measured in a service arrangement?
Against an evaluation set both parties agreed, run on a schedule, with the results reported. Quality commitments without an agreed measurement become arguments, and arguments in a service relationship are expensive.
03Who decides what the system does?
The client. A managed service operates the system and maintains its quality; changing what it decides, what it may do, or where the human boundary sits is the client's call and should be written that way.
04How should change be handled?
With a defined process for model updates, corpus changes, and prompt adjustments, including who approves what and how evaluation is run. Unmanaged change is how quality declines invisibly under a service arrangement.
05Why do exit provisions matter?
Because they keep the arrangement honest. A client who could leave â with the evaluation set, the decision records, and the documentation â is a client the provider serves rather than holds.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.