Playbook ┬╖ 6 minute read
How to Decommission an AI System Properly
Decommissioning an AI system properly means confirming nothing depends on it, deciding what happens to its data and records, communicating to everyone affected, closing provider arrangements, and recording that it was retired. Systems abandoned rather than retired keep processing data nobody watches.
Decommissioning is the lifecycle step organisations skip. Systems get abandoned rather than retired, which leaves them running, holding credentials, and processing data nobody watches. This playbook covers doing it properly, drawing on FISTA Solutions' AI enablement work.
When is this worth doing?
When a system has been replaced, when its use case no longer exists, when usage has fallen below the cost of operating it, or when it can no longer meet requirements it is subject to.
Also when nobody will own it. An unowned system in production is a liability, and retiring it deliberately is better than letting it decay.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Confirm the decision | Owner, reason, and date agreed |
| 2. Find every dependency | Logs, users, and an announced date |
| 3. Decide on the data | Delete, retain, or migrate |
| 4. Communicate | Everyone affected, with alternatives |
| 5. Shut down in stages | Disable, observe, then remove |
| 6. Close and record | Contracts, credentials, inventory |
Step 1 тАФ Confirm the decision and the reason
Get agreement from whoever owns the system and whoever depends on it, with a written reason and a target date.
The reason matters for later. Systems retired for cost reasons sometimes need reviving when the use case returns, and knowing why it went helps decide whether to rebuild or restore.
Where the system was serving a real need, identify the alternative before announcing. Retiring something people rely on without a replacement produces the same shadow usage that vendor consolidation causes.
Step 2 тАФ Find every dependency
Check what calls it in logs, ask the teams that use it, and announce a shutdown date to see who objects.
Each method catches different things. Logs find programmatic callers including ones nobody remembers building. Asking finds human workflows that route through it. The announcement finds the dependencies neither of the others surfaced, and it finds them while there is still time.
Allow a genuine window between announcement and shutdown. A week is not enough for someone on leave to notice.
Step 3 тАФ Decide what happens to the data
Every store the system touched needs a decision: the primary database, prompt and output logs, evaluation datasets, vector indexes, caches, and anything exported to analytics.
Default behaviour is not a decision. Systems shut down without one leave data in place indefinitely, which is a retention problem and a security surface.
Where personal data is involved, deletion has obligations attached and so does retention. Establish which applies before deleting anything, because both directions are irreversible in different ways. See AI and GDPR data subject rights.
Step 4 тАФ Communicate clearly and early
Tell everyone affected what is happening, when, why, and what to use instead.
Include the people whose workflow depends on it rather than only the technical owners. A system that disappears without warning from someone's daily process produces justified anger and a workaround you will not see.
Repeat the communication as the date approaches. Announcements made once, six weeks out, are forgotten by everyone who was not immediately affected.
Step 5 тАФ Shut down in stages
Disable first, observe, then remove.
A disabled system that can be re-enabled in minutes gives you a safe window to discover the dependency nobody found. Removing infrastructure immediately converts that discovery into an incident.
A week or two of disabled-but-restorable is usually enough. Watch for error rates in other systems during that window тАФ the failures caused by a missing dependency frequently appear somewhere unrelated.
Step 6 тАФ Close arrangements and record it
Cancel provider contracts, revoke credentials and API keys, remove service accounts, delete infrastructure, and stop the billing.
Credentials are the most commonly forgotten. An API key that still works months after a system was retired is an unmonitored access path, and nobody is watching for its use.
Then update the inventory with the retirement date and what happened to the data. An inventory that still lists retired systems is as misleading as one missing live ones. See what is an ai inventory.
What if the system must stay available read-only?
That is a different arrangement and should be named as such: an archive rather than a retired system.
Archives need an owner, a retention period, access controls, and a review date of their own. A system described as retired that is actually still serving requests to a few users is neither, and it will be missed by both the operational and the governance processes.
What about models and artefacts?
Fine-tuned models, embeddings, and evaluation sets are assets with their own handling questions.
A fine-tuned model may be expensive to reproduce and worth archiving, or it may embed data that should be deleted. Evaluation sets are frequently worth keeping regardless, because they represent accumulated knowledge about what the task required.
Decide explicitly rather than letting them persist in storage nobody reviews.
Who needs to be involved?
The system owner, someone from each dependent team, whoever holds the provider contract, and whoever maintains the inventory.
The contract holder is frequently different from the technical owner, which is how subscriptions outlive the systems they supported.
How long does it take?
Two to six weeks including the announcement window and the disabled-but-restorable period. Compressing it removes the safety margin that makes the discovery of a missed dependency survivable.
What are the common failure modes?
Abandoning rather than retiring. Finding dependencies only from logs. No data decision. Forgetting credentials and contracts. Removing infrastructure before observing. And not updating the inventory.
How do you know it worked?
No incidents caused by the shutdown, data handled according to an explicit decision, provider spend actually stopping, credentials revoked, and the inventory accurate afterwards.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
List every credential, API key, and service account the system uses. That list is the part most often left behind and the easiest to check now.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: dependencies found from logs, users, and an announced date rather than one source, data and credentials decided explicitly before anything is removed, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read what is an AI inventory.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why does decommissioning matter?
Because abandoned systems keep running. They continue processing data, holding credentials, incurring cost, and appearing in scope for assessments, and nobody is watching them because nobody owns them any more.
02How do you find dependencies?
From logs of what calls it, from the teams that use it, and by announcing a shutdown date and watching for objections. Each catches different dependencies, and none finds all of them alone.
03What happens to the data?
It needs a decision: what is deleted, what is retained and for how long, and where retained data will live. Prompts, outputs, logs, evaluation sets, and vector indexes all need covering, not just the primary database.
04What about records obligations?
They can outlive the system. Regulated contexts frequently require retaining decision records, validation evidence, and audit logs for defined periods, which means the data has to move somewhere before the system goes.
05What is the final step?
Recording the retirement in the inventory with the date and what happened to the data. An inventory listing systems that no longer exist is as misleading as one missing systems that do.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.