FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Leadership · 5 minute read

How to Respond to an AI Vendor Failure

AI vendors fail in five ways: outage, quality degradation, security breach, pricing shock, and discontinuation. Each needs a different response, and all are survivable if the company has a tested alternative, a gateway that makes switching a configuration change, and contract terms that were negotiated before the failure.

By FISTA Solutions· AI-Native Engineering Team·
How to Respond to an AI Vendor Failure article cover

AI vendor failures are not hypothetical: models are deprecated on notice periods measured in months, providers have outages, pricing changes, and startups disappear. The question is not whether one affects you but whether the response is a configuration change or a crisis. This guide covers the five failure types, the response to each, and how to reduce exposure.

What are the five failure types?

FailureWhat it looks likeFirst responsePreparation that helps
OutageRequests failing or timing outFail over or degrade gracefullyTested alternative; gateway failover
Quality degradationNo errors; outputs quietly worseWithdraw autonomy; evaluate; switch if neededScheduled evaluation; monitoring
Security breachProvider discloses an incidentAssess exposure; rotate credentials; consider suspensionData minimization; breach notice terms
Pricing shockMaterial increase or model changeQuantify; check terms; reroute volumePrice protection; alternatives
DiscontinuationModel deprecated or vendor shuts downExecute migration; recover assetsMigration plan; exit terms; export

How should an outage be handled?

Mechanically, if the preparation exists. The gateway fails over to a validated alternative model; agents continue with a possible quality difference that has been measured in advance. Where no alternative is validated, agents degrade to human handling for the affected processes, which requires that the manual path still exists and that people know it.

The executive decision that matters was made months earlier: whether a second model was evaluated against the same test set and kept warm. The AI gateways explained for executives piece covers the layer that makes failover a configuration change.

Why is quality degradation the hardest failure?

Because nothing breaks. A provider updates a model version, and the agent continues to run, returning outputs that are subtly worse: more hedging, different formatting, weaker reasoning on edge cases. There is no alert, no error rate spike, and no ticket. It surfaces weeks later as a rise in exceptions or a customer complaint.

The only reliable control is scheduled evaluation on production samples, with alerts on pass-rate movement. When degradation is detected: withdraw autonomy for the affected action classes immediately, run the evaluation against alternatives, switch if one passes, and raise it with the provider with the evidence, which is a considerably stronger conversation than a subjective complaint. The AI evaluation explained for executives piece covers the discipline.

What if the vendor is breached?

Treat it as your incident, because your data may be involved. Establish what data of yours the vendor held or processed and over what period; assess your own notification obligations, which may be triggered regardless of the vendor's; rotate credentials and revoke access; review what the vendor's integration could reach in your systems; and consider suspending use while facts are established. Contractual breach notice and liability terms determine what you are owed. Consult counsel; this is general guidance, not legal advice.

How should a pricing shock be handled?

Quantify first: the effect on cost per task for each affected agent and on total run cost at current and projected volume. Then check the contract for notice periods and price protection. Then evaluate alternatives on your own cases, and route price-sensitive volume accordingly. The model routing explained for executives piece covers the mechanism that makes selective rerouting possible.

A company that can move volume negotiates differently. That position is created by the gateway and the evaluation set, not by the negotiation itself.

What about discontinuation?

For a model deprecation, execute the migration plan: validate the alternative against the evaluation set, switch at the gateway, revalidate affected agents, and document the change for any regulated context. Notice periods are usually months, which is adequate if the plan exists and inadequate if it does not. The model deprecation risk management guide covers the planning.

For a vendor shutdown, act immediately on asset recovery: data, prompts, configurations, connector definitions, and evaluation assets, under the exit terms, while the company is still responding. Distressed vendors stop supporting exports before they stop operating.

What should happen after any failure?

A review that asks not only "how do we replace this vendor" but "why were we this exposed." The durable responses are structural: a gateway, a tested alternative, evaluation that detects degradation, exit terms in contracts, and concentration limits in the risk appetite. Replacing one single point of failure with another is the common mistake. The chief risk officer's guide to AI agents covers the concentration view.

What should executives ask now?

  • Which providers are load-bearing, and for which agents?
  • Do we have a validated alternative, and when was it last tested?
  • Would we detect quality degradation, and how quickly?
  • What do our contracts say about notice, price protection, breach, and exit?
  • Could we recover our prompts, configurations, and data tomorrow?

How can FISTA Solutions help?

FISTA Solutions builds AI agents behind a gateway with validated alternative models, scheduled evaluation that detects degradation, and exportable assets, and works with executive and risk teams through its AI enablement practice on concentration limits, contract requirements, and migration plans. Since 2017, FISTA has delivered 150+ projects for 50+ companies across 12+ countries, with a 99.9% uptime record on production systems.

To assess your exposure before a provider forces the question, talk to FISTA on WhatsApp, or read LLM vendor lock-in.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What should a company do during an AI provider outage?

Fail over to the tested alternative if one exists, degrade agents gracefully to human handling where it does not, communicate to affected internal and external users, and record the duration and impact for the service level claim. The decision that matters was made earlier: whether an alternative was validated in advance.

02How do you detect and respond to AI quality degradation?

Detect through scheduled evaluation on production samples and monitoring of exception and agreement rates; a quality collapse produces no error message. Respond by withdrawing autonomy for affected action classes, switching models if the alternative passes, and raising it with the provider with your evidence.

03What if an AI vendor suffers a security breach?

Treat it as your incident: determine what data of yours was exposed, assess notification obligations, revoke and rotate credentials, review what the vendor's systems could access, and consider suspending use pending facts. Contractual breach notice terms determine what you are owed; consult counsel.

04How should a company handle an AI price increase?

Quantify the effect on cost per task, check contractual notice and price-protection terms, evaluate alternatives on your own cases, and route price-sensitive volume elsewhere if the numbers justify it. Companies with a gateway and a tested alternative negotiate from a different position than those without.

05What if an AI vendor discontinues a model or shuts down?

Execute the migration plan that should already exist: validate the alternative against your evaluation set, switch at the gateway, and revalidate affected agents. For a vendor shutdown, recover data, prompts, and configurations under your exit terms immediately, while the company still responds to requests.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project