FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook · 6 minute read

How to Run an AI Proof of Value That Decides Something

A proof of value answers a specific question with real data and ends in a decision. Proofs of concept demonstrate that something is possible, which is rarely in doubt, and they produce enthusiasm rather than the evidence a funding decision needs.

By FISTA Solutions· AI-Native Engineering Team·
How to Run an AI Proof of Value That Decides Something article cover

Proofs of concept demonstrate that something is possible, which is rarely the question anyone actually has. A proof of value answers whether it is worth doing here, and ends in a decision. This playbook covers running one, drawing on FISTA Solutions' AI enablement work.

When is this worth doing?

Before committing significant budget to an AI initiative whose feasibility on your data and process is genuinely uncertain.

It is not worth doing where the approach is well established and the uncertainty is about appetite rather than feasibility. In that case the proof of value is theatre and the real question is a funding conversation.

What does the sequence look like?

StepPurpose
1. Frame the questionSpecific and answerable
2. Set the decision criteriaBefore starting
3. Get real dataIncluding the hard cases
4. Measure the current processThe comparison
5. Build the smallest testNot a product
6. Decide and recordIncluding negatives

Step 1 — Frame a specific question

Write the question the exercise will answer: can this approach classify these documents accurately enough to reduce review time, on our data.

Specific questions are answerable. Broad ones — can AI help our support team — produce demonstrations and no conclusion.

The question should be one where the answer could plausibly be no. Questions whose answer everyone already assumes do not need an exercise.

Step 2 — Set the decision criteria first

What quality threshold, at what cost, within what response time, would justify proceeding — and what result would mean stopping.

Writing these before starting is what makes the conclusion readable rather than arguable. Without them, the result gets interpreted by whoever has the strongest view, and the exercise settles nothing.

Anchor the quality threshold to the current process rather than to a round number. See how to set ai quality thresholds.

Step 3 — Get real data, including the hard cases

Real documents, real questions, real edge cases, with the access arrangements sorted before the clock starts.

Curated data hides exactly the problems the exercise exists to surface. A proof of value run on clean examples produces a positive result and no information.

Data access is frequently the longest lead time, particularly in regulated environments. Start it first and be honest about whether the timebox begins before or after it arrives.

Step 4 — Measure the current process

Time per case, error rate, and volume, for the process as it works today.

Without this there is nothing to compare against, and the conclusion becomes a judgement about whether the outputs looked good. That judgement favours whichever system writes more fluently.

The measurement also frequently changes the framing: teams sometimes discover the current process is faster or more accurate than assumed, which is useful before spending rather than after.

Step 5 — Build the smallest thing that answers the question

Not a product. A script, a notebook, or a rough interface that processes real cases and produces measurable results.

Proof-of-value effort spent on interface polish is effort not spent on the question, and a polished demonstration creates expectations about a delivery timeline that has not been estimated.

Resist the pull towards making it look finished. The output is a decision, not a system.

Step 6 — Decide and record, including negatives

Compare the result against the criteria, make the decision, and write down what was learned.

Record negatives clearly. A proof of value establishing that an approach will not work has saved the delivery cost of discovering it later, and organisations that treat that as failure stop receiving honest results.

Keep the artefacts: the evaluation set, the data preparation, and the measurement. Those are reusable whichever way the decision goes. See how to kill an ai project.

What if the result is ambiguous?

That is common and it is a result. Say what was established, what remains uncertain, and what would resolve it.

The honest options are extending with a specific question, narrowing the scope to the part that worked, or stopping. What does not work is proceeding to delivery on an ambiguous result, which converts the uncertainty into a delivery risk.

Ambiguity frequently means the question was too broad. Narrowing it and re-running is cheaper than delivering into uncertainty.

Who should see the output?

Whoever will decide, with the evidence rather than a summary of the enthusiasm.

Proof-of-value outputs presented as demonstrations get funded on impression. Outputs presented as measurement against criteria get funded on evidence, and the second holds up better when delivery is harder than the demonstration suggested.

Show the failures alongside the successes. Decision-makers who see both trust the result.

Who needs to be involved?

An engineer who can build quickly, someone who can judge correctness in the domain, and the person who will make the funding decision.

The decision-maker's involvement in setting the criteria is what makes the conclusion binding. Criteria set without them get renegotiated at the end.

How long does it take?

Two to six weeks including data access, which is frequently the longest part. Exercises running longer have usually become delivery projects without anyone deciding.

What are the common failure modes?

Demonstrating rather than answering. Curated data. No criteria. No baseline. Polishing the interface. And treating a negative result as a failure.

How do you know it worked?

A decision made on evidence against pre-set criteria, reusable artefacts kept, and a team willing to report a negative result next time.

What does it cost?

Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.

Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.

What should you do first?

Write the question the exercise will answer and what result would mean stopping. If you cannot write the second sentence, the exercise is a demonstration.

How FISTA Solutions helps

FISTA Solutions runs this work alongside client teams rather than around them: proofs framed as questions with decision criteria set in advance, run on real data including the cases that are difficult, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.

To run this with support, message FISTA on WhatsApp, or read how to estimate an AI project.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What is the difference from a proof of concept?

A proof of concept shows something is possible, which is rarely the question. A proof of value establishes whether it is worth doing here, on this data, for this process, at a cost that makes sense.

02Why use real data?

Because curated data hides the problems. Real documents, real questions, and real edge cases reveal where the approach struggles, and that is the information a decision needs.

03What decision criteria should be set?

What quality threshold, at what cost, within what time, would justify proceeding — and what result would mean stopping. Both, written down before starting, so the conclusion is read rather than argued.

04How long should it run?

Two to six weeks. Long enough to use real data and reach a defensible conclusion, short enough that the organisation has not committed to the outcome by the time it finishes.

05Is a negative result a failure?

No. A proof of value that establishes the approach will not work has saved the cost of finding out during delivery, which is the point of doing it. Treating negatives as failures is how organisations stop getting honest ones.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project