FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook ¡ 6 minute read

How to Run an AI Maturity Assessment Worth Acting On

An AI maturity assessment is worth doing when it is scored against evidence rather than opinion and finishes with a sequenced plan rather than a score. Assessments based on self-reported questionnaires reliably overstate capability and produce charts nobody acts on.

By FISTA Solutions¡ AI-Native Engineering Team¡
How to Run an AI Maturity Assessment Worth Acting On article cover

Maturity assessments usually produce a chart and no change, because they are scored on what people say rather than on what exists. This playbook covers assessing against evidence and finishing with a plan, drawing on FISTA Solutions' AI enablement work.

When is this worth doing?

When an organisation wants to know where it actually stands before committing to a programme, or when several AI efforts are underway with inconsistent quality.

It is not worth doing as an annual ritual. Assessments run on a calendar with no decision attached produce documents, and the second one is read even less than the first.

What does the sequence look like?

StepPurpose
1. Define the dimensionsBroader than technology
2. Ask for artefactsNot descriptions
3. Sample real systemsRather than surveying opinion
4. Score against evidenceWhat exists, not what is planned
5. Identify the blocking gapsTwo or three, not twenty
6. Produce a sequenced planOwners and dates

Step 1 — Define the dimensions honestly

Delivery capability, operational practice, data readiness, evaluation and quality, governance and compliance, and adoption.

Assessments covering only technology mislead, because most organisations' constraint is somewhere else: review capacity, data ownership, or the absence of anyone accountable for a system in production.

Keep the dimension count small. Frameworks with twenty dimensions produce assessments that take longer than the remediation would have.

Step 2 — Ask for artefacts, not descriptions

For each dimension, request something concrete: the inventory, an evaluation result with a date, a change record, an incident postmortem, a decision record, a runbook.

What exists is the assessment. A team that describes a rigorous evaluation practice and cannot produce a result from the last month has told you the score.

This is also faster than interviewing. Twenty minutes looking at artefacts is worth an hour of discussion about intentions.

Step 3 — Sample real systems

Pick three or four systems across different teams and assess those specifically rather than surveying the organisation's opinion of itself.

Organisational averages hide the variance that matters. One team with good practice and three without is a different situation from uniform mediocrity, and it calls for a different plan.

Include at least one system nobody nominated. Self-selected examples are the best ones.

Step 4 — Score against evidence

Rate each dimension on what you saw rather than on what was described, with the artefact recorded as the basis.

Be willing to score low. An assessment that finds everything adequate has either measured a genuinely mature organisation or has not looked hard enough, and the second is considerably more common.

Record the evidence alongside the score. That makes the assessment defensible and makes reassessment meaningful later.

Step 5 — Identify the two or three blocking gaps

Most organisations have many gaps and are blocked by a small number.

A team without evaluation infrastructure cannot improve any system safely, which blocks everything downstream. A team without an inventory cannot govern anything. Those are blocking; a missing prompt library is not.

Say which gaps block the others. A plan addressing twenty gaps in parallel achieves less than one addressing three in sequence.

Step 6 — Produce a sequenced plan

Owners, dates, and an order, addressing the blocking gaps first.

The assessment's value is entirely in this output. A score without a plan is information nobody is obliged to act on, and it will be quoted in a slide next year with nothing having changed.

Agree the plan with the people who will do it before publishing the assessment. Plans imposed by an assessment are implemented reluctantly.

Should you use a published maturity model?

As a checklist rather than as a framework to be scored against.

Published models are useful for making sure you did not forget a dimension. They are less useful as a scoring scheme, because the levels are generic and the distance between them varies enormously by organisation.

What matters is the evidence and the plan, not the level number.

What about comparing against other organisations?

Interesting and rarely actionable. Benchmarks against an industry average tell you where you sit and not what to do.

The useful comparison is against your own capability six months ago and against what your own systems need. An organisation whose systems require evaluation infrastructure needs it regardless of what peers have.

Who needs to be involved?

An assessor who did not build the systems, the owners of the systems being sampled, and someone with the authority to fund the plan.

Self-assessment produces optimistic results reliably. An outside perspective — internal audit, another team, or an external party — is what makes the evidence-based approach work.

How long does it take?

One to two weeks for a focused assessment of three or four systems, including the plan. Assessments running longer are usually surveying opinion rather than examining evidence.

What are the common failure modes?

Self-reported questionnaires. Surveying opinion rather than sampling systems. Twenty dimensions. Scoring everything adequate. Producing a chart. And a plan with no owners.

How do you know it worked?

A score supported by artefacts, agreement on which gaps block the others, and a plan with owners that is underway within a month.

What does it cost?

Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.

Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.

What should you do first?

Ask one team to show you their evaluation results and their inventory. Whether those exist tells you more than a questionnaire will.

What about organisations with no AI in production?

Assess readiness rather than maturity. The useful questions are different: is there a candidate workflow with a measurable baseline, does the data exist, is there anyone who could own a system in production, and is there a sanctioned path to the tooling.

Those four answers tell an organisation whether it is ready to start, which is a more useful output than a low score across dimensions that do not yet apply. Scoring an organisation against operational practices for systems it does not have produces a chart of zeroes and no insight.

How FISTA Solutions helps

FISTA Solutions runs this work alongside client teams rather than around them: assessment scored against artefacts rather than self-reported practice, finishing with a sequenced plan that names the gaps blocking the others, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.

To run this with support, message FISTA on WhatsApp, or read how to run an AI governance review.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why do maturity assessments fail?

Because they are scored on self-reported answers. Teams describe intentions and best cases, which produces a flattering picture and a plan addressing gaps that are not the real ones.

02What should be scored against?

Artefacts. Ask to see the inventory, an evaluation result, a change record, an incident postmortem, and a decision record. What exists is the score; what people describe is a different measurement.

03What dimensions matter?

Delivery capability, operational practice, data readiness, evaluation and quality, governance, and adoption. Assessments covering only technology mislead, because most organisations' real constraint is operational or organisational rather than technical, and a technology- only view will not surface it.

04What should the output be?

A sequenced plan with owners and dates, addressing the two or three gaps that block the most. A radar chart is a summary of the assessment rather than a reason to have done it.

05How often should it run?

Annually at most, or when something material changes such as a new regulatory obligation or a significant shift in the portfolio. More frequent assessment measures the assessment process rather than capability, which moves slowly and does not need quarterly inspection.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project