FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Trends ¡ 5 minute read

The Next Five Years of AI Engineering: A Grounded View

Predictions about AI usually overshoot capability and undershoot operational difficulty. The likely direction is better tooling, harder operational problems as systems take on work with real consequence, and a discipline that increasingly resembles ordinary engineering with probabilistic components sitting inside it.

By FISTA Solutions¡ AI-Native Engineering Team¡
The Next Five Years of AI Engineering: A Grounded View article cover

Predictions about AI usually overshoot capability and undershoot operational difficulty. This is a grounded view of the engineering discipline's direction, drawn from FISTA Solutions' AI agents and AI enablement delivery work.

What is likely to change?

Six directions with reasonable confidence.

DirectionConfidence
Evaluation becomes standard toolingHigh
Cost engineering becomes a disciplineHigh
Governance becomes assumedHigh
Small models take more production workHigh
Agent operations becomes a roleModerate
Interoperability standards settleModerate

Why will evaluation standardise first?

Because every team needs it and everyone is building it separately.

That is the classic precondition for tooling to consolidate. Case management, scoring, regression detection, and pipeline integration are the same problems in every organisation, differing only in the cases and criteria.

Expect this to follow the path that testing frameworks took: bespoke, then libraries, then standard infrastructure nobody thinks about. See why evaluation is the new moat.

Why does cost engineering become a discipline?

Because the spend is large, variable, and optimisable.

Routing between models, caching, prompt compression, batching, and capacity planning are all specialist work with measurable returns. At sufficient scale, a person doing this full time pays for themselves repeatedly.

That is how performance engineering and cloud cost management became roles, and the pattern is the same. See how to reduce AI costs.

What happens to governance?

It becomes assumed rather than negotiated.

Audit logging, retention, disclosure, and inventory are moving from things a mature organisation does to things every organisation is expected to do. Regulation is part of that; customer expectation is the larger part.

The practical effect is shared infrastructure and standard patterns, because per-project governance does not scale. See the compliance layer of AI.

Why do small models grow their share?

Because most production tasks are narrow and the economics are decisive.

As routing infrastructure matures and fine-tuning becomes routine, the default architecture becomes several models with a router rather than one model for everything. The frontier model handles the hard cases; small models handle volume.

On-device inference extends this, particularly where privacy or latency dominates. See the quiet rise of small models.

What does agent operations look like?

A role concerned with what deployed agents are doing and whether it is right.

Monitoring trajectories, tuning permissions, reviewing escalations, investigating incidents, and managing the exception queue are ongoing responsibilities that do not fit neatly into existing roles.

This is the least certain prediction, because it depends on agent adoption reaching a scale that justifies dedicated staffing. In organisations that get there, the role seems inevitable. See how to staff an AI support rotation.

What stays stubbornly the same?

The organisational difficulty.

Data access approvals, unclear process ownership, security review capacity, and change management resistance are not technology problems and will not be solved by better technology.

Organisations that address them will move faster than their technical sophistication would predict, and the reverse is equally true. See the real bottleneck in enterprise AI.

What is the counter-argument?

The counter is that a sufficiently large capability jump could invalidate this entire framing, making the operational disciplines unnecessary. That is possible and it has been predicted before each previous release. The disciplines described here have become more necessary as capability has grown, not less, which is the better guide.

What does this change for engineering teams?

It means investing in the durable parts: evaluation, observability, cost attribution, and validation. Those have survived every model generation so far and there is no reason to expect otherwise.

It also means treating AI work as ordinary engineering with an unusual component, rather than as a separate discipline with its own rules.

What does this change for buyers?

It means preferring vendors investing in operational surfaces over those competing on capability, because capability converges and operations do not.

And assuming your requirements will grow: audit, cost visibility, and export will all matter more in three years than they do now.

What should leaders do about it now?

Fund the durable capabilities — evaluation, governance infrastructure, cost attribution — rather than chasing the capability frontier.

And fix the organisational bottlenecks, which are the constraint regardless of what the technology does.

What would change this view?

Reliability improving to the point where evaluation and review become unnecessary for consequential decisions. That would remove much of the operational burden described here.

Nothing in the current trajectory suggests it, and accountability requirements would likely persist even if accuracy did. But it is the assumption to watch.

How will you know if this is happening?

Watch for evaluation tooling consolidating, for cost engineering job titles appearing, and for governance questions arriving before technical ones in procurement. Each confirms the direction.

How FISTA Solutions reads this

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: investment concentrated in the durable layers — evaluation, observability, validation, cost attribution — that have survived every model generation so far, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To discuss what this means for your roadmap, message FISTA on WhatsApp, or read why evaluation is the new moat.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What is most likely to improve?

Tooling. Evaluation platforms, observability, and orchestration will mature into standard infrastructure the way testing and deployment tooling did, which removes work teams currently build themselves.

02What gets harder?

Operations. As systems take on more consequential work, the cost of failure rises, which raises the bar on monitoring, governance, review design, and incident response.

03Will model capability keep improving?

Almost certainly, and the practical effect is to widen the set of viable use cases rather than to remove the engineering around them. Capability has not historically reduced operational work.

04What new disciplines emerge?

Cost engineering as a named specialism, evaluation design as a distinct skill, and agent operations as a role — the equivalent of what site reliability became for distributed systems.

05What stays the same?

That most of the difficulty is organisational: data access, process ownership, and change management. No capability improvement addresses those.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project