FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Trends · 5 minute read

The Professionalization of Prompt Work Into Real Engineering

Prompt work is becoming ordinary engineering. Prompts determine system behaviour, so they need version control, review, tests, and named ownership — the same practices every other kind of behaviour-defining configuration eventually acquired, usually after an incident made the absence obvious.

By FISTA Solutions· AI-Native Engineering Team·
The Professionalization of Prompt Work Into Real Engineering article cover

Prompts determine system behaviour, which makes them code, and the practices around them are catching up to that. This piece covers the shift, drawing on FISTA Solutions' AI agents engineering work.

What does professional practice look like?

Six practices, none of them novel.

PracticeWhy it applies
Version controlBehaviour changes need history
Code reviewA prompt change is a behaviour change
Evaluation on every changeRegressions are otherwise invisible
Named ownershipPrevents unexplained accumulation
Deployment through the pipelineConsole edits bypass control
Documented intentFuture readers need the why

Why did prompts escape these practices?

Because they look like text rather than like code.

A prompt is written in prose, edited in a text box, and adjusted by whoever notices a problem. None of that resembles software, so none of the software practices were applied.

The consequence is prompts containing layers of instruction added over time by different people for reasons nobody recorded. Removing any line risks breaking something, so nothing is removed and they grow. See prompt review checklist.

What does review catch?

Instructions that conflict, that leak sensitive information, or that will not survive a model change.

Conflicting instructions are common in accumulated prompts — one section says be concise, another added later asks for thorough explanation. The model resolves the conflict unpredictably.

Review also catches prompts containing customer data, internal identifiers, or credentials, which happens more than teams expect because prompts are drafted against real examples.

What does testing involve?

An evaluation run comparing the new version against the old across a representative suite.

A prompt change made to fix one complaint frequently degrades other cases. Without measurement, the team ships an improvement that is net negative and discovers it through a different complaint weeks later.

This is the practice that most distinguishes teams with evaluation infrastructure. Without a suite, prompt changes are guesses. See how to build an agent evaluation harness.

Why does storage location matter?

Because prompts outside the repository are outside change control.

A prompt edited in a vendor console has no history, no review, no link to a ticket, and no relationship to the deployed version of anything else. Reproducing a past behaviour becomes impossible.

Keep prompts in version control, deploy them with the code, and log which version produced each output. That last point is what makes incident investigation possible. See how to set up AI change control.

What does ownership prevent?

Accumulation without pruning.

An owned prompt gets reviewed periodically, with lines removed when the reason for them has passed. An unowned one only grows, because adding is safe and removing is not.

Ownership also means someone can answer why an instruction is there, which is the question that makes the difference between a maintainable prompt and an artefact everyone is afraid of.

Is the specialist role going away?

It is being absorbed, which is what happens to most specialisms that matter.

The skill of expressing a task precisely, anticipating failure modes, and structuring output remains valuable. It is becoming part of what an AI engineer does rather than a separate job title.

That mirrors how database tuning, front-end performance, and build engineering evolved: specialist periods followed by absorption into general practice, with deep experts remaining for the hardest cases. See the new shape of engineering teams.

What is the counter-argument?

The counter is that heavy process around prompts slows the rapid iteration that makes them useful, and during exploration that is right. The practices here apply to production prompts — the ones customers depend on — not to experimentation, where friction is genuinely counterproductive.

What does this change for engineering teams?

It means prompt files in the repository, prompt changes in pull requests, and evaluation as a pipeline stage. None of that is new infrastructure; it is existing infrastructure applied to a new artefact.

It also means prompt versions logged with outputs, so a quality question can be traced to a specific change.

What does this change for buyers?

It means asking vendors how their prompts are managed and tested, because a vendor editing prompts in production without evaluation is changing your system's behaviour without measurement.

And asking whether you can see the prompts at all, which some products do not permit.

What should leaders do about it now?

Require prompts in version control and evaluation on every change. Those two rules eliminate most of the unexplained behaviour changes teams report.

Then assign ownership, because an unowned prompt only accumulates.

What changes with agents?

More prompts, with more interaction between them. An agent has a system prompt, tool descriptions, and step-level instructions, all of which shape behaviour and can conflict.

Tool descriptions in particular are prompts that determine whether the agent uses a capability correctly, and they are rarely reviewed as such. See how to build an AI agent.

How will you know if this is happening?

Watch for prompts edited in consoles, for behaviour changes nobody can trace, and for instructions nobody can explain. Each indicates the practices have not caught up.

How FISTA Solutions reads this

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: prompts kept in version control, reviewed as behaviour changes, and evaluated on every revision with the version logged against each output, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To discuss what this means for your roadmap, message FISTA on WhatsApp, or read how to set up AI change control.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why treat prompts as code?

Because they determine what the system does. A change to a prompt can alter behaviour as much as a code change, so it deserves the same version control, review, and testing.

02What goes wrong without this?

Prompts edited in consoles with no history, instructions accumulating that nobody can explain, and behaviour changes nobody can trace. Then a regression appears and there is no diff to examine.

03What does testing a prompt mean?

Running an evaluation suite against it and comparing results to the previous version. A change that improves one case and degrades three is only visible with systematic measurement.

04Is the prompt engineer role disappearing?

It is merging into engineering. The skill remains valuable; it is becoming part of building AI systems rather than a separate job, in the same way database tuning did.

05Where should prompts live?

In version control alongside the code that uses them, deployed through the same pipeline. Prompts in a vendor console or a document are outside change control by definition.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project