Playbook ¡ 6 minute read
How to Version Prompts and Models So Changes Are Traceable
AI systems change through prompts, model versions, retrieval corpora, and configuration, and most teams version only the code. Versioning all of them, recording which combination produced each output, and pinning provider models is what makes behaviour changes traceable and genuinely reversible.
AI systems change through prompts, model versions, and retrieval corpora, and most teams version only the code. That leaves the majority of behaviour changes untracked. This playbook covers versioning all of it, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
From the start of any AI system intended for production, and immediately for any system already there. Retrofitting is possible and loses the history.
It is particularly urgent for systems subject to validation, model risk, or change control requirements, where an unversioned behaviour change is a control failure rather than an inconvenience.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Enumerate what changes behaviour | More than you think |
| 2. Version prompts as artefacts | With review, not in a database row |
| 3. Pin model versions | No unpinned provider calls |
| 4. Version the corpus | Snapshots or change logs |
| 5. Record the combination per output | For reproducibility |
| 6. Make rollback routine | Practised, not documented |
Step 1 â Enumerate what changes behaviour
Prompts, system instructions, model version, temperature and sampling settings, retrieval corpus, chunking and embedding configuration, tool definitions, and thresholds.
Any of those changing alters what the system does. Teams that version code and treat the rest as configuration have most of their behaviour surface untracked.
Write the list for your own system. It is usually longer than expected, and the items nobody thought of are where unexplained changes come from.
Step 2 â Version prompts as artefacts
Keep prompts in version control alongside the code, with review, rather than in a database row someone can edit.
Editable prompts feel convenient and remove the audit trail entirely. A prompt changed in a production database at three in the afternoon is a behaviour change with no record, no review, and no rollback.
Where runtime editing is genuinely needed, version the entries and record who changed what and when. See what is a prompt template.
Step 3 â Pin provider model versions
Call a specific model version rather than an alias that follows the latest release.
Unpinned calls mean your system's behaviour can change without a deployment, which breaks evaluation evidence, change control, and incident attribution simultaneously. A team debugging a regression that turns out to have been a provider update has lost days to an avoidable problem.
Update deliberately: pin the new version, run the evaluation, review the differences, and deploy as a change. That is a slightly slower process and a far more controllable one.
Step 4 â Version the retrieval corpus
Snapshot the corpus or maintain a change log with timestamps, so you can say which state produced a given answer.
Corpus changes are behaviour changes. A document added, removed, or edited alters answers, and a system that cannot say what its knowledge base contained yesterday cannot explain yesterday's answer.
Full snapshots are expensive for large corpora; a change log with document versions and timestamps is usually sufficient and much cheaper.
Step 5 â Record the combination per output
Every logged output should carry the prompt version, model version, corpus version, and configuration in force.
That is what makes an output reproducible. Without it, investigating a failure means guessing which of several things changed, and the guess is frequently wrong.
It is also what makes evaluation results meaningful over time. A score from three months ago describes a system you can only identify if the versions were recorded. See what is a regression suite for ai.
Step 6 â Make rollback routine
Rolling back a prompt, model version, or configuration should be a routine operation that someone has done recently.
Rollback paths that exist only in documentation fail when used, because configurations drift and nobody noticed. Exercising it during ordinary operations makes it reliable during incidents.
Include the corpus in the rollback story. A retrieval system whose corpus changed badly needs a way back, and that is harder to arrange after the fact.
How does this interact with experimentation?
It makes experiments interpretable. An A/B test or shadow comparison between two versions is only meaningful if you can say precisely what differed.
Systems without version discipline produce experiments whose results cannot be attributed, which wastes the experiment and, worse, produces confident conclusions from a comparison nobody can characterise.
What about fine-tuned models?
They need versioning too: the base model, the training data version, the training configuration, and the resulting artefact.
A fine-tuned model is a build output, and treating it as one â reproducible from recorded inputs â is what allows it to be rebuilt, compared, and audited. Models that exist as a file nobody can reproduce are a dependency with no source.
Who needs to be involved?
An engineer to implement it, and whoever owns change control or validation where those apply.
It is largely an engineering discipline, and the regulatory stakeholders should know it exists because it is what their requirements depend on.
How long does it take?
One to two weeks to implement in a system of moderate size. Retrofitting into a large estate takes longer and is worth doing incrementally, starting with model pinning.
What are the common failure modes?
Versioning only code. Unpinned model calls. Prompts editable in production. Unversioned corpora. Outputs logged without versions. And rollback that has never been exercised.
How do you know it worked?
Any output reproducible from its recorded versions, regressions attributable to a specific change, and rollback performed without drama when needed.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Check whether any of your production calls use an unpinned model alias. That is the single change with the largest immediate benefit.
What about configuration that lives in a dashboard?
Provider consoles and vendor dashboards frequently let someone change a setting that alters behaviour, with no record reaching your version control. Treat those as production configuration: restrict who can change them, and mirror the values into your own records so a discrepancy is visible.
Systems where a support engineer can change a temperature setting in a console have an unversioned behaviour surface, and it will eventually explain a regression nobody could attribute.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: prompts, models, corpora and configuration all versioned with the combination recorded per output, provider versions pinned so behaviour cannot shift silently, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to set up AI change control.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What needs versioning in an AI system?
Prompts, model versions, retrieval corpora, configuration such as temperature and thresholds, and tool definitions. Any of them changing alters behaviour, and versioning only the code leaves most changes untracked.
02Why pin model versions?
Because providers update models, and an unpinned call means behaviour can change without any deployment on your side. That breaks evaluation evidence, change control, and the ability to attribute a regression to anything.
03Why version the corpus?
Because a retrieval system whose knowledge base changed has changed. Documents added, removed, or edited alter answers, and a system that cannot say which corpus state produced an answer cannot explain it.
04What should be recorded per output?
The prompt version, model version, corpus version, and configuration in force at the time. That combination is what makes an output reproducible and a regression attributable to a specific change rather than to a guess between several things that moved at once.
05How does this relate to evaluation?
Directly. An evaluation result recorded without versions attached describes a system nobody can identify afterwards. Results should carry the full combination so a comparison made three months later is measuring what you think it is measuring rather than an unknown configuration.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.