Governance · 5 minute read
AI and SOC 2 Type 2: Evidence Over a Period
SOC 2 Type 2 tests whether controls operated effectively across a period, so AI systems become an evidence problem rather than a design one. The criteria they touch hardest are change management, logical access, vendor management, and monitoring, with model providers usually treated as subservice organisations.
A Type 2 examination tests whether controls operated effectively across a period rather than whether they were designed well at a point. For AI systems that turns compliance into an evidence problem, and the evidence has to exist throughout. This guide covers what auditors examine, drawing on FISTA Solutions' AI enablement work. This article is general guidance, not legal advice.
Which criteria do AI systems touch?
Mostly the ones you already have, applied to components that behave differently.
| Area | What AI adds |
|---|---|
| Change management | Model version updates are changes |
| Logical access | Model outputs are data with access controls |
| Vendor management | Model providers as subservice organisations |
| Monitoring | Quality drift alongside availability |
| Confidentiality | Prompts and outputs in scope |
| Incident management | AI failures routed through the process |
Why do model updates matter so much?
Because a model version change alters system behaviour, which is precisely what change management exists to control.
Systems that adopt provider updates automatically have changes occurring without records. Auditors treat that as a control gap, and the fix — pinning versions and updating deliberately with evaluation evidence — is the same thing good engineering practice would ask for anyway. See what is a regression suite for ai.
How are model providers treated?
Usually as subservice organisations, which gives two options: carve them out of the report while describing the complementary controls customers should assume, or include them with supporting evidence such as their own attestation.
The choice affects what your report says and what customers read into it. Make it deliberately with your auditor rather than discovering it in the draft.
What does operating effectiveness actually look like?
Evidence that the control ran consistently across the whole period.
A change record for every model update, access reviews performed on schedule, monitoring alerts triaged with outcomes, and vendor reviews completed. A control that operated for ten months and lapsed for two produces an exception, and exceptions in a Type 2 report are visible to every customer who reads it.
What evidence do you need?
Change records covering every model version update, access review evidence across the period, monitoring and alerting records with triage outcomes, vendor and subservice assessments, and logs showing who accessed model outputs.
If that evidence exists as a by-product of how systems are built and operated, you are in good shape. If it exists only as documents written for a review, you are not, and the difference is visible to anyone who looks carefully.
How does this change engineering practice?
It pushes version pinning, change records, and output access logging into the build. All three are modest engineering tasks that are impossible to reconstruct after the period has passed.
The practical discipline is treating the model as a dependency with a version, like any library: pinned, updated through a pull request with evaluation results attached, and recorded. Teams that do this produce audit evidence as a by-product.
How does it interact with other regimes?
Usually more than expected. The same system can attract questions from a data protection authority, a sector supervisor, and a general AI regulator, each starting from a different premise and arriving at overlapping requirements.
One evidence base mapped to several requirements answers all of them. Separate programmes produce separate documents describing the same systems, and inconsistencies between them are themselves a finding.
What does compliance cost?
Mostly the cost of good engineering practice: evaluation, documentation, logging, and oversight design. Built into a project, the incremental cost is modest and much of it is work the system needed anyway.
Retrofitted onto a live system it becomes a project, performed under a deadline you did not choose, on something people already depend on. See AI compliance audit cost.
What are the common mistakes?
Adopting provider model updates automatically. Treating outputs as ephemeral rather than as accessed data. Discovering the subservice treatment in the draft report. And building evidence at the end of the period rather than throughout.
Who owns this internally?
The function that owns the systems, with legal and compliance support. Ownership by compliance alone produces documents describing systems nobody changed; ownership by engineering alone produces good practice with no one accountable for the interpretation.
Name a person per system rather than a committee. Committees review; people decide.
What should you ask a supplier?
What documentation they provide about capabilities and limitations, what evaluation evidence they share, how they handle personal data, where processing happens, and what happens to your prompts and outputs.
Suppliers who have prepared answer those quickly. Suppliers who have not take weeks, and that delay is itself information about how the relationship will run.
How do you keep this current?
Assign someone to watch the sources that actually bind you rather than general commentary. Record what was checked and when, so the next review starts from a known point.
Rules in this area change, and a position taken eighteen months ago and never revisited is a risk in itself.
How does this compare with a Type 1 report?
A Type 1 tests design at a point in time, which an organisation can prepare for in weeks. A Type 2 tests operation over months, which cannot be prepared for retrospectively.
That is why the practical advice for AI systems is to put change records, access logging, and monitoring in place before the observation period starts. Nothing you do afterwards creates evidence for a period that has already passed.
What should you do first?
Check whether you can produce a change record for every model version change in the last quarter. If you cannot, that gap is what an examination will find.
How FISTA Solutions helps
FISTA Solutions builds AI systems so the evidence exists when it is needed: model versions pinned and updated through recorded changes with evaluation evidence attached, access to model outputs logged, evaluation results dated and versioned, oversight designed structurally rather than asserted in policy, and documentation produced during the build rather than reconstructed afterwards. Delivery runs through AI enablement, AI agents, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries.
To align a system with these requirements, message FISTA on WhatsApp, or read AI and ISO 27001.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Which criteria do AI systems touch hardest?
Change management, logical access, vendor and subservice organisation management, and monitoring. Confidentiality and privacy criteria apply where the relevant categories are in scope. This is general guidance, not legal advice.
02Are model updates changes?
Yes. A model version change alters system behaviour, which is exactly what change management exists to control. Systems that adopt provider updates automatically have a change occurring without a record, which auditors treat as a control gap.
03How are model providers treated?
Usually as subservice organisations, which means either carving them out with complementary controls described, or including them with supporting evidence such as their own attestation. The choice affects the report and should be made deliberately.
04What does operating effectiveness look like?
Evidence that the control ran consistently across the period: change records for every model update, access reviews performed on schedule, monitoring alerts triaged, and vendor reviews completed. Gaps in the period are exceptions.
05What evidence should you keep?
Change records covering model version updates, access review evidence across the period, monitoring and alerting records with triage, vendor and subservice assessments, and logs showing who accessed model outputs.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.