FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Whitepaper · 8 minute read

The First 100 Days of an AI Program: A Whitepaper

In the first 100 days, decide the thesis and evidence standard in week one, choose two outcomes with owners and measure their baselines by week four, build the minimum platform alongside the first agent, reach supervised production by week twelve, and hold the first evidence review with real numbers before day 100.

By FISTA Solutions· AI-Native Engineering Team·
The First 100 Days of an AI Program: A Whitepaper article cover

The first 100 days set the pattern for everything after. A program that spends them on discovery, vendor selection, and strategy documents arrives at day 100 with slides; one that spends them on decisions, a baseline, and a first build arrives with numbers. This whitepaper gives a week-by-week plan and the evidence that should exist at the end of it.

Why does the first 100 days matter disproportionately?

Because it establishes what the organization believes the program is. If the first visible output is a strategy document, the program is understood as a planning exercise and attracts planning behaviors. If it is an agent in production with a measured result, the program is understood as delivery, and business units start bringing candidate processes.

It also establishes the evidence standard in practice rather than on paper. The first time an executive asks for a pass rate instead of accepting a demo, the standard becomes real. The AI-native leadership playbook whitepaper covers the habits this period forms.

Weeks 1–2: decisions

Five decisions, made rather than researched:

DecisionOutputOwner
The thesisOne page: which outcomes change, by how much, by whenCEO or accountable executive
The evidence standardWhat counts as proof before scaling or granting autonomyAccountable executive
Committed outcomesTwo or three named processes with business ownersExecutive team
BoundariesWhat will not be automated, and the policy linesExecutive team with risk and legal
AccountabilityWho runs the program; who owns each outcomeCEO

These do not require discovery. A leadership team knows which processes have volume and pain; the scoring can be done in a session. The how to choose your first AI agent guide covers the criteria, and the how to run an AI executive offsite guide covers doing it in one day.

Weeks 2–4: baselines and specification

Nothing is built until the baselines exist. For each committed outcome: the current cost per unit, cycle time, error or exception rate, and volume, measured from systems rather than estimated, with a named person responsible for the figure.

In parallel, the specification for the first agent: what correct behavior is for each class of input, what the agent does, what it escalates, and what it must never do. This is written with the people who do the work, because they know the exceptions. The spec-driven development practice covers the discipline.

By week four: two baselines, one specification, and the start of an evaluation set built from real cases.

Weeks 3–8: platform minimum and build

The platform minimum is built alongside the first agent, not before it: a model gateway with logging and cost attribution, per-agent identity and permissions, an evaluation harness, and run tracing. That is all. Connectors are built for what the first agent needs, and they become the first entries in a reusable catalog.

Building the platform as a separate project is the most common way to spend a quarter with nothing to show; building an agent with no platform is the most common way to arrive at agent three with three incompatible stacks. The CIO's guide to AI and agentic AI covers the layers.

The build itself proceeds against the specification and the evaluation set, with the pass rate as the release criterion.

Weeks 8–12: supervised production

The agent goes live with human review of every consequential action. During this period:

  • Agreement rates are measured: how often the reviewer accepts the agent's proposed action.
  • Exceptions are examined: what falls outside the specification, and why.
  • The evaluation set grows with every failure.
  • The monitoring is proven: does drift detection work, do alerts reach someone, does the kill switch function?
  • The human role is refined: is the handoff giving reviewers what they need?

The AI agent lifecycle explained for executives piece covers the stages and their gates.

Weeks 6–14: the rhythm, started early

The monthly evidence review starts before there is much to report, in its final format. The first one takes twenty minutes and covers a baseline, a specification, and a build status. The second covers the first production numbers. By the third, the format is established and nobody questions it.

Starting the rhythm when it is easy is what makes it survive when reporting becomes uncomfortable. The agentic operating review whitepaper covers the full cadence system.

Weeks 12–14: second outcome and inventory

The second outcome enters specification, deliberately reusing the first's connectors where possible so the cost difference is visible and quotable. The inventory is established with the first agent's entry: owner, permissions, systems touched, risk tier, evaluation status.

What should exist at day 100?

ArtifactWhy it matters
One agent in supervised production, real volumeProves the company can ship
A baseline-to-actual comparisonProves the result
An evaluation set and pass rateProves quality is measured
A minimum platform in useProves the second agent will be cheaper
An inventory with owner and risk tierProves governance exists
A monthly review held at least twice, same formatProves the rhythm is real
A second outcome specifiedProves momentum

Five of the seven are artifacts rather than opinions, which is the point: at day 100 the program's status should be checkable rather than asserted.

What should be avoided?

A broad discovery exercise. Six weeks of interviews produces a long list and no evidence. Score candidates in a session instead.

A technology selection process. Consumes the quarter and matters less than the specification and the data. Choose a credible provider, keep it replaceable behind the gateway, and move.

A platform project with no agent attached. Produces infrastructure nobody has validated against a real need.

The hardest or most visible use case. Both are poor first choices for reasons covered in the how to choose your first AI agent guide.

Any build without a baseline. The result will be unprovable, and the program will spend its second quarter arguing about whether anything improved.

What does the team look like during this period?

Small, and mostly borrowed. A typical first-100-days team is one accountable executive at perhaps a day a week, two business owners at a day or two a week each during specification and supervised deployment, one or two engineers full time, and a partner or contractor where internal capacity is short. Hiring a team first is a common and expensive detour: the roles that matter become clear during the first build, and recruiting against a real deployment produces better hires than recruiting against a job description written in the abstract.

Two staffing decisions do matter early. Someone must own the platform, even part time, or it becomes nobody's work and fragments. And the business owners must be genuinely available; an owner who cannot attend specification sessions is an owner in name, and the deployment will drift toward what engineers assumed the process was. The how to interview an AI leader guide covers hiring once the shape of the need is clear.

How should the first 100 days be communicated?

Twice, to different audiences, and honestly. Internally, the affected teams should know before the build starts what is changing, what the agent will do, what they will do instead, and what happens to the capacity, including where that is undecided. The specification work makes this natural, because the people doing the process are in the room. The how to communicate AI changes to employees guide covers the message.

To the executive team and board, the communication should set the expectation that day 100 produces evidence rather than transformation: one process, measured, with a rhythm established. Programs that promise broad change in the first quarter spend the second explaining why it did not happen, which costs more credibility than the modest promise would have.

What if the 100 days do not go to plan?

The common failures and their diagnoses: data not accessible (the process was chosen badly, or the access problem was known and ignored); owner unavailable (the wrong process, and it should be swapped rather than persisted with); specification impossible (experienced people disagree on correct behavior, which is a process problem rather than an AI problem and worth solving anyway); pass rate below threshold (either the threshold was wrong or the approach is, and evaluation makes that distinction possible).

None of these is fatal at day 100 if they are named. What is fatal is arriving at day 100 with none of the seven artifacts and a plan to have them next quarter.

What should executives ask at day 100?

  • Is there an agent in production carrying real volume?
  • What was the baseline, and what is the current number?
  • What is the pass rate, on how many cases?
  • Did the monthly review happen twice in the same format?
  • Is the second outcome specified, and will it reuse anything?
  • What did we decide not to automate, and is it written down?

How can FISTA Solutions help?

FISTA Solutions runs first-100-day engagements through its AI enablement practice: the decision session, baseline measurement, specification with the people who do the work, the platform minimum built alongside the first AI agent, supervised deployment, and the operating rhythm installed. Its forward deployed engineers work inside client teams so the capability and the artifacts stay with the company. Since 2017, FISTA has delivered 150+ projects for 50+ companies across 12+ countries; clients report efficiency gains of up to 47% on automated processes.

To start a program that has numbers at day 100, talk to FISTA on WhatsApp, or read how to lead an AI transformation for the wider sequence.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What should be decided in the first week of an AI program?

The thesis (which outcomes change and by how much), the evidence standard (what counts as proof), the two or three committed outcomes with named business owners, the boundaries (what will not be automated), and who is accountable for the program. Discovery can run in parallel; these decisions should not wait for it.

02How quickly can a first AI agent reach production?

About twelve weeks from decision to supervised production for a well-chosen process with an available owner and accessible data. Longer usually indicates the process was too complex, the data was not reachable, or the platform was being built without that being acknowledged in the plan.

03Should the platform be built before the first agent?

Alongside it, never before. A platform built in isolation accumulates features nobody needs and delays evidence by months. A platform built with the first agent contains exactly what an agent requires: gateway, identity, evaluation harness, and tracing.

04What should exist by day 100 of an AI program?

One agent in supervised production with real volume, a baseline-to-actual comparison, an evaluation set and pass rate, a minimum platform, an inventory with owners and risk tiers, a monthly evidence review that has met at least twice, and a second outcome specified.

05What should a new AI program avoid in its first 100 days?

A broad discovery exercise that delays the first build; a technology selection process that consumes the quarter; a platform project with no agent attached; the hardest or most visible use case; and any deployment without a measured baseline.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project