FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Leadership ¡ 5 minute read

The VP of Engineering's Guide to AI and Agentic AI

A VP of Engineering delivers agentic AI by organizing a small platform team plus embedded engineers, running a delivery process where evaluation is the release gate, adopting coding agents under review policy, operating production agents with on-call, monitoring, and kill switches, and reporting pass rates, incidents, and cycle time rather than demos.

By FISTA Solutions¡ AI-Native Engineering Team¡
The VP of Engineering's Guide to AI and Agentic AI article cover

The VP of Engineering is where AI strategy becomes shipped software that someone has to support at three in the morning. This guide covers what changes in team structure, delivery process, tooling, and operations when the systems being built are probabilistic agents that take actions, and what to report so that leadership sees evidence rather than demos.

Why does agentic AI change engineering management?

Three things are different. First, the definition of done changes: a feature is not done when the code is merged but when it passes an evaluation set at a rate the business accepts. Second, the work is cross-functional in a new way: the hardest part of an agent is the process it serves, which lives with business owners, not engineers. Third, operations never stop: models change, data drifts, and an agent that passed last month can fail this month without a code change.

FISTA's agentic SDLC whitepaper sets out the full delivery model; this guide focuses on the management decisions.

How should the organization be structured?

UnitOwnsStaffed byMeasured on
Platform teamGateway, agent identity, connectors, evaluation harness, observability, inventorySmall, senior, applied AI and infrastructure engineersAdoption, reliability, time to first agent for a new team
Embedded engineersAgents built inside business or product teamsEngineers paired with process ownersEvaluation pass rates, production outcomes
Business ownersOutcomes, supervision levels, exception handlingOperations, product, or functional leadersBaseline-to-actual results
On-call and operationsMonitoring, incidents, runbooks, kill switchesRotation across platform and embedded engineersTime to detection, time to resolution

This is the forward deployed engineering model applied inside a company: shared infrastructure, engineers close to the work, and named owners. A central AI lab that builds agents for other teams tends to produce prototypes the receiving teams cannot own.

What does the delivery process look like?

  1. Specify. Write what correct behavior is for each class of input, including when to escalate. See spec-driven development.
  2. Build the evaluation set from real cases before building the agent.
  3. Develop against the set. Prompts, tools, retrieval, and model choice are tuned to raise the pass rate.
  4. Review everything. Hand-written and generated code alike, with attention to the failure modes agents produce.
  5. Gate release on the pass rate, with no regression against the baseline.
  6. Deploy under supervision. Human review of consequential actions until evidence supports releasing it.
  7. Monitor and feed back. Every production failure becomes an evaluation case.

The evaluation-driven development whitepaper details the harness and the metrics.

How should coding agents be adopted?

Coding agents raise output substantially. The bottleneck moves to specification, review, and testing, and management has to move with it. A written policy should define scope (what agents may change autonomously and what requires a human author), review requirements, security boundaries (no secrets, no production configuration), and measurement. Invest in specifications and test coverage, because those are what agents optimize against. Measure cycle time, defect escape rate, and review load, never lines of code. FISTA's how to adopt AI coding agents safely and measuring AI developer productivity guides cover the policy and the metrics.

What operations do production agents need?

Agents are production services with an extra failure mode: silent quality drift. Operations should include an on-call rotation trained on agent behavior; monitoring for exception rates, latency, cost, and quality against the evaluation standard; runbooks for common failures; a tested kill switch per agent; scheduled evaluation runs in production; and change control for prompts, tools, and model versions with the same rigor as code. The AI agent runbook template and kill switch design guides are practical starting points.

What should the VP of Engineering report?

Monthly, in the same format: evaluation pass rates and trends per agent; agents in production and their autonomy levels; incidents with severity, root cause, and time to detection; cycle time from specification to production; platform adoption across teams; model spend and cost per task; and coding-agent metrics before and after adoption. These numbers survive scrutiny from the CFO and the board because they are measured, not asserted.

What should the first 90 days look like?

  1. Weeks 1–3: stand up the minimum platform: gateway, per-agent identity, an evaluation harness, and tracing. Do not build agents yet.
  2. Weeks 3–6: pair one embedded engineer with one business owner on one agent. Write the specification and the evaluation set first.
  3. Weeks 6–10: ship under full supervision. Establish on-call, monitoring, and the kill switch before launch, not after.
  4. Weeks 10–13: review evidence, release supervision where it is justified, publish the first monthly report, and start the second agent on the same platform.

What are the common management mistakes?

Approving builds without an evaluation set; letting prompt iteration substitute for specification; staffing a central lab instead of embedding; treating agents as done at launch; adopting coding agents without review policy; and reporting demos instead of pass rates. Each is avoidable with the structure above.

How can FISTA Solutions help a VP of Engineering?

FISTA Solutions embeds forward deployed engineers inside engineering teams to build production agents with evaluation, observability, and operations designed in, and its Applied division helps stand up the platform team, the evaluation harness, and the coding-agent policy described here. Since 2017, FISTA has delivered 150+ projects for 50+ companies across 12+ countries, with a 99.9% uptime record on production systems.

If your team is shipping prototypes that do not survive production, talk to FISTA on WhatsApp about a delivery-process review, or read the AI team structure guide first.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01How should an engineering organization be structured for AI agents?

A small platform team owns the model gateway, agent identity, connectors, evaluation harness, and observability. Engineers embedded in business or product teams build agents on that platform with the people who own the process. Each agent has a business owner for outcomes and a technical owner for the system. Avoid a central lab that builds for others.

02What does the delivery process look like for AI agents?

Specify correct behavior first, build the evaluation set from real cases, develop against it, review generated and hand-written code alike, gate release on the pass rate with no regression, deploy under human supervision, monitor in production, and feed every failure back into the evaluation set. The loop is continuous because models and data change.

03How should engineering adopt AI coding agents?

Under a written policy covering scope (what agents may change), review (all generated code reviewed with attention to plausible-but-wrong logic), security (no secrets or production configuration), and measurement (cycle time, defect escape rate, review load). Invest in specifications and test coverage, because those are what coding agents work from.

04What operations do production AI agents need?

An on-call rotation that understands agent behavior, monitoring for exception rates, latency, cost, and quality drift, runbooks for common failures, a tested kill switch per agent, scheduled evaluation runs in production, and a change-control process for prompts, tools, and model versions with the same rigor as code deployments.

05What metrics should a VP of Engineering report for AI?

Evaluation pass rates and trends per agent, agents in production and their autonomy levels, incidents with severity and time to detection, cycle time from specification to production, platform adoption across teams, model spend and cost per task, and for coding agents, cycle time and defect rates before and after adoption.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project