FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Whitepaper ┬╖ 9 minute read

Agent Orchestration at Scale: An Engineering Whitepaper

Running many agents requires orchestration chosen deliberately, workflows where the path is known and autonomy only where it is not, explicit state that survives handoffs, failure containment so one agent cannot cascade, per-agent identity and permissions, and a registry that records what exists. Without those, an agent estate becomes ungovernable within a year.

By FISTA Solutions┬╖ AI-Native Engineering Team┬╖
Agent Orchestration at Scale: An Engineering Whitepaper article cover

The first production agent is an engineering achievement. The fiftieth is an operating problem that no amount of individual engineering quality solves. Organisations reach that point faster than they expect, because agents proliferate once the platform exists, and they arrive at a recognisable state: nobody knows how many agents are running, several hold broader permissions than anyone intended, one team's agent calls another's through an undocumented interface, and an incident in one produces symptoms in three. This whitepaper sets out the patterns and platform that prevent it. It draws on FISTA Solutions' AI agents delivery and complements the multi-agent orchestration patterns whitepaper and the agentic AI governance whitepaper.

What orchestration patterns exist, and when does each fit?

PatternControl flowFits whenCost of misuse
Deterministic workflowCode decides the sequencePath is known; variation at defined pointsNone; this is the default
Workflow with agent stepsCode sequences, agents handle stepsSteps need judgement, sequence does notLow
Supervisor and workersOne agent delegates to specialistsTask decomposition varies by inputModerate; supervisor becomes bottleneck
Autonomous planningAgent determines its own sequencePath genuinely cannot be enumeratedHigh; unpredictable cost and behaviour
Peer collaborationAgents negotiate among themselvesRarely justified in enterprise settingsVery high; hard to debug or govern

The discipline that matters is preferring the top of this table. Autonomous orchestration is more interesting and almost always less appropriate: enterprise processes usually have known paths, and encoding them in code produces a system that is cheaper, faster, testable, and explicable to an auditor.

A useful test: if a competent analyst could draw the process on a whiteboard, it should be a workflow. Autonomy earns its place where the analyst would say it depends on what we find.

How should state be managed?

Externally, explicitly, and in a structure both the orchestrator and the agents can read. The failure pattern is carrying state in the conversation history passed between steps, which grows unbounded, loses structure, mixes the agent's reasoning with the facts of the case, and makes resumption after a failure impossible.

An explicit state object holds the case identifier, the goal, the facts established, the decisions made with their reasons, the steps completed, and what remains. Each agent reads what it needs and writes what it established. Context is rendered from state at each step rather than accumulated.

The practical tests are whether a workflow can be paused at any step and resumed by a different process, and whether a human can read the state and understand where the case stands. Systems that pass both are debuggable; systems that fail them are not.

What does a handoff contract require?

The same discipline as any service interface. Specifically: the data passed and its schema; the outcome expected and its schema; the timeout; the behaviour on failure, refusal, or timeout; idempotency guarantees, since retries happen; and the owner of each side.

Handoffs without contracts produce the characteristic multi-agent failure: agent A passes work to agent B, B fails or declines, A has no handling for that case and proceeds as though the work completed, and the error surfaces three steps later in a form that gives no indication of its origin.

Where a handoff crosses team boundaries, the contract should be versioned and treated as a published interface, because the team on the other side will change their agent without telling you.

How is failure contained?

By assuming any agent can fail in any way and designing so that no single failure propagates. The mechanisms:

Bounded blast radius. Each agent's permissions limited to its role, so a malfunction cannot reach beyond it.

Circuit breakers. Repeated failures from a downstream agent stop calls to it rather than retrying indefinitely.

Timeouts at every boundary, with defined behaviour when they expire.

Step and depth limits, so recursion or delegation loops terminate.

Cost caps per workflow instance, which catch runaway behaviour that latency monitoring misses.

Graceful degradation, where a workflow can complete a reduced version of its task or escalate to a person rather than failing entirely.

Independent kill switches, so one agent can be disabled without stopping the estate.

Cost caps deserve emphasis: in agent systems the most common runaway is not an outage but a loop that completes successfully at each step while consuming budget. See ai agent guardrails.

Why does each agent need its own identity?

Three reasons that all become acute at scale.

Least privilege. Agents have different jobs and should hold different permissions. A shared service account means every agent holds the union of all permissions, which means the blast radius of any compromise is the whole estate.

Attribution. When an incorrect record change is discovered, the question is which agent did it and under whose instruction. Shared credentials make that unanswerable.

Revocation. Containing a compromised or malfunctioning agent should not require disabling every agent that shares its credentials.

Practically this means agents as first-class identities in the identity provider, with scoped credentials, short-lived tokens, and logged actions attributable to both the agent and the human whose request initiated the chain. That second attribution matters: when agent A invokes agent B on behalf of a user, B must act with the user's entitlements, not with A's. See the agent identity and access control whitepaper.

What does the registry hold?

Every agent in the estate with its purpose, business and technical owners, the systems and data it can access, its tools and their permissions, its service levels, its current version and model, its evaluation status and last score, its dependencies on other agents, and its kill path.

This is the governance minimum, and organisations discover its absence during their first significant incident, when the questions are which agents touch this system, who owns them, and what else will break if we disable this one.

The registry should be generated from the platform rather than maintained by hand, because a manually maintained registry is wrong within a month. See what is an agent registry.

What does the shared platform provide?

The components every agent needs, built once: a model gateway with version pinning and cost attribution; tool definitions with permission scoping; the state store; tracing across agent boundaries with a correlation identifier that survives handoffs; the evaluation harness; guardrail libraries; the registry; and the kill infrastructure.

The economic argument is that the tenth agent should cost a fraction of the first. The governance argument is stronger: an estate built on a shared platform can answer questions about itself, while an estate of independently built agents cannot, regardless of how well each was engineered.

How is observability handled across agents?

With correlation identifiers that survive every handoff, so a single user request that traverses four agents produces one trace rather than four disconnected ones. Each span records the agent, its version, the model, the tools called, the state read and written, the cost, and the latency.

Without this, diagnosing a multi-agent failure means manually correlating logs across systems by timestamp, which is slow and frequently impossible. With it, the question of which agent introduced the wrong value is a query. See the AI observability whitepaper.

How does evaluation work for a multi-agent system?

At both levels. Each agent has its own reference set and is evaluated independently, which is where most regressions are caught. The composed workflow is evaluated end to end against business outcomes, which catches the failures that only appear in combination, such as an agent whose output is individually correct but poorly formed for the next step.

End-to-end evaluation is more expensive and slower, so the practical pattern runs agent-level evaluation on every change and workflow-level evaluation on a schedule and before significant releases.

What does maturity look like?

LevelOrchestrationStateIdentityGovernance
InitialAd hoc, often autonomousIn conversation historyShared credentialsNone
ManagedWorkflows for known pathsExternal but informalPer-agent credentialsSpreadsheet inventory
DefinedPatterns chosen deliberatelyExplicit schema, resumableIdentity provider integratedGenerated registry
OptimisedContracts versioned, containment testedState machine with auditDelegated user entitlementsRegistry drives access review

What goes wrong?

Autonomy chosen for processes with known paths. State in transcripts. Handoffs without contracts. Shared service accounts. No correlation identifiers, so incidents cannot be traced. Cost caps absent, so a loop becomes an invoice. Agents calling each other through undocumented interfaces. And a registry that does not exist until the first incident demands one.

How does an organisation get from one agent to an estate?

Deliberately, by building the platform components when the second agent is proposed rather than the tenth. The first agent legitimately hard-codes its model access, logging, and state handling. The second is the decision point: extract those into shared components, or accept that every subsequent agent will reinvent them slightly differently.

The sequence that works extracts the gateway and tracing first, because they are needed by everything and are cheapest to retrofit early. The state store and handoff contracts follow when the first multi-step workflow appears. Identity and the registry come next, before the estate exceeds what one person can hold in their head, which is roughly five agents.

Organisations that defer this until the estate is large face a migration rather than a build, and migrations of production agents are expensive because each one must be re-evaluated after the change.

Who owns the platform?

A small platform team with a product mindset, serving the teams that build agents rather than building all the agents itself. The distinction matters: a central team that builds everything becomes the bottleneck and the reason shadow agents appear; a central team that provides the gateway, state store, tracing, evaluation harness, and registry lets functional teams build safely at their own pace.

The platform team's product is the golden path: the documented, supported way to build an agent that gets you registry entry, identity, logging, evaluation, and kill switch by default. When that path is genuinely easier than building around it, governance stops being enforcement and becomes the obvious choice.

How FISTA Solutions delivers this

FISTA Solutions builds agent estates on a shared platform with deterministic orchestration by default, explicit resumable state, versioned handoff contracts, per-agent identity with delegated entitlements, containment tested rather than assumed, and a generated registry that makes the estate governable, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To scale from one agent to many without losing control, message FISTA on WhatsApp, or read the multi-agent orchestration patterns whitepaper.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01When should orchestration be a workflow rather than autonomous?

Whenever the path is known. Deterministic workflows are cheaper, faster, testable, and debuggable, and most enterprise processes have known paths with variation only at specific decision points. Autonomy belongs where the sequence genuinely cannot be enumerated in advance.

02How should state be managed across agents?

Externally and explicitly, in a store that both the orchestrator and the agents read and write, rather than passed through conversation history. State carried in transcripts grows unbounded, loses structure, mixes reasoning with facts, and makes resumption after a failure impossible, which is the test that matters.

03What does a handoff contract contain?

The data passed and its schema, the expected outcome, the timeout, what happens on failure or refusal, and who owns each side. Handoffs without contracts produce silent failures where one agent assumes another completed work that never started.

04Why does each agent need its own identity?

For least privilege, attribution, and revocation. Shared credentials mean every agent holds the union of all permissions, no action can be attributed to a specific agent, and containing a compromised agent requires disabling all of them.

05What belongs in an agent registry?

Every agent with its purpose, owner, permissions, tools, data access, service levels, evaluation status, and current version. Without it, nobody can answer what agents exist, what they can do, or who is accountable, which makes both incidents and audits unmanageable.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.

Start a project