Hiring ¡ 5 minute read
How to Hire Agent Engineers: Signals, Tests and Scope
Agent engineers build systems where a model plans and acts through tools, which makes the work distributed systems engineering with a non-deterministic component. Test for failure design, evaluation discipline, and authority boundaries rather than prompt writing, because those are what separate shipped agents from demos.
Agent engineering is distributed systems work with a non-deterministic planner attached. Hiring for prompt skill alone produces convincing demos that fail in production. This guide covers what to test, drawing on FISTA Solutions' AI agents work.
What makes agent engineering different?
The planner is non-deterministic and the actions are real. A model deciding which tool to call, with which arguments, can take a wrong action against a live system.
| Concern | Conventional service | Agent |
|---|---|---|
| Control flow | Coded explicitly | Decided at runtime |
| Failure modes | Known set | Open-ended |
| Retry safety | Design choice | Requirement |
| Authority | Fixed per endpoint | Inherited by the planner |
| Evaluation | Input to output | Whole trajectory |
What should you test in an interview?
Failure design. Ask three questions: what happens when a tool call times out after the action succeeded, what happens when the model calls the same tool repeatedly, and what happens when it produces a plausible but wrong plan.
Those cover most production failures. Candidates with answers have shipped; candidates who talk about prompt improvements have not.
Why does tool authority matter so much?
Because an agent inherits whatever permissions its tools have, and it can be steered by content it reads.
Untrusted input reaching a model with write access is a confused deputy problem, and it must be addressed in design rather than in instructions. Ask how they scoped tool permissions. See what is a confused deputy attack.
How should agents be evaluated?
On trajectories as well as outcomes. An agent reaching the right answer through three unnecessary expensive tool calls is not working correctly.
Ask what they measured. Task completion, tool call efficiency, cost per task, and unsafe action rate are the useful signals. Outcome-only evaluation hides problems until the bill or an incident arrives.
Why does idempotency come up constantly?
Because agents retry, and retries against systems of record create duplicates. Every tool that writes needs to be safe to call twice.
Candidates who have not thought about this will build agents that occasionally double-charge or double-order, and the failure will be intermittent and hard to trace. See what is idempotency in ai agents.
What about human oversight?
Ask where they put a human in the loop and why. Good answers tie the checkpoint to consequence: actions that are expensive, irreversible, or externally visible get review; routine reads do not.
Blanket approval on everything defeats the purpose; approval on nothing defeats the safety argument.
How do you handle context limits?
Ask how they managed context across a long task. Summarisation, external memory, and scoping each step are the practical approaches, and each has failure modes.
Agents that lose earlier context mid-task produce confident wrong actions, which is worse than failing.
What about cost control?
Agents can loop, and loops are billed. Ask what limits they set: maximum steps, cost ceilings per task, and circuit breakers on repeated identical calls.
Candidates without limits have either not run agents at scale or have had an expensive month.
When do you not need an agent?
When the task is a fixed sequence. A workflow with known steps should be a workflow â cheaper, faster, and easier to reason about.
Ask candidates when they chose not to build an agent. The answer indicates judgement rather than enthusiasm.
Contract, staff augmentation, or permanent hire?
Augmentation suits building the first production agent, where experience compresses the learning curve substantially. Permanent hiring suits organisations building a portfolio of them.
What are the common hiring mistakes?
Screening on prompt writing. Ignoring tool authority. Skipping trajectory evaluation. And building agents for fixed workflows.
How do you onboard them well?
Give them the systems the agent will act on, the permission model, and the failure tolerance for each action. Those three define the design space.
What does good look like after 90 days?
One agent in production with scoped tool permissions, idempotent writes, step and cost limits, trajectory evaluation running, and a defined human checkpoint on consequential actions.
What should be measured?
Task completion rate, unsafe or incorrect action rate, cost per completed task, and escalation appropriateness.
What should you do first?
List the actions an agent would take and mark which are reversible. That list determines the authority design and the oversight points.
What background do good candidates come from?
Backend and distributed systems engineering, more often than machine learning. The hard parts of the job are retry semantics, permission design, observability across a non-deterministic control flow, and deciding what a system may do unsupervised â all of which are engineering problems.
Model familiarity is genuinely useful and quickly learned. Ask a backend engineer how they would design the same system and the answer is frequently better than one from a candidate whose experience is entirely in prompting.
How do you debug an agent?
Ask directly. Useful answers involve recorded trajectories, replay against the same inputs, and tracing that spans model calls and tool calls together. Candidates who debug by reading logs of final outputs cannot diagnose why a plan went wrong, which is where the interesting failures are.
How FISTA Solutions helps
FISTA Solutions builds production agents through AI agents, AI enablement, and forward deployed engineers: tool permissions scoped to least privilege, writes made idempotent because agents retry, step and cost limits enforced, trajectory evaluation running continuously, and human checkpoints placed by consequence rather than by blanket rule. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To build a production agent, message FISTA on WhatsApp, or read what is an agent loop.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What makes agent engineering different?
The planner is non-deterministic and the actions are real. A model deciding which tool to call and with what arguments can take a wrong action against a live system, which makes authority boundaries and failure handling more important than the prompt.
02What should be tested in an interview?
Failure design. Ask what happens when a tool call times out after the action succeeded, when the model calls the same tool repeatedly, and when it produces a plausible but wrong plan. Those three cover most production failures.
03Why does tool authority matter?
Because an agent inherits whatever permissions its tools have, and it can be steered by content it reads. Untrusted input reaching a model with write access is a confused deputy problem, and it must be addressed in design rather than in prompts.
04How should agents be evaluated?
On trajectories as well as outcomes. An agent that reaches the right answer through three unnecessary and expensive tool calls is not working correctly, and outcome-only evaluation hides that until the bill or an incident arrives.
05When do you not need an agent?
When the task is a fixed sequence. A workflow with known steps should be a workflow, which is cheaper, faster, and easier to reason about. Agents earn their cost where the sequence genuinely varies by case.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.