Playbook ┬╖ 6 minute read
How to Build a Xero AI Agent
A Xero agent connects through the OAuth 2.0 API with tenant-scoped tokens, handles rate limits deliberately, and delivers most in bank reconciliation coding, bill capture and approval, and month-end preparation. Accounting practices gain most because one build serves many client organisations under separate tenant connections.
Xero has one of the better APIs in accounting software, which removes most of the integration friction and leaves the interesting problems: which workflows justify the effort, how to serve many client organisations from one build without mixing their data, and how to keep an automated ledger defensible. This guide covers those, drawing on FISTA Solutions' AI agents delivery in finance operations. It complements ai in accounting firms and how to build an invoice processing agent. This article is general guidance, not legal or accounting advice.
How does the connection model work?
Through OAuth 2.0, with tokens scoped per connected organisation. An application connects to a tenant, receives access and refresh tokens, and uses them to act within that organisation only.
For a practice, this means one application with many tenant connections, each with its own tokens, each representing a client relationship that can be granted and revoked. The engineering implications are concrete: tokens stored encrypted per tenant, refreshed proactively rather than on failure, revocation handled cleanly when a client disconnects, and every API call explicitly scoped so a coding error cannot cross organisations.
What do rate limits require?
Design rather than reaction. Xero enforces call limits per tenant and per application, and an agent that iterates transaction records synchronously will hit them quickly, particularly for a practice acting across many tenants at month end.
| Concern | Naive approach | What works |
|---|---|---|
| Reference data | Fetch per transaction | Cache per tenant with refresh |
| Bulk coding | Loop synchronously | Queue with concurrency control |
| Limit responses | Retry immediately | Exponential backoff with jitter |
| Multi-tenant month end | Process all tenants at once | Stagger with fair scheduling |
| Token refresh | Refresh on failure | Refresh proactively before expiry |
Which workflows deliver most?
Bank reconciliation coding. The highest-volume task in most Xero organisations. The agent proposes the account, tax treatment, and tracking categories from the transaction description, amount, and the organisation's own history with that payee, with confidence indicated. High-confidence items are confirmed in bulk; ambiguous ones get attention.
Bill capture and approval. Extraction from supplier invoices, coding proposal, duplicate detection, and routing to the approver, with the bill created in Xero once approved.
Month-end preparation. Reconciliation assembly, detection of accounts whose movement is unusual for the period, draft accruals from recurring patterns, and preparation of the questions the accountant will ask the client.
Advisory support for practices. Drafting management commentary from the numbers, identifying clients whose metrics have moved materially, and preparing meeting packs, which converts compliance work into advisory conversations.
Why do practices gain disproportionately?
Because the economics change with tenant count. A single small business cannot justify custom agent development against its bookkeeping volume. A practice with a hundred client organisations builds once and applies it across all of them, with each client's data isolated by tenant.
That also changes what is worth building. Practice-wide, it becomes worth investing in coding accuracy, because a percentage point across a hundred clients is real staff time, and in advisory preparation, because it changes what the practice can sell.
How is tenant isolation enforced?
As a hard boundary, in code rather than by convention. Every query, cache key, retrieval index, and stored artifact carries the tenant identifier, and the data access layer refuses to execute without one. Tests should attempt cross-tenant access and expect failure.
This is the most serious risk in the design. An agent that surfaces one client's supplier data to another has caused a confidentiality breach with professional and regulatory consequences for the practice, and the fact that it was a caching defect is not a defence. See the multi-tenant AI architecture whitepaper.
What controls keep the ledger defensible?
Coding proposals reviewed before posting, with the review interface making low-confidence items obvious. A reference on each transaction identifying agent involvement. An external log holding the proposal, its reasoning, the confidence, the reviewer, and the outcome. And clear client communication about what is automated, since the client's accountant remains responsible for the accounts.
For practices, professional obligations apply: the work is still the practice's, and an error coded by an agent is the practice's error.
How is it evaluated?
Against real transactions with the coding the practice actually applied, segmented by organisation type and transaction category. Measure coding accuracy, the proportion accepted without change, and the error rate on accepted proposals. Track it per tenant as well as in aggregate, because a client whose business differs from the norm may be coded badly while the average looks healthy.
What does the build sequence look like?
Connection and tenant management first, including token lifecycle and isolation tests, because everything else depends on it. Then reference data caching and the coding proposal engine, evaluated offline against history before anything is live. Then the review interface, which determines adoption. Then bill capture. Then month-end and advisory support.
A realistic first delivery for a practice runs eight to ten weeks, with the coding engine live for a pilot group of clients before the full base.
How should clients be told what is automated?
Explicitly, in the engagement letter and in conversation, because the practice remains professionally responsible for work an agent performed. Clients rarely object to automation that improves turnaround; they object to discovering it after an error.
The disclosure that works states which tasks are automated, that a qualified person reviews the output before it affects the accounts, how the client's data is handled including that it is not used to train external models, and where the client's own approval is still required. Practices that publish this find it accelerates rather than complicates client conversations, because it reads as capability rather than cost-cutting.
What does the review interface need to do?
Determine adoption, which is the outcome most builds underweight. A reviewer working through a month of bank transactions needs the proposed coding, the confidence, the reason, and the ability to accept or change it without leaving the screen or waiting for a page load.
Two design decisions matter more than the rest. Bulk confirmation for high-confidence items, so a hundred routine transactions are cleared in one action rather than a hundred. And clear visual separation of the items that need thought, so attention lands where the value is. Interfaces that present every proposal identically make the reviewer do the triage the agent was supposed to do.
What happens when Xero changes?
Regularly, and the agent must tolerate it. Xero deprecates and versions API endpoints, adds fields, and changes behaviour, and a practice-wide agent that breaks silently across a hundred tenants at month end is a serious operational event.
The protections are version pinning where the API supports it, monitoring that alerts on changed response shapes rather than only on errors, a canary tenant that processes first so problems surface on one client rather than all of them, and a staged rollout for agent changes of the practice's own.
How FISTA Solutions helps
FISTA Solutions builds Xero agents with tenant isolation enforced in the data layer, rate-limit-aware processing, history-based coding proposals with confidence routing, review interfaces designed for speed, and audit logging that satisfies professional obligations, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To automate bookkeeping across a client base, message FISTA on WhatsApp, or read ai in accounting firms.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How does an agent connect to Xero?
Through the OAuth 2.0 flow, obtaining tenant-scoped access and refresh tokens per connected organisation. Tokens are stored encrypted per tenant, refreshed proactively before expiry, and revoked cleanly when a client disconnects, which accounting practices handle frequently as client relationships change.
02What rate limits apply and how are they handled?
Xero enforces per-tenant and per-application call limits, so agents batch requests, cache reference data such as accounts and contacts, back off with jitter on limit responses, and queue bulk work rather than iterating records synchronously during a user interaction.
03Which Xero workflows pay back first?
Bank reconciliation coding, where the agent proposes account and tracking category from transaction history and payee patterns; bill capture with coding and approval routing; and month-end preparation including reconciliation assembly and detection of accounts whose movement is unusual for the period.
04Why do accounting practices benefit most?
Because one build serves every client organisation through separate tenant connections, so the engineering amortises across the client base while each client's data stays isolated. The economics that do not work for a single small business work well for a practice with a hundred clients.
05What is the most serious risk to design against?
Cross-tenant data leakage. An agent serving multiple client organisations must scope every query, cache key, and retrieval index by tenant, because surfacing one client's data to another is a confidentiality breach with professional consequences for the practice, not merely a defect.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.