Whitepaper · 9 minute read
Agent Memory Architecture: Designing What Agents Remember
Agent memory should be designed as distinct layers rather than one growing transcript: working state for the current task, episodic memory of past interactions, semantic memory of durable facts, and a governed user profile. Each has its own write rules, retention, access controls, and recall evaluation. Undifferentiated memory produces cost, confusion, and privacy exposure.
Memory is the feature teams add when an agent feels unhelpfully forgetful, and it is where more agent systems go wrong than anywhere except retrieval. The naive implementation, appending everything to a growing history and retrieving by similarity, produces an agent that is expensive, slow, occasionally confused by its own past, and holding personal data nobody can account for. This whitepaper sets out a layered architecture with explicit rules for what is written, kept, recalled, and forgotten. It draws on FISTA Solutions' production AI agents work and complements what is agent memory and how to build long term memory for ai agents.
What are the layers, and why separate them?
| Layer | Holds | Lifetime | Written when | Read when |
|---|---|---|---|---|
| Working state | Goal, plan, intermediate results, decisions in this task | The task | Every step, structured | Every step |
| Episodic | What happened in past interactions | Weeks to months | Interaction ends, if notable | Similar situation arises |
| Semantic | Durable facts about entities: accounts, products, people | Until superseded | A fact is resolved and confirmed | The entity is in scope |
| Profile | Stable preferences and settings | Until changed or deleted | User states or confirms a preference | Every interaction with that user |
They are separated because their rules differ in every respect. Working state is high-frequency, structured, and discarded when the task ends. Episodic memory is selective and time-bounded. Semantic memory must be superseded rather than appended, or contradictions accumulate. Profile data is the most privacy-sensitive and the most consequential when wrong.
Systems that collapse these into one store inherit the strictest requirement of each with the discipline of none: they retain everything forever, retrieve by similarity across categories, and cannot answer what they know about a person.
Why is working state the layer teams underbuild?
Because it looks like conversation history, and it is not. Working state is the agent's structured understanding of the current task: what it is trying to achieve, what it has established, what it has tried, what remains. Held as structured data and re-rendered into context each step, it stays compact and consistent. Held as an accumulating transcript, it grows linearly with steps, carries the agent's earlier confusion forward, and pushes relevant material out of the window.
The practical test is whether the agent can be interrupted at step seven, have its context rebuilt from state, and continue correctly. Systems that can do this are debuggable and resumable; systems that cannot are neither. See what is a scratchpad in ai agents.
What should be written to memory, and when?
Only what will change future behaviour. Concretely: preferences the user stated, facts resolved during the interaction that were not previously known, decisions made and their reasons, outcomes including failures, and corrections the user gave.
That excludes most of what happens. An agent that writes every turn produces a store that is expensive, noisy to search, and impossible to audit. The write decision belongs in the specification with explicit criteria, and it is worth implementing as an explicit step the agent takes rather than a side effect of the conversation loop, so it can be logged, reviewed, and evaluated.
Two further rules prevent common damage. Write with provenance: what was written, when, from which interaction, and whether it was stated by the user or inferred. And write with confidence: inferred facts are marked as such, so recall can weight them differently and decay them over time.
How should memory be superseded rather than accumulated?
Semantic memory holds facts that change. A customer's shipping address, a company's primary contact, an account's service tier, a stated preference: each can be updated, and the update must replace rather than join the old value. Systems that append produce retrieval that surfaces both, leaving the model to guess.
Practical mechanics: key semantic facts by entity and attribute so a write to the same key supersedes; keep the prior value in history for audit but exclude it from recall; and record the time and source of the change. For inferred facts, apply decay so that confidence falls with age unless reconfirmed, which prevents a single offhand comment from shaping an agent's behaviour indefinitely.
What does forgetting require?
Design, not cleanup. Each layer gets a retention period tied to its purpose and to legal requirements. Working state is discarded at task end. Episodic memory expires on a schedule. Semantic facts persist until superseded or until the underlying relationship ends. Profile data persists until changed or deleted by the user.
Deletion paths must be real: when a user exercises a deletion right, the system must remove their data from every layer including derived embeddings and any caches, and must be able to demonstrate it. Systems that embed memories into a vector index without a deletion path create a compliance problem that is expensive to remediate later. See what is data residency and ai data retention checklist.
How is recall governed and scoped?
Memory recall is subject to the same access control as any other data. A support agent serving one customer must not recall facts about another. In multi-tenant systems the tenant boundary is absolute, and recall queries must be scoped before search rather than filtered after, so cross-tenant material cannot influence ranking or leak through timing.
Within a tenant, scope still matters: an agent handling a billing question does not need the customer's support history about a product defect, and surfacing it wastes context and occasionally embarrasses. Recall scoping by task type is both a quality and a privacy control.
How is recall evaluated?
Separately from generation, with its own reference set. For each situation in the set, record which memories should be recalled and which should not. Then measure recall precision, the rate of stale or superseded memories surfacing, and the rate of out-of-scope recall.
The metric that matters most is harmful recall: the agent bringing up something irrelevant, outdated, or sensitive. Customers notice this far more than they notice an agent failing to remember something, because a wrong memory is an error with a human face on it. Evaluation practice is in the AI evaluation and testing whitepaper.
How does memory interact with retrieval?
They are different systems with different guarantees and should not share a store. Retrieval serves organisational knowledge: documents, policies, product information, curated and owned by a content function. Memory serves interaction history: what happened with this user or account, written by the agent itself.
Mixing them produces two failures. Agent-written memories pollute the knowledge corpus, so an agent's earlier mistaken inference becomes a retrievable fact. And knowledge documents surface in memory recall as if they were things that happened. Keep the stores, the write paths, and the evaluation separate, and assemble both into context deliberately. See what is rag.
What does this cost, and how is it bounded?
Memory costs storage, embedding compute at write, retrieval compute at read, and context tokens at every interaction that recalls. Unbounded memory therefore raises cost on every request forever, which is why bounds are an architectural requirement rather than an optimisation.
Practical bounds: a cap on memories recalled per interaction, a relevance threshold below which nothing is recalled, summarisation of episodic memory into semantic facts so volume does not grow linearly with interactions, and retention expiry. Measure memory's share of context tokens and cost per interaction, and treat growth in either as a defect. See ai inference cost.
What does a memory specification contain?
Layer definitions with the data each holds; write criteria per layer, expressed precisely enough to evaluate; retention and supersession rules; recall scoping rules by task type; access control boundaries including tenancy; privacy classification of each field with consent and deletion handling; bounds on recall volume and context share; and the recall evaluation set. Reviewed by security and privacy before deployment, this document answers the questions that otherwise surface during an audit.
What goes wrong most often?
One undifferentiated store. Writing every turn. Appending rather than superseding, so agents hold contradictory facts. No deletion path through embeddings. Recall scoped after search rather than before, which is the cross-tenant leak waiting to happen. Memory and knowledge in the same index. No recall evaluation, so nobody notices the agent has been surfacing a stale preference for months. And unbounded growth, which shows up first as rising cost and later as declining quality.
How does this differ for consumer versus enterprise agents?
Consumer agents carry heavier privacy obligations around profile data, stronger expectations of deletion, and more sensitivity to harmful recall in front of a person who did not ask for it. Enterprise agents operating on accounts rather than individuals carry tenancy isolation as the dominant risk and usually need longer semantic retention because account facts change slowly and matter commercially. Both need the same layered structure; the retention periods and the privacy controls differ.
How should a team introduce memory to an existing agent?
Not all at once. The sequence that works starts with working state, because it improves reliability on multi-step tasks immediately and carries no privacy weight: replace the accumulating transcript with structured state, confirm the agent can be interrupted and resumed, and measure step count and task success.
Profile memory comes next, limited to preferences the user has explicitly stated, with a visible way for them to see and change what is stored. Explicit statement is the safe boundary: agents that infer preferences from behaviour and act on them surprise people, and the surprise is rarely pleasant.
Semantic memory follows, scoped to a single entity type such as accounts, with supersession and provenance from the first write. Episodic memory comes last and is the easiest to over-build; many agents never need it, because the semantic facts extracted from past interactions are more useful than the interactions themselves.
At each stage the recall evaluation set grows, and the harmful-recall rate is watched. A team that adds all four layers in one release has no way to attribute a quality change to any of them.
How FISTA Solutions delivers this
FISTA Solutions designs agent memory as separate governed layers with explicit write criteria, supersession, retention and deletion paths, scoped recall, and recall evaluation, so agents are helpful without becoming expensive or a privacy liability, delivered through AI enablement, production AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To give agents memory you can govern, message FISTA on WhatsApp, or read how to build long term memory for ai agents.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What are the layers of agent memory?
Working state for the current task, holding goals, intermediate results, and decisions; episodic memory of prior interactions; semantic memory of durable facts learned about entities; and a user or account profile holding stable preferences. Each has different write criteria, retention, and access rules.
02What should an agent write to memory?
Only what will change future behaviour: stated preferences, resolved facts, decisions and their reasons, and outcomes. Writing every utterance produces a transcript that is expensive to store, noisy to retrieve, and impossible to govern. Write decisions need explicit criteria in the specification.
03How should memory be forgotten?
By design, with retention periods per layer, superseding rules so newer facts replace older ones rather than coexisting, confidence decay for inferred information, and deletion paths that satisfy privacy rights. An agent that recalls a preference the customer changed last year causes real harm.
04How is memory recall evaluated?
With a reference set of situations paired with the memories that should and should not be recalled, measuring recall precision and the rate of harmful recall, meaning stale, irrelevant, or out-of-scope memories surfacing. Recall quality is measured separately from generation quality.
05What privacy obligations attach to agent memory?
Memory is stored personal data when it concerns individuals, so purpose limitation, retention limits, access control, deletion rights, and disclosure obligations apply, and cross-user or cross-tenant leakage is a serious incident. Confirm specifics with counsel.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.