Memory is what separates a useful AI agent from an expensive autocomplete tool. Without memory, an agent forgets every conversation the moment it ends, cannot learn from past interactions, and cannot build understanding of a user or task across sessions. With the right memory architecture, an agent accumulates context that makes it progressively more useful over time. But “memory” in AI agents covers several distinct concepts that serve different purposes and have different implementation requirements. Getting the architecture right means choosing the right memory type for each use case rather than bolting on a single generic store.
The Four Types of Agent Memory
Agent memory systems typically distinguish four types, borrowed from cognitive science. Sensory memory holds raw, immediate inputs — the current conversation turn, newly received documents, fresh tool outputs. It is volatile and corresponds roughly to the model’s context window: once the context is full or the session ends, sensory memory is lost. Short-term memory (also called working memory) holds information relevant to the current task — intermediate results, the current state of a multi-step problem, facts established earlier in the current conversation. It typically lives in the context window or a session-scoped store. Long-term memory persists across sessions — user preferences, learned facts about the domain, successful strategies from past interactions. It requires external storage (a database or vector store) and explicit retrieval. Procedural memory encodes how to do things — the agent’s learned behaviours, successful tool-use patterns, verified workflows for recurring task types. It is often stored as examples or instructions that are retrieved and added to the context when relevant.
Context Window as Working Memory
The model’s context window is the most immediate form of agent memory. Everything in the context is available to the model at inference time with zero latency and no retrieval overhead. For single-session agents, the context window is often sufficient — a 200k token context (as in Claude 3.5, GPT-4o, and most 2025+ models) can hold hours of conversation, many retrieved documents, and extensive tool outputs before overflow becomes a concern. The limitation is cost (more tokens = higher API cost) and performance (very long contexts have slightly degraded attention over distant content). Context window management is the first memory architecture decision: what gets kept in full context, what gets summarised, and what gets offloaded to external storage and retrieved on demand.
Episodic Memory: Recalling Past Interactions
Episodic memory stores records of past interactions — previous conversations, past task outcomes, earlier decisions and their results. For a coding assistant, episodic memory might store the fact that a user consistently prefers functional patterns over object-oriented ones, or that a particular approach to a class of problem succeeded three weeks ago. For a customer support agent, it stores the history of past interactions with each user, which issues were resolved and how. Episodic memory is typically implemented as a vector database: each past interaction is embedded and stored, and at the start of a new session, the most relevant past episodes are retrieved based on semantic similarity to the current task and injected into the context. The retrieval step is the key engineering challenge — selecting which memories are relevant without overwhelming the context with noise.
Semantic Memory: Stored Knowledge
Semantic memory holds factual knowledge about the world, the user, or the domain that does not fit neatly into a specific past interaction. User preferences (“prefers bullet points over prose,” “based in Sydney, uses Australian English”), domain facts (“the company’s API rate limit is 1000 requests per minute”), and entity attributes (“this customer is on the enterprise plan, which includes…”) all belong in semantic memory. Unlike episodic memory, which stores experiences, semantic memory stores distilled facts. Implementation options range from a simple key-value store (for discrete facts) to a structured relational database (for entities with multiple attributes) to a knowledge graph (for facts with complex relationships between entities). The retrieval pattern is typically lookup by entity or keyword rather than semantic similarity search, which makes structured databases often more appropriate than vector stores for semantic memory.
Procedural Memory: Learned How-To
Procedural memory encodes patterns for how to accomplish recurring tasks — not specific past episodes but generalised playbooks derived from experience. For an agent that frequently books travel, procedural memory might store the sequence of steps that consistently succeeds for a particular booking site. For a coding agent, it might store verified patterns for common tasks (“to add a new endpoint in this codebase, always follow these three steps”). Procedural memory is often implemented as a retrieval-augmented prompt library: when the agent is faced with a task that matches a known procedure, the relevant procedure is retrieved and injected as additional context or instructions. The challenge is maintaining the library — keeping procedures current, removing ones that no longer work, and detecting when a current task is similar enough to a stored procedure to warrant retrieval.
Figure 1 — Agent memory architecture: four types and their storage patterns
Memory Consolidation: From Short to Long Term
A key architectural question is how information moves from the current session into long-term storage. One approach is explicit consolidation: at the end of each session, run a summarisation step that extracts key facts, preferences, and outcomes from the conversation and writes them to the long-term store. Another is continuous consolidation: as the conversation proceeds, a background process identifies noteworthy statements (“the user said they prefer X,” “this approach worked on problem Y”) and writes them immediately. The explicit approach is simpler and introduces less noise; the continuous approach captures more information but requires filtering logic to avoid storing everything. For most production agent systems, explicit end-of-session consolidation with a small LLM running a summarisation and fact-extraction prompt is the most practical starting point.
Retrieval-Augmented Memory
The dominant implementation pattern for long-term agent memory is retrieval-augmented: store information in an external system, retrieve the most relevant pieces at the start of each session or turn, and inject them into the context. The retrieval step is the critical engineering challenge. For episodic and procedural memories, semantic similarity search in a vector database (pgvector, Pinecone, Weaviate, Chroma) is the standard approach — embed both the stored memory and the current context, find nearest neighbours, inject the top results. For semantic memories (discrete facts), structured queries against a relational database or key-value store are more precise than vector search, which can retrieve semantically similar but factually different results. Many production agents use a hybrid: vector search for episodic and procedural memories, structured lookup for semantic facts.
Memory Decay and Relevance Management
Long-term memory systems accumulate stale information — preferences that have changed, facts that are no longer true, strategies that worked once but have since been superseded. Without management, the memory store becomes a source of misleading context that reduces rather than improves agent performance. Several approaches address this: recency weighting (more recent memories score higher in retrieval), explicit expiration (memories older than N days are automatically downweighted or removed), contradiction detection (when a new fact contradicts a stored one, flag the stored one for review), and confidence scoring (memories written from explicit user statements score higher than those inferred from behaviour). Production memory systems typically combine recency weighting with periodic pruning — automatically removing memories that have not been retrieved in a long time on the assumption that irrelevant memories do not get retrieved.
Practical Implementation Choices
For a new agent project, the practical question is how much memory infrastructure to build. Start with the context window — it is sufficient for single-session agents and requires no external infrastructure. Add a simple key-value store for user preferences and semantic facts when you need cross-session personalisation — Redis or even a JSON file works for small deployments. Add a vector database for episodic memory only when you have evidence that past interaction recall meaningfully improves agent performance on your use case; the operational overhead of maintaining and querying a vector store is non-trivial. Libraries like Mem0, Zep, and LangChain Memory provide higher-level abstractions over these storage systems, handling the serialisation, retrieval, and injection into context automatically — worth evaluating before building custom memory infrastructure.
Memory in Multi-Agent Systems
In multi-agent architectures, memory architecture becomes more complex because multiple agents may need access to shared state. Approaches range from a shared memory pool (all agents read from and write to the same store, with appropriate locking), to message-passing (agents communicate state via messages, with no shared external store), to hierarchical memory (each agent has its own working memory, with a coordinator managing shared long-term state). The right choice depends on how much state needs to be shared across agents and how concurrency is handled. For most multi-agent systems, a shared semantic memory store for global facts plus agent-local context windows for working state is the most practical starting architecture — complex enough to handle coordination, simple enough to debug when things go wrong.
Memory architecture is the layer of agent design that most directly determines whether the system feels intelligent and useful over time or merely reactive in the moment. Get the taxonomy right first — understand which of the four memory types your use case actually needs — then choose the simplest implementation that meets each requirement. Most production agents need working memory (the context window, already provided) plus a simple semantic store (user preferences, domain facts). Episodic and procedural memory add significant value for personalisation and task efficiency but require more infrastructure and maintenance. Build what you need rather than what memory architecture papers describe.