Agent Memory & State
A chatbot that forgets everything after each message is a toy. A production agent needs to remember who you are, what you asked last week, and what it already tried that didn't work. Memory turns a stateless API call into a persistent collaborator.
Three layers, each with a different lifespan and purpose. The art is deciding what goes in each layer โ and what gets dropped.
Without memory, every conversation starts from zero.
A stateless LLM call is a blank slate. Every request carries the full context in the prompt โ system instructions, chat history, retrieved documents. When the response comes back, all that context is discarded. The next message starts over.
This works for one-shot tasks. It breaks for anything that spans multiple turns: a debugging session that lasts an hour, a project that spans a week, a customer relationship that spans months. Memory is the layer that lets an agent pick up where it left off without re-explaining everything from scratch.
Not all memory is the same. Each type solves a different problem.
Conversation Memory
The last N messages in the current session. Simple sliding window. Old messages fall off. This is the "what were we just talking about" layer. Every chatbot has this by default โ it's just the chat history array.
Episodic Memory
Specific past interactions stored for later retrieval. "Remember when we debugged the auth system last month and you suggested using JWTs?" Stored in a vector database as embeddings, retrieved by semantic similarity when the current conversation needs that context.
Semantic Memory
Facts and knowledge extracted from conversations. "The user prefers TypeScript over Python" or "The production database is PostgreSQL 16." These are distilled truths, not raw conversation logs. Usually stored as key-value pairs or structured profiles.
Procedural Memory
Learned patterns of behavior. "When the user says 'ship it,' they mean: run tests, build, deploy to staging, wait for CI, then promote to production." The agent learns multi-step workflows from repeated interactions and stores them as reusable routines.
Three building blocks, combined differently for each memory type.
Vector Databases
Embed conversations into vectors, store them, retrieve by semantic similarity. When the current conversation mentions "auth bug," the vector DB surfaces past conversations about authentication issues. Tools: pgvector, Pinecone, Chroma, Qdrant.
Summarization
Periodically ask the model to summarize the conversation so far. Store the summary as compressed context. When the conversation gets too long for the context window, replace the raw history with the summary. Loses detail but preserves the gist.
Structured Profiles
Extract facts into a structured user profile: name, preferences, project details, tech stack. Update fields as new information appears. Query the profile before each session starts. This is the "who are you and what are we working on" layer.
Deciding what to keep and what to drop is the hardest problem in agent memory.
Context windows have limits. You cannot store every conversation forever. At some point, something has to be dropped, summarized, or archived. The question is what โ and when.
Keep everything. Raw conversation logs grow unbounded. Retrieval quality degrades as the database fills with noise. Old, irrelevant conversations pollute searches with false matches.
Keep facts, not transcripts. Extract what was learned from each conversation and discard the raw log. "User runs PostgreSQL 16" is 30 bytes. The 5,000-token conversation that established it can be dropped.
A common pattern: summarize conversations into compressed form after they end. Store those summaries, not the raw transcripts. When memory retrieval runs, it searches summaries first. Only pull the full transcript if the summary indicates high relevance.
Memory systems that survive real-world usage.
Hybrid Retrieval
Combine vector search (semantic similarity) with keyword search (exact matches) and recency ranking (newer conversations score higher). Pure vector search misses exact names and dates. Pure keyword search misses conceptual connections. Combined, they don't miss much.
Tiered Storage
Hot memory (last 24 hours): full transcripts in context window. Warm memory (last 30 days): summarized conversations in vector DB. Cold memory (older): heavily compressed summaries in cheap object storage. Retrieval checks tiers in order, stops at first hit.
Intent-Based Retrieval
Don't search memory on every message. Trigger retrieval only when the current message references the past: "remember when," "like we discussed," "same issue as." Classify the intent first, then decide whether to pay the retrieval cost.
User-Controlled Forget
Give users explicit control: "forget this conversation," "forget everything about project X," "don't remember anything I said in the last hour." Memory without user-facing forget controls is a privacy liability. Every memory system needs a delete path.
"The difference between a chatbot and an agent is memory. Everything else is just an API call."
Store facts, not transcripts. A structured fact like "user prefers TypeScript" is 30 bytes and infinitely reusable. The 5,000-token conversation that established it is noise.
Every memory system needs a delete path before it goes to production. Users must be able to forget โ and the system must actually forget, not just hide.