โ† Back to ConceptsVerified: 2026-07-30
Agent MemoryPersistence Across Sessions

Agent Memory & State

A chatbot that forgets everything after each message is a toy. A production agent needs to remember who you are, what you asked last week, and what it already tried that didn't work. Memory turns a stateless API call into a persistent collaborator.

CONVERSATION BUFFER โ€” last 10-20 messages, fed into every requestImmediate context. The agent's short-term working set. Lost when the session ends.WORKING MEMORY โ€” session state: current task, active tools, known facts, user preferencesSurvives across turns within a session. Reset when the task completes.LONG-TERM MEMORY โ€” vector DB + summaries: past conversations, learned facts, user profilePersists across sessions. What the agent knows about you when you come back next week.

Three layers, each with a different lifespan and purpose. The art is deciding what goes in each layer โ€” and what gets dropped.

01Why It Matters

Without memory, every conversation starts from zero.

A stateless LLM call is a blank slate. Every request carries the full context in the prompt โ€” system instructions, chat history, retrieved documents. When the response comes back, all that context is discarded. The next message starts over.

This works for one-shot tasks. It breaks for anything that spans multiple turns: a debugging session that lasts an hour, a project that spans a week, a customer relationship that spans months. Memory is the layer that lets an agent pick up where it left off without re-explaining everything from scratch.

02Types of Memory

Not all memory is the same. Each type solves a different problem.

Type 1

Conversation Memory

The last N messages in the current session. Simple sliding window. Old messages fall off. This is the "what were we just talking about" layer. Every chatbot has this by default โ€” it's just the chat history array.

Type 2

Episodic Memory

Specific past interactions stored for later retrieval. "Remember when we debugged the auth system last month and you suggested using JWTs?" Stored in a vector database as embeddings, retrieved by semantic similarity when the current conversation needs that context.

Type 3

Semantic Memory

Facts and knowledge extracted from conversations. "The user prefers TypeScript over Python" or "The production database is PostgreSQL 16." These are distilled truths, not raw conversation logs. Usually stored as key-value pairs or structured profiles.

Type 4

Procedural Memory

Learned patterns of behavior. "When the user says 'ship it,' they mean: run tests, build, deploy to staging, wait for CI, then promote to production." The agent learns multi-step workflows from repeated interactions and stores them as reusable routines.

03How It's Built

Three building blocks, combined differently for each memory type.

Block 1

Vector Databases

Embed conversations into vectors, store them, retrieve by semantic similarity. When the current conversation mentions "auth bug," the vector DB surfaces past conversations about authentication issues. Tools: pgvector, Pinecone, Chroma, Qdrant.

Block 2

Summarization

Periodically ask the model to summarize the conversation so far. Store the summary as compressed context. When the conversation gets too long for the context window, replace the raw history with the summary. Loses detail but preserves the gist.

Block 3

Structured Profiles

Extract facts into a structured user profile: name, preferences, project details, tech stack. Update fields as new information appears. Query the profile before each session starts. This is the "who are you and what are we working on" layer.

04The Hard Part

Deciding what to keep and what to drop is the hardest problem in agent memory.

Context windows have limits. You cannot store every conversation forever. At some point, something has to be dropped, summarized, or archived. The question is what โ€” and when.

Don't

Keep everything. Raw conversation logs grow unbounded. Retrieval quality degrades as the database fills with noise. Old, irrelevant conversations pollute searches with false matches.

Do

Keep facts, not transcripts. Extract what was learned from each conversation and discard the raw log. "User runs PostgreSQL 16" is 30 bytes. The 5,000-token conversation that established it can be dropped.

A common pattern: summarize conversations into compressed form after they end. Store those summaries, not the raw transcripts. When memory retrieval runs, it searches summaries first. Only pull the full transcript if the summary indicates high relevance.

05Production Patterns

Memory systems that survive real-world usage.

Pattern 1

Hybrid Retrieval

Combine vector search (semantic similarity) with keyword search (exact matches) and recency ranking (newer conversations score higher). Pure vector search misses exact names and dates. Pure keyword search misses conceptual connections. Combined, they don't miss much.

Pattern 2

Tiered Storage

Hot memory (last 24 hours): full transcripts in context window. Warm memory (last 30 days): summarized conversations in vector DB. Cold memory (older): heavily compressed summaries in cheap object storage. Retrieval checks tiers in order, stops at first hit.

Pattern 3

Intent-Based Retrieval

Don't search memory on every message. Trigger retrieval only when the current message references the past: "remember when," "like we discussed," "same issue as." Classify the intent first, then decide whether to pay the retrieval cost.

Pattern 4

User-Controlled Forget

Give users explicit control: "forget this conversation," "forget everything about project X," "don't remember anything I said in the last hour." Memory without user-facing forget controls is a privacy liability. Every memory system needs a delete path.

"The difference between a chatbot and an agent is memory. Everything else is just an API call."

01

Store facts, not transcripts. A structured fact like "user prefers TypeScript" is 30 bytes and infinitely reusable. The 5,000-token conversation that established it is noise.

02

Every memory system needs a delete path before it goes to production. Users must be able to forget โ€” and the system must actually forget, not just hide.