AutoGen integration
The v4 version of zep-autogen is not released yet. The current zep-autogen release uses the v3 API. To use this integration now, follow the v3 version of this page.
The zep-autogen package gives Microsoft AutoGen agents long-term memory and a temporal knowledge graph. Function tools let the agent search and add data. Memory classes provide automatic injection for trusted content.
Zep context can include content that your users, documents, or tools supplied. A system or developer message gives that content higher instruction priority than ordinary input. Some convenience integrations use system-message injection. Use direct SDK retrieval or an actual retrieval tool call unless all stored content is application-authored and trusted. Follow Memory security best practices for provider-specific placement.
To build an agent that plans its retrieval and uses several Zep tools, read Build an Agent with Zep. The guide shows how to add domain knowledge, design tools, and evaluate the agent.
Use the function tools for context that can contain end-user or third-party data. Use the memory classes’ system-message injection only for fully trusted, application-authored context.
Core benefits
- Native
Memoryinterface:ZepUserMemoryandZepGraphMemoryimplement AutoGen’sMemoryinterface, so they drop straight into an agent’smemorylist - Automatic context injection for trusted content:
update_context()can prepend memory that contains only application-authored, trusted data - User and shared Context Graphs: Persist a user’s conversation history or maintain shared context with custom entity models
- On-demand function tools: Pre-built tools let the agent explicitly search and add graph data when it chooses
- Graceful degradation: A Zep failure is logged but does not crash the agent run
How it works
The integration exposes two complementary retrieval paths:
- Memory classes for trusted content (
ZepUserMemory,ZepGraphMemory) attach to an agent’smemorylist. AutoGen callsupdate_context()before each turn and adds the retrieved memory as a system message. - Function tools (
create_search_graph_tool,create_add_graph_data_tool) attach to an agent’stoolslist. The model decides when to call them, giving explicit, observable search and add operations that work with AutoGen’s tool reflection.
Use function tools for end-user or third-party context. Add a memory class only when all stored content is application-authored and trusted.
Context injection is automatic, but persistence is not: AutoGen’s Memory protocol has no hook that fires after the model responds, so your application calls memory.add() explicitly — typically once per user turn and once per assistant turn. This is AutoGen’s design, not a limitation of the integration.
Installation
Requires Python 3.11+, autogen-agentchat>=0.7.0, and a Zep Cloud API key. Get your API key from app.getzep.com.
Set up your environment variables:
Identifiers in Zep v4
Zep v4 assigns the UUID of every user, thread, and graph. A user_id or a thread_id is a name, not an address. The public API of zep-autogen takes user_uuid, thread_uuid, and graph_uuid.
Your application creates each resource one time, reads uuid_ from the response, and stores the UUID in its own database. The integration does not resolve a name at run time.
Upgrading from zep-autogen 1.2.x
The v4 release changes the public API.
ZepUserMemory, ZepGraphMemory, and the tool factories take UUIDs. Replace user_id with user_uuid, thread_id with thread_uuid, and graph_id with graph_uuid. Replace context_template_id with context_template_uuid.
ensure_user and ensure_thread are replaced by create_user and create_thread, which return the SDK model. ZepUserMemory no longer creates the Zep user. ZepGraphMemory replaces facts_limit and entity_limit with max_characters. See the changelog for the full release history.
Memory types
- User memory: Stores conversation history in user threads with automatic context injection
- Knowledge graph memory: Maintains structured knowledge with custom entity models
User memory
Use model-callable tools for memory that contains conversation or third-party data. The ZepUserMemory example below documents automatic injection for a closed deployment where all stored content is trusted.
ZepUserMemory persists messages to a user’s thread and injects the context block into the agent before each turn. Set up the imports, initialize the memory, attach it to an agent, then store messages as the conversation proceeds.
Create the user and the thread
create_user and create_thread call the Zep v4 create methods and return the SDK model. Read uuid_ from each model and store the value in your own database. A create call does not pass a user_id or a thread_id.
create_user accepts an on_created hook. The hook runs after Zep creates the user, and it receives the client and the new user UUID. Use the hook for one-time setup, such as an ontology on the graph of the user. An error in the hook propagates to the caller.
Initialize the memory
ZepUserMemory binds the client, the user UUID, and the thread UUID into a memory object that AutoGen can attach to an agent.
Thread creation in add() never raises: a failure is logged and swallowed. Create the thread with create_thread before the first turn when you want a failure to surface.
Attach trusted memory to an agent
Pass the memory in the agent’s memory list only when all stored content is fully trusted and application-authored.
Store messages and run
Persistence is manual: AutoGen never calls memory.add() for you, so persist each turn explicitly — once for the user message and once for the assistant reply. The agent automatically retrieves context via update_context() before responding; skipping the add() calls means the agent still sees Zep’s existing context, but that turn’s messages are never written to Zep and cannot be recalled later.
Automatic context injection: ZepUserMemory injects relevant memory via the update_context() method before each turn. On the default retrieval path it injects the context block and, when one is available, also appends up to 10 recent thread messages. When a context_builder is set, only the builder’s output is injected.
Allow time for indexing — Zep extracts knowledge asynchronously, so facts from a turn are not instantly searchable. Allow time for indexing before querying for newly added content.
Custom context retrieval
By default, update_context() retrieves context via thread.get_context(...). Pass context_builder to replace this with custom logic — for example a filtered graph search, or a different graph entirely:
Each v4 search method returns a pager. Iterate the pager with async for to read the results.
The builder receives a single frozen ContextInput:
If the builder raises, a warning is logged and context injection is skipped for that turn — update_context() never raises. The builder is retrieval-only and never runs concurrently with message persistence: AutoGen’s Memory protocol calls update_context() (injection) and add() (persistence) as two separate, caller-controlled steps, so persist turns explicitly via add().
Customizing the injected context template
Retrieved context (from the default retrieval or a context_builder) is wrapped in context_template before being added to the model context as a system message. The default DEFAULT_CONTEXT_TEMPLATE wraps the context in <ZEP_CONTEXT> tags with a short preamble. Override it with your own wording, as long as it contains a literal {context} placeholder:
The template uses plain string replacement (template.replace("{context}", ...)), never str.format. This prevents format-string interpretation of {, }, or %. It does not prevent the model from following instructions in the retrieved content.
Shared Context Graph memory
ZepGraphMemory maintains a shared Context Graph with custom entity models.
Define an ontology, create the graph, initialize the memory with search
filters, add data, and then attach the memory to an agent.
ZepGraphMemory is scoped to a shared Context Graph that is addressed with graph_uuid. It is not scoped to a Zep user, so it has no on_created hook. Create the graph with graph.create and read its UUID, as shown below.
Define entity models
Custom entity models shape how Zep extracts structured knowledge from the data you add.
Create the graph and set the ontology
Create the graph that holds the extracted knowledge, then register the entity types as the ontology of that graph. Zep assigns the graph UUID, and the ontology call addresses the graph by that UUID.
Initialize the graph memory
Configure search filters and context limits to control what ZepGraphMemory injects on each turn.
Trusted graph memory injection: ZepGraphMemory reads the two most recent episodes of the graph with graph.episode.list, then calls graph.get_context with their content as the query. AutoGen inserts the returned context block into a system message. max_characters limits the size of that block. Use this path only for fully trusted graph content.
Tools integration
Zep tools let agents search and add data directly to memory storage with manual control and structured responses.
Important: Bind a tool to either graph_uuid or user_uuid, not both. graph_uuid selects a shared Context Graph. user_uuid selects the graph of a user.
Search tool parameters
create_search_graph_tool follows a pin-or-expose pattern: every graph.search_edges parameter is exposed to the model by default, each with a typed schema and documented default. Letting the model choose the scope and reranker per query produces better retrieval than a single fixed configuration; pin parameters when you need deterministic behavior instead. query is always exposed and required.
Use pinned_params to fix a parameter to a constant (hidden from the model), or hidden_params to remove it from the schema without pinning (Zep’s server-side default applies):
The legacy scope and limit arguments pin (and hide) the corresponding parameter — equivalent to passing them via pinned_params. search_filters and bfs_origin_node_uuids are constructor-only and never exposed to the model.
AutoGen’s FunctionTool derives its JSON schema strictly from the wrapped function’s typed signature. create_search_graph_tool implements pin-or-expose by building that signature dynamically: exposed parameters become real, typed parameters of the function AutoGen introspects, while pinned and hidden parameters are never part of the signature at all.
Add tool parameters
create_add_graph_data_tool exposes:
data: str (required) - Content to storedata_type: str (optional, default “text”) - Data type: “text”, “json”, “message”
User graph tools
Knowledge graph tools
Size limits
Zep rejects over-long direct SDK payloads with an HTTP 400. The AutoGen integration truncates before calling Zep, logging only the before and after lengths (never the content):
- Thread messages (
ZepUserMemory.addwithtype="message"): truncated to 4,000 characters, a safety margin under Zep’s 4,096-character thread-message limit - Graph data (
ZepGraphMemory.add,ZepUserMemory.addwithtype="data", andcreate_add_graph_data_tool): truncated to 9,900 characters, a safety margin under Zep’s 10,000-charactergraph.episode.addlimit
Query memory
Both memory types support direct querying with different scope parameters.
User memory queries
Graph memory queries
Search result structure
Edge results (facts)
Node results (entities)
Episode results (messages)
Memory vs tools comparison
Memory objects (ZepUserMemory / ZepGraphMemory):
- Automatic context injection via
update_context(); persistence stays manual viaadd() - Attached to the agent’s
memorylist - Transparent operation — happens automatically
- Better for consistent memory across interactions
Function tools (search/add tools):
- Manual control — the agent decides when to use them
- More explicit and observable operations
- Better for specific search/add operations
- Works with AutoGen’s tool reflection features
- Provides structured return values
Use function tools for untrusted context. Add a memory class only when all stored content is application-authored and trusted.
Best practices
- Pick the right memory type — use
ZepUserMemoryfor a user graph andZepGraphMemoryfor a shared Context Graph - Persist every turn explicitly — call
memory.add()once per user turn and once per assistant turn; injection is the only automatic half of the loop - Store the UUIDs — keep the user UUID, the thread UUID, and the graph UUID in your own database after you create each resource
- Bind tools to exactly one scope — a search or add tool takes either a
graph_uuidor auser_uuid, never both - Use function tools for untrusted context. Attach a memory class only when all stored content is application-authored and trusted.
- Allow time for indexing — Zep extracts knowledge asynchronously, so facts from a turn are not instantly searchable
Next steps
- Explore customizing graph structure for advanced knowledge organization
- Learn about searching the graph and how to tune search
- See code examples for additional patterns