> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs-beta.getzep.com/v4/autogen-memory/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-beta.getzep.com/_mcp/server. # AutoGen integration > **Note** > > The v4 version of `zep-autogen` is not released yet. The current `zep-autogen` release uses the v3 API. To use this integration now, follow the [v3 version of this page](/v3/autogen-memory). The `zep-autogen` package gives [Microsoft AutoGen](https://github.com/microsoft/autogen) agents long-term memory and a temporal knowledge graph. Function tools let the agent search and add data. Memory classes provide automatic injection for trusted content. > **Keep retrieved context out of privileged instructions** > > Zep context can include content that your users, documents, or tools supplied. A system or developer message gives that content higher instruction priority than ordinary input. Some convenience integrations use system-message injection. Use direct SDK retrieval or an actual retrieval tool call unless all stored content is application-authored and trusted. Follow [Memory security best practices](/memory-security) for provider-specific placement. > **Build an agent with Zep tools** > > To build an agent that plans its retrieval and uses several Zep tools, read [Build an Agent with Zep](/build-an-agent-with-zep). The guide shows how to add domain knowledge, design tools, and evaluate the agent. Use the function tools for context that can contain end-user or third-party data. Use the memory classes' system-message injection only for fully trusted, application-authored context. ## Core benefits * **Native `Memory` interface**: `ZepUserMemory` and `ZepGraphMemory` implement AutoGen's `Memory` interface, so they drop straight into an agent's `memory` list * **Automatic context injection for trusted content**: `update_context()` can prepend memory that contains only application-authored, trusted data * **User and shared Context Graphs**: Persist a user's conversation history or maintain shared context with custom entity models * **On-demand function tools**: Pre-built tools let the agent explicitly search and add graph data when it chooses * **Graceful degradation**: A Zep failure is logged but does not crash the agent run ## How it works The integration exposes two complementary retrieval paths: * **Memory classes for trusted content** (`ZepUserMemory`, `ZepGraphMemory`) attach to an agent's `memory` list. AutoGen calls `update_context()` before each turn and adds the retrieved memory as a system message. * **Function tools** (`create_search_graph_tool`, `create_add_graph_data_tool`) attach to an agent's `tools` list. The model decides when to call them, giving explicit, observable search and add operations that work with AutoGen's tool reflection. Use function tools for end-user or third-party context. Add a memory class only when all stored content is application-authored and trusted. Context injection is automatic, but persistence is not: AutoGen's `Memory` protocol has no hook that fires after the model responds, so your application calls `memory.add()` explicitly — typically once per user turn and once per assistant turn. This is AutoGen's design, not a limitation of the integration. ## Installation **`pip`** ```bash pip pip install zep-autogen autogen-core autogen-agentchat ``` **`uv`** ```bash uv uv add zep-autogen autogen-core autogen-agentchat ``` **`poetry`** ```bash poetry poetry add zep-autogen autogen-core autogen-agentchat ``` > **Info** > > Requires Python 3.11+, `autogen-agentchat>=0.7.0`, and a Zep Cloud API key. Get your API key from [app.getzep.com](https://app.getzep.com). Set up your environment variables: ```bash export ZEP_API_KEY="your-zep-api-key" export OPENAI_API_KEY="your-openai-api-key" ``` ## Identifiers in Zep v4 Zep v4 assigns the UUID of every user, thread, and graph. A `user_id` or a `thread_id` is a name, not an address. The public API of `zep-autogen` takes `user_uuid`, `thread_uuid`, and `graph_uuid`. Your application creates each resource one time, reads `uuid_` from the response, and stores the UUID in its own database. The integration does not resolve a name at run time. #### Upgrading from zep-autogen 1.2.x The v4 release changes the public API. > **Warning** > > `ZepUserMemory`, `ZepGraphMemory`, and the tool factories take UUIDs. Replace `user_id` with `user_uuid`, `thread_id` with `thread_uuid`, and `graph_id` with `graph_uuid`. Replace `context_template_id` with `context_template_uuid`. `ensure_user` and `ensure_thread` are replaced by `create_user` and `create_thread`, which return the SDK model. `ZepUserMemory` no longer creates the Zep user. `ZepGraphMemory` replaces `facts_limit` and `entity_limit` with `max_characters`. See the [changelog](https://github.com/getzep/zep/blob/main/integrations/autogen/python/CHANGELOG.md) for the full release history. ## Memory types * **User memory**: Stores conversation history in [user threads](/users) with automatic context injection * **Knowledge graph memory**: Maintains structured knowledge with [custom entity models](/customizing-graph-structure) ## User memory Use model-callable tools for memory that contains conversation or third-party data. The `ZepUserMemory` example below documents automatic injection for a closed deployment where all stored content is trusted. `ZepUserMemory` persists messages to a user's thread and injects the context block into the agent before each turn. Set up the imports, initialize the memory, attach it to an agent, then store messages as the conversation proceeds. ### Import dependencies ```python import os import asyncio from autogen_agentchat.agents import AssistantAgent from autogen_ext.models.openai import OpenAIChatCompletionClient from autogen_core.memory import MemoryContent, MemoryMimeType from zep_cloud.client import AsyncZep from zep_autogen import ZepUserMemory, create_thread, create_user ``` ### Create the user and the thread `create_user` and `create_thread` call the Zep v4 create methods and return the SDK model. Read `uuid_` from each model and store the value in your own database. A create call does not pass a `user_id` or a `thread_id`. ```python zep_client = AsyncZep(api_key=os.environ.get("ZEP_API_KEY")) user = await create_user( zep_client, first_name="Alice", email="alice@example.com", ) user_uuid = str(user.uuid_) thread = await create_thread(zep_client, user_uuid=user_uuid) thread_uuid = str(thread.uuid_) ``` `create_user` accepts an `on_created` hook. The hook runs after Zep creates the user, and it receives the client and the new user UUID. Use the hook for one-time setup, such as an ontology on the graph of the user. An error in the hook propagates to the caller. ### Initialize the memory `ZepUserMemory` binds the client, the user UUID, and the thread UUID into a memory object that AutoGen can attach to an agent. ```python memory = ZepUserMemory( client=zep_client, user_uuid=user_uuid, thread_uuid=thread_uuid, ) ``` | Parameter | Description | | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | `client` | An initialized `AsyncZep` instance (required) | | `user_uuid` | UUID of the Zep user (required) | | `thread_uuid` | UUID of the Zep thread; `add()` creates a thread when the value is omitted | | `graph_uuid` | UUID of the graph that receives data with `type="data"`; defaults to the graph of the user | | `context_template_uuid` | UUID of the Zep context template used to render the retrieved context block; ignored when `context_builder` is set | | `context_builder` | Async callable replacing the default context retrieval in `update_context()` — see [custom context retrieval](#custom-context-retrieval) | | `context_template` | Template wrapping injected context; defaults to `DEFAULT_CONTEXT_TEMPLATE` | Thread creation in `add()` never raises: a failure is logged and swallowed. Create the thread with `create_thread` before the first turn when you want a failure to surface. ### Attach trusted memory to an agent Pass the memory in the agent's `memory` list only when all stored content is fully trusted and application-authored. ```python # Trusted-only path: AutoGen inserts this memory into a system message. agent = AssistantAgent( name="MemoryAwareAssistant", model_client=OpenAIChatCompletionClient( model="gpt-5.6-terra", api_key=os.environ.get("OPENAI_API_KEY") ), memory=[memory], system_message="You are a helpful assistant with persistent memory." ) ``` ### Store messages and run Persistence is manual: AutoGen never calls `memory.add()` for you, so persist each turn explicitly — once for the user message and once for the assistant reply. The agent automatically retrieves context via `update_context()` before responding; skipping the `add()` calls means the agent still sees Zep's existing context, but that turn's messages are never written to Zep and cannot be recalled later. ```python # Helper function to store messages with proper metadata async def add_message(message: str, role: str, name: str = None): """Store a message in Zep memory following AutoGen standards.""" metadata = {"type": "message", "role": role} if name: metadata["name"] = name await memory.add(MemoryContent( content=message, mime_type=MemoryMimeType.TEXT, metadata=metadata )) # Example conversation with memory persistence user_message = "My name is Alice and I love hiking in the mountains." print(f"User: {user_message}") # Store user message await add_message(user_message, "user", "Alice") # Run agent - it will automatically retrieve context via update_context() response = await agent.run(task=user_message) agent_response = response.messages[-1].content print(f"Agent: {agent_response}") # Store agent response await add_message(agent_response, "assistant") ``` > **Info** > > **Automatic context injection**: `ZepUserMemory` injects relevant memory via the `update_context()` method before each turn. On the default retrieval path it injects the context block and, when one is available, also appends up to 10 recent thread messages. When a `context_builder` is set, only the builder's output is injected. > **Note** > > **Allow time for indexing** — Zep extracts knowledge asynchronously, so facts from a turn are not instantly searchable. Allow time for indexing before querying for newly added content. ## Custom context retrieval By default, `update_context()` retrieves context via `thread.get_context(...)`. Pass `context_builder` to replace this with custom logic — for example a filtered graph search, or a different graph entirely: ```python from zep_autogen.memory import ContextInput async def my_builder(ctx: ContextInput) -> str | None: pager = await ctx.zep.graph.search_edges( ctx.graph_uuid, query=ctx.user_message, limit=10, ) facts = [edge.fact async for edge in pager if edge.fact] if not facts: return None return "\n".join(facts) memory = ZepUserMemory( client=zep_client, user_uuid=user_uuid, thread_uuid=thread_uuid, context_builder=my_builder, ) ``` Each v4 search method returns a pager. Iterate the pager with `async for` to read the results. The builder receives a single frozen `ContextInput`: | Field | Description | | --------------- | ------------------------------------------------------------------------------ | | `zep` | The `AsyncZep` client in use by this memory instance | | `user_uuid` | The UUID of the Zep user the memory is scoped to | | `thread_uuid` | The UUID of the Zep thread the memory records the conversation in | | `graph_uuid` | The UUID of the graph the memory searches | | `user_message` | The last user-role message's text from the model context (`""` if none) | | `model_context` | The AutoGen `ChatCompletionContext` passed to `update_context()` for this call | If the builder raises, a warning is logged and context injection is skipped for that turn — `update_context()` never raises. The builder is retrieval-only and never runs concurrently with message persistence: AutoGen's `Memory` protocol calls `update_context()` (injection) and `add()` (persistence) as two separate, caller-controlled steps, so persist turns explicitly via `add()`. ### Customizing the injected context template Retrieved context (from the default retrieval or a `context_builder`) is wrapped in `context_template` before being added to the model context as a system message. The default `DEFAULT_CONTEXT_TEMPLATE` wraps the context in `` tags with a short preamble. Override it with your own wording, as long as it contains a literal `{context}` placeholder: ```python memory = ZepUserMemory( client=zep_client, user_uuid=user_uuid, thread_uuid=thread_uuid, context_template="Relevant background:\n{context}", ) ``` The template uses plain string replacement (`template.replace("{context}", ...)`), never `str.format`. This prevents format-string interpretation of `{`, `}`, or `%`. It does not prevent the model from following instructions in the retrieved content. ## Shared Context Graph memory `ZepGraphMemory` maintains a shared Context Graph with custom entity models. Define an ontology, create the graph, initialize the memory with search filters, add data, and then attach the memory to an agent. `ZepGraphMemory` is scoped to a shared Context Graph that is addressed with `graph_uuid`. It is not scoped to a Zep user, so it has no `on_created` hook. Create the graph with `graph.create` and read its UUID, as shown below. ### Define entity models Custom entity models shape how Zep extracts structured knowledge from the data you add. ```python from zep_autogen.graph_memory import ZepGraphMemory from zep_cloud import EntityProperty, EntityType ENTITY_TYPES = [ EntityType( name="ProgrammingLanguage", description="A programming language entity.", properties=[ EntityProperty( name="paradigm", type="text", description="programming paradigm, such as object-oriented or functional", ), EntityProperty( name="use_case", type="text", description="primary use cases for this language", ), ], ), EntityType( name="Framework", description="A software framework or library.", properties=[ EntityProperty( name="language", type="text", description="the programming language this framework is built for", ), EntityProperty( name="purpose", type="text", description="primary purpose of this framework", ), ], ), ] ``` ### Create the graph and set the ontology Create the graph that holds the extracted knowledge, then register the entity types as the ontology of that graph. Zep assigns the graph UUID, and the ontology call addresses the graph by that UUID. ```python from zep_cloud import SearchFilters graph = await zep_client.graph.create(name="Programming Knowledge Graph") graph_uuid = str(graph.uuid_) await zep_client.graph.set_ontology(graph_uuid, entity_types=ENTITY_TYPES) ``` ### Initialize the graph memory Configure search filters and context limits to control what `ZepGraphMemory` injects on each turn. ```python # Create graph memory with search configuration graph_memory = ZepGraphMemory( client=zep_client, graph_uuid=graph_uuid, search_filters=SearchFilters( node_labels=["ProgrammingLanguage", "Framework"] ), max_characters=4000, # Optional budget for the injected context block ) ``` ### Add data and wait for indexing Knowledge extraction is asynchronous, so allow time for indexing before the data is searchable. ```python # Add structured knowledge await graph_memory.add(MemoryContent( content="Python is excellent for data science and AI development", mime_type=MemoryMimeType.TEXT, metadata={"type": "data"} # "data" stores in graph, "message" stores as episode )) # Wait for graph processing (required) print("Waiting for graph indexing...") await asyncio.sleep(30) # Allow time for knowledge extraction ``` ### Attach trusted graph memory to an agent Pass graph memory in the agent's `memory` list only when all graph content is application-authored and trusted. Use `create_search_graph_tool` for end-user or third-party graph content. ```python # Trusted-only path: AutoGen inserts graph memory into a system message. agent = AssistantAgent( name="GraphMemoryAssistant", model_client=OpenAIChatCompletionClient(model="gpt-5.6-terra"), memory=[graph_memory], system_message="You are a technical assistant with programming knowledge." ) ``` > **Info** > > **Trusted graph memory injection**: `ZepGraphMemory` reads the two most recent episodes of the graph with `graph.episode.list`, then calls `graph.get_context` with their content as the query. AutoGen inserts the returned context block into a system message. `max_characters` limits the size of that block. Use this path only for fully trusted graph content. ## Tools integration Zep tools let agents search and add data directly to memory storage with manual control and structured responses. > **Warning** > > **Important**: Bind a tool to either `graph_uuid` or `user_uuid`, not both. `graph_uuid` selects a shared Context Graph. `user_uuid` selects the graph of a user. ### Search tool parameters `create_search_graph_tool` follows a **pin-or-expose** pattern: every `graph.search_edges` parameter is exposed to the model by default, each with a typed schema and documented default. Letting the model choose the scope and reranker per query produces better retrieval than a single fixed configuration; pin parameters when you need deterministic behavior instead. `query` is always exposed and required. | Parameter | Default | Description | | ------------------ | --------- | -------------------------------------------------------------------------------- | | `scope` | `"edges"` | One of `edges`, `nodes`, `episodes`, `observations`, `thread_summaries`, `auto` | | `reranker` | `"rrf"` | One of `rrf`, `mmr`, `node_distance`, `episode_mentions`, `cross_encoder` | | `limit` | `10` | Maximum number of results; the integration clamps the value to the range 1 to 50 | | `mmr_lambda` | `None` | Diversity (0.0) vs. relevance (1.0) balance; only used when `reranker="mmr"` | | `center_node_uuid` | `None` | Center node for `reranker="node_distance"` | Use `pinned_params` to fix a parameter to a constant (hidden from the model), or `hidden_params` to remove it from the schema without pinning (Zep's server-side default applies): ```python # Model chooses scope/reranker/limit/mmr_lambda/center_node_uuid freely (default) tool = create_search_graph_tool(zep_client, user_uuid=user_uuid) # Pin scope to "nodes" and limit to 5 — hidden from the model, always sent as given tool = create_search_graph_tool( zep_client, user_uuid=user_uuid, pinned_params={"scope": "nodes", "limit": 5} ) # Hide mmr_lambda from the schema without pinning it — Zep's own default applies tool = create_search_graph_tool(zep_client, user_uuid=user_uuid, hidden_params={"mmr_lambda"}) ``` The legacy `scope` and `limit` arguments pin (and hide) the corresponding parameter — equivalent to passing them via `pinned_params`. `search_filters` and `bfs_origin_node_uuids` are constructor-only and never exposed to the model. > **Note** > > AutoGen's `FunctionTool` derives its JSON schema strictly from the wrapped function's typed signature. `create_search_graph_tool` implements pin-or-expose by building that signature dynamically: exposed parameters become real, typed parameters of the function AutoGen introspects, while pinned and hidden parameters are never part of the signature at all. ### Add tool parameters `create_add_graph_data_tool` exposes: * `data`: str (required) - Content to store * `data_type`: str (optional, default "text") - Data type: "text", "json", "message" ### User graph tools ```python from zep_autogen import create_search_graph_tool, create_add_graph_data_tool # Create tools bound to the graph of the user search_tool = create_search_graph_tool(zep_client, user_uuid=user_uuid) add_tool = create_add_graph_data_tool(zep_client, user_uuid=user_uuid) # Agent with user graph tools agent = AssistantAgent( name="UserKnowledgeAssistant", model_client=OpenAIChatCompletionClient(model="gpt-5.6-terra"), tools=[search_tool, add_tool], system_message="You can search and add data to the user's knowledge graph.", reflect_on_tool_use=True # Enables tool usage reflection ) ``` ### Knowledge graph tools ```python # Create tools bound to the knowledge graph search_tool = create_search_graph_tool(zep_client, graph_uuid=graph_uuid) add_tool = create_add_graph_data_tool(zep_client, graph_uuid=graph_uuid) # Agent with knowledge graph tools agent = AssistantAgent( name="KnowledgeGraphAssistant", model_client=OpenAIChatCompletionClient(model="gpt-5.6-terra"), tools=[search_tool, add_tool], system_message="You can search and add data to the knowledge graph.", reflect_on_tool_use=True ) ``` ## Size limits Zep rejects over-long direct SDK payloads with an HTTP 400. The AutoGen integration truncates before calling Zep, logging only the before and after lengths (never the content): * **Thread messages** (`ZepUserMemory.add` with `type="message"`): truncated to 4,000 characters, a safety margin under Zep's 4,096-character thread-message limit * **Graph data** (`ZepGraphMemory.add`, `ZepUserMemory.add` with `type="data"`, and `create_add_graph_data_tool`): truncated to 9,900 characters, a safety margin under Zep's 10,000-character `graph.episode.add` limit ## Query memory Both memory types support direct querying with different scope parameters. ### User memory queries ```python # Query user conversation history results = await memory.query("What does Alice like?", limit=5) # Process different result types for result in results.results: content = result.content metadata = result.metadata if 'edge_name' in metadata: # Fact/relationship result print(f"Fact: {content}") print(f"Relationship: {metadata['edge_name']}") print(f"Valid: {metadata.get('valid_at', 'N/A')} - {metadata.get('invalid_at', 'present')}") elif 'node_name' in metadata: # Entity result print(f"Entity: {metadata['node_name']}") print(f"Summary: {content}") else: # Episode/message result print(f"Message: {content}") print(f"Role: {metadata.get('episode_role', 'unknown')}") print(f"Source: {metadata.get('source')}\n") ``` ### Graph memory queries ```python # Query knowledge graph with scope control facts_results = await graph_memory.query( "Python frameworks", limit=10, ) print(f"Found {len(facts_results.results)} facts about Python frameworks:") for result in facts_results.results: print(f"- {result.content}") entities_results = await graph_memory.query( "programming languages", limit=5, ) print(f"\nFound {len(entities_results.results)} programming language entities:") for result in entities_results.results: entity_name = result.metadata.get('node_name', 'Unknown') print(f"- {entity_name}: {result.content}") ``` ### Search result structure #### Edge results (facts) ```json { "content": "fact text", "metadata": { "source": "graph" | "user_graph", "edge_name": "relationship_name", "edge_attributes": {...}, "created_at": "timestamp", "valid_at": "timestamp", "invalid_at": "timestamp", "expired_at": "timestamp" } } ``` #### Node results (entities) ```json { "content": "entity_name:\n entity_summary", "metadata": { "source": "graph" | "user_graph", "node_name": "entity_name", "node_attributes": {...}, "created_at": "timestamp" } } ``` #### Episode results (messages) ```json { "content": "episode_content", "metadata": { "source": "graph" | "user_graph", "episode_type": "source_type", "episode_role": "role_type", "episode_name": "role_name", "created_at": "timestamp" } } ``` ## Memory vs tools comparison > **Note** > > **Memory objects** (`ZepUserMemory` / `ZepGraphMemory`): > > * Automatic context injection via `update_context()`; persistence stays manual via `add()` > * Attached to the agent's `memory` list > * Transparent operation — happens automatically > * Better for consistent memory across interactions > > **Function tools** (search/add tools): > > * Manual control — the agent decides when to use them > * More explicit and observable operations > * Better for specific search/add operations > * Works with AutoGen's tool reflection features > * Provides structured return values > > Use function tools for untrusted context. Add a memory class only when all stored content is application-authored and trusted. ## Best practices * **Pick the right memory type** — use `ZepUserMemory` for a user graph and `ZepGraphMemory` for a shared Context Graph * **Persist every turn explicitly** — call `memory.add()` once per user turn and once per assistant turn; injection is the only automatic half of the loop * **Store the UUIDs** — keep the user UUID, the thread UUID, and the graph UUID in your own database after you create each resource * **Bind tools to exactly one scope** — a search or add tool takes either a `graph_uuid` or a `user_uuid`, never both * **Use function tools for untrusted context.** Attach a memory class only when all stored content is application-authored and trusted. * **Allow time for indexing** — Zep extracts knowledge asynchronously, so facts from a turn are not instantly searchable ## Next steps * Explore [customizing graph structure](/customizing-graph-structure) for advanced knowledge organization * Learn about [searching the graph](/searching-the-graph) and how to tune search * See [code examples](https://github.com/getzep/zep/tree/main/integrations/autogen/python/examples) for additional patterns > Add long-term agent memory to AutoGen agents