Pydantic AI integration
The v4 version of zep-pydantic-ai is not released yet. The current zep-pydantic-ai release uses the v3 API. To use this integration now, follow the v3 version of this page.
Pydantic AI agents using Zep gain long-term memory backed by a temporal knowledge graph. The zep-pydantic-ai package persists conversation turns and adds a model-callable graph-search tool. Its native capability can inject trusted content into the model prompt.
Zep context can include content that your users, documents, or tools supplied. A system or developer message gives that content higher instruction priority than ordinary input. Some convenience integrations use system-message injection. Use direct SDK retrieval or an actual retrieval tool call unless all stored content is application-authored and trusted. Follow Memory security best practices for provider-specific placement.
To build an agent that plans its retrieval and uses several Zep tools, read Build an Agent with Zep. The guide shows how to add domain knowledge, design tools, and evaluate the agent.
Core benefits
- Native Pydantic AI capabilities:
zep_capabilities(deps)bundles the currentProcessHistoryhistory-processor hook with aHooks(after_run=...)hook — not a deprecated kwarg - Automatic assistant persistence: The bundled
after_runhook persists the assistant’s reply when the run completes, so no manual persistence call is needed - Single round-trip: Persists the user turn and retrieves context in one
add_messagescall - Correct under tool calls: Dedupes per run (keyed by
RunContext.run_id), so a run that makes tool calls records the turn exactly once - Pin-or-expose graph search: A model-callable tool over
graph.search_edges— every search parameter is model-exposed by default, or pinned/hidden per deployment - Out-of-band provisioning:
create_user/create_threadcreate the Zep resources up front and return the UUIDs that the application stores - Graceful degradation: A Zep failure on the turn path is logged but never crashes the agent run
How it works
The integration plugs into Pydantic AI through three components:
ZepDeps— a dataclass used as the agent’sdeps_type. It carries the Zep client, the user UUID, the thread UUID, the optional graph UUID, and optional context-building configuration. Construct one per conversation and pass it toagent.run(..., deps=deps); the history processor, theafter_runhook, and the search tool all reach it throughRunContext.deps.zep_capabilities(deps)registers memory for fully trusted, application-authored context. It returnsProcessHistory(zep_history_processor)and aHooks(after_run=...)hook. The processor persists the latest user message withthread.add_messages(return_context=True)and prepends the context block as a system message. The hook persists the assistant reply when the run completes. Use this path only when all stored content is application-authored and trusted.create_zep_search_tool— a factory returning a model-callablepydantic_ai.Toolover the v4 graph search methods (for examplegraph.search_edges). The model decides when to search the knowledge graph and, by default, which search parameters to use.
Because ProcessHistory fires once per model request (not once per run), the history processor dedupes per run, keyed by RunContext.run_id: it persists and retrieves on the first model request of a run and replays the cached context on later requests within that same run, so tool-calling runs never create duplicate episodes.
Installation
Requires Python 3.11+, pydantic-ai>=1.107,<2, and a Zep Cloud API key. Get your API key from app.getzep.com.
Zep v4 addresses a user, a thread, and a graph by a server-generated UUID. A v4 create call accepts no client-chosen name, and the server rejects a request that sends one. Your application creates the user and the thread one time, reads the UUIDs from the responses, stores them in its own database, and passes them to the integration. The integration does not resolve a name at run time.
Set up your environment variables:
Upgrading to the Zep v4 SDK
The package targets Zep v4 only. ZepDeps takes user_uuid, thread_uuid, and the optional graph_uuid in place of user_id and thread_id. ensure_user and ensure_thread are replaced by create_user and create_thread, which return the created model; the integration no longer creates a resource on the turn path. The search tool takes graph_uuid in place of graph_id. See the migration guide and the package changelog.
Upgrading from zep-pydantic-ai 0.1.x
Two changes can require code updates: create_zep_search_tool returns a pydantic_ai.Tool rather than a bare function — code that invoked the return value directly should call tool.function(ctx, query=..., **kwargs) — and the default injected context wording follows the canonical DEFAULT_CONTEXT_TEMPLATE; pass context_template=... on ZepDeps to keep custom wording. See the package changelog for the full list of changes.
Capability usage with trusted context
zep_capabilities(deps) inserts context into a system message. Use this pattern only for fully trusted, application-authored context. For other context, register create_zep_search_tool so retrieval uses an actual tool call. You can also retrieve context with the Zep SDK and place it through your provider’s documented data channel.
When you use the bundled capability, pass the same ZepDeps to each run. Both sides of every turn are persisted automatically. zep_capabilities(deps) closes over one ZepDeps instance, so construct the Agent inside your per-conversation setup rather than sharing it across users.
Explicit control over persistence
To control exactly when the assistant’s reply reaches Zep, register the history processor directly and call persist_run yourself after the run completes:
persist_run sends only assistant text — tool-call and tool-return scaffolding is skipped — so Zep records one clean assistant message per turn. It is not needed when the agent uses zep_capabilities(deps).
On-demand graph search
Beyond the automatic context injection, create_zep_search_tool() returns a model-callable pydantic_ai.Tool over the v4 graph search methods; pass it directly in tools=[...]. The model decides when to look up specific facts, entities, or prior episodes, and the tool returns a formatted text summary of the matching results. By default it searches the graph that graph_uuid on ZepDeps names; pass graph_uuid=... to the factory to target a shared Context Graph.
Every search parameter (scope, reranker, limit, mmr_lambda, center_node_uuid) is exposed to the model in the tool’s JSON schema by default, with documented defaults. Two constructor arguments override this per deployment: pinned_params fixes a parameter to a constant value and hides it from the schema, and hidden_params hides a parameter without pinning it, so Zep’s server-side default applies:
The scope, reranker, and limit constructor arguments are back-compat aliases that pin (and hide) those parameters; prefer pinned_params in new code. search_filters and bfs_origin_node_uuids are constructor-only — their complex shapes are not exposed to the model.
Memory vs tools
The integration combines two retrieval paths on the same agent:
Injection grounds each turn with cross-session context; the search tool lets the model actively dig for specific details.
Provisioning
create_user and create_thread create the Zep user and the thread out of band, before the first turn. Each helper returns the v4 model, so the application reads the UUIDs from the response and stores them:
A create call sends no client-chosen name, because the server generates the UUID. After the user exists, you can configure per-user resources against user.graph_uuid — a custom ontology, custom extraction instructions, or user summary instructions; see customizing graph structure for the available options.
The integration does not create a user or a thread on the turn path, and it does not look a name up at run time. It expects the stored UUIDs on ZepDeps. A Zep failure on the turn path is logged and degrades to no memory rather than breaking the run.
Custom context building
Set context_builder on ZepDeps to replace the default context retrieval with custom logic — for example, searching a different graph, applying filters, or combining multiple sources:
ContextInput is a frozen dataclass bundling zep (the AsyncZep client), user_uuid, thread_uuid, graph_uuid, user_message, and run_context (the Pydantic AI RunContext for the turn). Returning None skips injection for that turn.
When context_builder is set, message persistence (add_messages without return_context) and the builder run concurrently, with per-side failure isolation:
- If the builder raises, a warning is logged and context injection is skipped for that turn — persistence still completes.
- If persistence raises, a warning is logged and the turn is not marked as persisted (so it retries on the next model request) — a successful builder result is still injected.
Context template
context_template on ZepDeps controls how retrieved context is wrapped before injection. It must contain a literal {context} placeholder. Plain string replacement prevents format-string interpretation of {, }, or %. It does not prevent the model from following instructions in the retrieved content.
The default is DEFAULT_CONTEXT_TEMPLATE, an explicit <ZEP_CONTEXT>...</ZEP_CONTEXT> block with canonical wording shared across Zep’s framework integrations.
Configuration options
ZepDeps
create_zep_search_tool
Constructor arguments (returns a pydantic_ai.Tool[ZepDeps]):
Model-exposed search parameters (when not pinned or hidden), with their defaults:
Best practices
- Construct one
ZepDepsper conversation and reuse a singleAsyncZepclient across runs - Pass real names so Zep can anchor the user’s identity node in the graph
- Use
create_zep_search_toolfor untrusted context. Usezep_capabilities(deps)only when all stored context is application-authored and trusted. - Create the user and the thread in your onboarding flow with
create_user/create_thread, and store the returned UUIDs in your own database - Allow time for indexing — Zep extracts knowledge asynchronously, so facts from a turn are not instantly searchable
Next steps
- Explore customizing graph structure for advanced knowledge organization
- Learn about searching the graph and how to tune search
- See code examples for additional patterns