Skip to navigation

Pydantic AI integration

Add long-term agent memory to Pydantic AI agents

The v4 version of zep-pydantic-ai is not released yet. The current zep-pydantic-ai release uses the v3 API. To use this integration now, follow the v3 version of this page.

Pydantic AI agents using Zep gain long-term memory backed by a temporal knowledge graph. The zep-pydantic-ai package persists conversation turns and adds a model-callable graph-search tool. Its native capability can inject trusted content into the model prompt.

Keep retrieved context out of privileged instructions

Zep context can include content that your users, documents, or tools supplied. A system or developer message gives that content higher instruction priority than ordinary input. Some convenience integrations use system-message injection. Use direct SDK retrieval or an actual retrieval tool call unless all stored content is application-authored and trusted. Follow Memory security best practices for provider-specific placement.

Build an agent with Zep tools

To build an agent that plans its retrieval and uses several Zep tools, read Build an Agent with Zep. The guide shows how to add domain knowledge, design tools, and evaluate the agent.

Core benefits

  • Native Pydantic AI capabilities: zep_capabilities(deps) bundles the current ProcessHistory history-processor hook with a Hooks(after_run=...) hook — not a deprecated kwarg
  • Automatic assistant persistence: The bundled after_run hook persists the assistant’s reply when the run completes, so no manual persistence call is needed
  • Single round-trip: Persists the user turn and retrieves context in one add_messages call
  • Correct under tool calls: Dedupes per run (keyed by RunContext.run_id), so a run that makes tool calls records the turn exactly once
  • Pin-or-expose graph search: A model-callable tool over graph.search_edges — every search parameter is model-exposed by default, or pinned/hidden per deployment
  • Out-of-band provisioning: create_user / create_thread create the Zep resources up front and return the UUIDs that the application stores
  • Graceful degradation: A Zep failure on the turn path is logged but never crashes the agent run

How it works

The integration plugs into Pydantic AI through three components:

  • ZepDeps — a dataclass used as the agent’s deps_type. It carries the Zep client, the user UUID, the thread UUID, the optional graph UUID, and optional context-building configuration. Construct one per conversation and pass it to agent.run(..., deps=deps); the history processor, the after_run hook, and the search tool all reach it through RunContext.deps.
  • zep_capabilities(deps) registers memory for fully trusted, application-authored context. It returns ProcessHistory(zep_history_processor) and a Hooks(after_run=...) hook. The processor persists the latest user message with thread.add_messages(return_context=True) and prepends the context block as a system message. The hook persists the assistant reply when the run completes. Use this path only when all stored content is application-authored and trusted.
  • create_zep_search_tool — a factory returning a model-callable pydantic_ai.Tool over the v4 graph search methods (for example graph.search_edges). The model decides when to search the knowledge graph and, by default, which search parameters to use.

Because ProcessHistory fires once per model request (not once per run), the history processor dedupes per run, keyed by RunContext.run_id: it persists and retrieves on the first model request of a run and replays the cached context on later requests within that same run, so tool-calling runs never create duplicate episodes.

Installation

pip install zep-pydantic-ai

Requires Python 3.11+, pydantic-ai>=1.107,<2, and a Zep Cloud API key. Get your API key from app.getzep.com.

Zep v4 addresses a user, a thread, and a graph by a server-generated UUID. A v4 create call accepts no client-chosen name, and the server rejects a request that sends one. Your application creates the user and the thread one time, reads the UUIDs from the responses, stores them in its own database, and passes them to the integration. The integration does not resolve a name at run time.

Set up your environment variables:

export ZEP_API_KEY="your-zep-api-key"
export OPENAI_API_KEY="your-openai-api-key"

The package targets Zep v4 only. ZepDeps takes user_uuid, thread_uuid, and the optional graph_uuid in place of user_id and thread_id. ensure_user and ensure_thread are replaced by create_user and create_thread, which return the created model; the integration no longer creates a resource on the turn path. The search tool takes graph_uuid in place of graph_id. See the migration guide and the package changelog.

Two changes can require code updates: create_zep_search_tool returns a pydantic_ai.Tool rather than a bare function — code that invoked the return value directly should call tool.function(ctx, query=..., **kwargs) — and the default injected context wording follows the canonical DEFAULT_CONTEXT_TEMPLATE; pass context_template=... on ZepDeps to keep custom wording. See the package changelog for the full list of changes.

Capability usage with trusted context

zep_capabilities(deps) inserts context into a system message. Use this pattern only for fully trusted, application-authored context. For other context, register create_zep_search_tool so retrieval uses an actual tool call. You can also retrieve context with the Zep SDK and place it through your provider’s documented data channel.

When you use the bundled capability, pass the same ZepDeps to each run. Both sides of every turn are persisted automatically. zep_capabilities(deps) closes over one ZepDeps instance, so construct the Agent inside your per-conversation setup rather than sharing it across users.

Python
import asyncio
from pydantic_ai import Agent
from zep_cloud.client import AsyncZep
from zep_pydantic_ai import (
ZepDeps,
create_thread,
create_user,
create_zep_search_tool,
zep_capabilities,
)
zep = AsyncZep(api_key="your-zep-api-key")
async def main() -> None:
# Create the resources one time, then store the UUIDs in your own database.
user = await create_user(zep, first_name="Jane", last_name="Smith")
thread = await create_thread(zep, user_uuid=user.uuid_)
deps = ZepDeps(
client=zep,
user_uuid=user.uuid_,
thread_uuid=thread.uuid_,
graph_uuid=user.graph_uuid,
first_name="Jane",
last_name="Smith",
)
agent = Agent(
"openai:gpt-5.6-terra",
deps_type=ZepDeps,
# Trusted-only path: this capability inserts stored context into a system message.
capabilities=zep_capabilities(deps),
tools=[create_zep_search_tool()],
instructions="You are a helpful assistant with long-term memory.",
)
result = await agent.run("What did I tell you about my project?", deps=deps)
print(result.output)
# The user turn and the assistant's reply are both already persisted.
asyncio.run(main())

Explicit control over persistence

To control exactly when the assistant’s reply reaches Zep, register the history processor directly and call persist_run yourself after the run completes:

Python
from pydantic_ai import Agent
from pydantic_ai.capabilities import ProcessHistory
from zep_pydantic_ai import ZepDeps, persist_run, zep_history_processor
agent = Agent(
"openai:gpt-5.6-terra",
deps_type=ZepDeps,
capabilities=[ProcessHistory(zep_history_processor)],
instructions="You are a helpful assistant with long-term memory.",
)
result = await agent.run("What did I tell you about my project?", deps=deps)
# Persist the assistant's reply (the user turn was already persisted).
await persist_run(deps, result.new_messages())

persist_run sends only assistant text — tool-call and tool-return scaffolding is skipped — so Zep records one clean assistant message per turn. It is not needed when the agent uses zep_capabilities(deps).

Beyond the automatic context injection, create_zep_search_tool() returns a model-callable pydantic_ai.Tool over the v4 graph search methods; pass it directly in tools=[...]. The model decides when to look up specific facts, entities, or prior episodes, and the tool returns a formatted text summary of the matching results. By default it searches the graph that graph_uuid on ZepDeps names; pass graph_uuid=... to the factory to target a shared Context Graph.

Every search parameter (scope, reranker, limit, mmr_lambda, center_node_uuid) is exposed to the model in the tool’s JSON schema by default, with documented defaults. Two constructor arguments override this per deployment: pinned_params fixes a parameter to a constant value and hides it from the schema, and hidden_params hides a parameter without pinning it, so Zep’s server-side default applies:

# Model chooses scope/reranker/limit/mmr_lambda/center_node_uuid freely.
tool = create_zep_search_tool()
# Pin scope to "nodes" and limit to 5 — hidden from the model, always sent.
tool = create_zep_search_tool(pinned_params={"scope": "nodes", "limit": 5})
# Hide mmr_lambda from the schema; Zep applies its own default when omitted.
tool = create_zep_search_tool(hidden_params={"mmr_lambda"})

The scope, reranker, and limit constructor arguments are back-compat aliases that pin (and hide) those parameters; prefer pinned_params in new code. search_filters and bfs_origin_node_uuids are constructor-only — their complex shapes are not exposed to the model.

Memory vs tools

The integration combines two retrieval paths on the same agent:

PathHowWhen it fires
Automatic injectionzep_history_processor (included in zep_capabilities(deps))Before every model request — prepends the context block
On-demand searchcreate_zep_search_tool()When the model chooses to call it for a specific lookup

Injection grounds each turn with cross-session context; the search tool lets the model actively dig for specific details.

Provisioning

create_user and create_thread create the Zep user and the thread out of band, before the first turn. Each helper returns the v4 model, so the application reads the UUIDs from the response and stores them:

from zep_pydantic_ai import create_thread, create_user
user = await create_user(
zep,
first_name="Jane",
last_name="Smith",
)
thread = await create_thread(zep, user_uuid=user.uuid_)
# Store these three values in your own database.
user_uuid = user.uuid_
graph_uuid = user.graph_uuid
thread_uuid = thread.uuid_

A create call sends no client-chosen name, because the server generates the UUID. After the user exists, you can configure per-user resources against user.graph_uuid — a custom ontology, custom extraction instructions, or user summary instructions; see customizing graph structure for the available options.

The integration does not create a user or a thread on the turn path, and it does not look a name up at run time. It expects the stored UUIDs on ZepDeps. A Zep failure on the turn path is logged and degrades to no memory rather than breaking the run.

Custom context building

Set context_builder on ZepDeps to replace the default context retrieval with custom logic — for example, searching a different graph, applying filters, or combining multiple sources:

from zep_pydantic_ai import ContextInput, ZepDeps
async def my_builder(ctx: ContextInput) -> str | None:
if ctx.graph_uuid is None:
return None
pager = await ctx.zep.graph.search_edges(
ctx.graph_uuid,
query=ctx.user_message,
)
edges = pager.items or []
if not edges:
return None
return "\n".join(edge.fact for edge in edges if edge.fact)
deps = ZepDeps(
client=zep,
user_uuid=user_uuid,
thread_uuid=thread_uuid,
graph_uuid=graph_uuid,
context_builder=my_builder,
)

ContextInput is a frozen dataclass bundling zep (the AsyncZep client), user_uuid, thread_uuid, graph_uuid, user_message, and run_context (the Pydantic AI RunContext for the turn). Returning None skips injection for that turn.

When context_builder is set, message persistence (add_messages without return_context) and the builder run concurrently, with per-side failure isolation:

  • If the builder raises, a warning is logged and context injection is skipped for that turn — persistence still completes.
  • If persistence raises, a warning is logged and the turn is not marked as persisted (so it retries on the next model request) — a successful builder result is still injected.

Context template

context_template on ZepDeps controls how retrieved context is wrapped before injection. It must contain a literal {context} placeholder. Plain string replacement prevents format-string interpretation of {, }, or %. It does not prevent the model from following instructions in the retrieved content.

deps = ZepDeps(
client=zep,
user_uuid=user_uuid,
thread_uuid=thread_uuid,
context_template="Relevant memory:\n{context}",
)

The default is DEFAULT_CONTEXT_TEMPLATE, an explicit <ZEP_CONTEXT>...</ZEP_CONTEXT> block with canonical wording shared across Zep’s framework integrations.

Configuration options

ZepDeps

FieldTypeRequiredDefaultDescription
clientAsyncZepYes—Initialized Zep async client (caller owns its lifecycle)
user_uuidstrYes—UUID of the Zep user (one user graph)
thread_uuidstrYes—UUID of the Zep thread for the conversation
graph_uuidstr | NoneNoNoneUUID of the user’s graph, from user.graph_uuid; the search tool and a custom context builder use it
first_namestrNoNoneUser first name (recommended; anchors the user node)
last_namestrNoNoneUser last name
user_namestrNoNoneDisplay name for persisted user messages (defaults to first + last)
assistant_namestrNo"Assistant"Display name for persisted assistant messages
ignore_roleslist[str]NoNoneRoles to exclude from graph ingestion
context_builderContextBuilder | NoneNoNoneCustom async context-retrieval callable (see custom context building)
context_templatestrNoDEFAULT_CONTEXT_TEMPLATETemplate wrapping injected context; must contain a literal {context} placeholder

create_zep_search_tool

Constructor arguments (returns a pydantic_ai.Tool[ZepDeps]):

ParameterTypeDefaultDescription
graph_uuidstr | NoneNoneUUID of a shared Context Graph to search; when unset, the tool uses graph_uuid on ZepDeps
pinned_paramsdict[str, Any] | NoneNoneFix a search parameter to a value; hidden from the model schema
hidden_paramsset[str] | NoneNoneHide a search parameter from the schema without pinning (Zep’s server-side default applies)
search_filtersdict[str, Any] | NoneNoneConstructor-only Zep search filters (node_labels, edge_types, etc.)
bfs_origin_node_uuidslist[str] | NoneNoneConstructor-only node UUIDs for BFS seeding
namestr"zep_search"Tool name exposed to the model
descriptionstrBuilt-in descriptionTool description exposed to the model
scopeScope | NoneNoneBack-compat alias for pinned_params={"scope": scope}
rerankerReranker | NoneNoneBack-compat alias for pinned_params={"reranker": reranker}
limitint | NoneNoneBack-compat alias for pinned_params={"limit": limit}

Model-exposed search parameters (when not pinned or hidden), with their defaults:

ParameterTypeDefaultDescription
scope"edges" | "nodes" | "episodes" | "observations" | "thread_summaries" | "auto""edges"What to search
reranker"rrf" | "mmr" | "node_distance" | "episode_mentions" | "cross_encoder""rrf"Result ordering (ignored for scope="auto")
limitint10Maximum results (clamped to Zep’s ceiling of 50)
mmr_lambdafloat—Diversity/relevance balance; only used when reranker="mmr"
center_node_uuidstr—Center node for reranker="node_distance"

Best practices

  • Construct one ZepDeps per conversation and reuse a single AsyncZep client across runs
  • Pass real names so Zep can anchor the user’s identity node in the graph
  • Use create_zep_search_tool for untrusted context. Use zep_capabilities(deps) only when all stored context is application-authored and trusted.
  • Create the user and the thread in your onboarding flow with create_user / create_thread, and store the returned UUIDs in your own database
  • Allow time for indexing — Zep extracts knowledge asynchronously, so facts from a turn are not instantly searchable

Next steps