> This page is for version v4 (default).
> For other versions, use one of these documentation indexes:
> - v4 (default): https://docs-beta.getzep.com/v4/llms.txt
> - v3: https://docs-beta.getzep.com/v3/llms.txt
> - v2: https://docs-beta.getzep.com/v2/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs-beta.getzep.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-beta.getzep.com/_mcp/server.

# Pydantic AI integration

> **Note**
>
> The v4 version of `zep-pydantic-ai` is not released yet. The current `zep-pydantic-ai` release uses the v3 API. To use this integration now, follow the [v3 version of this page](/v3/pydantic-ai-memory).

[Pydantic AI](https://ai.pydantic.dev) agents using Zep gain long-term memory backed by a temporal knowledge graph. The `zep-pydantic-ai` package persists conversation turns and adds a model-callable graph-search tool. Its native capability can inject trusted content into the model prompt.

> **Keep retrieved context out of privileged instructions**
>
> Zep context can include content that your users, documents, or tools supplied. A system or developer message gives that content higher instruction priority than ordinary input. Some convenience integrations use system-message injection. Use direct SDK retrieval or an actual retrieval tool call unless all stored content is application-authored and trusted. Follow [Memory security best practices](/memory-security) for provider-specific placement.

> **Build an agent with Zep tools**
>
> To build an agent that plans its retrieval and uses several Zep tools, read [Build an Agent with Zep](/build-an-agent-with-zep). The guide shows how to add domain knowledge, design tools, and evaluate the agent.

## Core benefits

* **Native Pydantic AI capabilities**: `zep_capabilities(deps)` bundles the current `ProcessHistory` history-processor hook with a `Hooks(after_run=...)` hook — not a deprecated kwarg
* **Automatic assistant persistence**: The bundled `after_run` hook persists the assistant's reply when the run completes, so no manual persistence call is needed
* **Single round-trip**: Persists the user turn and retrieves context in one `add_messages` call
* **Correct under tool calls**: Dedupes per run (keyed by `RunContext.run_id`), so a run that makes tool calls records the turn exactly once
* **Pin-or-expose graph search**: A model-callable tool over `graph.search_edges` — every search parameter is model-exposed by default, or pinned/hidden per deployment
* **Out-of-band provisioning**: `create_user` / `create_thread` create the Zep resources up front and return the UUIDs that the application stores
* **Graceful degradation**: A Zep failure on the turn path is logged but never crashes the agent run

## How it works

The integration plugs into Pydantic AI through three components:

* **`ZepDeps`** — a dataclass used as the agent's `deps_type`. It carries the Zep client, the user UUID, the thread UUID, the optional graph UUID, and optional context-building configuration. Construct one per conversation and pass it to `agent.run(..., deps=deps)`; the history processor, the `after_run` hook, and the search tool all reach it through `RunContext.deps`.
* **`zep_capabilities(deps)`** registers memory for fully trusted, application-authored context. It returns `ProcessHistory(zep_history_processor)` and a `Hooks(after_run=...)` hook. The processor persists the latest user message with `thread.add_messages(return_context=True)` and prepends the context block as a system message. The hook persists the assistant reply when the run completes. Use this path only when all stored content is application-authored and trusted.
* **`create_zep_search_tool`** — a factory returning a model-callable `pydantic_ai.Tool` over the v4 graph search methods (for example `graph.search_edges`). The model decides when to search the knowledge graph and, by default, which search parameters to use.

Because `ProcessHistory` fires once per model request (not once per run), the history processor dedupes per run, keyed by `RunContext.run_id`: it persists and retrieves on the first model request of a run and replays the cached context on later requests within that same run, so tool-calling runs never create duplicate episodes.

## Installation

```bash
pip install zep-pydantic-ai
```

> **Info**
>
> Requires Python 3.11+, `pydantic-ai>=1.107,<2`, and a Zep Cloud API key. Get your API key from [app.getzep.com](https://app.getzep.com).

Zep v4 addresses a user, a thread, and a graph by a server-generated UUID. A v4 create call accepts no client-chosen name, and the server rejects a request that sends one. Your application creates the user and the thread one time, reads the UUIDs from the responses, stores them in its own database, and passes them to the integration. The integration does not resolve a name at run time.

Set up your environment variables:

```bash
export ZEP_API_KEY="your-zep-api-key"
export OPENAI_API_KEY="your-openai-api-key"
```

#### Upgrading to the Zep v4 SDK

The package targets Zep v4 only. `ZepDeps` takes `user_uuid`, `thread_uuid`, and the optional `graph_uuid` in place of `user_id` and `thread_id`. `ensure_user` and `ensure_thread` are replaced by `create_user` and `create_thread`, which return the created model; the integration no longer creates a resource on the turn path. The search tool takes `graph_uuid` in place of `graph_id`. See the [migration guide](/migrating-from-v3) and the [package changelog](https://github.com/getzep/zep/blob/main/integrations/pydantic-ai/python/CHANGELOG.md).

#### Upgrading from zep-pydantic-ai 0.1.x

Two changes can require code updates: `create_zep_search_tool` returns a `pydantic_ai.Tool` rather than a bare function — code that invoked the return value directly should call `tool.function(ctx, query=..., **kwargs)` — and the default injected context wording follows the canonical `DEFAULT_CONTEXT_TEMPLATE`; pass `context_template=...` on `ZepDeps` to keep custom wording. See the [package changelog](https://github.com/getzep/zep/blob/main/integrations/pydantic-ai/python/CHANGELOG.md) for the full list of changes.

## Capability usage with trusted context

`zep_capabilities(deps)` inserts context into a system message. Use this pattern only for fully trusted, application-authored context. For other context, register `create_zep_search_tool` so retrieval uses an actual tool call. You can also retrieve context with the Zep SDK and place it through your provider's documented data channel.

When you use the bundled capability, pass the same `ZepDeps` to each run. Both sides of every turn are persisted automatically. `zep_capabilities(deps)` closes over one `ZepDeps` instance, so construct the `Agent` inside your per-conversation setup rather than sharing it across users.

**`Python`**

```python Python
import asyncio
from pydantic_ai import Agent
from zep_cloud.client import AsyncZep
from zep_pydantic_ai import (
    ZepDeps,
    create_thread,
    create_user,
    create_zep_search_tool,
    zep_capabilities,
)

zep = AsyncZep(api_key="your-zep-api-key")

async def main() -> None:
    # Create the resources one time, then store the UUIDs in your own database.
    user = await create_user(zep, first_name="Jane", last_name="Smith")
    thread = await create_thread(zep, user_uuid=user.uuid_)

    deps = ZepDeps(
        client=zep,
        user_uuid=user.uuid_,
        thread_uuid=thread.uuid_,
        graph_uuid=user.graph_uuid,
        first_name="Jane",
        last_name="Smith",
    )

    agent = Agent(
        "openai:gpt-5.6-terra",
        deps_type=ZepDeps,
        # Trusted-only path: this capability inserts stored context into a system message.
        capabilities=zep_capabilities(deps),
        tools=[create_zep_search_tool()],
        instructions="You are a helpful assistant with long-term memory.",
    )

    result = await agent.run("What did I tell you about my project?", deps=deps)
    print(result.output)
    # The user turn and the assistant's reply are both already persisted.

asyncio.run(main())
```

### Explicit control over persistence

To control exactly when the assistant's reply reaches Zep, register the history processor directly and call `persist_run` yourself after the run completes:

**`Python`**

```python Python
from pydantic_ai import Agent
from pydantic_ai.capabilities import ProcessHistory
from zep_pydantic_ai import ZepDeps, persist_run, zep_history_processor

agent = Agent(
    "openai:gpt-5.6-terra",
    deps_type=ZepDeps,
    capabilities=[ProcessHistory(zep_history_processor)],
    instructions="You are a helpful assistant with long-term memory.",
)

result = await agent.run("What did I tell you about my project?", deps=deps)
# Persist the assistant's reply (the user turn was already persisted).
await persist_run(deps, result.new_messages())
```

`persist_run` sends only assistant text — tool-call and tool-return scaffolding is skipped — so Zep records one clean assistant message per turn. It is not needed when the agent uses `zep_capabilities(deps)`.

## On-demand graph search

Beyond the automatic context injection, `create_zep_search_tool()` returns a model-callable `pydantic_ai.Tool` over the v4 graph search methods; pass it directly in `tools=[...]`. The model decides when to look up specific facts, entities, or prior episodes, and the tool returns a formatted text summary of the matching results. By default it searches the graph that `graph_uuid` on `ZepDeps` names; pass `graph_uuid=...` to the factory to target a shared Context Graph.

Every search parameter (`scope`, `reranker`, `limit`, `mmr_lambda`, `center_node_uuid`) is exposed to the model in the tool's JSON schema by default, with documented defaults. Two constructor arguments override this per deployment: `pinned_params` fixes a parameter to a constant value and hides it from the schema, and `hidden_params` hides a parameter without pinning it, so Zep's server-side default applies:

```python
# Model chooses scope/reranker/limit/mmr_lambda/center_node_uuid freely.
tool = create_zep_search_tool()

# Pin scope to "nodes" and limit to 5 — hidden from the model, always sent.
tool = create_zep_search_tool(pinned_params={"scope": "nodes", "limit": 5})

# Hide mmr_lambda from the schema; Zep applies its own default when omitted.
tool = create_zep_search_tool(hidden_params={"mmr_lambda"})
```

The `scope`, `reranker`, and `limit` constructor arguments are back-compat aliases that pin (and hide) those parameters; prefer `pinned_params` in new code. `search_filters` and `bfs_origin_node_uuids` are constructor-only — their complex shapes are not exposed to the model.

## Memory vs tools

The integration combines two retrieval paths on the same agent:

| Path                    | How                                                            | When it fires                                           |
| ----------------------- | -------------------------------------------------------------- | ------------------------------------------------------- |
| **Automatic injection** | `zep_history_processor` (included in `zep_capabilities(deps)`) | Before every model request — prepends the context block |
| **On-demand search**    | `create_zep_search_tool()`                                     | When the model chooses to call it for a specific lookup |

Injection grounds each turn with cross-session context; the search tool lets the model actively dig for specific details.

## Provisioning

`create_user` and `create_thread` create the Zep user and the thread out of band, before the first turn. Each helper returns the v4 model, so the application reads the UUIDs from the response and stores them:

```python
from zep_pydantic_ai import create_thread, create_user

user = await create_user(
    zep,
    first_name="Jane",
    last_name="Smith",
    email="jane@example.com",
)
thread = await create_thread(zep, user_uuid=user.uuid_)

# Store these three values in your own database.
user_uuid = user.uuid_
graph_uuid = user.graph_uuid
thread_uuid = thread.uuid_
```

A create call sends no client-chosen name, because the server generates the UUID. After the user exists, you can configure per-user resources against `user.graph_uuid` — a custom ontology, custom extraction instructions, or user summary instructions; see [customizing graph structure](/customizing-graph-structure) for the available options.

The integration does not create a user or a thread on the turn path, and it does not look a name up at run time. It expects the stored UUIDs on `ZepDeps`. A Zep failure on the turn path is logged and degrades to no memory rather than breaking the run.

## Custom context building

Set `context_builder` on `ZepDeps` to replace the default context retrieval with custom logic — for example, searching a different graph, applying filters, or combining multiple sources:

```python
from zep_pydantic_ai import ContextInput, ZepDeps

async def my_builder(ctx: ContextInput) -> str | None:
    if ctx.graph_uuid is None:
        return None
    pager = await ctx.zep.graph.search_edges(
        ctx.graph_uuid,
        query=ctx.user_message,
    )
    edges = pager.items or []
    if not edges:
        return None
    return "\n".join(edge.fact for edge in edges if edge.fact)

deps = ZepDeps(
    client=zep,
    user_uuid=user_uuid,
    thread_uuid=thread_uuid,
    graph_uuid=graph_uuid,
    context_builder=my_builder,
)
```

`ContextInput` is a frozen dataclass bundling `zep` (the `AsyncZep` client), `user_uuid`, `thread_uuid`, `graph_uuid`, `user_message`, and `run_context` (the Pydantic AI `RunContext` for the turn). Returning `None` skips injection for that turn.

When `context_builder` is set, message persistence (`add_messages` without `return_context`) and the builder run concurrently, with per-side failure isolation:

* If the builder raises, a warning is logged and context injection is skipped for that turn — persistence still completes.
* If persistence raises, a warning is logged and the turn is not marked as persisted (so it retries on the next model request) — a successful builder result is still injected.

## Context template

`context_template` on `ZepDeps` controls how retrieved context is wrapped before injection. It must contain a literal `{context}` placeholder. Plain string replacement prevents format-string interpretation of `{`, `}`, or `%`. It does not prevent the model from following instructions in the retrieved content.

```python
deps = ZepDeps(
    client=zep,
    user_uuid=user_uuid,
    thread_uuid=thread_uuid,
    context_template="Relevant memory:\n{context}",
)
```

The default is `DEFAULT_CONTEXT_TEMPLATE`, an explicit `<ZEP_CONTEXT>...</ZEP_CONTEXT>` block with canonical wording shared across Zep's framework integrations.

## Configuration options

### ZepDeps

| Field              | Type                     | Required | Default                    | Description                                                                                           |
| ------------------ | ------------------------ | -------- | -------------------------- | ----------------------------------------------------------------------------------------------------- |
| `client`           | `AsyncZep`               | Yes      | —                          | Initialized Zep async client (caller owns its lifecycle)                                              |
| `user_uuid`        | `str`                    | Yes      | —                          | UUID of the Zep user (one user graph)                                                                 |
| `thread_uuid`      | `str`                    | Yes      | —                          | UUID of the Zep thread for the conversation                                                           |
| `graph_uuid`       | `str \| None`            | No       | `None`                     | UUID of the user's graph, from `user.graph_uuid`; the search tool and a custom context builder use it |
| `first_name`       | `str`                    | No       | `None`                     | User first name (recommended; anchors the user node)                                                  |
| `last_name`        | `str`                    | No       | `None`                     | User last name                                                                                        |
| `user_name`        | `str`                    | No       | `None`                     | Display name for persisted user messages (defaults to first + last)                                   |
| `assistant_name`   | `str`                    | No       | `"Assistant"`              | Display name for persisted assistant messages                                                         |
| `ignore_roles`     | `list[str]`              | No       | `None`                     | Roles to exclude from graph ingestion                                                                 |
| `context_builder`  | `ContextBuilder \| None` | No       | `None`                     | Custom async context-retrieval callable (see [custom context building](#custom-context-building))     |
| `context_template` | `str`                    | No       | `DEFAULT_CONTEXT_TEMPLATE` | Template wrapping injected context; must contain a literal `{context}` placeholder                    |

### create\_zep\_search\_tool

Constructor arguments (returns a `pydantic_ai.Tool[ZepDeps]`):

| Parameter               | Type                     | Default              | Description                                                                                   |
| ----------------------- | ------------------------ | -------------------- | --------------------------------------------------------------------------------------------- |
| `graph_uuid`            | `str \| None`            | `None`               | UUID of a shared Context Graph to search; when unset, the tool uses `graph_uuid` on `ZepDeps` |
| `pinned_params`         | `dict[str, Any] \| None` | `None`               | Fix a search parameter to a value; hidden from the model schema                               |
| `hidden_params`         | `set[str] \| None`       | `None`               | Hide a search parameter from the schema without pinning (Zep's server-side default applies)   |
| `search_filters`        | `dict[str, Any] \| None` | `None`               | Constructor-only Zep search filters (`node_labels`, `edge_types`, etc.)                       |
| `bfs_origin_node_uuids` | `list[str] \| None`      | `None`               | Constructor-only node UUIDs for BFS seeding                                                   |
| `name`                  | `str`                    | `"zep_search"`       | Tool name exposed to the model                                                                |
| `description`           | `str`                    | Built-in description | Tool description exposed to the model                                                         |
| `scope`                 | `Scope \| None`          | `None`               | Back-compat alias for `pinned_params={"scope": scope}`                                        |
| `reranker`              | `Reranker \| None`       | `None`               | Back-compat alias for `pinned_params={"reranker": reranker}`                                  |
| `limit`                 | `int \| None`            | `None`               | Back-compat alias for `pinned_params={"limit": limit}`                                        |

Model-exposed search parameters (when not pinned or hidden), with their defaults:

| Parameter          | Type                                                                                 | Default   | Description                                                  |
| ------------------ | ------------------------------------------------------------------------------------ | --------- | ------------------------------------------------------------ |
| `scope`            | `"edges" \| "nodes" \| "episodes" \| "observations" \| "thread_summaries" \| "auto"` | `"edges"` | What to search                                               |
| `reranker`         | `"rrf" \| "mmr" \| "node_distance" \| "episode_mentions" \| "cross_encoder"`         | `"rrf"`   | Result ordering (ignored for `scope="auto"`)                 |
| `limit`            | `int`                                                                                | `10`      | Maximum results (clamped to Zep's ceiling of 50)             |
| `mmr_lambda`       | `float`                                                                              | —         | Diversity/relevance balance; only used when `reranker="mmr"` |
| `center_node_uuid` | `str`                                                                                | —         | Center node for `reranker="node_distance"`                   |

## Best practices

* **Construct one `ZepDeps` per conversation** and reuse a single `AsyncZep` client across runs
* **Pass real names** so Zep can anchor the user's identity node in the graph
* **Use `create_zep_search_tool` for untrusted context.** Use `zep_capabilities(deps)` only when all stored context is application-authored and trusted.
* **Create the user and the thread in your onboarding flow** with `create_user` / `create_thread`, and store the returned UUIDs in your own database
* **Allow time for indexing** — Zep extracts knowledge asynchronously, so facts from a turn are not instantly searchable

## Next steps

* Explore [customizing graph structure](/customizing-graph-structure) for advanced knowledge organization
* Learn about [searching the graph](/searching-the-graph) and how to tune search
* See [code examples](https://github.com/getzep/zep/tree/main/integrations/pydantic-ai/python/examples) for additional patterns