> This page is for version v4 (default).
> For other versions, use one of these documentation indexes:
> - v4 (default): https://docs-beta.getzep.com/v4/llms.txt
> - v3: https://docs-beta.getzep.com/v3/llms.txt
> - v2: https://docs-beta.getzep.com/v2/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs-beta.getzep.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-beta.getzep.com/_mcp/server.

# Performance guide

> Reduce application-side latency when you use Zep

This guide describes application-side methods that reduce latency when you use Zep.

Retrieval latency depends on the graph size, query, filters, result limits, deployment, and application network path. Measure end-to-end latency with representative data and requests. The methods below reduce avoidable application-side latency.

## Reuse the Zep SDK Client

The Zep SDK client maintains an HTTP connection pool. Reuse the client to avoid creating a connection for each request:

* Create a single client instance and reuse it across your application
* Avoid creating new client instances for each request or function
* Consider implementing a client singleton pattern in your application
* For serverless environments, initialize the client outside the handler function

## Optimizing Context Operations

The `thread.add_messages` and `thread.get_context` methods are optimized for conversational messages and low-latency retrieval. For optimal performance:

* Use `graph.episode.add` for larger documents, tool outputs, or business data (up to 10,000 characters per call)
* [Chunk large documents](/chunking-large-documents) before adding them to the graph
* Remove unnecessary metadata or content before persistence
* For bulk document ingestion, use [`zep-ingest`](/zep-ingest) or the [Batch API](/adding-batch-data) rather than parallel `graph.episode.add` calls

```python
# Recommended for conversations
zep_client.thread.add_messages(
    zep_thread_uuid,
    messages=[
        {
            "role": "user",
            "name": "Alice",
            "content": "What's the weather like today?"
        }
    ]
)

# Recommended for large documents
zep_client.graph.episode.add(
    graph_uuid,             # The graph_uuid of the user, or of a shared graph
    data=document_content,  # Your chunked document content
    type="text"             # Can be "text", "message", or "json"
)
```

### Get the Context Block sooner

You can request the Context Block directly in the response to the `thread.add_messages()` call.
This optimization eliminates the need for a separate `thread.get_context()` call.
Read more about our [Context Block](/retrieving-context#zeps-context-block).

In this scenario you can pass in the `return_context=True` flag to the `thread.add_messages()` method.
Zep will perform a user graph search right after persisting the data and return the context relevant to the recently added messages.

**`Python`**

```python Python
memory_response = await zep_client.thread.add_messages(
    zep_thread_uuid,
    messages=messages,
    return_context=True
)

context = memory_response.context
```

**`TypeScript`**

```typescript TypeScript
const memoryResponse = await zepClient.thread.addMessages(zepThreadUuid, {
    messages: messages,
    returnContext: true
});

const context = memoryResponse.context;
```

**`Go`**

```go Go
memoryResponse, err := zepClient.Thread.AddMessages(
    context.TODO(),
    zepThreadUUID,
    &zep.AddMessagesRequest{
        Messages:      messages,
        ReturnContext: zep.Bool(true),
    },
)
if err != nil {
    // handle error
}
contextBlock := memoryResponse.Context
```

> **Tip**
>
> Read more in the [Thread SDK Reference](/sdk-reference/thread/add-messages)

### Searching the Graph Sooner

Use [graph search](/searching-the-graph) and [Advanced Context Block construction](/advanced-context-block-construction) when you need custom parameters. You can search the graph while another request adds data.

```python
import asyncio
from zep_cloud.client import AsyncZep

client = AsyncZep(api_key="your-zep-api-key")

async def add_and_retrieve_from_zep(zep_thread_uuid, user_graph_uuid, messages):
    # Concatenate message content to create query string
    query = " ".join([msg.content for msg in messages])

    # Execute all operations concurrently
    add_result, edges_result, nodes_result = await asyncio.gather(
        client.thread.add_messages(
            zep_thread_uuid,
            messages=messages
        ),
        client.graph.search_edges(
            user_graph_uuid,
            query=query
        ),
        client.graph.search_nodes(
            user_graph_uuid,
            query=query
        )
    )

    return add_result, edges_result, nodes_result
```

Construct a custom Context Block from the search results. [Advanced Context Block construction](/advanced-context-block-construction) provides examples.

## Optimizing Search Queries

Zep uses hybrid retrieval combining semantic (vector) similarity, BM25 full-text search, and graph traversal in a single ranked result. For optimal performance:

* Keep your queries concise.
* Longer queries may not improve search quality and will increase latency
* Consider breaking down complex searches into smaller, focused queries
* Use specific, contextual queries rather than generic ones

Best practices for search:

* Keep search queries concise and specific
* Structure queries to target relevant information
* Use natural language queries for better semantic matching
* Consider the scope of your search (graphs versus user graphs)

```python
# Recommended - concise query
results = await zep_client.graph.search_edges(
    graph_uuid,  # The graph_uuid of the user, or of a shared graph
    query="project requirements discussion"
)

# Not accepted - query exceeds the public limit
results = await zep_client.graph.search_edges(
    graph_uuid,
    query="very long text with multiple paragraphs..."  # A long query increases latency
)
```

## Warming the User Cache

Use `graph.warm` with the `graph_uuid` of the user before a likely retrieval to prepare the user's graph data. For example, call it when a user logs in to your service or opens your application.

**`Python`**

```python Python
# Warm the user's cache when they log in
client.graph.warm(user_graph_uuid)
```

**`TypeScript`**

```typescript TypeScript
// Warm the user's cache when they log in
await client.graph.warm(userGraphUuid);
```

**`Go`**

```go Go
// Warm the user's cache when they log in
_, err := client.Graph.Warm(context.TODO(), userGraphUUID)
if err != nil {
    log.Printf("Error warming user cache: %v", err)
}
```

> **Tip**
>
> Read more in the [Graph SDK Reference](/sdk-reference/graph/warm)

## Summary

* Reuse Zep SDK client instances to optimize connection management
* Use appropriate methods for different types of content (`thread.add_messages` for conversations, `graph.episode.add` for large documents)
* Keep search queries focused and under the token limit for optimal performance
* Warm the user cache when users log in or open your app for faster retrieval