> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs-beta.getzep.com/v4/performance/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-beta.getzep.com/_mcp/server. # Performance guide > Reduce application-side latency when you use Zep This guide describes application-side methods that reduce latency when you use Zep. Retrieval latency depends on the graph size, query, filters, result limits, deployment, and application network path. Measure end-to-end latency with representative data and requests. The methods below reduce avoidable application-side latency. ## Reuse the Zep SDK Client The Zep SDK client maintains an HTTP connection pool. Reuse the client to avoid creating a connection for each request: * Create a single client instance and reuse it across your application * Avoid creating new client instances for each request or function * Consider implementing a client singleton pattern in your application * For serverless environments, initialize the client outside the handler function ## Optimizing Context Operations The `thread.add_messages` and `thread.get_context` methods are optimized for conversational messages and low-latency retrieval. For optimal performance: * Use `graph.episode.add` for larger documents, tool outputs, or business data (up to 10,000 characters per call) * [Chunk large documents](/chunking-large-documents) before adding them to the graph * Remove unnecessary metadata or content before persistence * For bulk document ingestion, use [`zep-ingest`](/zep-ingest) or the [Batch API](/adding-batch-data) rather than parallel `graph.episode.add` calls ```python # Recommended for conversations zep_client.thread.add_messages( zep_thread_uuid, messages=[ { "role": "user", "name": "Alice", "content": "What's the weather like today?" } ] ) # Recommended for large documents zep_client.graph.episode.add( graph_uuid, # The graph_uuid of the user, or of a shared graph data=document_content, # Your chunked document content type="text" # Can be "text", "message", or "json" ) ``` ### Get the Context Block sooner You can request the Context Block directly in the response to the `thread.add_messages()` call. This optimization eliminates the need for a separate `thread.get_context()` call. Read more about our [Context Block](/retrieving-context#zeps-context-block). In this scenario you can pass in the `return_context=True` flag to the `thread.add_messages()` method. Zep will perform a user graph search right after persisting the data and return the context relevant to the recently added messages. **`Python`** ```python Python memory_response = await zep_client.thread.add_messages( zep_thread_uuid, messages=messages, return_context=True ) context = memory_response.context ``` **`TypeScript`** ```typescript TypeScript const memoryResponse = await zepClient.thread.addMessages(zepThreadUuid, { messages: messages, returnContext: true }); const context = memoryResponse.context; ``` **`Go`** ```go Go memoryResponse, err := zepClient.Thread.AddMessages( context.TODO(), zepThreadUUID, &zep.AddMessagesRequest{ Messages: messages, ReturnContext: zep.Bool(true), }, ) if err != nil { // handle error } contextBlock := memoryResponse.Context ``` > **Tip** > > Read more in the [Thread SDK Reference](/sdk-reference/thread/add-messages) ### Searching the Graph Sooner Use [graph search](/searching-the-graph) and [Advanced Context Block construction](/advanced-context-block-construction) when you need custom parameters. You can search the graph while another request adds data. ```python import asyncio from zep_cloud.client import AsyncZep client = AsyncZep(api_key="your-zep-api-key") async def add_and_retrieve_from_zep(zep_thread_uuid, user_graph_uuid, messages): # Concatenate message content to create query string query = " ".join([msg.content for msg in messages]) # Execute all operations concurrently add_result, edges_result, nodes_result = await asyncio.gather( client.thread.add_messages( zep_thread_uuid, messages=messages ), client.graph.search_edges( user_graph_uuid, query=query ), client.graph.search_nodes( user_graph_uuid, query=query ) ) return add_result, edges_result, nodes_result ``` Construct a custom Context Block from the search results. [Advanced Context Block construction](/advanced-context-block-construction) provides examples. ## Optimizing Search Queries Zep uses hybrid retrieval combining semantic (vector) similarity, BM25 full-text search, and graph traversal in a single ranked result. For optimal performance: * Keep your queries concise. * Longer queries may not improve search quality and will increase latency * Consider breaking down complex searches into smaller, focused queries * Use specific, contextual queries rather than generic ones Best practices for search: * Keep search queries concise and specific * Structure queries to target relevant information * Use natural language queries for better semantic matching * Consider the scope of your search (graphs versus user graphs) ```python # Recommended - concise query results = await zep_client.graph.search_edges( graph_uuid, # The graph_uuid of the user, or of a shared graph query="project requirements discussion" ) # Not accepted - query exceeds the public limit results = await zep_client.graph.search_edges( graph_uuid, query="very long text with multiple paragraphs..." # A long query increases latency ) ``` ## Warming the User Cache Use `graph.warm` with the `graph_uuid` of the user before a likely retrieval to prepare the user's graph data. For example, call it when a user logs in to your service or opens your application. **`Python`** ```python Python # Warm the user's cache when they log in client.graph.warm(user_graph_uuid) ``` **`TypeScript`** ```typescript TypeScript // Warm the user's cache when they log in await client.graph.warm(userGraphUuid); ``` **`Go`** ```go Go // Warm the user's cache when they log in _, err := client.Graph.Warm(context.TODO(), userGraphUUID) if err != nil { log.Printf("Error warming user cache: %v", err) } ``` > **Tip** > > Read more in the [Graph SDK Reference](/sdk-reference/graph/warm) ## Summary * Reuse Zep SDK client instances to optimize connection management * Use appropriate methods for different types of content (`thread.add_messages` for conversations, `graph.episode.add` for large documents) * Keep search queries focused and under the token limit for optimal performance * Warm the user cache when users log in or open your app for faster retrieval > Reduce avoidable application latency when you ingest and retrieve context