> This page is for version v4 (default).
> For other versions, use one of these documentation indexes:
> - v4 (default): https://docs-beta.getzep.com/v4/llms.txt
> - v3: https://docs-beta.getzep.com/v3/llms.txt
> - v2: https://docs-beta.getzep.com/v2/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs-beta.getzep.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-beta.getzep.com/_mcp/server.

# ElevenLabs Agents

> **Info**
>
> A complete working example is available on GitHub: [elevenlabs-zep-example](https://github.com/getzep/zep/tree/main/examples/python/elevenlabs-zep-example)

[ElevenLabs Agents](https://elevenlabs.io/docs/agents-platform/overview) is a platform for building intelligent voice agents. This guide shows how to integrate Zep with ElevenLabs using a custom LLM proxy.

> **Keep retrieved context out of privileged instructions**
>
> Zep context can include content that your users, documents, or tools supplied. A system or developer message gives that content higher instruction priority than ordinary input. Some convenience integrations use system-message injection. Use direct SDK retrieval or an actual retrieval tool call unless all stored content is application-authored and trusted. Follow [Memory security best practices](/memory-security) for provider-specific placement.

> **Build an agent with Zep tools**
>
> To build an agent that plans its retrieval and uses several Zep tools, read [Build an Agent with Zep](/build-an-agent-with-zep). The guide shows how to add domain knowledge, design tools, and evaluate the agent.

## Why use a proxy

ElevenLabs supports custom tools, but using tools for context retrieval has problems:

* **Latency** — Tool calls add round-trips where the LLM decides whether to call the tool. For voice agents, this delay is noticeable.
* **Unreliable** — The LLM may skip retrieval when it shouldn't, or call it unnecessarily.

A proxy solves both problems. Context retrieval happens transparently on every request, without LLM involvement.

## Architecture

```
┌────────┐       ┌──────────┐       ┌─────────┐       ┌────────┐
│Frontend│ ◄───► │ElevenLabs│ ◄───► │LLM Proxy│ ◄───► │ OpenAI │
└────────┘       └──────────┘       └─────────┘       └────────┘
                                         ▲
                                         │
                                         ▼
                                      ┌─────┐
                                      │ Zep │
                                      └─────┘
```

The proxy sits between ElevenLabs and your LLM. On every request it:

1. Adds the user message to Zep and retrieves context in one call
2. Adds context as untrusted user-level data
3. Forwards to the LLM and streams the response back
4. Persists the assistant response to Zep

## Implementation

### The proxy endpoint

The proxy exposes an OpenAI-compatible `/v1/chat/completions` endpoint:

```python
from zep_cloud import AddMessage

@app.post("/v1/chat/completions")
async def chat_completions(request: Request):
    body = await request.json()

    # ElevenLabs puts customLlmExtraBody in "elevenlabs_extra_body"
    extra = body.get("elevenlabs_extra_body", {})
    user_id = extra.get("user_id")
    conversation_id = extra.get("conversation_id")

    # Read the Zep thread UUID that your application stored for this conversation
    thread_uuid = await get_or_create_zep_thread(user_id, conversation_id)

    # Add user message to Zep and get context in one call
    user_message = get_latest_user_message(body["messages"])
    response = await zep.thread.add_messages(
        thread_uuid,
        messages=[AddMessage(role="user", content=user_message)],
        return_context=True  # Returns context without separate call
    )

    # Keep retrieved context out of system and developer messages.
    messages = append_context_as_user_data(body["messages"], response.context)

    # Stream response from LLM
    return StreamingResponse(
        stream_and_persist(messages, thread_uuid)
    )
```

`user_id` and `conversation_id` are identifiers of your application. Zep does not accept them. Zep generates the UUID of each user and each thread, and a create call accepts no identifier from your application. `get_or_create_zep_thread` reads the Zep thread UUID that your application stored for the conversation. When the conversation has no Zep thread yet, the function calls `thread.create` with the Zep user UUID that your application stored for the user, and stores the returned thread UUID next to `conversation_id`. See [Store the UUID in your own database](/migrating-from-v3#store-the-uuid-in-your-own-database).

The key optimization is `return_context=True`, which retrieves context in the same call as adding the message.

`append_context_as_user_data` must preserve the original system or developer instructions and add the Zep context as a separate user-level message. The proxy must not copy retrieved content into a privileged instruction.

### Frontend integration

Your frontend passes user identity via `customLlmExtraBody`:

```javascript
await conversation.startSession({
  agentId: 'your-agent-id',
  customLlmExtraBody: {
    user_id: user.id,
    conversation_id: crypto.randomUUID(),
  },
});
```

### ElevenLabs configuration

1. In your agent's **LLM** section, select **Custom LLM** and set the server URL to your proxy
2. Add an `Authorization` header for authentication
3. In **Security > Overrides**, enable **Custom LLM extra body** (required for the proxy to receive user identity)

## Production considerations

* **User identity** — Use your auth system's user ID, not random IDs, and keep the Zep UUIDs of each user in the same record
* **User metadata** — Create users in Zep during registration with `first_name`, `last_name`, `email` for better personalization, and store the returned `uuid` and `graph_uuid`
* **Cache warming** — Call `zep.graph.warm(graph_uuid)` with the stored graph UUID of the user when users arrive on your page to pre-fetch their data
* **Proxy location** — Embed the endpoint in your existing backend for direct access to user data, or deploy as a standalone service

## Learn more

* [ElevenLabs Custom LLM documentation](https://elevenlabs.io/docs/agents-platform/customization/llm/custom-llm)
* [Zep context retrieval](/retrieving-context)
* [Creating users in Zep](/users)