Skip to navigation

ElevenLabs Agents

Add persistent context to ElevenLabs voice agents using a custom LLM proxy.

A complete working example is available on GitHub: elevenlabs-zep-example

ElevenLabs Agents is a platform for building intelligent voice agents. This guide shows how to integrate Zep with ElevenLabs using a custom LLM proxy.

Keep retrieved context out of privileged instructions

Zep context can include content that your users, documents, or tools supplied. A system or developer message gives that content higher instruction priority than ordinary input. Some convenience integrations use system-message injection. Use direct SDK retrieval or an actual retrieval tool call unless all stored content is application-authored and trusted. Follow Memory security best practices for provider-specific placement.

Build an agent with Zep tools

To build an agent that plans its retrieval and uses several Zep tools, read Build an Agent with Zep. The guide shows how to add domain knowledge, design tools, and evaluate the agent.

Why use a proxy

ElevenLabs supports custom tools, but using tools for context retrieval has problems:

  • Latency — Tool calls add round-trips where the LLM decides whether to call the tool. For voice agents, this delay is noticeable.
  • Unreliable — The LLM may skip retrieval when it shouldn’t, or call it unnecessarily.

A proxy solves both problems. Context retrieval happens transparently on every request, without LLM involvement.

Architecture

┌────────┐ ┌──────────┐ ┌─────────┐ ┌────────┐
│Frontend│ ◄───► │ElevenLabs│ ◄───► │LLM Proxy│ ◄───► │ OpenAI │
└────────┘ └──────────┘ └─────────┘ └────────┘
▲
│
▼
┌─────┐
│ Zep │
└─────┘

The proxy sits between ElevenLabs and your LLM. On every request it:

  1. Adds the user message to Zep and retrieves context in one call
  2. Adds context as untrusted user-level data
  3. Forwards to the LLM and streams the response back
  4. Persists the assistant response to Zep

Implementation

The proxy endpoint

The proxy exposes an OpenAI-compatible /v1/chat/completions endpoint:

from zep_cloud import AddMessage
@app.post("/v1/chat/completions")
async def chat_completions(request: Request):
body = await request.json()
# ElevenLabs puts customLlmExtraBody in "elevenlabs_extra_body"
extra = body.get("elevenlabs_extra_body", {})
user_id = extra.get("user_id")
conversation_id = extra.get("conversation_id")
# Read the Zep thread UUID that your application stored for this conversation
thread_uuid = await get_or_create_zep_thread(user_id, conversation_id)
# Add user message to Zep and get context in one call
user_message = get_latest_user_message(body["messages"])
response = await zep.thread.add_messages(
thread_uuid,
messages=[AddMessage(role="user", content=user_message)],
return_context=True # Returns context without separate call
)
# Keep retrieved context out of system and developer messages.
messages = append_context_as_user_data(body["messages"], response.context)
# Stream response from LLM
return StreamingResponse(
stream_and_persist(messages, thread_uuid)
)

user_id and conversation_id are identifiers of your application. Zep does not accept them. Zep generates the UUID of each user and each thread, and a create call accepts no identifier from your application. get_or_create_zep_thread reads the Zep thread UUID that your application stored for the conversation. When the conversation has no Zep thread yet, the function calls thread.create with the Zep user UUID that your application stored for the user, and stores the returned thread UUID next to conversation_id. See Store the UUID in your own database.

The key optimization is return_context=True, which retrieves context in the same call as adding the message.

append_context_as_user_data must preserve the original system or developer instructions and add the Zep context as a separate user-level message. The proxy must not copy retrieved content into a privileged instruction.

Frontend integration

Your frontend passes user identity via customLlmExtraBody:

await conversation.startSession({
agentId: 'your-agent-id',
customLlmExtraBody: {
user_id: user.id,
conversation_id: crypto.randomUUID(),
},
});

ElevenLabs configuration

  1. In your agent’s LLM section, select Custom LLM and set the server URL to your proxy
  2. Add an Authorization header for authentication
  3. In Security > Overrides, enable Custom LLM extra body (required for the proxy to receive user identity)

Production considerations

  • User identity — Use your auth system’s user ID, not random IDs, and keep the Zep UUIDs of each user in the same record
  • User metadata — Create users in Zep during registration with first_name, last_name, email for better personalization, and store the returned uuid and graph_uuid
  • Cache warming — Call zep.graph.warm(graph_uuid) with the stored graph UUID of the user when users arrive on your page to pre-fetch their data
  • Proxy location — Embed the endpoint in your existing backend for direct access to user data, or deploy as a standalone service

Learn more