> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs-beta.getzep.com/v4/livekit-memory/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-beta.getzep.com/_mcp/server. # LiveKit integration > **Note** > > The v4 version of `zep-livekit` is not released yet. The current `zep-livekit` release uses the v3 API. To use this integration now, follow the [v3 version of this page](/v3/livekit-memory). The `zep-livekit` package adds long-term agent memory to [LiveKit](https://docs.livekit.io/agents/) voice agents. It wraps LiveKit's `Agent` so that completed conversation turns are persisted to Zep and relevant context is injected before each response. Choose between [user thread memory](/users) or structured [knowledge graph memory](/graph-overview). > **Keep retrieved context out of privileged instructions** > > Zep context can include content that your users, documents, or tools supplied. A system or developer message gives that content higher instruction priority than ordinary input. Some convenience integrations use system-message injection. Use direct SDK retrieval or an actual retrieval tool call unless all stored content is application-authored and trusted. Follow [Memory security best practices](/memory-security) for provider-specific placement. > **Build an agent with Zep tools** > > To build an agent that plans its retrieval and uses several Zep tools, read [Build an Agent with Zep](/build-an-agent-with-zep). The guide shows how to add domain knowledge, design tools, and evaluate the agent. The wrapper inserts context into a system message. For end-user or third-party context, retrieve data with the Zep SDK and place it through your model provider's documented data channel. ## Core benefits * **Persistent voice memory**: Each completed turn is stored in Zep and contributes to the user's temporal knowledge graph * **Automatic context injection for trusted content**: Relevant context can be added as a system message when it contains only application-authored, trusted data * **Two access patterns**: `ZepUserAgent` for thread-based conversation memory, `ZepGraphAgent` for direct knowledge graph access * **Drop-in replacement**: Both classes subclass LiveKit's `Agent` and accept all standard `Agent` parameters ## How it works LiveKit's `AgentSession` owns the audio pipeline — speech-to-text, voice activity detection, turn detection, and text-to-speech. Zep does not touch audio. Instead, the Zep agent hooks into LiveKit's turn lifecycle and runs a write-then-read cycle on each completed user turn: 1. **Persist the turn** — when LiveKit fires `on_user_turn_completed`, the user message is written to Zep (a thread for `ZepUserAgent`, the graph for `ZepGraphAgent`). Assistant responses are captured separately via the `conversation_item_added` session event. 2. **Retrieve context** — `ZepUserAgent` folds persistence and retrieval into a single `thread.add_messages(..., return_context=True)` round-trip; `ZepGraphAgent` writes the message to the graph with `graph.episode.add`, then calls `graph.get_context`, which assembles the context block on the server. 3. **Inject context** — the retrieved context is wrapped in a [context template](#customizing-retrieved-context) and added to the turn as a system message, so the LLM's next response is grounded in prior conversation. > **Info** > > **Allow time for indexing**: Turns are ingested and knowledge is extracted asynchronously, so facts from the current turn are not searchable within that same turn. Context retrieved on a given turn reflects knowledge extracted from earlier turns. ## Installation **`pip`** ```bash pip pip install zep-livekit "livekit-agents[openai,silero]>=1.0.0" ``` **`uv`** ```bash uv uv add zep-livekit "livekit-agents[openai,silero]>=1.0.0" ``` **`poetry`** ```bash poetry poetry add zep-livekit "livekit-agents[openai,silero]>=1.0.0" ``` > **Info** > > Requires Python 3.11+, `zep-livekit`, LiveKit Agents v1.0+ (not v0.x), and a Zep Cloud API key. The examples use the v1.0 `AgentSession` API. Get your API key from [app.getzep.com](https://app.getzep.com). Set up your environment variables: ```bash export ZEP_API_KEY="your-zep-api-key" export OPENAI_API_KEY="your-openai-api-key" export LIVEKIT_URL="your-livekit-url" export LIVEKIT_API_KEY="your-livekit-api-key" export LIVEKIT_API_SECRET="your-livekit-api-secret" ``` > **Info** > > `LIVEKIT_URL`, `LIVEKIT_API_KEY`, and `LIVEKIT_API_SECRET` come from a [LiveKit Cloud](https://cloud.livekit.io) project or a self-hosted LiveKit server. They configure the LiveKit infrastructure your agent connects to and are unrelated to Zep. #### Upgrading from earlier versions of zep-livekit `zep-livekit` 0.3.0 targets the Zep v4 SDK. Zep v4 addresses every user, thread, and graph by a server-generated UUID, so the integration's constructor arguments change: `ZepUserAgent` takes `user_uuid` and `thread_uuid`, `ZepGraphAgent` takes `graph_uuid`, and `create_graph_search_tool` takes a required `graph_uuid`. The `ensure_user`/`ensure_thread` helpers are replaced by `create_user`/`create_thread`, which return the created `User` and `Thread` — read `user.uuid_`, `user.graph_uuid`, and `thread.uuid_` from the responses and store the UUIDs in your own database. The lazy per-agent provisioning parameters (`first_name`, `last_name`, `email`, `on_created`) and the per-scope retrieval limits on `ZepGraphAgent` (`facts_limit`, `entity_limit`, `episode_limit`, `reranker`) are removed; `ZepGraphAgent` now calls `graph.get_context` and takes `max_characters`. See the [package changelog](https://github.com/getzep/zep/blob/main/integrations/livekit/python/CHANGELOG.md) for the full list of changes. ## Agent types * **`ZepUserAgent`**: Uses [user threads](/users) for conversation memory with automatic context injection * **`ZepGraphAgent`**: Reads and writes a [knowledge graph](/graph-overview), optionally shaped by [custom entity models](/customizing-graph-structure) ## Identity and isolation Zep v4 addresses every user, thread, and graph by a server-generated UUID. A create call takes no client-chosen name, so the integration cannot resolve a name at run time, and it makes no lookup call. Instead, your application creates each resource one time, stores the UUID that Zep returns in its own database, and passes the stored UUID to the agent. Map the Zep user UUID to the stable identity in your auth system so a returning user's memory accumulates across sessions. Create a new thread (or graph) for each room when you want per-session isolation while still attributing every session to the same long-lived user. ```python # Create the resources one time — during signup or the first session — # and store the returned UUIDs against your own identifiers. user = await zep_client.user.create(first_name="Alice") thread = await zep_client.thread.create(user_uuid=user.uuid_) user_uuid = user.uuid_ thread_uuid = thread.uuid_ # Persist user_uuid and thread_uuid in your database. ``` On later sessions, load the stored UUIDs from your database and pass them to the agent. Do not create a new user or thread per session — that fragments a returning user's history and prevents Zep from accumulating long-term memory for that person. ## Provisioning users and threads The `create_user` and `create_thread` helpers wrap `user.create` and `thread.create`. Each call creates a new resource and returns the SDK response object — read the `uuid_` field and store it. The helpers raise on failure. Call them before the first turn — for example during account or session onboarding — so misconfiguration surfaces loudly. The optional `on_created` hook on `create_user` fires right after the user is created — use it to seed initial facts, set custom instructions, or configure an ontology. **`Python`** ```python Python from zep_livekit import create_thread, create_user async def seed_new_user(zep_client, user_uuid: str) -> None: """Runs right after the user is created.""" ... user = await create_user( zep_client, first_name="Alice", on_created=seed_new_user, ) user_uuid = user.uuid_ thread = await create_thread(zep_client, user_uuid=user_uuid) thread_uuid = thread.uuid_ ``` `ZepGraphAgent` does not accept `on_created`. It is scoped to a shared Context Graph that is addressed with `graph_uuid`, not to a Zep user. Therefore, there is no user-created event to use. Passing `on_created` raises `TypeError`. ## User memory agent for trusted deployments `ZepUserAgent` inserts retrieved context into a system message. Use this wrapper only when all stored content is fully trusted and application-authored. `ZepUserAgent` stores each turn in a Zep thread and injects a context block before the next response. **`Python`** ```python Python import logging import os from livekit import agents from livekit.agents import AutoSubscribe from livekit.plugins import openai, silero from zep_cloud.client import AsyncZep from zep_livekit import ZepUserAgent async def entrypoint(ctx: agents.JobContext): zep_client = AsyncZep(api_key=os.environ.get("ZEP_API_KEY")) # Load the UUIDs you created at onboarding and stored in your database. user_uuid = load_user_uuid(ctx) # your lookup thread_uuid = load_thread_uuid(ctx) # your lookup # Subscribe to audio only — a voice agent has no use for video tracks await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY) # AgentSession owns the audio pipeline (STT, VAD, turn detection, TTS) session = agents.AgentSession( stt=openai.STT(), llm=openai.LLM(model="gpt-5.6-terra"), tts=openai.TTS(), vad=silero.VAD.load(), ) # Trusted-only path: ZepUserAgent inserts stored context into a system message. agent = ZepUserAgent( zep_client=zep_client, user_uuid=user_uuid, thread_uuid=thread_uuid, user_message_name="Alice", assistant_message_name="Assistant", instructions="You are a helpful voice assistant with long-term memory. " "Reference details from previous conversations naturally.", ) await session.start(agent=agent, room=ctx.room) logging.info("Voice assistant with Zep memory is running") if __name__ == "__main__": agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint)) ``` > **Info** > > **Automatic memory integration**: `ZepUserAgent` captures each voice turn and injects relevant context from previous conversations, enabling continuity across sessions without manual memory management. ### ZepUserAgent configuration `ZepUserAgent` accepts the following parameters in addition to all standard LiveKit `Agent` parameters (`stt`, `llm`, `tts`, `instructions`, `tools`, `chat_ctx`, etc.): | Parameter | Description | | ------------------------ | --------------------------------------------------------------------------------------------------------------------- | | `zep_client` | Initialized `AsyncZep` client | | `user_uuid` | UUID of the Zep user (memory isolation boundary) | | `thread_uuid` | UUID of the Zep thread (conversation continuity) | | `user_message_name` | Optional name attributed to user messages in Zep | | `assistant_message_name` | Optional name attributed to assistant messages in Zep | | `context_builder` | Async callable replacing the built-in retrieval — see [customizing retrieved context](#customizing-retrieved-context) | | `context_template` | Template wrapping injected context (default: `DEFAULT_CONTEXT_TEMPLATE`) | ## Knowledge graph agent for trusted deployments `ZepGraphAgent` also inserts retrieved context into a system message. Use this wrapper only for an application-authored graph whose content is fully trusted. `ZepGraphAgent` writes each turn directly to a shared Context Graph and retrieves a prompt-ready context block with `graph.get_context`, which assembles the relevant facts, entities, and episodes on the server. You can optionally shape the graph with custom entity models. **`Python`** ```python Python import os from livekit import agents from livekit.agents import AutoSubscribe from livekit.plugins import openai, silero from pydantic import Field from zep_cloud import SearchFilters from zep_cloud.client import AsyncZep from zep_cloud.ontology import EdgeModel, EntityModel, EntityText, build_ontology from zep_cloud.types import EdgeSourceTarget from zep_livekit import ZepGraphAgent class Person(EntityModel): """A person entity for voice interactions.""" role: EntityText = Field(description="person's role or profession", default=None) interests: EntityText = Field(description="topics the person is interested in", default=None) class Topic(EntityModel): """A conversation topic or subject.""" category: EntityText = Field(description="category of the topic", default=None) importance: EntityText = Field(description="importance to the user", default=None) async def entrypoint(ctx: agents.JobContext): zep_client = AsyncZep(api_key=os.environ.get("ZEP_API_KEY")) # Load the graph UUID you created at onboarding and stored in your # database. For a shared graph, create it one time first: # graph = await zep_client.graph.create(name="LiveKit Knowledge Graph") # graph_uuid = graph.uuid_ graph_uuid = load_graph_uuid(ctx) # your lookup # Optional: define a custom ontology for structured extraction on # this graph. entity_types, edge_types = build_ontology( entities={"Person": Person, "Topic": Topic} ) await zep_client.graph.set_ontology( graph_uuid, entity_types=entity_types, edge_types=edge_types, ) # Subscribe to audio only — a voice agent has no use for video tracks await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY) session = agents.AgentSession( stt=openai.STT(), llm=openai.LLM(model="gpt-5.6-terra"), tts=openai.TTS(), vad=silero.VAD.load(), ) agent = ZepGraphAgent( zep_client=zep_client, graph_uuid=graph_uuid, max_characters=4000, # Limit the size of the retrieved context block search_filters=SearchFilters(node_labels=["Person"]), # Constrain to Person entities instructions="You are a knowledgeable voice assistant. Use the provided " "context about entities and facts to give informed responses.", ) await session.start(agent=agent, room=ctx.room) if __name__ == "__main__": agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint)) ``` > **Info** > > **Search filters**: The `search_filters` parameter constrains which results the agent retrieves. Use `node_labels` to filter by entity types defined in your ontology. > **Info** > > **Graph memory context**: `ZepGraphAgent` writes each turn to the graph and injects relevant facts, entities, and episodes as context, grounding responses in prior conversations. ### ZepGraphAgent configuration `ZepGraphAgent` accepts the following parameters in addition to all standard LiveKit `Agent` parameters: | Parameter | Description | | ------------------ | --------------------------------------------------------------------------------------------------------------------- | | `zep_client` | Initialized `AsyncZep` client | | `graph_uuid` | UUID of the graph used for knowledge storage | | `user_name` | Optional name prefixed to stored messages for attribution | | `max_characters` | Maximum length of the retrieved context block | | `search_filters` | Optional `SearchFilters` applied to context retrieval | | `context_builder` | Async callable replacing the built-in retrieval — see [customizing retrieved context](#customizing-retrieved-context) | | `context_template` | Template wrapping injected context (default: `DEFAULT_CONTEXT_TEMPLATE`) | `ZepGraphAgent` has no `on_created` parameter — it is graph-scoped, with no Zep user to provision. Passing `on_created` raises `TypeError`. ## Customizing retrieved context Both agents wrap injected context in `DEFAULT_CONTEXT_TEMPLATE`, the shared `...` block, before adding it as a system message. Override with `context_template`: it must contain a literal `{context}` placeholder. Plain string replacement prevents format-string interpretation of `{`, `}`, or `%`. It does not prevent the model from following instructions in the retrieved content. **`Python`** ```python Python agent = ZepUserAgent( zep_client=zep_client, user_uuid=user_uuid, thread_uuid=thread_uuid, context_template="Known facts about the user:\n{context}", ) ``` To replace the retrieval logic itself — a filtered graph search, a different graph, or multi-source context assembly — pass `context_builder`. On `ZepUserAgent` the builder is an async callable receiving a frozen `ContextInput` (`zep`, `user_uuid`, `thread_uuid`, `user_message`, `session`) and returning the context string, or `None` to skip injection: **`Python`** ```python Python from zep_livekit import ContextInput async def my_builder(ctx: ContextInput) -> str | None: user = await ctx.zep.user.get(ctx.user_uuid) results = await ctx.zep.graph.search_edges( user.graph_uuid, # the user's personal graph query=ctx.user_message, ) edges = results.items or [] if not edges: return None return "\n".join(edge.fact for edge in edges) agent = ZepUserAgent( zep_client=zep_client, user_uuid=user_uuid, thread_uuid=thread_uuid, context_builder=my_builder, ) ``` When `context_builder` is set on `ZepUserAgent`, message persistence and the builder run concurrently for lower latency, with per-side failure isolation: a builder error is logged and skips injection for that turn but does not stop persistence, and a persistence error is logged but a successful builder result is still injected. `ZepGraphAgent` takes the analogous `context_builder` typed as `GraphContextBuilder`, receiving a `GraphContextInput` (`zep`, `graph_uuid`, `user_message`, `session`). Setting it fully replaces the built-in `graph.get_context` call rather than running concurrently with anything — graph message persistence happens independently, earlier in the turn. ## Graph search tool In addition to the context injected automatically every turn, `create_graph_search_tool` builds a model-callable LiveKit function tool (via `function_tool(raw_schema=...)`, returning a `RawFunctionTool`) that lets the agent search a Zep graph on demand. The `graph_uuid` parameter is required: pass the UUID of a shared standalone graph, or the `graph_uuid` of a user object to search that user's personal graph. Register it through the standard `tools=[...]` parameter: **`Python`** ```python Python from zep_livekit import ZepUserAgent, create_graph_search_tool # graph_uuid is a stored standalone graph UUID, or the graph_uuid field of a # stored user object to search that user's personal graph. search_tool = create_graph_search_tool(zep_client, graph_uuid=graph_uuid) agent = ZepUserAgent( zep_client=zep_client, user_uuid=user_uuid, thread_uuid=thread_uuid, tools=[search_tool], instructions="...", ) ``` The tool exposes the search parameters to the model by default: | Parameter | Values | Default | | ------------------ | ------------------------------------------------------------------------ | ------------------ | | `query` | Natural language search query (required) | — | | `scope` | `edges`, `nodes`, `episodes`, `observations`, `thread_summaries`, `auto` | `edges` | | `reranker` | `rrf`, `mmr`, `node_distance`, `episode_mentions`, `cross_encoder` | `rrf` | | `limit` | Maximum results | `10` | | `mmr_lambda` | Diversity/relevance balance for the `mmr` reranker | omitted when unset | | `center_node_uuid` | Center node for `node_distance` reranking | omitted when unset | Each `scope` maps to a dedicated search method (`graph.search_edges`, `graph.search_nodes`, `graph.search_episodes`, `graph.search_observations`, or `graph.search_thread_summaries`); `auto` calls `graph.get_context`. Use `pinned_params` to fix a parameter to a constant value (hidden from the model, always sent), or `hidden_params` to hide a parameter without pinning it (Zep's server-side default applies). `search_filters` and `bfs_origin_node_uuids` are constructor-only and never exposed to the model. **`Python`** ```python Python search_tool = create_graph_search_tool( zep_client, graph_uuid=graph_uuid, pinned_params={"scope": "edges", "limit": 5}, hidden_params={"center_node_uuid"}, ) ``` Zep failures are caught and returned as an error string to the model — the tool never raises into the voice session. ## Size limits * Zep rejects direct thread-message payloads over 4,096 characters; LiveKit agents truncate message content to 4,000 characters before writing it, logging lengths only and never content. * Zep rejects direct `graph.episode.add` payloads over 10,000 characters; LiveKit graph agents truncate graph payloads to 9,900 characters before calling `graph.episode.add`. ## Best practices * **Store the UUIDs in your database** — create each user, thread, and graph one time and persist the returned `uuid_` against your own identifiers, so a returning user's memory accumulates instead of fragmenting across sessions * **Scope sessions with a new thread or graph** — create one thread (or graph) per room when you want per-session isolation, keeping the user UUID constant * **Use SDK or tool retrieval for untrusted context.** The wrapper's automatic retrieval enters a system message. * **Allow time for indexing** — Zep extracts knowledge asynchronously, so facts from a turn are not instantly searchable ## Next steps * Explore [customizing graph structure](/customizing-graph-structure) for advanced knowledge organization * Learn about [searching the graph](/searching-the-graph) and how to tune search > Add long-term agent memory to LiveKit voice agents