Skip to navigation

LiveKit integration

Add long-term agent memory to LiveKit voice agents

The v4 version of zep-livekit is not released yet. The current zep-livekit release uses the v3 API. To use this integration now, follow the v3 version of this page.

The zep-livekit package adds long-term agent memory to LiveKit voice agents. It wraps LiveKit’s Agent so that completed conversation turns are persisted to Zep and relevant context is injected before each response. Choose between user thread memory or structured knowledge graph memory.

Keep retrieved context out of privileged instructions

Zep context can include content that your users, documents, or tools supplied. A system or developer message gives that content higher instruction priority than ordinary input. Some convenience integrations use system-message injection. Use direct SDK retrieval or an actual retrieval tool call unless all stored content is application-authored and trusted. Follow Memory security best practices for provider-specific placement.

Build an agent with Zep tools

To build an agent that plans its retrieval and uses several Zep tools, read Build an Agent with Zep. The guide shows how to add domain knowledge, design tools, and evaluate the agent.

The wrapper inserts context into a system message. For end-user or third-party context, retrieve data with the Zep SDK and place it through your model provider’s documented data channel.

Core benefits

  • Persistent voice memory: Each completed turn is stored in Zep and contributes to the user’s temporal knowledge graph
  • Automatic context injection for trusted content: Relevant context can be added as a system message when it contains only application-authored, trusted data
  • Two access patterns: ZepUserAgent for thread-based conversation memory, ZepGraphAgent for direct knowledge graph access
  • Drop-in replacement: Both classes subclass LiveKit’s Agent and accept all standard Agent parameters

How it works

LiveKit’s AgentSession owns the audio pipeline — speech-to-text, voice activity detection, turn detection, and text-to-speech. Zep does not touch audio. Instead, the Zep agent hooks into LiveKit’s turn lifecycle and runs a write-then-read cycle on each completed user turn:

  1. Persist the turn — when LiveKit fires on_user_turn_completed, the user message is written to Zep (a thread for ZepUserAgent, the graph for ZepGraphAgent). Assistant responses are captured separately via the conversation_item_added session event.
  2. Retrieve context — ZepUserAgent folds persistence and retrieval into a single thread.add_messages(..., return_context=True) round-trip; ZepGraphAgent writes the message to the graph with graph.episode.add, then calls graph.get_context, which assembles the context block on the server.
  3. Inject context — the retrieved context is wrapped in a context template and added to the turn as a system message, so the LLM’s next response is grounded in prior conversation.

Allow time for indexing: Turns are ingested and knowledge is extracted asynchronously, so facts from the current turn are not searchable within that same turn. Context retrieved on a given turn reflects knowledge extracted from earlier turns.

Installation

pip install zep-livekit "livekit-agents[openai,silero]>=1.0.0"

Requires Python 3.11+, zep-livekit, LiveKit Agents v1.0+ (not v0.x), and a Zep Cloud API key. The examples use the v1.0 AgentSession API. Get your API key from app.getzep.com.

Set up your environment variables:

export ZEP_API_KEY="your-zep-api-key"
export OPENAI_API_KEY="your-openai-api-key"
export LIVEKIT_URL="your-livekit-url"
export LIVEKIT_API_KEY="your-livekit-api-key"
export LIVEKIT_API_SECRET="your-livekit-api-secret"

LIVEKIT_URL, LIVEKIT_API_KEY, and LIVEKIT_API_SECRET come from a LiveKit Cloud project or a self-hosted LiveKit server. They configure the LiveKit infrastructure your agent connects to and are unrelated to Zep.

zep-livekit 0.3.0 targets the Zep v4 SDK. Zep v4 addresses every user, thread, and graph by a server-generated UUID, so the integration’s constructor arguments change: ZepUserAgent takes user_uuid and thread_uuid, ZepGraphAgent takes graph_uuid, and create_graph_search_tool takes a required graph_uuid. The ensure_user/ensure_thread helpers are replaced by create_user/create_thread, which return the created User and Thread — read user.uuid_, user.graph_uuid, and thread.uuid_ from the responses and store the UUIDs in your own database. The lazy per-agent provisioning parameters (first_name, last_name, email, on_created) and the per-scope retrieval limits on ZepGraphAgent (facts_limit, entity_limit, episode_limit, reranker) are removed; ZepGraphAgent now calls graph.get_context and takes max_characters.

See the package changelog for the full list of changes.

Agent types

Identity and isolation

Zep v4 addresses every user, thread, and graph by a server-generated UUID. A create call takes no client-chosen name, so the integration cannot resolve a name at run time, and it makes no lookup call. Instead, your application creates each resource one time, stores the UUID that Zep returns in its own database, and passes the stored UUID to the agent.

Map the Zep user UUID to the stable identity in your auth system so a returning user’s memory accumulates across sessions. Create a new thread (or graph) for each room when you want per-session isolation while still attributing every session to the same long-lived user.

# Create the resources one time — during signup or the first session —
# and store the returned UUIDs against your own identifiers.
user = await zep_client.user.create(first_name="Alice")
thread = await zep_client.thread.create(user_uuid=user.uuid_)
user_uuid = user.uuid_
thread_uuid = thread.uuid_
# Persist user_uuid and thread_uuid in your database.

On later sessions, load the stored UUIDs from your database and pass them to the agent. Do not create a new user or thread per session — that fragments a returning user’s history and prevents Zep from accumulating long-term memory for that person.

Provisioning users and threads

The create_user and create_thread helpers wrap user.create and thread.create. Each call creates a new resource and returns the SDK response object — read the uuid_ field and store it. The helpers raise on failure. Call them before the first turn — for example during account or session onboarding — so misconfiguration surfaces loudly.

The optional on_created hook on create_user fires right after the user is created — use it to seed initial facts, set custom instructions, or configure an ontology.

Python
from zep_livekit import create_thread, create_user
async def seed_new_user(zep_client, user_uuid: str) -> None:
"""Runs right after the user is created."""
...
user = await create_user(
zep_client,
first_name="Alice",
on_created=seed_new_user,
)
user_uuid = user.uuid_
thread = await create_thread(zep_client, user_uuid=user_uuid)
thread_uuid = thread.uuid_

ZepGraphAgent does not accept on_created. It is scoped to a shared Context Graph that is addressed with graph_uuid, not to a Zep user. Therefore, there is no user-created event to use. Passing on_created raises TypeError.

User memory agent for trusted deployments

ZepUserAgent inserts retrieved context into a system message. Use this wrapper only when all stored content is fully trusted and application-authored.

ZepUserAgent stores each turn in a Zep thread and injects a context block before the next response.

Python
import logging
import os
from livekit import agents
from livekit.agents import AutoSubscribe
from livekit.plugins import openai, silero
from zep_cloud.client import AsyncZep
from zep_livekit import ZepUserAgent
async def entrypoint(ctx: agents.JobContext):
zep_client = AsyncZep(api_key=os.environ.get("ZEP_API_KEY"))
# Load the UUIDs you created at onboarding and stored in your database.
user_uuid = load_user_uuid(ctx) # your lookup
thread_uuid = load_thread_uuid(ctx) # your lookup
# Subscribe to audio only — a voice agent has no use for video tracks
await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)
# AgentSession owns the audio pipeline (STT, VAD, turn detection, TTS)
session = agents.AgentSession(
stt=openai.STT(),
llm=openai.LLM(model="gpt-5.6-terra"),
tts=openai.TTS(),
vad=silero.VAD.load(),
)
# Trusted-only path: ZepUserAgent inserts stored context into a system message.
agent = ZepUserAgent(
zep_client=zep_client,
user_uuid=user_uuid,
thread_uuid=thread_uuid,
user_message_name="Alice",
assistant_message_name="Assistant",
instructions="You are a helpful voice assistant with long-term memory. "
"Reference details from previous conversations naturally.",
)
await session.start(agent=agent, room=ctx.room)
logging.info("Voice assistant with Zep memory is running")
if __name__ == "__main__":
agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint))

Automatic memory integration: ZepUserAgent captures each voice turn and injects relevant context from previous conversations, enabling continuity across sessions without manual memory management.

ZepUserAgent configuration

ZepUserAgent accepts the following parameters in addition to all standard LiveKit Agent parameters (stt, llm, tts, instructions, tools, chat_ctx, etc.):

ParameterDescription
zep_clientInitialized AsyncZep client
user_uuidUUID of the Zep user (memory isolation boundary)
thread_uuidUUID of the Zep thread (conversation continuity)
user_message_nameOptional name attributed to user messages in Zep
assistant_message_nameOptional name attributed to assistant messages in Zep
context_builderAsync callable replacing the built-in retrieval — see customizing retrieved context
context_templateTemplate wrapping injected context (default: DEFAULT_CONTEXT_TEMPLATE)

Knowledge graph agent for trusted deployments

ZepGraphAgent also inserts retrieved context into a system message. Use this wrapper only for an application-authored graph whose content is fully trusted.

ZepGraphAgent writes each turn directly to a shared Context Graph and retrieves a prompt-ready context block with graph.get_context, which assembles the relevant facts, entities, and episodes on the server. You can optionally shape the graph with custom entity models.

Python
import os
from livekit import agents
from livekit.agents import AutoSubscribe
from livekit.plugins import openai, silero
from pydantic import Field
from zep_cloud import SearchFilters
from zep_cloud.client import AsyncZep
from zep_cloud.ontology import EdgeModel, EntityModel, EntityText, build_ontology
from zep_cloud.types import EdgeSourceTarget
from zep_livekit import ZepGraphAgent
class Person(EntityModel):
"""A person entity for voice interactions."""
role: EntityText = Field(description="person's role or profession", default=None)
interests: EntityText = Field(description="topics the person is interested in", default=None)
class Topic(EntityModel):
"""A conversation topic or subject."""
category: EntityText = Field(description="category of the topic", default=None)
importance: EntityText = Field(description="importance to the user", default=None)
async def entrypoint(ctx: agents.JobContext):
zep_client = AsyncZep(api_key=os.environ.get("ZEP_API_KEY"))
# Load the graph UUID you created at onboarding and stored in your
# database. For a shared graph, create it one time first:
# graph = await zep_client.graph.create(name="LiveKit Knowledge Graph")
# graph_uuid = graph.uuid_
graph_uuid = load_graph_uuid(ctx) # your lookup
# Optional: define a custom ontology for structured extraction on
# this graph.
entity_types, edge_types = build_ontology(
entities={"Person": Person, "Topic": Topic}
)
await zep_client.graph.set_ontology(
graph_uuid,
entity_types=entity_types,
edge_types=edge_types,
)
# Subscribe to audio only — a voice agent has no use for video tracks
await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)
session = agents.AgentSession(
stt=openai.STT(),
llm=openai.LLM(model="gpt-5.6-terra"),
tts=openai.TTS(),
vad=silero.VAD.load(),
)
agent = ZepGraphAgent(
zep_client=zep_client,
graph_uuid=graph_uuid,
max_characters=4000, # Limit the size of the retrieved context block
search_filters=SearchFilters(node_labels=["Person"]), # Constrain to Person entities
instructions="You are a knowledgeable voice assistant. Use the provided "
"context about entities and facts to give informed responses.",
)
await session.start(agent=agent, room=ctx.room)
if __name__ == "__main__":
agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint))

Search filters: The search_filters parameter constrains which results the agent retrieves. Use node_labels to filter by entity types defined in your ontology.

Graph memory context: ZepGraphAgent writes each turn to the graph and injects relevant facts, entities, and episodes as context, grounding responses in prior conversations.

ZepGraphAgent configuration

ZepGraphAgent accepts the following parameters in addition to all standard LiveKit Agent parameters:

ParameterDescription
zep_clientInitialized AsyncZep client
graph_uuidUUID of the graph used for knowledge storage
user_nameOptional name prefixed to stored messages for attribution
max_charactersMaximum length of the retrieved context block
search_filtersOptional SearchFilters applied to context retrieval
context_builderAsync callable replacing the built-in retrieval — see customizing retrieved context
context_templateTemplate wrapping injected context (default: DEFAULT_CONTEXT_TEMPLATE)

ZepGraphAgent has no on_created parameter — it is graph-scoped, with no Zep user to provision. Passing on_created raises TypeError.

Customizing retrieved context

Both agents wrap injected context in DEFAULT_CONTEXT_TEMPLATE, the shared <ZEP_CONTEXT>...</ZEP_CONTEXT> block, before adding it as a system message. Override with context_template: it must contain a literal {context} placeholder. Plain string replacement prevents format-string interpretation of {, }, or %. It does not prevent the model from following instructions in the retrieved content.

Python
agent = ZepUserAgent(
zep_client=zep_client,
user_uuid=user_uuid,
thread_uuid=thread_uuid,
context_template="Known facts about the user:\n{context}",
)

To replace the retrieval logic itself — a filtered graph search, a different graph, or multi-source context assembly — pass context_builder. On ZepUserAgent the builder is an async callable receiving a frozen ContextInput (zep, user_uuid, thread_uuid, user_message, session) and returning the context string, or None to skip injection:

Python
from zep_livekit import ContextInput
async def my_builder(ctx: ContextInput) -> str | None:
user = await ctx.zep.user.get(ctx.user_uuid)
results = await ctx.zep.graph.search_edges(
user.graph_uuid, # the user's personal graph
query=ctx.user_message,
)
edges = results.items or []
if not edges:
return None
return "\n".join(edge.fact for edge in edges)
agent = ZepUserAgent(
zep_client=zep_client,
user_uuid=user_uuid,
thread_uuid=thread_uuid,
context_builder=my_builder,
)

When context_builder is set on ZepUserAgent, message persistence and the builder run concurrently for lower latency, with per-side failure isolation: a builder error is logged and skips injection for that turn but does not stop persistence, and a persistence error is logged but a successful builder result is still injected.

ZepGraphAgent takes the analogous context_builder typed as GraphContextBuilder, receiving a GraphContextInput (zep, graph_uuid, user_message, session). Setting it fully replaces the built-in graph.get_context call rather than running concurrently with anything — graph message persistence happens independently, earlier in the turn.

Graph search tool

In addition to the context injected automatically every turn, create_graph_search_tool builds a model-callable LiveKit function tool (via function_tool(raw_schema=...), returning a RawFunctionTool) that lets the agent search a Zep graph on demand. The graph_uuid parameter is required: pass the UUID of a shared standalone graph, or the graph_uuid of a user object to search that user’s personal graph. Register it through the standard tools=[...] parameter:

Python
from zep_livekit import ZepUserAgent, create_graph_search_tool
# graph_uuid is a stored standalone graph UUID, or the graph_uuid field of a
# stored user object to search that user's personal graph.
search_tool = create_graph_search_tool(zep_client, graph_uuid=graph_uuid)
agent = ZepUserAgent(
zep_client=zep_client,
user_uuid=user_uuid,
thread_uuid=thread_uuid,
tools=[search_tool],
instructions="...",
)

The tool exposes the search parameters to the model by default:

ParameterValuesDefault
queryNatural language search query (required)—
scopeedges, nodes, episodes, observations, thread_summaries, autoedges
rerankerrrf, mmr, node_distance, episode_mentions, cross_encoderrrf
limitMaximum results10
mmr_lambdaDiversity/relevance balance for the mmr rerankeromitted when unset
center_node_uuidCenter node for node_distance rerankingomitted when unset

Each scope maps to a dedicated search method (graph.search_edges, graph.search_nodes, graph.search_episodes, graph.search_observations, or graph.search_thread_summaries); auto calls graph.get_context.

Use pinned_params to fix a parameter to a constant value (hidden from the model, always sent), or hidden_params to hide a parameter without pinning it (Zep’s server-side default applies). search_filters and bfs_origin_node_uuids are constructor-only and never exposed to the model.

Python
search_tool = create_graph_search_tool(
zep_client,
graph_uuid=graph_uuid,
pinned_params={"scope": "edges", "limit": 5},
hidden_params={"center_node_uuid"},
)

Zep failures are caught and returned as an error string to the model — the tool never raises into the voice session.

Size limits

  • Zep rejects direct thread-message payloads over 4,096 characters; LiveKit agents truncate message content to 4,000 characters before writing it, logging lengths only and never content.
  • Zep rejects direct graph.episode.add payloads over 10,000 characters; LiveKit graph agents truncate graph payloads to 9,900 characters before calling graph.episode.add.

Best practices

  • Store the UUIDs in your database — create each user, thread, and graph one time and persist the returned uuid_ against your own identifiers, so a returning user’s memory accumulates instead of fragmenting across sessions
  • Scope sessions with a new thread or graph — create one thread (or graph) per room when you want per-session isolation, keeping the user UUID constant
  • Use SDK or tool retrieval for untrusted context. The wrapper’s automatic retrieval enters a system message.
  • Allow time for indexing — Zep extracts knowledge asynchronously, so facts from a turn are not instantly searchable

Next steps