Skip to navigation

Google ADK integration

Add persistent context and knowledge graphs to Google ADK agents

The v4 versions of zep-adk (Python), @getzep/zep-adk (TypeScript), and github.com/getzep/zep/integrations/adk/go (Go) are not released yet. The current releases use the v3 API. To use this integration now, follow the v3 version of this page.

Google’s Agent Development Kit (ADK) agents equipped with Zep’s context layer can maintain context across conversations and access personalized knowledge graphs. The zep-adk package provides real-time message persistence and automatic context injection for ADK agents, and ships for Python, TypeScript, and Go.

Keep retrieved context out of privileged instructions

Zep context can include content that your users, documents, or tools supplied. A system or developer message gives that content higher instruction priority than ordinary input. Some convenience integrations use system-message injection. Use direct SDK retrieval or an actual retrieval tool call unless all stored content is application-authored and trusted. Follow Memory security best practices for provider-specific placement.

Build an agent with Zep tools

To build an agent that plans its retrieval and uses several Zep tools, read Build an Agent with Zep. The guide shows how to add domain knowledge, design tools, and evaluate the agent.

Use ZepGraphSearchTool for context that can contain end-user or third-party data. Use the automatic context hook only for fully trusted, application-authored context.

Core benefits

  • Zero restructuring: Add Zep to an existing ADK agent without changing your agent architecture
  • Shared-agent architecture: One Agent definition serves all users. Per-user identity is resolved at runtime from ADK session state
  • Real-time persistence: Both user and assistant messages are persisted to Zep on every turn, not batched at session end
  • Automatic context injection for trusted content: Zep’s context block can be inserted before each response when it contains only application-authored, trusted data
  • Explicit provisioning: Helpers create Zep users and threads once, out of band, before the first turn
  • ADK-native memory service: Zep backs ADK’s built-in load_memory/preload_memory tools through a BaseMemoryService implementation

How it works

The integration hooks into ADK’s agent lifecycle to persist the user’s message and inject relevant context before each model call, then persist the assistant’s reply afterward. Each language exposes the same capabilities through its idiomatic ADK extension points:

CapabilityPythonTypeScriptGo
Context injection (per turn)ZepContextToolcreateZepBeforeModelCallback or ZepContextToolNewBeforeModelCallback
Assistant persistencecreate_after_model_callbackcreateZepAfterModelCallbackNewAfterModelCallback
Provisioningcreate_user / create_threadcreateUser / createThreadCreateUser / CreateThread
Custom context blockcontext_buildercontextBuilderWithContextBuilder
Injection templatecontext_templatecontextTemplateWithContextTemplate
Model-callable graph searchZepGraphSearchToolZepGraphSearchToolNewGraphSearchTool
ADK-native memory serviceZepMemoryServiceZepMemoryServiceNewMemoryService

In Python and TypeScript, ZepContextTool is a BaseTool that hooks ADK’s process_llm_request() lifecycle method — the same hook ADK’s own PreloadMemoryTool uses — and is never called by the model directly. In TypeScript, use either createZepBeforeModelCallback or ZepContextTool, not both: running both persists each user message twice. Go intentionally has no tool-based injection — callbacks are the idiomatic Go ADK hook.

On each turn the context hook resolves the user’s Zep identity, persists the user’s message, retrieves the relevant context block, and injects it into the model’s system instruction. Tool-loop continuations are skipped, so a turn is recorded in Zep exactly once. The turn path assumes the Zep user and thread already exist — provision them with the create_user/create_thread helpers (Python) before the first turn (see Provisioning users and threads). If persistence targets a user or thread that doesn’t exist, a warning naming the helpers is logged and the turn continues without Zep memory.

What gets persisted

Only the user’s message and the model’s final response are persisted to Zep on each turn. Intermediate model outputs — such as “thinking” text emitted alongside a tool call (e.g. “Let me look that up for you.”) — are not persisted. Tool calls and tool results are also excluded. This keeps the Zep thread clean: one user message and one assistant message per turn, reflecting the actual conversation rather than internal agent mechanics.

If the user message contains multiple text parts (e.g. text alongside an image), all text parts are joined. Non-text parts (images, files) are ignored — only text is sent to Zep. The Zep API rejects thread messages over 4,096 characters; the ADK integration truncates longer messages before persisting rather than dropping the turn.

Installation

pip install zep-adk

Requires a Zep Cloud API key — get yours from app.getzep.com — plus the ADK runtime for your language: Python 3.11+ with google-adk>=1.19.0,<3, Node.js 20+ with @google/adk (a ^1.2.0 peer dependency), or Go 1.25+ with google.golang.org/adk v1.4.0. The Go package is imported as zepadk "github.com/getzep/zep/integrations/adk/go".

Set up your Zep API key and Google API key:

export ZEP_API_KEY="your-zep-api-key"
export GOOGLE_API_KEY="your-google-api-key"

Versions zep-adk 0.3.0 (Python), @getzep/zep-adk 0.2.0 (TypeScript), and zepadk 0.2.0 (Go) replaced lazy in-band resource creation with explicit provisioning. If you’re upgrading:

  • Python: ensure_user/ensure_thread are replaced by create_user/create_thread, which return the created User and Thread. Read user.uuid_, user.graph_uuid, and thread.uuid_ from the responses and store them.
  • Python: The on_created hook and the UserSetupHook type are removed. Run one-time user setup after create_user returns.
  • Python: The zep_user_id and zep_thread_id session-state keys are replaced by zep_user_uuid, zep_thread_uuid, and zep_graph_uuid. ContextInput carries user_uuid and thread_uuid.
  • Python: ContextBuilder takes a single ContextInput argument instead of four positional arguments.
  • TypeScript: ZepResourceManager is removed — use createZepCallbacks, or share a TurnDedup instance via the dedup option.
  • TypeScript: the package targets the Zep v4 TypeScript SDK. ensureUser and ensureThread are replaced by createUser and createThread, which return the UUIDs that Zep generates. The callbacks and the tools take those UUIDs as userUuid and threadUuid, and ZepGraphSearchTool and ZepMemoryService take a graphUuid. The session-state keys are zep_user_uuid and zep_thread_uuid.
  • Go: the package targets the Zep v4 Go SDK, github.com/getzep/zep-go/v4. EnsureUser and EnsureThread are replaced by CreateUser and CreateThread, which return the UUIDs that Zep generates. The callbacks, the search tool, and the memory service take those UUIDs through WithThreadUUID, WithUserUUID, WithAfterThreadUUID, WithGraphUUID, and WithMemoryGraphUUID, or through the matching resolver option.
  • All languages: The zep_email session-state key is removed — pass email to create_user/createUser/CreateUser.

For the full list of changes, see the package CHANGELOGs in the zep-adk repository.

Automatic context for trusted deployments

The automatic callbacks insert retrieved context into the model’s system instruction. Use these examples only when all stored content is fully trusted and application-authored.

Whether you’re building a new agent or adding Zep to an existing one, the setup is the same: provision the Zep user and thread out of band, then wire up the context hook and the after-model callback. The Python example below shows the full runner flow; the TypeScript and Go tabs show the equivalent wiring.

import asyncio
import os
from uuid import uuid4
from google.adk.agents import Agent
from google.adk.runners import Runner
from google.adk.sessions import InMemorySessionService
from google.genai import types
from zep_cloud.client import AsyncZep
from zep_adk import ZepContextTool, create_after_model_callback, create_user, create_thread
async def main() -> None:
zep = AsyncZep(api_key=os.environ["ZEP_API_KEY"])
# One shared agent definition serves all users.
agent = Agent(
name="my_agent",
model="gemini-3.7-flash",
instruction="You are a helpful assistant with long-term memory.",
tools=[
ZepContextTool(
zep_client=zep,
ignore_roles=["assistant"],
),
],
after_model_callback=create_after_model_callback(
zep_client=zep,
assistant_name="my_agent",
ignore_roles=["assistant"],
),
)
session_service = InMemorySessionService()
runner = Runner(agent=agent, app_name="my_app", session_service=session_service)
# Provision the Zep user and thread before the first turn. Zep generates the UUIDs.
session_id = f"session-{uuid4().hex[:8]}"
user = await create_user(
zep,
first_name="Jane",
last_name="Smith",
)
thread = await create_thread(zep, user_uuid=user.uuid_)
# Store user.uuid_, user.graph_uuid, and thread.uuid_ in your own database,
# then put them into the ADK session state.
await session_service.create_session(
app_name="my_app",
user_id="user-123", # your own application identifier
session_id=session_id,
state={
"zep_user_uuid": user.uuid_,
"zep_thread_uuid": thread.uuid_,
"zep_graph_uuid": user.graph_uuid,
"zep_first_name": "Jane",
"zep_last_name": "Smith",
},
)
# Trusted-only path: context enters the model's system instruction.
content = types.Content(
role="user",
parts=[types.Part(text="Hi, I work at Acme Corp.")],
)
async for event in runner.run_async(
user_id="user-123",
session_id=session_id,
new_message=content,
):
if event.is_final_response() and event.content:
print(event.content.parts[0].text)
asyncio.run(main())

That’s it. Every user message is persisted to Zep, relevant context is injected into the LLM prompt, and assistant responses are captured — all automatically.

The ignore_roles parameter shown above excludes specific message roles from graph ingestion while still storing them in the thread history. This is useful when assistant messages don’t add meaningful knowledge to the graph — they’re preserved for conversation context but don’t create nodes or edges. Both ZepContextTool and create_after_model_callback accept ignore_roles (TypeScript: ignoreRoles). See Ignore assistant messages in the Zep docs for more detail.

Provisioning users and threads

create_user and create_thread (TypeScript: createUser/createThread, Go: CreateUser/CreateThread) are explicit provisioning helpers. Call them once — during onboarding, account creation, or before the first turn of a new conversation — before the agent runs. Each calls the Zep SDK’s create method directly. TypeScript and Go return the UUIDs that Zep generates: createUser returns { userUuid, graphUuid } and createThread returns { threadUuid, graphUuid }; CreateUser returns the UUID of the user and the UUID of the graph of the user, and CreateThread returns the UUID of the thread. Genuine failures (auth, network, 5xx) raise, so misconfiguration is caught immediately rather than silently swallowed.

Zep v4 addresses a user, a thread, and a graph by a server-generated UUID. A v4 create call takes no client-chosen identifier, thus a v4 resource has no name. In Python, create_user and create_thread accept no user_id and no thread_id. create_user returns the created User and create_thread returns the created Thread. Read user.uuid_, user.graph_uuid, and thread.uuid_, and store the UUIDs in your own database. Each Python call creates a new resource. The two Python helpers are not idempotent.

Two error philosophies apply, by design:

  • Provisioning fails loudly. create_user/create_thread raise on failures, so a misconfigured API key or network problem surfaces before the agent ever runs.
  • The turn path degrades gracefully. The callbacks and tools never raise a Zep error into the agent — failures are logged and the turn continues without Zep memory. If a persist call targets a user or thread that was never provisioned, the logged warning names create_user/create_thread (TypeScript: createUser/createThread).

Pass the user’s email to create_user (TypeScript: createUser) — the name and email on the Zep user profile are set at provisioning time, not through session state.

Identity and session state

Zep v4 addresses each user and each thread by a UUID that the server generates. Each language reads the Zep UUIDs as the following paragraphs describe. Zep’s knowledge graph is per-user, not per-thread — it accumulates knowledge across all of a user’s conversations, so when they start a new session they get context from everything Zep has learned about them.

In TypeScript, the integration takes the UUIDs that Zep generates. The construction options are userUuid and threadUuid, and the session-state keys are zep_user_uuid and zep_thread_uuid. The same precedence applies: the construction options take precedence over the session-state keys, which take precedence over the ADK userId/sessionId. The two ADK fields are used only when they hold Zep UUIDs. The integration makes no lookup call at run time, so the application resolves each UUID one time and stores it in its own database. The zep_first_name and zep_last_name session-state keys continue to set the author name on a persisted message.

In Go, the integration does not map ADK identifiers to Zep identifiers. Zep v4 addresses a user, a thread, and a graph by a UUID that the server generates, and an ADK user_id or session_id is not that UUID. The application supplies each UUID with WithUserUUID, WithThreadUUID, WithAfterThreadUUID, WithGraphUUID, and WithMemoryGraphUUID, or with the matching resolver option (WithUserUUIDResolver, WithThreadUUIDResolver, WithAfterThreadUUIDResolver, WithGraphUUIDResolver, and WithMemoryGraphUUIDResolver) when the UUID comes from session state or from your database. The integration makes no lookup call at run time. If no UUID is available, the callback or the tool logs an error and the turn continues without Zep memory. The zep_first_name and zep_last_name session-state keys continue to set the author name on a persisted message.

In Python, the keys hold UUIDs and they are not optional:

KeyRequiredDescription
zep_user_uuidYesThe UUID of the Zep user. The ADK user_id is the fallback, and it applies only when that field holds the Zep user UUID.
zep_thread_uuidYesThe UUID of the Zep thread. The ADK session_id is the fallback, and it applies only when that field holds the Zep thread UUID.
zep_graph_uuidNoThe UUID of the graph of the user. ZepGraphSearchTool and ZepMemoryService resolve it through user.get when it is absent.
zep_first_nameRecommendedUser’s first name.
zep_last_nameNoUser’s last name.

Advanced usage

Per-user setup

Configure per-user resources such as a custom ontology, custom extraction instructions, or user summary instructions one time, after the user is created. In Python, call your setup function directly after create_user returns. In TypeScript, pass the hook as onCreated; createUser creates a new user on each call, so the hook runs one time for that user and receives the UUID of the new user. Go has no hook: CreateUser creates a new user on each call, so the code that follows the call runs one time for that user.

from zep_cloud import CustomInstruction, EntityProperty, EntityType, UserInstruction
from zep_cloud.client import AsyncZep
from zep_adk import create_user
COMPANY = EntityType(
name="Company",
description="A company or organization the user is associated with.",
properties=[
EntityProperty(name="industry", description="The company's industry", type="text")
],
)
async def setup_user(zep_client: AsyncZep, user_uuid: str, graph_uuid: str) -> None:
"""Runs once, directly after the Zep user is created."""
# Set a custom ontology for this user's knowledge graph
await zep_client.graph.set_ontology(graph_uuid, entity_types=[COMPANY])
# Add custom extraction instructions
await zep_client.graph.set_instructions(
graph_uuid,
inherited=False,
instructions=[
CustomInstruction(
name="purchase_intent",
text="Extract product preferences and purchase intent.",
)
],
)
# Configure how user summaries are generated
await zep_client.user.set_summary_instructions(
user_uuid,
inherited=False,
instructions=[
UserInstruction(
name="work_focus",
text="Focus on the user's role, team, and active projects.",
)
],
)
user = await create_user(zep, first_name="Jane")
await setup_user(zep, user.uuid_, user.graph_uuid)

If the setup code raises an exception, the exception propagates. The user was still created, so a retry will not re-run the setup code. Keep the setup logic idempotent and re-run it directly to recover from a partial failure.

See custom ontology, custom instructions, and user summary instructions for details on each API.

Custom context builder

By default, the integration uses thread.add_messages(return_context=True) to persist the message and retrieve context in one call. The hook inserts this context into the system instruction, so use the default only for trusted application content.

For advanced scenarios — multi-graph searches, custom filtering, or combining multiple Zep API calls — you can provide a context builder: context_builder on ZepContextTool (Python), contextBuilder on createZepBeforeModelCallback, ZepContextTool, or createZepCallbacks (TypeScript), or WithContextBuilder on NewBeforeModelCallback (Go). The builder receives a single input object bundling everything it needs.

A custom builder does not change context placement. These hooks still insert the builder result into the system instruction. Use ZepMemoryService, load_memory, or ZepGraphSearchTool for end-user or third-party content.

import asyncio
from zep_adk import ZepContextTool, ContextInput
async def my_context_builder(ctx: ContextInput) -> str | None:
"""Custom context: combine user context with a targeted graph search."""
user = await ctx.zep.user.get(ctx.user_uuid)
user_context, edge_pager = await asyncio.gather(
ctx.zep.thread.get_context(ctx.thread_uuid),
ctx.zep.graph.search_edges(
user.graph_uuid,
query=ctx.user_message,
limit=10,
),
)
parts = []
if user_context and user_context.context:
parts.append(user_context.context)
facts = [e.fact for e in edge_pager.items or [] if e.fact]
if facts:
parts.append("Additional facts:\n" + "\n".join(f"- {f}" for f in facts))
return "\n\n".join(parts) if parts else None
tool = ZepContextTool(zep_client=zep, context_builder=my_context_builder)

When a builder is set, message persistence and context building run concurrently for lower latency, and each is isolated from the other’s failure: if the builder fails, a warning is logged and injection is skipped, but persistence still completes; if persistence fails, the turn is not marked as persisted (so it can be retried), but a successful builder result may still be injected. Return None (TypeScript: undefined) from the builder to skip injection for that turn without affecting persistence.

The Python type signature (importable from zep_adk):

ContextBuilder = Callable[[ContextInput], Awaitable[str | None]]
# ContextInput is a frozen dataclass with fields:
# zep — the AsyncZep client
# user_uuid — resolved Zep user UUID
# thread_uuid — resolved Zep thread UUID
# user_message — the user's latest message text
# tool_context — ADK session state / invocation metadata
# llm_request — the outgoing model request

TypeScript exports the equivalent ContextBuilder and ContextBuilderInput types; Go’s builder is func(ctx context.Context, in zepadk.ContextInput) (string, error).

See advanced context block construction and context templates for more on assembling custom context.

Injection template

The retrieved (or built) context block is wrapped in a template before it is injected into the system instruction. The default — DEFAULT_CONTEXT_TEMPLATE (Python and TypeScript) or DefaultContextTemplate (Go) — introduces the context and wraps it in <ZEP_CONTEXT> tags; the wording is identical across all three languages. Override it with context_template / contextTemplate / WithContextTemplate:

from zep_adk import ZepContextTool
tool = ZepContextTool(
zep_client=zep,
context_template="Relevant memory:\n{context}",
)

The template must contain a literal {context} placeholder. Plain string replacement prevents format-string interpretation of {, }, %, or $. It does not prevent the model from following instructions in the retrieved content. In Go, WithContextPrefix is deprecated in favor of WithContextTemplate.

Graph search tool

Use ZepGraphSearchTool (Go: NewGraphSearchTool) for end-user or third-party context. This model-callable tool preserves retrieval as an actual tool call.

from zep_adk import ZepGraphSearchTool, create_after_model_callback
agent = Agent(
name="my_agent",
model="gemini-3.7-flash",
instruction="...",
tools=[
ZepGraphSearchTool(zep_client=zep),
],
after_model_callback=create_after_model_callback(zep_client=zep),
)

In Python the tool resolves the user identity from session state, so the model only needs to provide a search query. In TypeScript the application gives the tool a graph UUID with graphUuid; when graphUuid is omitted, the tool reads the graph of the resolved userUuid one time with zep.user.get and caches it, and the model only provides a search query. In Go the application gives the tool a graph UUID with WithGraphUUID or WithGraphUUIDResolver, and the model only provides a search query. Unless pinned, the model can also choose the scope (edges, nodes, episodes, observations, thread_summaries, auto), the reranker (rrf, mmr, node_distance, episode_mentions, cross_encoder), limit, mmr_lambda, and center_node_uuid — see search parameters.

Pinning and hiding parameters

Every search parameter is independently in one of three states at construction time:

StateHow to set itEffect
Exposed (default)Omit the parameterAppears in the model’s tool schema with the default below; the model chooses a value per call.
PinnedPass a concrete value (e.g. scope="edges"; Go: WithToolSearchScope, WithToolReranker, …)Hidden from the model’s tool schema. Always used, even if the model would have chosen differently.
HiddenPass None/null (Go: WithHiddenParams)Hidden from the model’s tool schema and omitted from the search call entirely.

Defaults when exposed: scope="edges", reranker="rrf", limit=10; mmr_lambda and center_node_uuid have no default and are omitted unless the model supplies one. search_filters and bfs_origin_node_uuids (TypeScript: searchFilters/bfsOriginNodeUuids, Go: WithToolSearchFilters/WithToolBFSOriginNodeUUIDs) are always constructor-only — never exposed to the model, always applied to every search when set.

# Pin reranker and limit; hide mmr_lambda and center_node_uuid;
# the model still chooses scope per call.
ZepGraphSearchTool(
zep_client=zep,
reranker="cross_encoder", # pinned — hidden from the model
limit=5, # pinned
mmr_lambda=None, # hidden — omitted from every search
center_node_uuid=None, # hidden
search_filters={"node_labels": ["Person"]}, # constructor-only
bfs_origin_node_uuids=["node-uuid-1"], # constructor-only — seed BFS traversal
)

An invalid enum value sent by the model never reaches Zep and never crashes the agent: TypeScript falls back to the documented default and logs a warning; Go rejects it through ADK’s schema validation and surfaces a tool error the model can correct on its next call.

Shared documentation graph

To search a fixed graph that all users share (e.g. a documentation knowledge base), pass the UUID of that graph as graph_uuid (TypeScript: graphUuid, Go: WithGraphUUID). The tool will search that graph instead of the current user’s personal graph. Use distinct name and description values when combining multiple instances:

agent = Agent(
name="my_agent",
model="gemini-3.7-flash",
instruction="...",
tools=[
ZepGraphSearchTool(
zep_client=zep,
name="search_user_memory",
description="Search the user's knowledge graph for information from previous conversations, known facts, or general context about the user.",
),
ZepGraphSearchTool(
zep_client=zep,
name="search_docs",
description="Search the shared documentation knowledge base.",
graph_uuid="6f1a9a9c-2c3e-4d1a-9e4b-2b1a7a1f7c31",
),
],
after_model_callback=create_after_model_callback(zep_client=zep),
)

The model sees two distinct tools and chooses which to call based on the user’s query.

Memory service

All three packages implement ADK’s native memory extension point: ZepMemoryService (Python and TypeScript) implements BaseMemoryService, and NewMemoryService (Go) returns an ADK memory.Service. Registered on the Runner, it lets ADK’s built-in load_memory/preload_memory tools (Go: ToolContext.SearchMemory) search the calling user’s Zep graph whenever the model decides memory is relevant.

The two extension points have different security properties. ZepContextTool and the before-model callback add context to the system instruction. Use them only for fully trusted content. The memory service is model-initiated and returns an actual tool result. Use the memory service for end-user or third-party content.

from google.adk.agents import Agent
from google.adk.runners import Runner
from google.adk.tools import load_memory
from zep_cloud.client import AsyncZep
from zep_adk import ZepMemoryService
zep = AsyncZep(api_key=os.getenv("ZEP_API_KEY"))
agent = Agent(
name="my_agent",
model="gemini-3.7-flash",
instruction="You are a helpful assistant. Use load_memory to recall prior context when relevant.",
tools=[load_memory],
)
runner = Runner(
agent=agent,
app_name="my_app",
session_service=session_service,
memory_service=ZepMemoryService(zep=zep),
)

Each memory search runs the v4 search method of the configured scope (for example graph.search_edges) against the graph of the calling user. TypeScript and Go search the graph UUID that you configure, and TypeScript reads the graph of the session’s userUuid when graphUuid is omitted. The scope is configurable, and the same six scopes as the graph search tool apply (edges, nodes, episodes, observations, thread_summaries, auto). The service maps each result into an ADK memory entry. A Zep failure is logged and returns an empty result rather than raising into the agent, so a memory lookup can never break a turn.

add_session_to_memory (TypeScript: addSessionToMemory) is a deliberate no-op: Zep already ingests each turn live via the context tool and after-model callback, so flushing the full session again would persist the same conversation into the graph twice.

In TypeScript, the memory service requires the full Runner, not InMemoryRunner — only Runner’s RunnerConfig accepts a memoryService option. Wiring Runner directly means providing a sessionService yourself; an InMemorySessionService works for development.

Backfill strategy for existing users

If you have existing users with conversation history, you can backfill their data into Zep so they get rich context from day one. Use direct thread.add_messages calls for small or session-scale imports. For large historical imports, use the Batch API with thread_message items so Zep can process the backfill as an asynchronous job.

ID matching

Record the Zep UUIDs against your ADK identifiers. Zep generates a UUID for each user and each thread, so the backfill script cannot reuse your ADK identifiers. The script records each UUID against your own identifier:

  • User UUIDs must be stored against the user_id that you pass to ADK’s create_session(). Put the stored UUID in zep_user_uuid (TypeScript: the userUuid option or zep_user_uuid). This links live sessions to the correct knowledge graph. If a live session uses a different user UUID, the backfilled history is orphaned.
  • Thread UUIDs must be stored against the ADK session_id for each conversation. If a user continues an existing session after cutover, put the stored UUID in zep_thread_uuid. If a live session uses a different thread UUID, the conversation history is split — the continued thread won’t see the backfilled messages in its thread context.

Example small backfill script

This runs outside of ADK as a standalone script using the Zep Python SDK directly, with zep-adk’s provisioning helpers. The script creates one Zep user and one Zep thread for each source conversation, and it returns the UUIDs for your own database. It keeps each thread.add_messages call within Zep’s limits: at most 30 messages per request, and below the 4,096-character hard limit per message. The sample uses a 4,000-character safety margin, matching the other integration examples. Map source-system roles to Zep’s canonical roles before sending: user, assistant, system, function, tool, or user.

import asyncio
from zep_cloud.client import AsyncZep
from zep_cloud import AddMessage
from zep_adk import create_user, create_thread
zep = AsyncZep(api_key="your-zep-api-key")
MAX_MESSAGES_PER_CALL = 30
MAX_MESSAGE_CHARS = 4000
def truncate_for_zep(content: str) -> str:
return content[:MAX_MESSAGE_CHARS]
async def backfill_user(
user_id: str, # your own identifier, kept only for your records
first_name: str,
last_name: str,
conversations: list[dict], # list of {session_id, messages} dicts
) -> dict:
# 1. Create the user. Zep generates the UUID; store it in your database.
user = await create_user(zep, first_name=first_name, last_name=last_name)
# 2. Load each conversation into its own Zep thread.
thread_uuids = {}
for convo in conversations:
thread = await create_thread(zep, user_uuid=user.uuid_)
thread_uuids[convo["session_id"]] = thread.uuid_
messages = [
AddMessage(
role=msg["role"],
content=truncate_for_zep(msg["content"]),
name=f"{first_name} {last_name}" if msg["role"] == "user" else "Assistant",
)
for msg in convo["messages"]
]
for start in range(0, len(messages), MAX_MESSAGES_PER_CALL):
await zep.thread.add_messages(
thread.uuid_,
messages=messages[start : start + MAX_MESSAGES_PER_CALL],
)
print(f"Backfilled {len(conversations)} conversations for {user_id}")
return {"user_uuid": user.uuid_, "thread_uuids": thread_uuids}
async def main():
users = [
{
"user_id": "user-123", # same ID used in ADK sessions
"first_name": "Jane",
"last_name": "Smith",
"conversations": [
{
"session_id": "session-abc", # original ADK session ID
"messages": [
{"role": "user", "content": "I need help with my account settings."},
{"role": "assistant", "content": "I can help. What would you like to change?"},
{"role": "user", "content": "I want to enable two-factor authentication."},
{"role": "assistant", "content": "Go to Settings > Security > 2FA to enable it."},
],
},
],
},
]
for user in users:
await backfill_user(**user)
asyncio.run(main())

After backfilling, allow time for Zep to process the messages and build knowledge graphs. Zep processes messages asynchronously — the graph won’t be available instantly. For large backfills, prefer the Batch API over manual sleeps and direct SDK loops.

Transition gap

Messages created between the backfill and deployment are not in Zep. Use a dual-write period when you must preserve all thread messages.

After the backfill, write new messages to both Zep and the existing system until deployment completes.

Cutover checklist

  1. Run the backfill script.
  2. For trusted context, add ZepContextTool and the after-model callback.
  3. For untrusted context, add a model-callable graph search tool and the after-model callback.
  4. Include the Zep UUIDs, zep_first_name, and zep_last_name in create_session() calls.
  5. Call create_user and create_thread before the first turn, and store the returned UUIDs.
  6. Deploy the updated agent.

Next steps