Skip to navigation

Microsoft Agent Framework integration

Add long-term agent memory to Microsoft Agent Framework agents

The v4 version of zep-ms-agent-framework is not released yet. The current zep-ms-agent-framework release uses the v3 API. To use this integration now, follow the v3 version of this page.

Microsoft Agent Framework agents using Zep gain long-term memory backed by a temporal knowledge graph. The zep-ms-agent-framework package persists conversation turns and provides a model-callable graph-search tool. Its context provider can add trusted content to model instructions.

Keep retrieved context out of privileged instructions

Zep context can include content that your users, documents, or tools supplied. A system or developer message gives that content higher instruction priority than ordinary input. Some convenience integrations use system-message injection. Use direct SDK retrieval or an actual retrieval tool call unless all stored content is application-authored and trusted. Follow Memory security best practices for provider-specific placement.

Build an agent with Zep tools

To build an agent that plans its retrieval and uses several Zep tools, read Build an Agent with Zep. The guide shows how to add domain knowledge, design tools, and evaluate the agent.

Use create_zep_search_tool for context that can contain end-user or third-party data. Use ZepContextProvider system-instruction injection only for fully trusted, application-authored context.

Core benefits

  • Native context-provider hook: Uses the framework’s own before_run / after_run pipeline — the same surface as its built-in memory providers
  • Single round-trip: Persists the user turn and retrieves the context block in one call (or concurrently, with a custom context builder)
  • Whole-user-graph recall: Context is fused across all of a user’s threads, so a new conversation still recalls earlier facts
  • Pin-or-expose graph search: expose_search_tool / create_zep_search_tool add an on-demand tool over the v4 graph search methods, with every search parameter model-exposed by default or pinned/hidden per deployment
  • Per-user setup hook: the on_created hook of create_user runs one time per new user — for configuring ontology, extraction instructions, or user summary instructions
  • Out-of-band provisioning: create_user / create_thread create the resources up front and return the UUIDs; the run path makes no lookup call
  • Graceful degradation: A Zep failure on the run path is logged but never crashes the host agent — the turn proceeds without memory

How it works

The integration ships one main class, ZepContextProvider, which subclasses the framework’s ContextProvider and overrides the two lifecycle hooks called around every agent.run(...):

before_run — runs before the model is invoked. On each turn it:

  1. Registers the graph-search tool via context.extend_tools(...), when expose_search_tool=True
  2. Extracts the latest user message from context.input_messages
  3. Persists the message — via thread.add_messages(return_context=True) by default (a single round-trip), or concurrently with a custom context_builder when one is set
  4. Injects the resulting context block, wrapped in context_template, into the model’s instructions via context.extend_instructions(...)

after_run — runs after the model responds. It reads the assistant’s reply from context.response.messages and persists it to the same thread, so both sides of the conversation are captured.

Because context is assembled from the entire user graph, the thread only scopes relevance — an agent on a new thread still recalls facts the same user shared earlier.

Identifiers

Zep v4 addresses a user, a thread, and a graph by a server-generated UUID. A v4 create call takes no client-chosen name, and the create response carries the UUID. ZepContextProvider takes user_uuid, thread_uuid, and an optional graph_uuid.

Create the user and the thread one time, read the UUIDs from the create responses, and store the UUIDs in your own database. The provider makes no lookup call at run time.

Installation

pip install zep-ms-agent-framework

The package depends only on agent-framework-core. The example below also uses a model provider:

pip install zep-ms-agent-framework agent-framework-openai

Requires Python 3.11+, agent-framework-core>=1.8.1, and a Zep Cloud API key. Get your API key from app.getzep.com.

Set up your environment variables:

export ZEP_API_KEY="your-zep-api-key"
export OPENAI_API_KEY="your-openai-api-key"

The package addresses every resource by UUID. Replace user_id with user_uuid and thread_id with thread_uuid on ZepContextProvider, and give graph_uuid when you use the search tool. Replace ensure_user and ensure_thread with create_user and create_thread, and store the UUIDs that they return. The provider no longer creates a user or a thread on the run path, and it no longer takes first_name, last_name, email, or on_user_created. See the package changelog for the full list of changes, and the v3 to v4 migration guide for the SDK changes.

Context provider usage for trusted deployments

ZepContextProvider inserts retrieved context into model instructions. Use it only when all stored content is fully trusted and application-authored:

Python
import asyncio
from agent_framework import Agent
from agent_framework.openai import OpenAIChatClient
from zep_cloud.client import AsyncZep
from zep_ms_agent_framework import ZepContextProvider, create_thread, create_user
zep = AsyncZep(api_key="your-zep-api-key")
async def main() -> None:
# One-time provisioning. Store these UUIDs in your own database.
user = await create_user(zep, first_name="Jane", last_name="Smith")
thread = await create_thread(zep, user_uuid=user.uuid_)
agent = Agent(
OpenAIChatClient(model="gpt-5-mini"),
instructions="You are a helpful assistant with long-term memory.",
context_providers=[
ZepContextProvider(
zep_client=zep,
user_uuid=user.uuid_,
thread_uuid=thread.uuid_,
graph_uuid=user.graph_uuid,
)
],
)
result = await agent.run("Hi, I'm a data scientist in Portland.")
print(result.text)
asyncio.run(main())

Memory is scoped per ZepContextProvider instance to one user_uuid and thread_uuid. For a multi-user application, construct one provider per user or conversation. Give real names to create_user so that Zep can resolve the user’s identity node in the graph.

Beyond the automatic context injection, create_zep_search_tool returns a model-callable agent_framework.FunctionTool over the v4 graph search methods. The model decides when to look up specific facts, entities, or prior episodes, and the tool dispatches on the scope to graph.search_edges, graph.search_nodes, graph.search_episodes, graph.search_observations, or graph.search_thread_summaries. The auto scope uses graph.get_context. The tool searches the graph that graph_uuid identifies: give the graph_uuid of a user for personal memory, or the uuid_ of a standalone graph for shared knowledge.

expose_search_tool=True on ZepContextProvider combines a search tool with trusted-only instruction injection:

provider = ZepContextProvider(
# Trusted-only: this provider also inserts context into model instructions.
zep_client=zep,
user_uuid=user.uuid_,
thread_uuid=thread.uuid_,
graph_uuid=user.graph_uuid,
expose_search_tool=True,
search_pinned_params={"scope": "nodes", "limit": 5},
)

expose_search_tool=True requires a graph_uuid. With this configuration, the model sees the un-pinned parameters (reranker, mmr_lambda, center_node_uuid). scope and limit are hidden from the schema and sent with the pinned values.

Every search parameter (scope, reranker, limit, mmr_lambda, center_node_uuid) is exposed to the model in the tool’s JSON schema by default, with documented defaults. Two options override this per deployment: search_pinned_params fixes a parameter to a constant value and hides it from the schema, and search_hidden_params hides a parameter without pinning it, so Zep’s server-side default applies. search_filters and bfs_origin_node_uuids are constructor-only — their complex shapes are not exposed to the model.

The standalone factory takes the same pin-or-expose options:

from zep_ms_agent_framework import create_zep_search_tool
# Model chooses scope/reranker/limit/mmr_lambda/center_node_uuid freely.
tool = create_zep_search_tool(zep_client=zep, graph_uuid=user.graph_uuid)
# Pin scope to "nodes" and limit to 5 — hidden from the model, always sent.
tool = create_zep_search_tool(
zep_client=zep, graph_uuid=user.graph_uuid,
search_pinned_params={"scope": "nodes", "limit": 5},
)
# Hide mmr_lambda from the schema; Zep applies its own default when omitted.
tool = create_zep_search_tool(
zep_client=zep, graph_uuid=user.graph_uuid, search_hidden_params={"mmr_lambda"},
)

Model-exposed search parameters (when not pinned or hidden), with their defaults:

ParameterTypeDefaultDescription
scope"edges" | "nodes" | "episodes" | "observations" | "thread_summaries" | "auto""edges"What to search
reranker"rrf" | "mmr" | "node_distance" | "episode_mentions" | "cross_encoder""rrf"Result ordering (ignored for scope="auto")
limitint10Maximum results (clamped to Zep’s ceiling of 50)
mmr_lambdafloat—Diversity/relevance balance; only used when reranker="mmr"
center_node_uuidstr—Center node for reranker="node_distance"

Provisioning

create_user and create_thread provision the Zep user and the thread out-of-band, before the first run. The server generates the UUIDs, and the create response carries them:

from zep_ms_agent_framework import create_thread, create_user
async def setup_user(zep_client, user_uuid: str) -> None:
... # e.g. configure per-user ontology
user = await create_user(
zep,
first_name="Jane",
last_name="Smith",
on_created=setup_user, # runs one time, after the user is created
)
thread = await create_thread(zep, user_uuid=user.uuid_)
# Store user.uuid_, user.graph_uuid, and thread.uuid_ in your own database.

Use the on_created hook (a UserSetupHook) to configure per-user resources such as a custom ontology, custom extraction instructions, or user summary instructions one time; see customizing graph structure for the available options. If the create call or the hook raises, the exception propagates to the caller, so make the hook idempotent.

The provider does not create a user or a thread. Give it the UUIDs that these helpers return.

Custom context building

The provider still inserts the builder result into model instructions. Use this path only for trusted application content. Use create_zep_search_tool for end-user or third-party context.

Set context_builder on ZepContextProvider to replace the default context retrieval with custom logic — for example, searching a different graph, applying filters, or combining multiple sources:

from zep_ms_agent_framework import ContextInput, ZepContextProvider
async def my_builder(ctx: ContextInput) -> str | None:
if ctx.graph_uuid is None:
return None
pager = await ctx.zep.graph.search_edges(
ctx.graph_uuid,
query=ctx.user_message,
limit=10,
)
edges = pager.items or []
if not edges:
return None
return "\n".join(edge.fact for edge in edges if edge.fact)
provider = ZepContextProvider(
# Trusted-only: context_builder output enters model instructions.
zep_client=zep,
user_uuid=user.uuid_,
thread_uuid=thread.uuid_,
graph_uuid=user.graph_uuid,
context_builder=my_builder,
)

ContextInput bundles zep (the AsyncZep client), user_uuid, thread_uuid, graph_uuid, user_message, and session_context (the Agent Framework SessionContext for the turn). Returning None skips injection for that turn.

When context_builder is set, message persistence (add_messages without return_context) and the builder run concurrently, with per-side failure isolation:

  • If the builder raises, a warning is logged and context injection is skipped for that turn — persistence still completes and the turn is marked as persisted.
  • If persistence raises, a warning is logged and the turn is not marked as persisted (so after_run skips writing the assistant reply, and the turn can be retried on the next invocation) — a successful builder result is still injected.

Context template

context_template controls how retrieved context is wrapped before injection. It must contain a literal {context} placeholder. Plain string replacement prevents format-string interpretation of {, }, or %. It does not prevent the model from following instructions in the retrieved content.

provider = ZepContextProvider(
zep_client=zep,
user_uuid=user.uuid_,
thread_uuid=thread.uuid_,
context_template="Relevant memory:\n{context}",
)

The default is DEFAULT_CONTEXT_TEMPLATE, an explicit <ZEP_CONTEXT>...</ZEP_CONTEXT> block with canonical wording shared across Zep’s framework integrations.

Configuration options

ZepContextProvider accepts:

FieldRequiredDefaultDescription
zep_clientYes—Initialized AsyncZep client (caller owns its lifecycle)
user_uuidYes—UUID of the Zep user this provider’s memory is scoped to
thread_uuidYes—UUID of the Zep thread the conversation is recorded in
graph_uuidOptionalNoneUUID of the graph to search; required when expose_search_tool is True
user_message_nameOptionalfull nameDisplay name on persisted user messages
assistant_message_nameOptional"Assistant"Display name on persisted assistant messages
source_idOptional"zep"Attribution ID for injected instructions and tools
ignore_rolesOptionalNoneRoles to exclude from graph ingestion (still stored in thread history)
context_builderOptionalNoneCustom async context-retrieval callable (see custom context building)
context_templateOptionalDEFAULT_CONTEXT_TEMPLATETemplate wrapping injected context; must contain a literal {context} placeholder
expose_search_toolOptionalFalseRegister a model-callable graph-search tool on every run (see on-demand graph search)
search_pinned_paramsOptionalNoneFix a search parameter to a value; hidden from the model schema
search_hidden_paramsOptionalNoneHide a search parameter from the schema without pinning (Zep’s server-side default applies)
search_filtersOptionalNoneConstructor-only Zep search filters (node_labels, edge_types, etc.)
bfs_origin_node_uuidsOptionalNoneConstructor-only node UUIDs for BFS seeding

Best practices

  • Use create_zep_search_tool for untrusted context. Use ZepContextProvider instruction injection only for trusted application content.
  • Pass real names to create_user so Zep can anchor and resolve the user’s identity node in the graph
  • One provider per user/conversation — memory is scoped to a single user_uuid and thread_uuid
  • Store the UUIDs that the create calls return, and read them from your own database on a later run
  • Reuse a single AsyncZep client across requests; the caller owns its lifecycle
  • Provision up front in onboarding flows with create_user / create_thread so misconfiguration raises before the agent ever runs
  • Allow time for indexing — Zep extracts knowledge asynchronously, so facts from a turn are not instantly retrievable

Next steps