Skip to navigation

Mastra integration

Add long-term agent memory to Mastra agents with processors and tools

The v4 version of @getzep/zep-mastra is not released yet. The current @getzep/zep-mastra release uses the v3 API. To use this integration now, follow the v3 version of this page.

Keep retrieved context out of privileged instructions

Zep context can include content that your users, documents, or tools supplied. A system or developer message gives that content higher instruction priority than ordinary input. Some convenience integrations use system-message injection. Use direct SDK retrieval or an actual retrieval tool call unless all stored content is application-authored and trusted. Follow Memory security best practices for provider-specific placement.

Build an agent with Zep tools

To build an agent that plans its retrieval and uses several Zep tools, read Build an Agent with Zep. The guide shows how to add domain knowledge, design tools, and evaluate the agent.

Mastra agents using Zep gain long-term memory backed by a temporal knowledge graph. The @getzep/zep-mastra package provides two complementary surfaces:

  • Automatic memory for trusted content — createZepProcessors builds a ZepInputProcessor/ZepOutputProcessor pair that plugs into Mastra’s native inputProcessors/outputProcessors pipeline. The input processor inserts context into a system message.
  • Tools — createZepToolset builds zepRemember/zepSearch/zepContext tools that let the model decide when to persist or recall. Use retrieval tools for end-user or third-party context.

Core benefits

  • Automatic memory loop for trusted content: Processors can inject application-authored context and persist each completed turn
  • Model-in-the-loop tools: zepRemember, zepSearch, and zepContext drop straight into an Agent’s tools record
  • Per-call identity: Resolve graphUuid/threadUuid from Mastra’s requestContext so one processor or tool instance serves many end users
  • User and shared Context Graphs: Bind to a user’s personal graph or a shared knowledge base
  • Graceful degradation: A Zep outage is logged and surfaced as a non-fatal result — it never crashes the host agent

How it works

The processors sit on opposite sides of the model call:

  1. ZepInputProcessor runs before the model is called. It extracts the latest user message, retrieves a Zep context block (thread.getContext, or a custom contextBuilder), wraps it with contextTemplate/formatContext, and injects it as a system message.
  2. ZepOutputProcessor runs after the model responds. It persists the completed turn — the latest user message plus the assistant’s response — to the bound thread via a single thread.addMessages call. The assistant text persisted is the final step’s text; when generation ends mid-tool-loop (finishReason === "tool-calls"), the user message is still persisted.

Because injection and persistence sit on opposite sides of the model call, it’s safe to enable both processors together — they don’t interfere with each other. Every Zep call is wrapped: a missing threadUuid or any Zep failure degrades gracefully — messages pass through unchanged, a warning is logged — and the input processor never calls abort() or throws into the agent loop.

Zep is a temporal knowledge graph, not a row-oriented message store, so the package exposes Zep’s two real operations — persist and retrieve — through processors and tools rather than a MastraStorage adapter, which would require CRUD operations a temporal knowledge graph can’t honor faithfully.

Installation

npm install @getzep/zep-mastra @getzep/zep-cloud@preview @mastra/core

Requires Node.js 20+, @mastra/core>=1.42.0 (peer), @getzep/zep-cloud, and a Zep Cloud API key. Get your API key from app.getzep.com.

Set up your environment variables:

export ZEP_API_KEY="your-zep-api-key"
export OPENAI_API_KEY="your-openai-api-key"

Identifiers

Zep v4 addresses every user, thread, and graph by a server-generated UUID. A create call accepts no client-chosen identifier, so a new user, thread, or graph has no name.

  • createZepUserAndThread returns userUuid, graphUuid, and threadUuid. Store all three in your own database.
  • The processors take graphUuid and threadUuid. The tools take them on a ZepBinding.
  • The integration never calls lookup at run time. Resolve a v3 name to its UUID one time, then reuse the UUID.

The package now targets the Zep v4 SDK, and the public API takes UUIDs:

  • ensureZepUserAndThread is replaced by createZepUserAndThread, which returns { userUuid, graphUuid, threadUuid } or null. Zep v4 has no name-addressed create, so the function is not idempotent on a name.
  • ZepBinding takes graphUuid in place of userId/graphId, and ZepThreadBinding takes threadUuid in place of threadId.
  • The processors take graphUuid and threadUuid, and ResolvedZepIdentity returns the same two fields.
  • templateId becomes templateUuid, and searchFilters becomes filters.

See the CHANGELOG for the full migration notes.

Automatic memory for trusted deployments

createZepProcessors inserts retrieved context into a system message. Use this processor pair only when all stored content is fully trusted and application-authored:

import { ZepClient } from "@getzep/zep-cloud";
import { Agent } from "@mastra/core/agent";
import { createZepProcessors, createZepUserAndThread } from "@getzep/zep-mastra";
const client = new ZepClient({ apiKey: process.env.ZEP_API_KEY! });
// 1. Create the Zep user + thread before the first turn. Zep generates the UUIDs.
const identity = await createZepUserAndThread({
client,
firstName: "Jane",
lastName: "Smith",
});
if (!identity) throw new Error("Could not create the Zep user and thread.");
// 2. Trusted-only path: the input processor inserts context into a system message.
const { inputProcessor, outputProcessor } = createZepProcessors({
client,
graphUuid: identity.graphUuid,
threadUuid: identity.threadUuid,
});
// 3. Attach to a Mastra agent (id and name are both required by Mastra).
const agent = new Agent({
id: "memory-agent",
name: "Memory Agent",
instructions: "You have long-term memory about the user. Use it to personalize replies.",
model: "openai/gpt-5.6-terra",
inputProcessors: [inputProcessor],
outputProcessors: [outputProcessor],
});

Customizing context injection

By default the input processor retrieves the context block with thread.getContext and wraps it in DEFAULT_CONTEXT_TEMPLATE — the canonical <ZEP_CONTEXT> wrapper shared by Zep’s framework integrations. Three options change that, in increasing order of control:

const { inputProcessor, outputProcessor } = createZepProcessors({
client,
graphUuid,
threadUuid,
// Replace thread.getContext with your own retrieval:
contextBuilder: async ({ client, graphUuid, userMessage }) => {
if (!graphUuid) return undefined;
const result = await client.graph.searchEdges(graphUuid, {
body: { query: userMessage },
});
return result.data.map((e) => e.fact).join("\n");
},
// Or just customize the wrapping template (must contain a literal `{context}`):
contextTemplate: "Known facts about the user:\n{context}",
// Or fully take over formatting (wins over contextTemplate):
formatContext: (context) => `<memory>${context}</memory>`,
});
  • contextBuilder replaces the default retrieval with your own async function. It receives a ZepContextBuilderInput — the client, the resolved graphUuid/threadUuid, and the latest user message — and returns the context string (or undefined to inject nothing for that turn). The result still passes through the template or formatContext.
  • contextTemplate customizes the wrapping text. It must contain a literal {context} placeholder, replaced via literal string replacement (not a format string), so braces, %, and $ in the retrieved context are safe.
  • formatContext takes over formatting entirely and wins over contextTemplate.

Per-call identity

Pass resolveIdentity (a ZepIdentityResolver, sync or async) to resolve graphUuid/threadUuid per call from Mastra’s requestContext instead of binding a fixed identity at construction time — useful when a single processor instance serves many end users:

const { inputProcessor, outputProcessor } = createZepProcessors({
client,
resolveIdentity: (requestContext) => ({
graphUuid: (requestContext as { graphUuid?: string } | undefined)?.graphUuid,
threadUuid: (requestContext as { threadUuid?: string } | undefined)?.threadUuid,
}),
});

The same resolveIdentity option is accepted by createZepSearchTool, createZepRememberTool, and createZepContextTool (resolved from each tool call’s context.requestContext), and createZepToolset forwards it to all three tools.

If both a fixed graphUuid/threadUuid and resolveIdentity are set, resolveIdentity’s result wins for whichever fields it returns; any field it omits or resolves to undefined falls back to the constructor-bound value.

Provisioning with createZepUserAndThread

Zep requires the user and thread to exist before messages are added. Call createZepUserAndThread once, out-of-band, before the first turn, then store the returned UUIDs. Zep v4 addresses every later call by UUID, so this function is not idempotent on a name: a second call creates a second user. A failure (auth, network, 5xx) is logged at warn and reported as null, never thrown.

Pass onUserCreated (a ZepUserCreatedHook) to run one-time setup — per-user ontology, custom instructions, seeding — immediately after the user is created. The hook receives the userUuid:

const identity = await createZepUserAndThread({
client,
firstName: "Jane",
lastName: "Smith",
// Runs exactly once, immediately after the user is created —
// e.g. configure per-user summary instructions:
onUserCreated: async (client, userUuid) => {
await client.user.setSummaryInstructions(userUuid, {
instructions: [{ name: "diet", text: "Track the user's dietary preferences." }],
});
},
});
// identity: { userUuid, graphUuid, threadUuid } | null

The function sends no client-chosen identifier. The v4 create endpoints reject a userId or a threadId, and the returned UUIDs are the only addresses.

Tools

The toolset puts the model in the loop: the agent calls a tool when it decides to persist or recall. Use it standalone or alongside the processors.

import { ZepClient } from "@getzep/zep-cloud";
import { Agent } from "@mastra/core/agent";
import { createZepToolset, createZepUserAndThread } from "@getzep/zep-mastra";
const client = new ZepClient({ apiKey: process.env.ZEP_API_KEY! });
// 1. Create the Zep user + thread before the first turn.
const identity = await createZepUserAndThread({ client, firstName: "Jane", lastName: "Smith" });
if (!identity) throw new Error("Could not create the Zep user and thread.");
// 2. Build the tool set bound to that graph + thread.
const binding = { graphUuid: identity.graphUuid, threadUuid: identity.threadUuid };
const { zepRemember, zepSearch, zepContext } = createZepToolset({ client, binding });
// 3. Attach the tools to an Agent (id and name are both required by Mastra).
const agent = new Agent({
id: "memory-agent",
name: "Memory Agent",
instructions: "You have long-term memory. Store and recall user facts.",
model: "openai/gpt-5.6-terra",
tools: { zepRemember, zepSearch, zepContext },
});

The toolset provides three tools:

ToolZep operationWhat it does
zepRememberthread.addMessages / graph.episode.addPersists a message via thread.addMessages only when a role and a threadUuid are both present; otherwise the content is ingested as a fact via graph.episode.add
zepSearchgraph.searchEdges and the other scope methodsModel-callable search over the bound graph; each search parameter can be exposed to the model, pinned, or hidden
zepContextthread.getContextReturns the prompt-ready context block assembled from the whole user graph

Each tool is also exported as a standalone factory (createZepRememberTool, createZepSearchTool, createZepContextTool) for wiring a single tool with custom options.

Each tool has a typed input and output schema:

ToolInputOutput
zepRemembercontent (string); optional role; optional name{ stored: boolean, message: string }
zepSearchquery (string, 1–400 chars); optional scope, reranker, limit, mmrLambda, centerNodeUuid unless pinned or hidden{ facts: string[], found: boolean }
zepContextnone{ context: string, found: boolean }

zepSearch returns facts as extracted strings tailored to the search scope — edge facts, "name: summary" for entities, episode content, and so on — with found set to true when the result is non-empty.

Pin-or-expose search parameters

createZepSearchTool exposes each search parameter to the model by default, alongside the always-required query, so the model can tune its own searches per call:

ParameterModel-visible valuesDefault
scopeedges, nodes, episodes, observations, thread_summaries, autoedges
rerankerrrf, mmr, node_distance, episode_mentions, cross_encoderrrf
limitnumber (values above 50 are clamped to 50)10
mmrLambdanumberZep server default
centerNodeUuidstringZep server default

Each parameter is independently tri-state at construction time (ZepSearchPinnableParams):

  • pinnedParams fixes a parameter to a constant value: hidden from the model’s schema, always sent.
  • hiddenParams removes a parameter from the schema without pinning it: omitted from the search call entirely, so Zep’s own server default applies.
  • Omitted from both — exposed to the model with the documented default.
// Model only ever sees `query`; scope/reranker/limit are fixed.
createZepSearchTool({
client,
binding: { graphUuid },
pinnedParams: { scope: "edges", reranker: "rrf", limit: 10 },
});
// Hide mmrLambda/centerNodeUuid from the schema without fixing a value.
createZepSearchTool({
client,
binding: { graphUuid },
hiddenParams: new Set(["mmrLambda", "centerNodeUuid"]),
});

filters and bfsOriginNodeUuids are always constructor-only — never exposed to the model — and applied whenever set. The legacy scope/reranker/limit constructor arguments pin (and hide) their parameter, equivalent to the corresponding pinnedParams entry.

Binding: user graph vs shared Context Graph

Zep v4 addresses every graph by its server-generated UUID, so a user graph and a standalone graph are the same kind of address. Tools and processors are bound with graphUuid, and the thread-scoped surfaces are bound with threadUuid:

  • The graphUuid of a user graph is the graphUuid field of the User that user.create returns. A user graph is the home for personalized agent memory.
  • The graphUuid of a shared Context Graph is the uuid field of the Graph that graph.create returns. A shared Context Graph holds shared or domain knowledge such as a product knowledge base or runbooks. It has no user node and no user summary.
  • The threadUuid is the uuid field of the Thread that thread.create returns. Context retrieval and the zepContext tool need it. The thread scopes relevance; retrieval still spans the whole user graph.

If no graphUuid or no threadUuid can be resolved, tools and processors degrade gracefully instead of throwing.

Roles

zepRemember accepts an arbitrary role string and maps it onto Zep’s closed RoleType enum: user, assistant, system, tool, or function. Host-framework role names like human or ai are coerced safely; an unknown role is omitted. The mapper is exported as toRoleType.

Best practices

  • Use tools for untrusted context. Use processors only when all stored content is application-authored and trusted.
  • Call createZepUserAndThread once before the first turn, store the returned UUIDs, then reuse a single ZepClient
  • Pass real names so Zep can anchor the user’s identity node in the graph
  • Don’t read-after-write within a turn — Zep builds the graph asynchronously, so a just-stored fact is not instantly retrievable
  • Pass a custom logger to route Zep warnings into your logging stack

Next steps