Skip to navigation

Adding Business Data

Add structured and unstructured business data to Context Graphs

Requests to add data to the same graph are completed sequentially to ensure the graph is built correctly, and processing may be slow for large datasets. For large historical datasets, use batch ingestion.

Use graph.episode.add to add business records, documents, events, and communications to a Context Graph. Zep supports three data types: JSON, text, and message.

Choose the destination and source information

Pass the graph_uuid of the destination graph. Use the uuid of a graph that you created with graph.create when the context belongs to a shared customer account, project, product, organization, or business domain. Use the graph_uuid of a user from user.create when the context belongs to one application user. Zep gives each graph and user a UUID when you create it. Store the UUID in your application database next to your own identifier, and use the stored value on each call.

Use source_description to give a human-readable description of the source. Use metadata for source attributes that your application will use for search filters, traceability, or access policies.

import json
# The UUID that graph.create returned for the Acme account graph
zep_graph_uuid = stored_graph_uuid
result = client.graph.episode.add(
zep_graph_uuid,
type="json",
data=json.dumps({
"account_id": "acme",
"plan": "Enterprise",
"region": "eu-west",
}),
source_description="Account record from the CRM",
metadata={"source": "crm", "account_id": "acme"},
)

The source description and metadata do not establish that the source content is true. See Source traceability for the relationship between episodes and derived graph data.

The message type is ideal for adding data in the form of chat messages that are not directly associated with a Zep Thread’s chat history. This encompasses any communication with a designated speaker, such as emails or previous chat logs.

The text type is designed for raw text data without a specific speaker attribution. This category includes content from internal documents, wiki articles, or company handbooks. It’s important to note that Zep does not process text directly from links or files.

The JSON type may be used to add any JSON document to Zep. This may include REST API responses or JSON-formatted business data.

Adding Message Data

Here’s an example demonstrating how to add message data to the graph:

from zep_cloud.client import Zep
client = Zep(
api_key=API_KEY,
)
message = "Paul (user): I went to Eric Clapton concert last night"
# zep_graph_uuid is the graph_uuid of the user, or the uuid of a graph
result = client.graph.episode.add(
zep_graph_uuid,
type="message",
data=message,
)

Adding Text Data

Here’s an example demonstrating how to add text data to the graph:

from zep_cloud.client import Zep
client = Zep(
api_key=API_KEY,
)
# zep_graph_uuid is the graph_uuid of the user, or the uuid of a graph
result = client.graph.episode.add(
zep_graph_uuid,
type="text",
data="The user is an avid fan of Eric Clapton",
)

Adding JSON Data

Before ingesting JSON, name and contextualize each record so Zep builds an accurate graph from it.

Here’s an example demonstrating how to add JSON data to the graph:

from zep_cloud.client import Zep
import json
client = Zep(
api_key=API_KEY,
)
json_data = {"name": "Eric Clapton", "age": 78, "genre": "Rock"}
json_string = json.dumps(json_data)
# zep_graph_uuid is the graph_uuid of the user, or the uuid of a graph
result = client.graph.episode.add(
zep_graph_uuid,
type="json",
data=json_string,
)

Setting data timestamps

When adding data via the graph.episode.add method, you can provide the reference_time timestamp in RFC3339 format. The reference_time timestamp represents the time when the data was originally created. For messages, this would be when the message was originally sent. For events represented as JSON, this would be when the event occurred. Setting the reference_time timestamp ensures the user’s Context Graph has accurate temporal understanding of user history (since this time is used in our fact invalidation process).

from zep_cloud.client import Zep
import json
client = Zep(
api_key=API_KEY,
)
# Example: Adding a JSON event with its original timestamp
event_data = {"event": "purchase", "item": "laptop", "amount": 1299.99}
json_string = json.dumps(event_data)
result = client.graph.episode.add(
zep_graph_uuid,
type="json",
data=json_string,
reference_time="2025-06-01T13:11:12Z" # Time the event originally occurred
)

Grouping episodes with a document ID

Assign a document_id when extraction of a later episode needs earlier episodes. For example, an earlier episode can identify the subject of a pronoun.

Pass the same ID on each add in the group. See Documents.

Zep scopes prior-episode context to that ID, so extraction can resolve pronouns and other references against the earlier episodes. Zep also generates an incremental document summary.

Omit document_id for independent records that share only a folder, customer, or export. document_id is optional and 1 to 100 characters. Use a stable identifier from the source system.

Data Size Limit and Chunking

The graph.episode.add method has a data size limit of 10,000 characters when adding data to the graph. If you need to add a document which is more than 10,000 characters, see our Chunking Large Documents cookbook for a complete implementation with contextualized retrieval and best practices for chunking for Zep. Pass the same document_id on every chunk so Zep treats them as one source.

Episode metadata

You can attach key-value metadata to episodes when adding data. Metadata is useful for tagging episodes with their data source, category, priority, or other attributes. Zep projects this metadata onto every graph artifact derived from the episode, which is what lets you filter graph search results so that only facts derived from matching episodes are returned.

Metadata values must be scalars (string, number (int/float), or boolean) or non-empty arrays of such scalars, such as {"tags": ["red", "blue"]}. A maximum of 10 keys are allowed per episode. Nested objects are not supported, and empty arrays or arrays containing null are rejected.

from zep_cloud.client import Zep
client = Zep(
api_key=API_KEY,
)
result = client.graph.episode.add(
zep_graph_uuid,
type="text",
data="Patient blood glucose level was 95 mg/dL, within normal range.",
metadata={"source": "lab_report", "priority": 5, "reviewed": True},
)

Updating episode metadata

You can update an episode’s metadata after creation using merge semantics: new keys are added, existing keys are overwritten, and keys set to null are removed.

from zep_cloud.client import Zep
client = Zep(
api_key=API_KEY,
)
# Original metadata: {"source": "lab_report", "priority": 5, "reviewed": True}
updated_episode = client.graph.episode.update(
zep_graph_uuid,
episode_uuid,
metadata={"priority": 10, "department": "endocrinology", "reviewed": None},
)
# Result: {"source": "lab_report", "priority": 10, "department": "endocrinology"}
# "priority" was overwritten, "department" was added, "reviewed" was removed

Managing Your Data on the Graph

The graph.episode.add method returns the episode that was created when you added the data, and a task that tracks its processing. You can maintain a mapping between your data and its episode. You can then delete specific data from the graph with the delete episode method.