> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs-beta.getzep.com/v3/documents/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-beta.getzep.com/_mcp/server. # Documents ## Overview A document in Zep is a grouping of episodes, not a file type. Pages of a PDF, a Slack thread, and an email thread are all valid documents. Any sequence whose later episodes need the earlier ones as context qualifies. Assign a `document_id` when extraction of a later episode needs that prior context, most often to resolve a pronoun. Documents are the graph analogue of [threads](/threads) on user graphs. Zep scopes prior-episode context to the same `document_id`, which gives extraction the earlier episodes as context. Zep also generates incremental document summaries from them. `document_id` is optional. When present it must be 1 to 100 characters. Use a stable identifier from the source system, such as a file name and version or a conversation ID. The same `document_id` on different graphs identifies different documents. Pass `document_id` on [`graph.add`](/adding-business-data) and on batch `graph_episode` items. You can target a Context Graph with `graph_id` or a user graph with `user_id`. Listing episodes and document summaries takes `graph_id`. ## When to assign a document ID Zep loads earlier episodes with the same `document_id` when it extracts from a new episode. Share an ID across a group when a later episode is easier to interpret with the earlier ones in view. A pronoun is the clearest case. In these two chunks, `She` resolves to Alice only because the first chunk names her: ```text Alice joined Acme Corp as a designer. She reports to the product team in Austin. ``` Without a shared `document_id`, Zep extracts the second chunk alone and cannot resolve `She`. Other references behave the same way, including definite phrases such as `the company` or `that ticket`, a heading or preamble that only the first chunk carries, and facts stated once and assumed afterward. Use a `document_id` for any of these groups: * Chunks of one file, such as pages of a PDF or sections of a handbook * Messages in a Slack thread or email thread that you add with `graph.add` Do not assign a `document_id` only because records share a folder, customer, export, or other business grouping. Independent JSON records, unrelated tickets, and separate files in one directory are each self-contained, so extraction gains nothing from the others. Omit `document_id` for those episodes. Conversational messages on a user graph already get this grouping through a [thread](/threads). Use `document_id` on `graph.add` and batch `graph_episode` items. Do not pass it on `thread.add_messages`. ## Add episodes with a document ID Pass optional `document_id` when you add episodes, or when you append `graph_episode` batch items. Episodes that share a `document_id` on the same graph are associated together. > **Note** > > `document_id` requires `zep-cloud` 3.29.0 or later, or the matching TypeScript and Go packages. **`Python`** ```python Python from zep_cloud.client import Zep client = Zep(api_key="YOUR_API_KEY") client.graph.add( graph_id="graph_id", type="text", data="Alice joined Acme Corp as a designer.", document_id="handbook-v1", ) client.graph.add( graph_id="graph_id", type="text", data="She reports to the product team in Austin.", document_id="handbook-v1", ) ``` **`TypeScript`** ```typescript TypeScript import { ZepClient } from "@getzep/zep-cloud"; const client = new ZepClient({ apiKey: "YOUR_API_KEY" }); await client.graph.add({ graphId: "graph_id", type: "text", data: "Alice joined Acme Corp as a designer.", documentId: "handbook-v1", }); await client.graph.add({ graphId: "graph_id", type: "text", data: "She reports to the product team in Austin.", documentId: "handbook-v1", }); ``` **`Go`** ```go Go import ( "context" v3 "github.com/getzep/zep-go/v3" zepclient "github.com/getzep/zep-go/v3/client" "github.com/getzep/zep-go/v3/graph" "github.com/getzep/zep-go/v3/option" ) client := zepclient.NewClient(option.WithAPIKey("YOUR_API_KEY")) _, _ = client.Graph.Add(context.TODO(), &v3.AddDataRequest{ GraphID: v3.String("graph_id"), Type: v3.GraphDataTypeText, Data: "Alice joined Acme Corp as a designer.", DocumentID: v3.String("handbook-v1"), }) _, _ = client.Graph.Add(context.TODO(), &v3.AddDataRequest{ GraphID: v3.String("graph_id"), Type: v3.GraphDataTypeText, Data: "She reports to the product team in Austin.", DocumentID: v3.String("handbook-v1"), }) ``` Send each later episode with the same graph identifier and `document_id`. The [Batch API](/adding-batch-data) accepts the same optional `document_id` on `graph_episode` items. ## List episodes for a document **`Python`** ```python Python response = client.graph.get_episodes_for_document( "handbook-v1", graph_id="graph_id", ) episodes = response.episodes ``` **`TypeScript`** ```typescript TypeScript const response = await client.graph.getEpisodesForDocument("handbook-v1", { graphId: "graph_id", }); const episodes = response.episodes; ``` **`Go`** ```go Go response, err := client.Graph.GetEpisodesForDocument( context.TODO(), "handbook-v1", &v3.GraphGetEpisodesForDocumentRequest{GraphID: "graph_id"}, ) ``` > **Note** > > If no episodes have been associated with that `document_id`, the list is empty. A missing document is not an error. ## List document summaries for a graph Document summaries are generated asynchronously after episodes are associated. Listing returns one summary per document that has been summarized. **`Python`** ```python Python summaries = client.graph.document_summary.get_by_graph_id( graph_id="graph_id", ) for summary in summaries: print(summary.document_id, summary.summary) ``` **`TypeScript`** ```typescript TypeScript const summaries = await client.graph.documentSummary.getByGraphId("graph_id"); for (const summary of summaries) { console.log(summary.documentId, summary.summary); } ``` **`Go`** ```go Go summaries, err := client.Graph.DocumentSummary.GetByGraphID( context.TODO(), "graph_id", &graph.GraphDocumentSummariesRequest{}, ) ``` ## Related * [Threads](/threads) — message grouping on user graphs * [Thread summaries](/thread-summaries) — per-thread incremental summaries * [Adding business data](/adding-business-data) — `graph.add` fields, including `document_id` * [Prepare data for ingestion](/prepare-data-for-ingestion) — when to group chunks of one source * [Chunking large documents](/chunking-large-documents) — split sources that exceed the episode size limit > Group episodes on a Context Graph the way threads group messages