> This page is for version v4 (default).
> For other versions, use one of these documentation indexes:
> - v4 (default): https://docs-beta.getzep.com/v4/llms.txt
> - v3: https://docs-beta.getzep.com/v3/llms.txt
> - v2: https://docs-beta.getzep.com/v2/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs-beta.getzep.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-beta.getzep.com/_mcp/server.

# Documents

## Overview

A document in Zep is a grouping of episodes, not a file type. Pages of a PDF, a Slack thread, and an email thread are all valid documents. Any sequence whose later episodes need the earlier ones as context qualifies.

Assign a `document_id` when extraction of a later episode needs that prior context, most often to resolve a pronoun.

Documents are the graph analogue of [threads](/threads) on user graphs. Zep scopes prior-episode context to the same `document_id`, which gives extraction the earlier episodes as context. Zep also generates incremental document summaries from them.

`document_id` is optional. When present it must be 1 to 100 characters. Use a stable identifier from the source system, such as a file name and version or a conversation ID. The same `document_id` on different graphs identifies different documents.

Pass `document_id` on [`graph.episode.add`](/adding-business-data) and on batch `graph_episode` items. Each call takes a `graph_uuid`. Use the `uuid` of a graph from `graph.create`, or the `graph_uuid` of a user from `user.create`. Listing episodes and document summaries also takes `graph_uuid`.

## When to assign a document ID

Zep loads earlier episodes with the same `document_id` when it extracts from a new episode. Share an ID across a group when a later episode is easier to interpret with the earlier ones in view.

A pronoun is the clearest case. In these two chunks, `She` resolves to Alice only because the first chunk names her:

```text
Alice joined Acme Corp as a designer.
She reports to the product team in Austin.
```

Without a shared `document_id`, Zep extracts the second chunk alone and cannot resolve `She`.

Other references behave the same way, including definite phrases such as `the company` or `that ticket`, a heading or preamble that only the first chunk carries, and facts stated once and assumed afterward.

Use a `document_id` for any of these groups:

* Chunks of one file, such as pages of a PDF or sections of a handbook
* Messages in a Slack thread or email thread that you add with `graph.episode.add`

Do not assign a `document_id` only because records share a folder, customer, export, or other business grouping. Independent JSON records, unrelated tickets, and separate files in one directory are each self-contained, so extraction gains nothing from the others. Omit `document_id` for those episodes.

Conversational messages on a user graph already get this grouping through a [thread](/threads). Use `document_id` on `graph.episode.add` and batch `graph_episode` items. Do not pass it on `thread.add_messages`.

## Add episodes with a document ID

Pass optional `document_id` when you add episodes, or when you append `graph_episode` batch items. Episodes that share a `document_id` on the same graph are associated together.

**`Python`**

```python Python
from zep_cloud.client import Zep

client = Zep(api_key="YOUR_API_KEY")

# The UUID that graph.create returned, stored by your application
zep_graph_uuid = stored_graph_uuid

client.graph.episode.add(
    zep_graph_uuid,
    type="text",
    data="Alice joined Acme Corp as a designer.",
    document_id="handbook-v1",
)
client.graph.episode.add(
    zep_graph_uuid,
    type="text",
    data="She reports to the product team in Austin.",
    document_id="handbook-v1",
)
```

**`TypeScript`**

```typescript TypeScript
import { ZepClient } from "@getzep/zep-cloud";

const client = new ZepClient({ apiKey: "YOUR_API_KEY" });

// The UUID that graph.create returned, stored by your application
const zepGraphUuid = storedGraphUuid;

await client.graph.episode.add(zepGraphUuid, {
  type: "text",
  data: "Alice joined Acme Corp as a designer.",
  documentId: "handbook-v1",
});
await client.graph.episode.add(zepGraphUuid, {
  type: "text",
  data: "She reports to the product team in Austin.",
  documentId: "handbook-v1",
});
```

**`Go`**

```go Go
import (
    "context"

    zep "github.com/getzep/zep-go/v4"
    zepclient "github.com/getzep/zep-go/v4/client"
    "github.com/getzep/zep-go/v4/graph"
    "github.com/getzep/zep-go/v4/option"
)

client := zepclient.NewClient(option.WithAPIKey("YOUR_API_KEY"))

// The UUID that Graph.Create returned, stored by your application
zepGraphUUID := storedGraphUUID

_, _ = client.Graph.Episode.Add(context.TODO(), zepGraphUUID, &graph.AddEpisodeRequest{
    Type:       graph.V4AddEpisodeRequestTypeText.Ptr(),
    Data:       "Alice joined Acme Corp as a designer.",
    DocumentID: zep.String("handbook-v1"),
})
_, _ = client.Graph.Episode.Add(context.TODO(), zepGraphUUID, &graph.AddEpisodeRequest{
    Type:       graph.V4AddEpisodeRequestTypeText.Ptr(),
    Data:       "She reports to the product team in Austin.",
    DocumentID: zep.String("handbook-v1"),
})
```

Send each later episode with the same `graph_uuid` and `document_id`.

The [Batch API](/adding-batch-data) accepts the same optional `document_id` on `graph_episode` items.

## List episodes for a document

**`Python`**

```python Python
for episode in client.graph.episode.list_for_document(
    zep_graph_uuid,
    "handbook-v1",
):
    print(episode.uuid_, episode.content)
```

**`TypeScript`**

```typescript TypeScript
const episodes = await client.graph.episode.listForDocument(zepGraphUuid, "handbook-v1");
for await (const episode of episodes) {
  console.log(episode.uuid, episode.content);
}
```

**`Go`**

```go Go
page, err := client.Graph.Episode.ListForDocument(
    context.TODO(),
    zepGraphUUID,
    "handbook-v1",
    &graph.EpisodeListForDocumentRequest{},
)
if err != nil {
    return err
}
iter := page.Iterator()
for iter.Next(context.TODO()) {
    episode := iter.Current()
    fmt.Println(*episode.UUID, *episode.Content)
}
if err := iter.Err(); err != nil {
    return err
}
```

> **Note**
>
> If no episodes have been associated with that `document_id`, the list is empty. A missing document is not an error.

## List document summaries for a graph

Document summaries are generated asynchronously after episodes are associated. Listing returns one summary per document that has been summarized.

**`Python`**

```python Python
summaries = client.graph.document_summary.list(zep_graph_uuid)
for summary in summaries:
    print(summary.document_id, summary.summary)
```

**`TypeScript`**

```typescript TypeScript
const summaries = await client.graph.documentSummary.list(zepGraphUuid);
for await (const summary of summaries) {
  console.log(summary.documentId, summary.summary);
}
```

**`Go`**

```go Go
page, err := client.Graph.DocumentSummary.List(
    context.TODO(),
    zepGraphUUID,
    &graph.DocumentSummaryListRequest{},
)
if err != nil {
    return err
}
iter := page.Iterator()
for iter.Next(context.TODO()) {
    summary := iter.Current()
    fmt.Println(*summary.DocumentID, *summary.Summary)
}
if err := iter.Err(); err != nil {
    return err
}
```

## Related

* [Threads](/threads) — message grouping on user graphs
* [Thread summaries](/thread-summaries) — per-thread incremental summaries
* [Adding business data](/adding-business-data) — `graph.episode.add` fields, including `document_id`
* [Prepare data for ingestion](/prepare-data-for-ingestion) — when to group chunks of one source
* [Chunking large documents](/chunking-large-documents) — split sources that exceed the episode size limit