> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs-beta.getzep.com/v3/chunking-large-documents/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-beta.getzep.com/_mcp/server.
# Chunking Large Documents with Contextualized Retrieval
The `graph.add` endpoint has a 10,000-character limit per request. Split larger documents before ingestion.
This cookbook uses [contextualized retrieval](https://www.anthropic.com/news/contextual-retrieval). A large language model adds document context to each chunk before Zep ingests it.
This approach produces richer knowledge graphs with better entity and relationship extraction compared to naive chunking.
#### [Chunk documents with zep-ingest](/zep-ingest)
Use `ingest_documents` for paragraph-aware chunking, optional LLM context, preview, and submission.
> **Note**
>
> View the complete source code on GitHub: [Python](https://github.com/getzep/zep/tree/main/examples/python/chunking-example) | [TypeScript](https://github.com/getzep/zep/tree/main/examples/typescript/chunking-example) | [Go](https://github.com/getzep/zep/tree/main/examples/go/chunking-example)
## Overview
The ingestion pipeline follows these steps:
1. **Read the document** from a text file
2. **Chunk the document** into smaller pieces using paragraph-aware splitting
3. **Contextualize each chunk** using an LLM to add situational context
4. **Add each chunk to Zep** via `graph.add`
## Setup
Install the required dependencies:
**`Python`**
```bash Python
pip install zep-cloud openai python-dotenv
```
**`TypeScript`**
```bash TypeScript
npm install @getzep/zep-cloud openai dotenv
```
**`Go`**
```bash Go
go get github.com/getzep/zep-go/v3 github.com/sashabaranov/go-openai github.com/joho/godotenv
```
Initialize the clients:
**`Python`**
```python Python
import os
from openai import OpenAI
from zep_cloud.client import Zep
from dotenv import load_dotenv
load_dotenv()
openai_client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
zep_client = Zep(api_key=os.environ.get("ZEP_API_KEY"))
```
**`TypeScript`**
```typescript TypeScript
import { config } from "dotenv";
import { ZepClient } from "@getzep/zep-cloud";
import OpenAI from "openai";
config();
const openaiClient = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
const zepClient = new ZepClient({ apiKey: process.env.ZEP_API_KEY });
```
**`Go`**
```go Go
import (
"os"
"github.com/getzep/zep-go/v3"
zepclient "github.com/getzep/zep-go/v3/client"
"github.com/getzep/zep-go/v3/option"
"github.com/joho/godotenv"
openai "github.com/sashabaranov/go-openai"
)
godotenv.Load()
openaiClient := openai.NewClient(os.Getenv("OPENAI_API_KEY"))
zepClient := zepclient.NewClient(option.WithAPIKey(os.Getenv("ZEP_API_KEY")))
```
## Chunking the Document
> **Note**
>
> **Alternative chunking libraries:** If you prefer using an established library over the custom implementation below, consider [LangChain](https://docs.langchain.com/oss/python/integrations/splitters/index), [LlamaIndex](https://developers.llamaindex.ai/python/framework/module_guides/loading/node_parsers/), [Unstructured](https://docs.unstructured.io/open-source/core-functionality/chunking), or [Chonkie](https://docs.chonkie.ai).
The chunking algorithm splits text at paragraph boundaries first, then falls back to sentence boundaries for long paragraphs. This preserves semantic coherence better than fixed-size splitting.
**`Python`**
```python Python
import re
from typing import Generator
def chunk_document(
text: str,
chunk_size: int = 500,
chunk_overlap: int = 50
) -> Generator[tuple[int, str], None, None]:
"""
Split a document into chunks with configurable size and overlap.
Args:
text: The full document text
chunk_size: Maximum characters per chunk (default 500)
chunk_overlap: Characters to overlap between chunks for continuity
Yields:
Tuple of (chunk_index, chunk_text)
"""
if not text:
return
text = text.strip()
paragraphs = text.split('\n\n')
current_chunk = ""
chunk_index = 0
for paragraph in paragraphs:
paragraph = paragraph.strip()
if not paragraph:
continue
# If adding this paragraph exceeds chunk_size, yield current chunk
if len(current_chunk) + len(paragraph) + 2 > chunk_size:
if current_chunk:
yield (chunk_index, current_chunk.strip())
chunk_index += 1
# Start new chunk with overlap from previous
if chunk_overlap > 0 and len(current_chunk) > chunk_overlap:
overlap_text = current_chunk[-chunk_overlap:]
first_space = overlap_text.find(' ')
if first_space > 0:
overlap_text = overlap_text[first_space + 1:]
current_chunk = overlap_text + "\n\n"
else:
current_chunk = ""
# Handle single paragraphs longer than chunk_size
if len(paragraph) > chunk_size:
for sub_chunk in split_long_paragraph(paragraph, chunk_size, chunk_overlap):
yield (chunk_index, sub_chunk)
chunk_index += 1
current_chunk = ""
else:
current_chunk += paragraph
else:
if current_chunk:
current_chunk += "\n\n" + paragraph
else:
current_chunk = paragraph
# Yield final chunk
if current_chunk.strip():
yield (chunk_index, current_chunk.strip())
def split_long_paragraph(
paragraph: str,
chunk_size: int,
chunk_overlap: int
) -> Generator[str, None, None]:
"""Split a long paragraph by sentences."""
sentences = re.split(r'(?<=[.!?])\s+', paragraph)
current_chunk = ""
for sentence in sentences:
if len(current_chunk) + len(sentence) + 1 > chunk_size:
if current_chunk:
yield current_chunk.strip()
if chunk_overlap > 0:
overlap = current_chunk[-chunk_overlap:]
first_space = overlap.find(' ')
if first_space > 0:
current_chunk = overlap[first_space + 1:] + " "
else:
current_chunk = ""
else:
current_chunk = ""
current_chunk += sentence + " "
if current_chunk.strip():
yield current_chunk.strip()
```
**`TypeScript`**
```typescript TypeScript
const CHUNK_SIZE = 500;
const CHUNK_OVERLAP = 50;
function splitIntoSentences(text: string): string[] {
const sentences = text.match(/[^.!?]+[.!?]+[\s]*/g) || [text];
return sentences.map((s) => s.trim()).filter((s) => s.length > 0);
}
function* chunkDocument(
document: string,
chunkSize: number = CHUNK_SIZE,
chunkOverlap: number = CHUNK_OVERLAP
): Generator<[number, string]> {
const paragraphs = document.split(/\n\n+/).filter((p) => p.trim().length > 0);
let currentChunk = "";
let chunkIndex = 0;
for (const paragraph of paragraphs) {
const trimmedParagraph = paragraph.trim();
if (trimmedParagraph.length > chunkSize) {
// Yield current chunk if it exists
if (currentChunk.length > 0) {
yield [chunkIndex, currentChunk.trim()];
chunkIndex++;
currentChunk = currentChunk.slice(-chunkOverlap);
}
// Split long paragraph by sentences
const sentences = splitIntoSentences(trimmedParagraph);
for (const sentence of sentences) {
if (currentChunk.length + sentence.length + 1 > chunkSize) {
if (currentChunk.length > 0) {
yield [chunkIndex, currentChunk.trim()];
chunkIndex++;
currentChunk = currentChunk.slice(-chunkOverlap);
}
}
currentChunk = currentChunk.length > 0
? currentChunk + " " + sentence
: sentence;
}
} else {
if (currentChunk.length + trimmedParagraph.length + 2 > chunkSize) {
if (currentChunk.length > 0) {
yield [chunkIndex, currentChunk.trim()];
chunkIndex++;
currentChunk = currentChunk.slice(-chunkOverlap);
}
}
currentChunk = currentChunk.length > 0
? currentChunk + "\n\n" + trimmedParagraph
: trimmedParagraph;
}
}
if (currentChunk.trim().length > 0) {
yield [chunkIndex, currentChunk.trim()];
}
}
```
**`Go`**
```go Go
import (
"regexp"
"strings"
)
const (
ChunkSize = 500
ChunkOverlap = 50
)
func splitIntoSentences(text string) []string {
re := regexp.MustCompile(`([.!?]+)\s+`)
parts := re.Split(text, -1)
delimiters := re.FindAllString(text, -1)
var sentences []string
for i, part := range parts {
if part == "" {
continue
}
sentence := part
if i < len(delimiters) {
sentence += strings.TrimSpace(delimiters[i])
}
sentences = append(sentences, sentence)
}
return sentences
}
func getOverlapText(text string, overlapSize int) string {
if len(text) <= overlapSize {
return text
}
overlap := text[len(text)-overlapSize:]
spaceIdx := strings.Index(overlap, " ")
if spaceIdx > 0 && spaceIdx < len(overlap)/2 {
overlap = overlap[spaceIdx+1:]
}
return overlap
}
func chunkDocument(text string, chunkSize, chunkOverlap int) [][2]interface{} {
paragraphs := regexp.MustCompile(`\n\s*\n`).Split(text, -1)
var chunks [][2]interface{}
var currentChunk strings.Builder
chunkIndex := 0
for _, para := range paragraphs {
para = strings.TrimSpace(para)
if para == "" {
continue
}
if len(para) > chunkSize {
// Yield current chunk if exists
if currentChunk.Len() > 0 {
chunks = append(chunks, [2]interface{}{chunkIndex, currentChunk.String()})
chunkIndex++
overlapText := getOverlapText(currentChunk.String(), chunkOverlap)
currentChunk.Reset()
currentChunk.WriteString(overlapText)
}
// Split long paragraph by sentences
sentences := splitIntoSentences(para)
for _, sentence := range sentences {
sentence = strings.TrimSpace(sentence)
if currentChunk.Len()+len(sentence)+1 > chunkSize && currentChunk.Len() > 0 {
chunks = append(chunks, [2]interface{}{chunkIndex, currentChunk.String()})
chunkIndex++
overlapText := getOverlapText(currentChunk.String(), chunkOverlap)
currentChunk.Reset()
currentChunk.WriteString(overlapText)
}
if currentChunk.Len() > 0 {
currentChunk.WriteString(" ")
}
currentChunk.WriteString(sentence)
}
} else {
if currentChunk.Len()+len(para)+2 > chunkSize && currentChunk.Len() > 0 {
chunks = append(chunks, [2]interface{}{chunkIndex, currentChunk.String()})
chunkIndex++
overlapText := getOverlapText(currentChunk.String(), chunkOverlap)
currentChunk.Reset()
currentChunk.WriteString(overlapText)
}
if currentChunk.Len() > 0 {
currentChunk.WriteString("\n\n")
}
currentChunk.WriteString(para)
}
}
if currentChunk.Len() > 0 {
chunks = append(chunks, [2]interface{}{chunkIndex, currentChunk.String()})
}
return chunks
}
```
## Contextualizing Chunks
This is the key step that improves retrieval quality. For each chunk, we ask the LLM to generate a short context that situates it within the full document. This context is prepended to the chunk before adding to Zep.
> **Tip**
>
> **Cost optimization:** When contextualizing many chunks from the same document, use [prompt caching](https://platform.openai.com/docs/guides/prompt-caching) with a repeated user-level data prefix. Do not place document content in a system or developer message. Reusing the document tokens can reduce inference time and cost.
**`Python`**
```python Python
def contextualize_chunk(
openai_client: OpenAI,
full_document: str,
chunk: str
) -> str:
"""
Use OpenAI to generate context for a chunk within its document.
Args:
openai_client: Initialized OpenAI client
full_document: The complete document text
chunk: The specific chunk to contextualize
Returns:
The contextualized chunk (context prepended to original chunk)
"""
prompt = f"""
{full_document}
Here is the chunk we want to situate within the whole document:
{chunk}
Please give a short succinct context to situate this chunk within the overall document for the purposes of improving search retrieval of the chunk. If the document has a publication date, please include the date in your context. Answer only with the succinct context and nothing else."""
response = openai_client.chat.completions.create(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": prompt}],
max_completion_tokens=256
)
context = response.choices[0].message.content.strip()
# Combine context with original chunk
return f"{context}\n\n---\n\n{chunk}"
```
**`TypeScript`**
```typescript TypeScript
async function contextualizeChunk(
openai: OpenAI,
fullDocument: string,
chunk: string
): Promise {
const prompt = `
${fullDocument}
Here is the chunk we want to situate within the whole document:
${chunk}
Please give a short succinct context to situate this chunk within the overall document for the purposes of improving search retrieval of the chunk. If the document has a publication date, please include the date in your context. Answer only with the succinct context and nothing else.`;
const response = await openai.chat.completions.create({
model: "gpt-5.6-terra",
messages: [{ role: "user", content: prompt }],
max_completion_tokens: 256,
});
const context = response.choices[0]?.message?.content?.trim() || "";
// Combine context with original chunk
return `${context}\n\n---\n\n${chunk}`;
}
```
**`Go`**
```go Go
import (
"context"
"fmt"
openai "github.com/sashabaranov/go-openai"
)
func contextualizeChunk(
ctx context.Context,
client *openai.Client,
fullDocument string,
chunk string,
) (string, error) {
prompt := fmt.Sprintf(`
%s
Here is the chunk we want to situate within the whole document:
%s
Please give a short succinct context to situate this chunk within the overall document for the purposes of improving search retrieval of the chunk. If the document has a publication date, please include the date in your context. Answer only with the succinct context and nothing else.`, fullDocument, chunk)
resp, err := client.CreateChatCompletion(ctx, openai.ChatCompletionRequest{
Model: "gpt-5.6-terra",
Messages: []openai.ChatCompletionMessage{
{
Role: openai.ChatMessageRoleUser,
Content: prompt,
},
},
MaxCompletionTokens: 256,
})
if err != nil {
return "", fmt.Errorf("OpenAI API error: %w", err)
}
contextText := resp.Choices[0].Message.Content
// Combine context with original chunk
return fmt.Sprintf("%s\n\n---\n\n%s", contextText, chunk), nil
}
```
## Adding Chunks to Zep
Each contextualized chunk is added to the user's graph with `graph.add`. The method returns an episode that you can use to track ingestion.
These examples submit independent chunks. To group chunks of one source, pass the same `document_id` on each add. See [Documents](/documents).
**`Python`**
```python Python
def add_chunk_to_zep(
zep_client: Zep,
user_id: str,
chunk_data: str
) -> dict:
"""
Add a contextualized chunk to Zep's graph.
Args:
zep_client: Initialized Zep client
user_id: The user ID to add data to
chunk_data: The contextualized chunk text
Returns:
The episode response from Zep
"""
episode = zep_client.graph.add(
user_id=user_id,
type="text",
data=chunk_data,
)
return episode
```
**`TypeScript`**
```typescript TypeScript
async function addChunkToZep(
zepClient: ZepClient,
userId: string,
chunkData: string
): Promise {
await zepClient.graph.add({
userId,
type: "text",
data: chunkData,
});
}
```
**`Go`**
```go Go
import (
"context"
"github.com/getzep/zep-go/v3"
zepclient "github.com/getzep/zep-go/v3/client"
)
func addChunkToZep(
ctx context.Context,
zepClient *zepclient.Client,
userID string,
chunkData string,
) error {
dataType := zep.GraphDataTypeText
_, err := zepClient.Graph.Add(ctx, &zep.AddDataRequest{
UserID: zep.String(userID),
Type: &dataType,
Data: zep.String(chunkData),
})
return err
}
```
## Complete Ingestion Pipeline
Here's how to put it all together:
**`Python`**
```python Python
import os
def ingest_document(
openai_client: OpenAI,
zep_client: Zep,
document_path: str,
user_id: str,
chunk_size: int = 500,
chunk_overlap: int = 50
) -> dict:
"""
Ingest a document into Zep with contextualized retrieval.
Args:
openai_client: Initialized OpenAI client
zep_client: Initialized Zep client
document_path: Path to the text document
user_id: Zep user ID to add the document to
chunk_size: Maximum characters per chunk
chunk_overlap: Character overlap between chunks
Returns:
Summary statistics of the ingestion
"""
# Read document
with open(document_path, 'r', encoding='utf-8') as f:
full_document = f.read()
# If document fits in a single request, add directly
if len(full_document) <= 10000:
episode = zep_client.graph.add(
user_id=user_id,
type="text",
data=full_document,
)
return {"total_chunks": 1, "successful": 1, "episodes": [episode.uuid_]}
# Chunk the document
chunks = list(chunk_document(full_document, chunk_size, chunk_overlap))
stats = {"total_chunks": len(chunks), "successful": 0, "episodes": []}
for chunk_index, chunk_text in chunks:
# Contextualize the chunk
contextualized = contextualize_chunk(
openai_client,
full_document,
chunk_text
)
# Validate size after contextualization
if len(contextualized) > 10000:
# Truncate context if needed
excess = len(contextualized) - 10000
contextualized = contextualized[excess:]
# Add to Zep
episode = add_chunk_to_zep(zep_client, user_id, contextualized)
stats["successful"] += 1
stats["episodes"].append(episode.uuid_)
return stats
```
**`TypeScript`**
```typescript TypeScript
import * as fs from "fs";
interface IngestStats {
totalChunks: number;
successful: number;
}
async function ingestDocument(
openaiClient: OpenAI,
zepClient: ZepClient,
documentPath: string,
userId: string,
chunkSize: number = 500,
chunkOverlap: number = 50
): Promise {
// Read document
const fullDocument = fs.readFileSync(documentPath, "utf-8");
// If document fits in a single request, add directly
if (fullDocument.length <= 10000) {
await zepClient.graph.add({
userId,
type: "text",
data: fullDocument,
});
return { totalChunks: 1, successful: 1 };
}
// Chunk the document
const chunks = Array.from(chunkDocument(fullDocument, chunkSize, chunkOverlap));
const stats: IngestStats = { totalChunks: chunks.length, successful: 0 };
for (const [chunkIndex, chunkText] of chunks) {
// Contextualize the chunk
let contextualized = await contextualizeChunk(
openaiClient,
fullDocument,
chunkText
);
// Validate size after contextualization
if (contextualized.length > 10000) {
// Truncate context if needed
const excess = contextualized.length - 10000;
contextualized = contextualized.slice(excess);
}
// Add to Zep
await addChunkToZep(zepClient, userId, contextualized);
stats.successful++;
}
return stats;
}
```
**`Go`**
```go Go
import (
"context"
"fmt"
"os"
"github.com/getzep/zep-go/v3"
zepclient "github.com/getzep/zep-go/v3/client"
openai "github.com/sashabaranov/go-openai"
)
type IngestStats struct {
TotalChunks int
Successful int
}
func ingestDocument(
ctx context.Context,
openaiClient *openai.Client,
zepClient *zepclient.Client,
documentPath string,
userID string,
chunkSize int,
chunkOverlap int,
) (*IngestStats, error) {
// Read document
docContent, err := os.ReadFile(documentPath)
if err != nil {
return nil, fmt.Errorf("error reading document: %w", err)
}
fullDocument := string(docContent)
// If document fits in a single request, add directly
if len(fullDocument) <= 10000 {
dataType := zep.GraphDataTypeText
_, err := zepClient.Graph.Add(ctx, &zep.AddDataRequest{
UserID: zep.String(userID),
Type: &dataType,
Data: zep.String(fullDocument),
})
if err != nil {
return nil, err
}
return &IngestStats{TotalChunks: 1, Successful: 1}, nil
}
// Chunk the document
chunks := chunkDocument(fullDocument, chunkSize, chunkOverlap)
stats := &IngestStats{TotalChunks: len(chunks), Successful: 0}
for _, chunk := range chunks {
chunkText := chunk[1].(string)
// Contextualize the chunk
contextualized, err := contextualizeChunk(ctx, openaiClient, fullDocument, chunkText)
if err != nil {
continue
}
// Validate size after contextualization
if len(contextualized) > 10000 {
// Truncate context if needed
excess := len(contextualized) - 10000
contextualized = contextualized[excess:]
}
// Add to Zep
if err := addChunkToZep(ctx, zepClient, userID, contextualized); err != nil {
continue
}
stats.Successful++
}
return stats, nil
}
```
## Usage Example
**`Python`**
```python Python
# Ensure the user exists
user_id = "user123"
zep_client.user.add(user_id=user_id)
# Ingest a document
stats = ingest_document(
openai_client=openai_client,
zep_client=zep_client,
document_path="company_handbook.txt",
user_id=user_id,
chunk_size=500,
chunk_overlap=50
)
print(f"Ingested {stats['successful']} of {stats['total_chunks']} chunks")
```
**`TypeScript`**
```typescript TypeScript
// Ensure the user exists
const userId = "user123";
await zepClient.user.add({ userId });
// Ingest a document
const stats = await ingestDocument(
openaiClient,
zepClient,
"company_handbook.txt",
userId,
500,
50
);
console.log(`Ingested ${stats.successful} of ${stats.totalChunks} chunks`);
```
**`Go`**
```go Go
// Ensure the user exists
userID := "user123"
zepClient.User.Add(ctx, &zep.CreateUserRequest{
UserID: zep.String(userID),
})
// Ingest a document
stats, err := ingestDocument(
ctx,
openaiClient,
zepClient,
"company_handbook.txt",
userID,
500,
50,
)
if err != nil {
log.Fatalf("Failed to ingest document: %v", err)
}
fmt.Printf("Ingested %d of %d chunks\n", stats.Successful, stats.TotalChunks)
```
## Best practices
* **Chunk size**: Use 500 characters or less for optimal graph construction. Smaller chunks allow Zep to capture more granular entities and relationships.
* **Chunk overlap**: 50 characters helps maintain continuity between chunks without excessive redundancy.
* **Small chunks produce better graphs**: Zep can capture more entities and relationships from smaller, focused chunks. While the 10K character limit allows larger chunks, smaller chunks yield richer knowledge graphs.
## Further Reading
* [Documents](/documents) - Group chunks of the same source with `document_id`
* [Adding Business Data](/adding-business-data) - Learn about the `graph.add` endpoint and data types
* [Batch ingestion](/adding-batch-data) - For ingesting large historical datasets in a single batch
* [Performance best practices](/performance) - Data ingestion and retrieval guidance
> Ingest documents larger than 10,000 characters using semantic chunking and LLM-powered contextualization