> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs-beta.getzep.com/v4/adding-batch-data/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-beta.getzep.com/_mcp/server. # Batch ingestion > Load large historical datasets into Context Graphs with the Batch API, including business records, documents, archived conversations, and migration data. The Batch API loads large historical datasets — backfills, document collections, archived conversations, migrations from another system — into your Context Graphs. It is the fastest transport for bulk ingestion, and it does not compete with the live `graph.episode.add` and `thread.add_messages` traffic serving your agents. [`zep-ingest`](/zep-ingest) is the recommended wrapper around the Batch API: it handles preparation, ordering, and monitoring on top. Use the Batch API directly when you want to manage batches yourself or already have ingestion code. #### [Create an Ingestion Pipeline](/zep-ingest) Prepare and preview exports, submit through the Batch API, and monitor the complete import. ## Why use the Batch API Calling `graph.episode.add` or `thread.add_messages` once per item works for live data but becomes hard to manage at scale. Compared to issuing those calls one at a time, the Batch API gives you: * **Faster processing.** It ingests large datasets faster than the same operations sent one at a time. * **No interference with live traffic.** It is designed not to slow the real-time `graph.episode.add` and `thread.add_messages` ingestion serving your agents, so a large backfill can run alongside production. * **Progress you can watch.** Monitor each batch's status, item counts, and errors in the [batch dashboard](#viewing-batches-in-the-dashboard), or poll programmatically. * **One batch instead of many calls.** Group items into a single batch — splitting across batches when needed (see [Batch limits](#batch-limits)) — and hand it off to Zep to process as one job. ## Backfills For historical imports, [`zep-ingest`](/zep-ingest) is the recommended path: it prepares sources, preserves timestamps and order, submits through the Batch API when available, and monitors completion. Use the Batch API directly when you already manage batches yourself. **Within one graph:** submit all episodes without polling between adds. See [Submit many episodes, poll once](/check-data-ingestion-status#submit-many-episodes-poll-once) for ingestion order and what to poll on each path. **Across multiple graphs:** you do not need to wait for one graph to finish extraction before submitting to another. Create every destination and set ontology first — ontology is not retroactive — then submit episodes to every graph. After all submits are queued, wait once per graph when you need the data to be searchable. Extraction runs share account-level concurrency limits, so a very large parallel backfill will not run every graph at full speed simultaneously. [`zep-ingest`](/zep-ingest) handles one destination per call and does not parallelize across graphs for you; run separate imports concurrently when you backfill several graphs at once. ## How batches work A batch follows a three-step lifecycle: #### Create Create an empty batch with optional metadata and an optional `strict_ontology` flag. #### Add Add items to the batch across one or more `batch.add_items` calls. Items can be graph episodes or thread messages, and may target different graphs, users, or threads. #### Process Start processing. Zep returns immediately and processes the batch asynchronously. You can poll for progress or watch it in the dashboard. Items in a batch are grouped by destination graph and processed in the order they were added. Episodes and messages added through the Batch API are priced the same as those added through `graph.episode.add` or `thread.add_messages`. ### Batch limits * A single batch can contain up to **50,000 items**. * Each call to `batch.add_items` accepts up to **350 items**. To ingest more than 350 items, make multiple `batch.add_items` calls against the same batch UUID before calling `batch.process`. ### Strict ontology Set `strict_ontology` once on `batch.create`. Processing copies that value onto every graph episode and every thread message in the batch. You cannot set the flag per item. **`Python`** ```python Python batch = client.batch.create( strict_ontology=True, metadata={"description": "Customer support backfill"}, ) ``` **`TypeScript`** ```typescript TypeScript const batch = await client.batch.create({ strictOntology: true, metadata: { description: "Customer support backfill" }, }); ``` **`Go`** ```go Go batch, err := client.Batch.Create(ctx, &zep.CreateBatchRequest{ StrictOntology: zep.Bool(true), Metadata: map[string]any{ "description": "Customer support backfill", }, }) ``` See [Setting Strict Ontology](/customizing-graph-structure#setting-strict-ontology) for what the flag does. ## Quickstart The example below creates a batch, adds a mix of graph episodes and thread messages, starts processing, and polls until the batch finishes. Each item identifies its destination with a UUID that Zep returned when you created the resource: the `graph_uuid` of a user from `user.create`, the `uuid` of a graph from `graph.create`, or the `uuid` of a thread from `thread.create`. Store these UUIDs in your application database next to your own identifiers, and read them from there when you build the batch. **`Python`** ```python Python import time from zep_cloud.client import Zep from zep_cloud import BatchItemInput client = Zep(api_key=API_KEY) # UUIDs that your application stored when it created the resources alice_graph_uuid = stored_alice_graph_uuid # graph_uuid of the user Alice company_kb_graph_uuid = stored_kb_graph_uuid # uuid of the shared graph alice_thread_uuid = stored_alice_thread_uuid # uuid of the support thread # 1. Create the batch batch = client.batch.create( metadata={"description": "Customer support backfill"}, ) batch_uuid = batch.uuid_ # 2. Add items to the batch items = [ BatchItemInput( type="graph_episode", graph_uuid=alice_graph_uuid, data="Alice signed up for the Pro plan on 2024-06-15.", data_type="text", ), BatchItemInput( type="graph_episode", graph_uuid=company_kb_graph_uuid, data="Refund policy: orders may be refunded within 30 days of purchase.", data_type="text", ), BatchItemInput( type="thread_message", thread_uuid=alice_thread_uuid, content="My dashboard isn't loading.", role="user", name="Alice", ), ] client.batch.add_items(batch_uuid, items=items) # 3. Start processing client.batch.process(batch_uuid) # 4. Poll until the batch finishes TERMINAL_STATUSES = ("succeeded", "partial", "failed", "invalid", "canceled") while True: summary = client.batch.get(batch_uuid) if summary.status in TERMINAL_STATUSES: break progress = summary.progress or {} print( f"Status: {summary.status} " f"({progress.get('succeeded_items', 0)}/{progress.get('total_items', 0)} items)" ) time.sleep(5) print(f"Final status: {summary.status}") ``` **`TypeScript`** ```typescript TypeScript import { ZepClient, Zep } from "@getzep/zep-cloud"; const client = new ZepClient({ apiKey: API_KEY }); // UUIDs that your application stored when it created the resources const aliceGraphUuid = storedAliceGraphUuid; // graphUuid of the user Alice const companyKbGraphUuid = storedKbGraphUuid; // uuid of the shared graph const aliceThreadUuid = storedAliceThreadUuid; // uuid of the support thread // 1. Create the batch const batch = await client.batch.create({ metadata: { description: "Customer support backfill" }, }); const batchUuid = batch.uuid!; // 2. Add items to the batch const items: Zep.BatchItemInput[] = [ { type: "graph_episode", graphUuid: aliceGraphUuid, data: "Alice signed up for the Pro plan on 2024-06-15.", dataType: "text", }, { type: "graph_episode", graphUuid: companyKbGraphUuid, data: "Refund policy: orders may be refunded within 30 days of purchase.", dataType: "text", }, { type: "thread_message", threadUuid: aliceThreadUuid, content: "My dashboard isn't loading.", role: "user", name: "Alice", }, ]; await client.batch.addItems(batchUuid, { items }); // 3. Start processing await client.batch.process(batchUuid); // 4. Poll until the batch finishes const TERMINAL_STATUSES = ["succeeded", "partial", "failed", "invalid", "canceled"]; const sleep = (ms: number) => new Promise(r => setTimeout(r, ms)); let summary = await client.batch.get(batchUuid); while (!TERMINAL_STATUSES.includes(summary.status!)) { const progress = summary.progress ?? {}; console.log(`Status: ${summary.status} (${progress.succeeded_items ?? 0}/${progress.total_items ?? 0} items)`); await sleep(5000); summary = await client.batch.get(batchUuid); } console.log(`Final status: ${summary.status}`); ``` **`Go`** ```go Go import ( "context" "fmt" "time" zep "github.com/getzep/zep-go/v4" zepclient "github.com/getzep/zep-go/v4/client" "github.com/getzep/zep-go/v4/option" ) ctx := context.Background() client := zepclient.NewClient(option.WithAPIKey(apiKey)) // UUIDs that your application stored when it created the resources aliceGraphUUID := storedAliceGraphUUID // GraphUUID of the user Alice companyKBGraphUUID := storedKBGraphUUID // UUID of the shared graph aliceThreadUUID := storedAliceThreadUUID // UUID of the support thread // 1. Create the batch batch, err := client.Batch.Create(ctx, &zep.CreateBatchRequest{ Metadata: map[string]any{ "description": "Customer support backfill", }, }) if err != nil { panic(err) } batchUUID := *batch.UUID // 2. Add items to the batch items := []*zep.BatchItemInput{ { Type: zep.BatchItemInputTypeGraphEpisode, GraphUUID: zep.String(aliceGraphUUID), Data: zep.String("Alice signed up for the Pro plan on 2024-06-15."), DataType: zep.BatchItemInputDataTypeText.Ptr(), }, { Type: zep.BatchItemInputTypeGraphEpisode, GraphUUID: zep.String(companyKBGraphUUID), Data: zep.String("Refund policy: orders may be refunded within 30 days of purchase."), DataType: zep.BatchItemInputDataTypeText.Ptr(), }, { Type: zep.BatchItemInputTypeThreadMessage, ThreadUUID: zep.String(aliceThreadUUID), Content: zep.String("My dashboard isn't loading."), Role: zep.BatchItemInputRoleUser.Ptr(), Name: zep.String("Alice"), }, } _, err = client.Batch.AddItems(ctx, batchUUID, &zep.AddBatchItemsRequest{Items: items}) if err != nil { panic(err) } // 3. Start processing if _, err := client.Batch.Process(ctx, batchUUID); err != nil { panic(err) } // 4. Poll until the batch finishes terminalStatuses := map[string]bool{ "succeeded": true, "partial": true, "failed": true, "invalid": true, "canceled": true, } for { summary, err := client.Batch.Get(ctx, batchUUID) if err != nil { panic(err) } if terminalStatuses[*summary.Status] { fmt.Printf("Final status: %s\n", *summary.Status) break } fmt.Printf("Status: %s (%v/%v items)\n", *summary.Status, summary.Progress["succeeded_items"], summary.Progress["total_items"]) time.Sleep(5 * time.Second) } ``` ## Adding items to a batch Each item in a batch is one of two types: * **`graph_episode`** — equivalent to a single `graph.episode.add` call. Targets a graph or a user graph by `graph_uuid`. * **`thread_message`** — equivalent to one message inside a `thread.add_messages` call. Targets a thread by `thread_uuid`. The fields below mirror the equivalent fields on `graph.episode.add` and `thread.add_messages`. See [Adding business data](/adding-business-data) and [Adding messages](/adding-messages) for the underlying semantics. A batch item is a single SDK type covering both kinds, so every field is settable on every item. Fields that do not apply to an item's `type` are still validated, but not stored — `source_description` on a `thread_message` is rejected above 500 characters, and a valid value is discarded. ### Common fields | Field | Description | | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `type` | Required. `graph_episode` or `thread_message`. | | `metadata` | Optional. Up to 10 key-value pairs. See [Episode metadata](/adding-business-data#episode-metadata) for constraints and search filtering. | | `reference_time` | Optional. ISO 8601 timestamp marking when the original event occurred. Used by Zep's fact invalidation process for both item types. See [Setting timestamps](#setting-timestamps-on-batch-items). | ### Graph episode fields (`type: "graph_episode"`) | Field | Description | | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `data` | Required. The episode content. Subject to the same 10,000-character limit as `graph.episode.add`. | | `data_type` | Required. `text`, `json`, or `message`. | | `graph_uuid` | Required. The destination graph: the `uuid` of a graph, or the `graph_uuid` of a user. | | `source_description` | Optional. Human-readable description of where the episode came from. Maximum 500 characters. | | `document_id` | Optional. Groups this episode with others so extraction can resolve pronouns and other references against earlier episodes in that [document](/documents). Ignored on `thread_message` items. 1 to 100 characters. | ### Thread message fields (`type: "thread_message"`) | Field | Description | | ------------- | ------------------------------------------------------------------- | | `thread_uuid` | Required. The `uuid` of the destination thread. | | `content` | Required. The message body. | | `role` | Required. One of `user`, `assistant`, `system`, `function`, `tool`. | | `name` | Optional. Speaker name. | ## Setting timestamps on batch items Pass `reference_time` on each item to give Zep accurate temporal information for historical data. This is important for backfills — Zep uses these timestamps in its fact invalidation process to determine the `valid_at` and `invalid_at` values on extracted facts (edges). The `reference_time` value should be in RFC3339 format (e.g., `"2024-06-15T10:30:00Z"`). Both item types honor the value you supply: a `thread_message` item's `reference_time` dates the episode Zep extracts from that message, exactly as it does for a `graph_episode` item, so a batch backfill and a direct `thread.add_messages` call place the same message at the same point in the timeline. An item without a `reference_time` is dated at ingestion time. **`Python`** ```python Python from zep_cloud import BatchItemInput items = [ BatchItemInput( type="graph_episode", graph_uuid=alice_graph_uuid, data="Alice joined the engineering team as a senior developer.", data_type="text", reference_time="2024-06-15T10:30:00Z", ), BatchItemInput( type="graph_episode", graph_uuid=alice_graph_uuid, data="Alice was promoted to tech lead of the engineering team.", data_type="text", reference_time="2024-09-01T09:00:00Z", ), ] client.batch.add_items(batch_uuid, items=items) ``` **`TypeScript`** ```typescript TypeScript const items: Zep.BatchItemInput[] = [ { type: "graph_episode", graphUuid: aliceGraphUuid, data: "Alice joined the engineering team as a senior developer.", dataType: "text", referenceTime: "2024-06-15T10:30:00Z", }, { type: "graph_episode", graphUuid: aliceGraphUuid, data: "Alice was promoted to tech lead of the engineering team.", dataType: "text", referenceTime: "2024-09-01T09:00:00Z", }, ]; await client.batch.addItems(batchUuid, { items }); ``` **`Go`** ```go Go items := []*zep.BatchItemInput{ { Type: zep.BatchItemInputTypeGraphEpisode, GraphUUID: zep.String(aliceGraphUUID), Data: zep.String("Alice joined the engineering team as a senior developer."), DataType: zep.BatchItemInputDataTypeText.Ptr(), ReferenceTime: zep.String("2024-06-15T10:30:00Z"), }, { Type: zep.BatchItemInputTypeGraphEpisode, GraphUUID: zep.String(aliceGraphUUID), Data: zep.String("Alice was promoted to tech lead of the engineering team."), DataType: zep.BatchItemInputDataTypeText.Ptr(), ReferenceTime: zep.String("2024-09-01T09:00:00Z"), }, } client.Batch.AddItems(ctx, batchUUID, &zep.AddBatchItemsRequest{Items: items}) ``` ## Tracking progress Two methods report on a running or completed batch: * **`batch.get(batch_uuid)`** returns the whole batch, including a `progress` object with the current `stage` and the counts `total_items`, `processing_items`, `succeeded_items`, and `failed_items`. Before `batch.process` is called the batch is in `draft` and the `progress` counts are unpopulated; once processing starts the counts begin to update. * **`batch.list_items(batch_uuid)`** returns each item with its individual status (`pending`, `queued`, `processing`, `succeeded`, `failed`, `skipped`, `canceled`). When polling `batch.get`, a few-second interval (e.g., 5 seconds) is appropriate for small batches. For batches with thousands of items or more, polling becomes impractical. Subscribe to the [`ingest.batch.completed` webhook](/webhooks#batch-completion-payloads) for `succeeded` and `canceled` runs, and for `partial` runs that finish with canceled items. Zep does not send that webhook when a run ends as `failed`, or as `partial` because items failed. Poll `batch.get` for those two outcomes. The payload includes the `batch_id`, which is the `uuid` of the batch, so you can match it back to the batch you submitted. ### Batch statuses The `status` field on `Batch` is one of: | Status | Meaning | | ------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `draft` | The batch was just created. Items can still be added with `batch.add_items`. Processing has not started. Can be deleted. | | `invalid` | `batch.process` was called, but one or more items reference graphs, users, or threads that don't exist. The batch cannot proceed. Can be deleted. | | `queued` | `batch.process` was called and the batch is waiting for a worker. | | `processing` | A worker is actively processing the batch. | | `succeeded` | Terminal. No item failed or was canceled. Items can still be skipped, so check the item statuses with `batch.list_items` before you treat the import as complete. | | `partial` | Terminal. Some items succeeded and others failed or were canceled. Use `batch.list_items` to see each item status. | | `failed` | Terminal. The batch as a whole failed. | | `canceled` | Terminal. Every item was canceled because its target graph, user, or thread was deleted while the batch was in flight, so nothing was ingested. | Once a batch reaches `succeeded`, `partial`, `failed`, or `canceled`, no further state changes occur. `invalid` is also non-progressing — the batch never starts processing, but the state persists until you delete the batch. When polling, exit on any of `succeeded`, `partial`, `failed`, `canceled`, or `invalid`. ### Per-item statuses The `status` field on each `BatchItem` is one of: | Status | Meaning | | ------------ | ---------------------------------------------------------------------------------------------------------------------------- | | `pending` | The item has been added to the batch but processing has not started. | | `queued` | The item is queued for processing. | | `processing` | The item is currently being processed. | | `succeeded` | The item processed successfully. | | `failed` | The item failed to process. The `error` field on the item describes why. | | `skipped` | The item was skipped during processing — for example, a thread message whose role matches a configured `ignore_roles` value. | | `canceled` | The item was not processed because its target graph, user, or thread was deleted before the item finished. | **`Python`** ```python Python summary = client.batch.get(batch_uuid) progress = summary.progress or {} print(f"Status: {summary.status}") print(f"Progress: {progress.get('succeeded_items')}/{progress.get('total_items')}") # Inspect individual items for item in client.batch.list_items(batch_uuid, limit=50): print(item.uuid_, item.status) ``` **`TypeScript`** ```typescript TypeScript const summary = await client.batch.get(batchUuid); console.log(`Status: ${summary.status}`); console.log(`Progress: ${summary.progress?.succeeded_items}/${summary.progress?.total_items}`); // Inspect individual items const items = await client.batch.listItems(batchUuid, { limit: 50 }); for await (const item of items) { console.log(item.uuid, item.status); } ``` **`Go`** ```go Go summary, _ := client.Batch.Get(ctx, batchUUID) fmt.Printf("Status: %s\n", *summary.Status) fmt.Printf("Progress: %v/%v\n", summary.Progress["succeeded_items"], summary.Progress["total_items"]) // Inspect individual items page, _ := client.Batch.ListItems(ctx, batchUUID, &zep.BatchListItemsRequest{Limit: zep.Int(50)}) iter := page.Iterator() for iter.Next(ctx) { item := iter.Current() fmt.Println(*item.UUID, *item.Status) } ``` ## Listing and managing batches Use `batch.list` to enumerate batches in your project, optionally filtered by status. Use `batch.delete` to remove a `draft` or `invalid` batch. Once a batch has started processing, it cannot be deleted. `batch.list` uses cursor pagination, and the SDK pagers fetch the next pages for you. **`Python`** ```python Python # List recent batches for b in client.batch.list(limit=20): print(b.uuid_, b.status) # List only batches that are still being processed processing = client.batch.list(status="processing") # Delete a draft batch client.batch.delete(batch_uuid) ``` **`TypeScript`** ```typescript TypeScript // List recent batches const batches = await client.batch.list({ limit: 20 }); for await (const b of batches) { console.log(b.uuid, b.status); } // List only batches that are still being processed await client.batch.list({ status: "processing" }); // Delete a draft batch await client.batch.delete(batchUuid); ``` **`Go`** ```go Go // List recent batches page, _ := client.Batch.List(ctx, &zep.BatchListRequest{Limit: zep.Int(20)}) iter := page.Iterator() for iter.Next(ctx) { b := iter.Current() fmt.Println(*b.UUID, *b.Status) } // List only batches that are still being processed client.Batch.List(ctx, &zep.BatchListRequest{Status: zep.String("processing")}) // Delete a draft batch client.Batch.Delete(ctx, batchUUID) ``` ## Viewing batches in the dashboard The Zep web dashboard provides a batches view showing all batches in your project, their status, item counts, and processing progress. Click into a batch to inspect its individual items and any errors. You can delete a `draft` or `invalid` batch from the list or the batch detail page. Batches that have started processing cannot be deleted. ## Deprecated batch methods The v4 SDKs have no `graph.add_batch` or `thread.add_messages_batch` method. Use `batch.create`, `batch.add_items`, and `batch.process` with `graph_episode` and `thread_message` items, as this page shows. For code that uses the v3 methods, see [Migrating from v3](/migrating-from-v3). > Ingest large historical datasets into your Context Graphs with the Batch API