---
title: "Content ingestion"
description: "How Edge Assist ingests your content: the JSON index file format, chunking guidance, the hourly refresh cycle and per-site ingestion limits."
url: "https://edge.network/docs/assist/ingestion"
section: "Assist"
---

# Content ingestion

Assist answers only from content you give it. You publish a JSON index of your content at a URL; Assist polls it hourly and re-indexes whenever it changes. Your build pipeline is the source of truth.

![Assist ingestion tab with content sources and index status](/media/docs/control-assist-ingestion.svg)

## The index file

A single JSON document, deliberately small. Most site generators can emit it in a few dozen lines of build script. If an AI agent builds the index for you, hand it [SKILL.md](/assist/SKILL.md): this whole spec in a form an agent can fetch and follow directly.

```json
{
  "hash": "3f9c2a71b4e8d05c",
  "chunks": [
    {
      "url": "/pricing",
      "title": "Pricing",
      "section": "pricing",
      "text": "Storage is billed at $0.015 per GB per month. There are no egress fees within…"
    },
    {
      "url": "/docs/storage/buckets",
      "title": "Buckets",
      "section": "docs",
      "text": "A bucket is a flat namespace for objects. Names must be globally unique and…"
    }
  ]
}
```

| Field | Status | Notes |
|---|---|---|
| `hash` | Required | Any string that changes when the content changes (a content hash is ideal). Assist re-fetches the index hourly but only re-indexes when the hash differs. A changed hash also invalidates the answer cache. |
| `chunks[].url` | Required | The page the text came from, as a relative path (e.g. `/pricing`). Used for citations and search results. |
| `chunks[].title` | Recommended | The page title, shown in search results and citation chips. |
| `chunks[].section` | Optional | A coarse grouping label (e.g. `docs`, `pricing`, `blog`). Shown in search results. |
| `chunks[].text` | Required | The chunk content: plain text, no HTML. |

## Chunking guidance

A chunk is the unit of retrieval: when a visitor asks a question, Assist pulls the best-matching chunks into the answer. Good chunking is most of the quality:

- **Aim for ~1,400 characters per chunk**, split on sentence or paragraph boundaries. A chunk should read as a coherent passage on its own.
- **Skip boilerplate.** Navigation, footers, and cookie banners add noise, not answers. Very short fragments (under ~80 characters) aren't worth indexing.
- **Keep facts with their subject.** If a price or limit lives in a table, render it into the chunk as a sentence ("The free tier includes 250 questions per month") so retrieval can find it.
- **One page, many chunks.** Repeat the same `url`/`title` across a long page's chunks. Citations group by page automatically.

## The refresh cycle

1. Assist fetches your index URL **hourly**.
2. If the `hash` matches the last ingest, nothing happens.
3. If it changed, the site's search index is rebuilt from the new chunks and the [answer cache](/docs/assist/caching) is cleared, so stale answers can't outlive the content they came from.

Deployed something and don't want to wait? **Assist → Ingestion → Re-ingest now** forces a fetch immediately. The same page shows the index status: page and chunk counts, last refresh time, and the last ingest error if a fetch or parse failed.

> [!NOTE]
> Ingestion failures never take the assistant down. It keeps answering from the last good index until a fetch succeeds again.

## Ingestion limits

| Limit | Free | Paid |
|---|---|---|
| Pages (distinct URLs) per site | 500 | 10,000 |
| Chunks per site | 50,000 | 50,000 |

Knowledge bases beyond 10,000 documents are handled by [Expert Services](/solutions/expert-services).

## Other ingestion sources

The index URL is the most direct path for website content. For document collections that don't live on a website, [bucket ingestion](/docs/assist/bucket-ingestion) points Assist at an Edge Storage bucket of markdown, text, HTML, or PDF files, with private access control built in. **Crawl-based** ingestion (give Assist your site URL and it follows your sitemap itself, with no build step) is on the way. Every source feeds the same index and behaves identically downstream.

## Next steps

- [Answer caching](/docs/assist/caching) — What happens to answers when content changes
- [API reference](/docs/assist/api) — Search and ask over the ingested content
