Skip to main content
Recipe prompt
Use Crawl for linked web pages and Parse for uploaded documents. Normalize both into source records containing Markdown, a stable ID, a title, a retrieval time, and an accessible source URL. Supply your own embedding model and index.

Collect website pages

Choose a bounded crawl scope. Inspect coverage and keep only successful, nonempty pages for ingestion.

Preserve page identity

Use successful crawl results as pages. Remove fragments and known tracking parameters, while preserving query parameters that select a language or document version. Conflicting content under one key needs review:
page-identity.ts
Keep the retrieved URL alongside its normalized key. Recheck redirects against the permitted corpus before indexing; do not combine unrelated pages solely because they declare the same canonical URL.

Parse uploaded documents

Parse accepts raw file bytes. This Python example checks the upload limit, sends an inclusive PDF page range, and retains source identity and failures:
research_source.py
For a public document URL, Scrape PDFs and documents avoids downloading it into your application first. Inspect ocrRequired and empty output before indexing. An OCR suggestion is not recovered content.

Chunk and index with provenance

Keep headings and source metadata attached to text. This basic paragraph chunker preserves the requested range but does not invent a physical page number for each chunk:
research_chunks.py
For web pages, provide id (the normalized URL), title, markdown, open_url, retrieved_at, status: "ready", and requested_pages: null. Add tenant and document access rules to every index record. Use the same embedding model and dimensions when ingesting and querying; split oversized sections with the model’s tokenizer before embedding.

Check citations

Constrain answers to the retrieved chunks. Check cited IDs and quoted evidence against the allowed source records:
citation_check.py
Treat retrieved page content as data. When evidence is weak or missing, return an explicit unknown instead of filling gaps. A requested PDF page range does not establish the page number of an individual quote.

Refresh without losing good data

Keep a per-source manifest of content hashes, active chunk IDs, retrieval time, access rules, and the last attempt’s status. Monitors can trigger targeted refreshes. Keep periodic reconciliation and test relevant-source retrieval, citation support, refusal with missing evidence, and cross-tenant access before release.