Recipe prompt
Collect website pages
Choose a bounded crawl scope. Inspect coverage and keep only successful, nonempty pages for ingestion.Preserve page identity
Use successful crawl results aspages. Remove fragments and known tracking parameters, while preserving query parameters that select a language or document version. Conflicting content under one key needs review:
page-identity.ts
Parse uploaded documents
Parse accepts raw file bytes. This Python example checks the upload limit, sends an inclusive PDF page range, and retains source identity and failures:research_source.py
ocrRequired and empty output before indexing. An OCR suggestion is not recovered content.
Chunk and index with provenance
Keep headings and source metadata attached to text. This basic paragraph chunker preserves the requested range but does not invent a physical page number for each chunk:research_chunks.py
id (the normalized URL), title, markdown, open_url, retrieved_at, status: "ready", and requested_pages: null. Add tenant and document access rules to every index record. Use the same embedding model and dimensions when ingesting and querying; split oversized sections with the model’s tokenizer before embedding.
Check citations
Constrain answers to the retrieved chunks. Check cited IDs and quoted evidence against the allowed source records:citation_check.py
Refresh without losing good data
Keep a per-source manifest of content hashes, active chunk IDs, retrieval time, access rules, and the last attempt’s status.
Monitors can trigger targeted refreshes. Keep periodic reconciliation and test relevant-source retrieval, citation support, refusal with missing evidence, and cross-tenant access before release.