Document ingestion pipeline
The stages that turn an uploaded file into searchable, versioned evidence.


1. Intake and ownership
The upload route validates the authenticated owner, file type and request metadata before recording a document. Citely returns a document record and processes it asynchronously so the interface can show queued, processing, ready or failed states.
2. Extraction
A parser is selected for the file family. Extracted text is normalized into page-oriented content while the original file remains available to supported readers. Extraction metadata includes a version so later citations and annotations can be checked against the same text representation.
3. Structure and chunking
The pipeline divides extracted content into retrievable chunks while retaining source, page and character-position metadata. Page-relative ranges are assigned before indexing; this is what makes precise passage navigation possible after retrieval.
4. Embedding and indexes
Chunks are embedded and written to the vector index. A lexical representation is also retained for keyword matching. A document becomes ready only after its searchable representation is complete; failed work is reported instead of presenting an incomplete index as usable.
Format scope
The interface currently accepts PDF, text, Markdown, CSV, JSON, DOCX, PPTX, XLSX and common image extensions. An accepted extension describes intake configuration, not equal processing quality. Layout, scanned text, formulas, charts and complex Office structures can produce different results and should be checked in the reader.


