Documentation/Document ingestion pipeline
02 / PREPARE THE SOURCE

Document ingestion pipeline

The stages that turn an uploaded file into searchable, versioned evidence.

Citely document library with files and processing status.Citely document library with files and processing status.
Processing status remains visible until a document is ready.Sample documents · Select image to enlarge ↗
An upload is only the beginning. — Day 2 Problem poster from Citely’s Building in Public series.
Day 2 / ProblemAn upload is only the beginning.Open full-size poster ↗
Split the content. Keep its address. — Day 2 Mechanism poster from Citely’s Building in Public series.
Day 2 / MechanismSplit the content. Keep its address.Open full-size poster ↗
The way back starts at upload. — Day 2 Takeaway poster from Citely’s Building in Public series.
Day 2 / TakeawayThe way back starts at upload.Open full-size poster ↗
01

1. Intake and ownership

The upload route validates the authenticated owner, file type and request metadata before recording a document. Citely returns a document record and processes it asynchronously so the interface can show queued, processing, ready or failed states.

02

2. Extraction

A parser is selected for the file family. Extracted text is normalized into page-oriented content while the original file remains available to supported readers. Extraction metadata includes a version so later citations and annotations can be checked against the same text representation.

03

3. Structure and chunking

The pipeline divides extracted content into retrievable chunks while retaining source, page and character-position metadata. Page-relative ranges are assigned before indexing; this is what makes precise passage navigation possible after retrieval.

04

4. Embedding and indexes

Chunks are embedded and written to the vector index. A lexical representation is also retained for keyword matching. A document becomes ready only after its searchable representation is complete; failed work is reported instead of presenting an incomplete index as usable.

05

Format scope

The interface currently accepts PDF, text, Markdown, CSV, JSON, DOCX, PPTX, XLSX and common image extensions. An accepted extension describes intake configuration, not equal processing quality. Layout, scanned text, formulas, charts and complex Office structures can produce different results and should be checked in the reader.