Skip to content

Documents and indexing

The Documents area is the source library for a workspace. It combines file management with the indexing controls that make content available to research.

You can upload individual files, upload a complete folder structure, or drag files and folders into the document area. The library supports creating, renaming, and moving folders, plus table and tree views for collections of different sizes.

Use search, sorting, and status filters to find a file or monitor a large ingestion job. The table view can show file type, page count, size, status, progress, extraction method, chunk count, and quality information.

Status Meaning Available to research?
Queued Waiting for an indexing worker No
Processing Content is being extracted or indexed No
Completed Indexing finished successfully Yes, when enabled
Failed Processing did not complete No
Canceled The indexing job was stopped No

You can monitor active and previous jobs, cancel an active job, and reindex after changing the source or extraction settings.

PDFs can use the workspace strategy or a document-specific override:

  • Text layer reads text already embedded in a digital PDF.
  • OCR recognizes text from page images and is appropriate for scans.
  • Vision model interprets a full page when meaning depends on visual structure.
  • Adaptive starts with the text layer and escalates low-quality results through OCR and then the vision model.

Workspace managers can also configure table structure, formulas, code blocks, pictures, and OCR. More intensive extraction may improve difficult documents, but it takes more compute and time.

Disabling a document removes it from the searchable collection without deleting the file. Use this when a source should remain in the workspace but must not influence current research.

Enable it again to return it to the research collection.

Open a document to compare the source preview with its indexed Markdown. The analysis view can show page coverage, chunks, embedding coverage, and quality information while keeping the source and extracted content aligned.

Users with the required permission can correct a specific indexed Markdown chunk. Reindexing may replace manual chunk edits, so prefer correcting the source file when possible.

Before starting high-impact research:

  1. Confirm all required documents show Completed.
  2. Confirm those documents are enabled.
  3. Spot-check scanned pages, dense tables, formulas, and images.
  4. Compare important extracted passages with the original source.
  5. Reprocess low-quality documents with a more suitable extraction method.

See supported files for the complete format groups.