Documents and indexing
The Documents area is the source library for a workspace. It combines file management with the indexing controls that make content available to research.
Add and organize files
Section titled “Add and organize files”You can upload individual files, upload a complete folder structure, or drag files and folders into the document area. The library supports creating, renaming, and moving folders, plus table and tree views for collections of different sizes.
Use search, sorting, and status filters to find a file or monitor a large ingestion job. The table view can show file type, page count, size, status, progress, extraction method, chunk count, and quality information.
Understand indexing status
Section titled “Understand indexing status”| Status | Meaning | Available to research? |
|---|---|---|
| Queued | Waiting for an indexing worker | No |
| Processing | Content is being extracted or indexed | No |
| Completed | Indexing finished successfully | Yes, when enabled |
| Failed | Processing did not complete | No |
| Canceled | The indexing job was stopped | No |
You can monitor active and previous jobs, cancel an active job, and reindex after changing the source or extraction settings.
Choose an extraction approach
Section titled “Choose an extraction approach”PDFs can use the workspace strategy or a document-specific override:
- Text layer reads text already embedded in a digital PDF.
- OCR recognizes text from page images and is appropriate for scans.
- Vision model interprets a full page when meaning depends on visual structure.
- Adaptive starts with the text layer and escalates low-quality results through OCR and then the vision model.
Workspace managers can also configure table structure, formulas, code blocks, pictures, and OCR. More intensive extraction may improve difficult documents, but it takes more compute and time.
Enable or disable a document
Section titled “Enable or disable a document”Disabling a document removes it from the searchable collection without deleting the file. Use this when a source should remain in the workspace but must not influence current research.
Enable it again to return it to the research collection.
Inspect the result
Section titled “Inspect the result”Open a document to compare the source preview with its indexed Markdown. The analysis view can show page coverage, chunks, embedding coverage, and quality information while keeping the source and extracted content aligned.
Users with the required permission can correct a specific indexed Markdown chunk. Reindexing may replace manual chunk edits, so prefer correcting the source file when possible.
Quality checklist
Section titled “Quality checklist”Before starting high-impact research:
- Confirm all required documents show Completed.
- Confirm those documents are enabled.
- Spot-check scanned pages, dense tables, formulas, and images.
- Compare important extracted passages with the original source.
- Reprocess low-quality documents with a more suitable extraction method.
See supported files for the complete format groups.