Index R2 File
Index a single file from a Cloudflare R2 bucket into a collection.
Headers:
- Authorization: Bearer {api_key} - Captain API key for authentication
- X-Organization-ID: Organization UUID
Args:
collection_name: Name of the collection (path parameter)
body: R2 file configuration with file_uri (r2://bucket/path/to/file.pdf)
Returns:
{ job_id, status: "pending" }
Path parameters
collection_name
Headers
Request
This endpoint expects an object.
access_key_id
account_id
bucket_name
file_uri
R2 object URI in the format r2://bucket-name/path/to/file.pdf
processing_type
Document processing type. 'advanced' uses agentic OCR with AI-enhanced extraction for complex layouts, tables, figures, charts, and documents containing images. 'basic' provides reliable OCR optimized for general document indexing and high-volume processing.
secret_access_key
custom_metadata
Custom metadata to attach to all chunks from this file. Keys must be strings. Values: str, int, float, bool, or List[str].
jurisdiction
overwrite_existing
When true, files that already exist in the collection will be deleted and re-indexed with the latest changes. Requires skip_existing=false. Setting both to true returns a 400 error.
parsing_script
Relative path to a JS parsing script for JSON files (e.g. 'research/paper-parser'). When provided, .json files are processed through a sandboxed V8 isolate. Without this, .json files are indexed as raw text.
skip_existing
When true, files already indexed in the collection are skipped and will not be re-indexed with incoming changes. When false, all incoming files are indexed regardless of whether they already exist.
mask_pii
When true, detected PII (emails, phone numbers, SSNs, credit cards, names, and locations) is masked in the parsed content before it is embedded and stored — replaced with entity tags like <PERSON> and <EMAIL_ADDRESS>. For images (including images embedded in PDFs), PII text visible in the image is also pixel-redacted. Opt-in; defaults to false, which leaves content unchanged.
transcription_language
AWS Transcribe language code for the spoken audio (e.g. 'es-US', 'pt-BR'). Omit to auto-detect per file. Video and audio files only. Supported codes: https://docs.aws.amazon.com/transcribe/latest/dg/supported-languages.html
Response
Successful Response
job_id
status
Errors
400
Bad Request Error