> ## Documentation Index
> Fetch the complete documentation index at: https://docs.canary.effectiveai.app/llms.txt
> Use this file to discover all available pages before exploring further.

# Read and cite document evidence

> Read existing text progressively, resume after interruption, and retain honest provenance for original downloads.

Start with a file ID from [search](/guides/search-filings#search-document-text) or
the [file inventory](/api-reference/filings/list-stored-filing-files). Use
[connection setup](/guides/index#connect-once) for authentication and request
conventions.

```http theme={null}
GET /api/v2/filings/{filingId}/files?limit=20
```

The response contains `files` and `nextCursor`. Select a returned `files[].fileId`
for the content requests below.

Follow inventory cursors until null if you need every caller-visible stored file.
Select a document using its role, name, and research relevance, then verify its
contents. Stored filenames and search passages alone do not establish the full
meaning of a form or policy.

## Read existing extracted text

```http theme={null}
GET /api/v2/files/{fileId}/text?maxBytes=16384
```

An abbreviated response fragment:

```json theme={null}
{
  "text": "Policy wording...",
  "textRevision": "<revision>",
  "versionId": null,
  "location": { "startOffsetBytes": 0, "endOffsetBytes": 17 },
  "nextCursor": null
}
```

A success returns `text`, `textRevision`, byte `location`, and `nextCursor`, with
the file ID and a nullable source `versionId`. See the
[text reference](/api-reference/files/read-a-bounded-chunk-of-existing-extracted-text)
for the schema and size bounds. Byte offsets count UTF-8 bytes, not characters;
a chunk can be shorter than `maxBytes` to preserve a complete character.

This reads existing extraction only. A downloadable file or indexed search match
need not have readable stored text. `409 text_unavailable` permits a fallback to
the original download; an available empty extraction instead succeeds with empty
text and a null cursor. `404 not_found` is not the same as missing extraction.
See [recovery](/guides/pagination-and-recovery) for access and environment failures.

## Save a checkpoint and resume

Save the file ID, `maxBytes`, `textRevision`, and the cursor used to request a
chunk alongside its text, byte offsets, and returned `nextCursor`. Store the chunk
before advancing the checkpoint. When `nextCursor` is non-null, resume using it:

```http theme={null}
GET /api/v2/files/{fileId}/text?maxBytes=16384&cursor={nextCursor}
```

Keep the same file and size. If a response is lost, retry the same input request;
deduplicate saved chunks by file ID, revision, and starting byte offset. Returned
cursor strings may differ on retry. Stop when `nextCursor` is null.

Cursors expire 24 hours after the first chunk; continuation does not extend their
lifetime, and every request checks access again. On `invalid_cursor`, restart the
intended text traversal separately. Never append a new revision to old chunks.
See [pagination and recovery](/guides/pagination-and-recovery) for failure handling.

## Download the original

Use [the download operation](/api-reference/files/download-one-immutable-file-version)
when you need layout, original pages, or a file without readable extraction.
Use the inventory's `versionId` when present. If it is null, omit that parameter
to pin the current original version.

```http theme={null}
GET /api/v2/files/{fileId}/download?versionId={versionId}
```

The redirect identifies the selected version:

```http theme={null}
HTTP/1.1 302 Found
File-Version-Id: <opaque-version>
Location: <short-lived-download-url>
```

Follow the signed URL without forwarding API authorization to the storage host.
After the download completes, inspect the actual bytes. Keep `File-Version-Id`
for citations and retries; the signed URL expires and is not a durable reference.
To retry, request the same file and saved version ID for a fresh redirect. Do not
silently substitute current bytes when an exact version is unavailable. Use
bounded concurrency if downloading multiple files.

## Cite what you actually read

| Evidence | Keep | Do not infer |
| - | - | - |
| Search passage | Filing/source reference, file ID, passage and available page metadata | A match to today's original bytes when evidence `versionId` is null |
| API text | File ID, `textRevision`, start/end byte offsets, observation time | Original-file version or verified source-page coordinates |
| Downloaded original | File ID, returned `File-Version-Id`, observed pages, optional local digest | That search or extracted text came from this same version |

The text operation currently returns source `versionId: null`; its revision
identifies extracted text only. Embedded page markers do not establish verified
page mapping. For a PDF citation, inspect the actual downloaded page and record
that original's version. A local PDF-to-text conversion derives from those local
bytes, but it is not an API `textRevision`.

For example: “Read page 1 of the downloaded form, file ID X, version V; search
evidence version was unknown.” That supports a source-grounded statement without
inventing a version relationship. Keep [extracted research data](/guides/filing-research-data)
and your interpretation separate from the original evidence.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.