files and nextCursor. Select a returned files[].fileId
for the content requests below.
Follow inventory cursors until null if you need every caller-visible stored file.
Select a document using its role, name, and research relevance, then verify its
contents. Stored filenames and search passages alone do not establish the full
meaning of a form or policy.
Read existing extracted text
text, textRevision, byte location, and nextCursor, with
the file ID and a nullable source versionId. See the
text reference
for the schema and size bounds. Byte offsets count UTF-8 bytes, not characters;
a chunk can be shorter than maxBytes to preserve a complete character.
This reads existing extraction only. A downloadable file or indexed search match
need not have readable stored text. 409 text_unavailable permits a fallback to
the original download; an available empty extraction instead succeeds with empty
text and a null cursor. 404 not_found is not the same as missing extraction.
See recovery for access and environment failures.
Save a checkpoint and resume
Save the file ID,maxBytes, textRevision, and the cursor used to request a
chunk alongside its text, byte offsets, and returned nextCursor. Store the chunk
before advancing the checkpoint. When nextCursor is non-null, resume using it:
nextCursor is null.
Cursors expire 24 hours after the first chunk; continuation does not extend their
lifetime, and every request checks access again. On invalid_cursor, restart the
intended text traversal separately. Never append a new revision to old chunks.
See pagination and recovery for failure handling.
Download the original
Use the download operation when you need layout, original pages, or a file without readable extraction. Use the inventory’sversionId when present. If it is null, omit that parameter
to pin the current original version.
File-Version-Id
for citations and retries; the signed URL expires and is not a durable reference.
To retry, request the same file and saved version ID for a fresh redirect. Do not
silently substitute current bytes when an exact version is unavailable. Use
bounded concurrency if downloading multiple files.
Cite what you actually read
The text operation currently returns source
versionId: null; its revision
identifies extracted text only. Embedded page markers do not establish verified
page mapping. For a PDF citation, inspect the actual downloaded page and record
that original’s version. A local PDF-to-text conversion derives from those local
bytes, but it is not an API textRevision.
For example: “Read page 1 of the downloaded form, file ID X, version V; search
evidence version was unknown.” That supports a source-grounded statement without
inventing a version relationship. Keep extracted research data
and your interpretation separate from the original evidence.