Skip to content

feat(read): add document paging and retain PDF text coverage - #3345

Merged
bobleer merged 2 commits into
GCWing:mainfrom
bobleer:bob/read-document-coverage
Oct 10, 2026
Merged

bobleer merged 2 commits into
GCWing:mainfrom
bobleer:bob/read-document-coverage

Conversation

@bobleer

@bobleer bobleer commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Read now accepts optional pages="1-3,7" so an agent can narrow a large document after seeing truncated output. It also preserves readable PDF text alongside explicit OCR gaps instead of rejecting the whole document.

  • Select original PDF pages, PPTX slides, or XLSX visible sheets in source order. DOCX and other reflowable/legacy formats expose explicitly labeled 1,600-character text chunks; these are not printed page numbers.
  • Return page_kind, page_count, normalized selected_pages, selected_page_count, and next_page. Existing offset/limit/tail then address Markdown lines within the same selection; next_offset guides continuation. Omitted pages preserves full extraction.
  • Scope PDF extraction status and OCR gaps to selected pages while retaining original page numbers and the whole-document page count. Unselected pages are not claimed as read.
  • Upgrade anydoc from 0.1.6 to 0.2.4. PDF uses its existing offline pdf-inspector page API; Office selection filters the in-memory OOXML manifest and preserves relationships, styles, shared strings and resources. Normalized selections participate in cache keys.
  • Distinguish empty extraction from an empty source file and line clipping from further line windows. Simplify Read/Edit and shared/Cowork file guidance, removing duplicated rules and incorrect local-workspace assumptions.

Type and Areas

Type: feature, bug fix, dependency update, prompt refinement, regression tests.

Areas: portable tool-runtime document/read primitives, Core Read/Edit tools, built-in Agent prompts.

Motivation / Impact

A PDF with readable pages 1 and 3 and scanned page 2 yields usable text plus an explicit page-2 gap. Reading pages="3" returns that source page and reports coverage for that selection. For a large spreadsheet, the agent can select one visible worksheet and then page through its Markdown lines.

The existing request shape, CSV source default, permissions and feature-off behavior remain supported. Invalid ranges and incompatible source/text requests fail explicitly. Extraction stays offline behind document-read; no OCR, model download, host-specific IO or new remote protocol is introduced.

Selection limits the extracted representation, not source transfer: the 64 MiB source limit and parser budgets still apply, and PDF/DOCX parsing may process the full document. The retained selected Markdown limit remains 16 MiB. Images, diagrams and hidden spreadsheet content are not claimed as extracted.

Verification

All 14 GitHub PR checks passed for cf8e3f977d7cfeb62e53ec849ebf1709c6adffdb (CI run). The first macOS CLI attempt hit child-process exit timeouts; the same commit passed on rerun after the full local CLI contract target also passed.

Local focused checks:

  • cargo test --locked -p tool-runtime --features document-read --lib -- fs::document fs::read_file — 31 passed.
  • cargo test --locked -p openbitfun-core --no-default-features --features agent-runtime,git,document-read --lib -- file_read_tool::tests file_edit_tool::tests prompt_builder::prompt_builder_impl::tests — 45 passed.
  • cargo test --locked -p openbitfun-core --no-default-features --features agent-runtime,git --lib file_read_tool::tests — 10 passed.
  • cargo test --locked -p openbitfun-agent-content --test prompt_catalog_contracts — 8 passed.
  • cargo test --locked -p openbitfun-cli --test cli_command_contracts — 40 passed on macOS after investigating CI subprocess-exit timeouts; the focused 8 streaming cases also passed.
  • pnpm run fmt:rs
  • pnpm run check:core-boundaries — passed.
  • git diff --check — passed.
  • node scripts/check-git-object-sizes.mjs --base upstream/main --head HEAD — 20 changed Git blobs passed.

Reviewer Notes

AI-assisted implementation with focused local regression testing. Fixtures cover text/mixed/scanned PDFs, disjoint/out-of-range selections, original PDF numbering and selection-scoped OCR, native slide order, hidden sheets, UTF-16/Strict OOXML, custom main-part locations, shared strings/assets, selection-specific cache results, Unicode/long-paragraph chunks, line/tail continuation, legacy requests and feature-off behavior. No model A/B evaluation was performed.

Remote workspace byte IO, selected PDF/text-chunk reads and missing-provider failure are exercised through a bound provider mock, including the bounded transfer path and no local fallback. Real SSH, Remote Control, Peer Device Mode and Detached Dispatch end-to-end runs were not performed locally.

Checklist

  • This PR is focused and does not include secrets, temporary prompts, generated scratch files, or unrelated artifacts.
  • Relevant verification is recorded above, or skipped checks are explained.
  • User-facing strings, docs, and locales are updated where applicable.

@bobleer bobleer changed the title fix(read): retain PDF text and report extraction gaps feat(read): add document paging and retain PDF text coverage Oct 10, 2026
@bobleer
bobleer merged commit 3957ef6 into GCWing:main Oct 10, 2026
25 of 27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant