frompdf is a simple CLI tool for extracting structured text from PDFs.
The initial and primary output format is Markdown, with additional output formats planned. frompdf uses robust heuristics to detect paragraphs, headings at various levels, block quotes, and other document features. Running headers and footers are detected and removed, while page numbers can optionally be exported as metadata.
These heuristics can never be perfect, but they should often provide a useful approximation of the actual document content — one that is more useful than plain text extraction for RAG and similar workflows, or for turning read-only PDFs into editable structured text.
While ML-based alternatives such as Docling may handle some details better, they are slower and have considerably higher computational overhead. frompdf's heuristics will not get every detail right, but they are fast, robust, and easy to run locally.
frompdf is released under the permissive MIT License. This can make it easier to use, modify, and integrate than tools based on PyMuPDF, which is available under the GNU AGPL or a commercial license.
The current version provides one command:
frompdf file.pdf- extract Markdown from a PDF
For an input named file.pdf, the command writes:
file.md- Markdown output
With diagnostic options, it can also write:
file-lines.csv- extracted line records with page, block, geometry, font size, and weight datafile-pagenos.csv- raw PDF page numbers mapped to detected or safely inferred visible page labels
If file.md already exists, frompdf renames it to file.md.bak before
writing the new output.
frompdf currently detects and serializes:
- paragraphs
- headings, based mostly on font size plus a document-relative font-weight boost
- block quotes, based on indentation
- repeated headers and footers, which are removed from the Markdown output
- visible page numbers found in removed headers or footers
- optional page-boundary markers embedded in the Markdown output
- document-aware unhyphenation of words split across lines within a block
- normalization of unspaced en and em dashes split across lines within a block
- inline superscripts and subscripts, preserved as
<sup>and<sub>tags
The internal block model tracks the raw PDF page number and, when available, the visible page number for each block.
- Python 3.11 or newer
pdftext, installed through this package's dependencies
Install the latest published version with pip:
pip install frompdfFor an isolated command-line installation, use uv:
uv tool install frompdfAlternatively, use pipx:
pipx install frompdfAll three options make the frompdf command available in your environment.
From a checkout of this repository, install the package in editable mode:
pip install -e .This makes the frompdf command use your local source code, including any
changes you make.
To use pipx with a local checkout, install in editable mode as well:
pipx install -e .Convert a PDF:
frompdf ./document.pdfExample output:
document.md written
Use -m or its long form --page-markers to embed page boundaries in the
Markdown itself:
frompdf -m ./document.pdfEach source page starts with a marker such as <<PAGE:7>>, containing its raw
PDF page number. When a different visible page label was detected or safely
inferred, the marker includes that label too, as in <<PAGE:7|LABEL:9>>.
The LABEL: field distinguishes the visible label from the raw page number,
including for compound visible labels such as <<PAGE:2|LABEL:51:2>>.
Markers are placed after Markdown prefixes such as ## for headings and
> for block quotes. Collected notes retain their source pages, so markers
can move backward in page order and can occur inside a continued note.
Detected footnotes are collected in endnote-style "Notes" sections, with their original numbers. A single uninterrupted numbering sequence is placed at the document end under H2. When numbering restarts, each group prefers a following heading at the level of the shallowest section it covers. This keeps notes spanning several sections out of individual subsections.
A boundary before the next group starts takes priority over a later major heading. Nearby restarts can also place old notes before a mid-page heading above their footnote area. Without a following heading, notes go at the end.
Generated headings default to H2 for one Notes section or H3 for several. Before a deeper heading, Notes matches that level so the following section does not become its child (H4 before H4, for example). A trailing group among several uses at least the level of the sections it covers. Groups sharing a destination share one Notes heading, with separate lists to preserve restarts.
Choose a different section title with --notes-title:
frompdf --notes-title Anmerkungen -m ./document.pdfWrite diagnostic CSV files with --dump-lines and --dump-pagenos:
frompdf --dump-lines --dump-pagenos ./document.pdfExample output:
document-lines.csv written
document-pagenos.csv written
document.md written
--dump-lines writes the extracted line records. --dump-pagenos always
writes one mapping row per PDF page. Missing visible labels are written as
?; gaps are filled when surrounding Arabic or Roman page labels determine
an unambiguous value. The page-marker and diagnostic options are independent;
they can be combined or used separately, and --page-markers and
--dump-pagenos use the same page-number mapping.
Small, raised or lowered characters in running text retain their position
using inline HTML tags, for example text<sup>2</sup> and
GaV<sub>4</sub>S<sub>8</sub>. Rendering these tags requires a Markdown viewer
that allows HTML. Detected footnote labels are still written as ordinary
numbered list markers, such as 2. .
Detection uses character size and position relative to nearby ordinary text,
including when a PDF reports placeholder font sizes or mixes scripts into
ordinary spans. Recognition is conservative; standalone scripts and complex
stacked mathematical notation may remain plain text. Diagnostic line CSVs
retain plain text and include scripts ranges with decoded character offsets.
PDFs do not contain document structure directly, so most higher-level structure has to be inferred. Current limitations include:
- heading detection is heuristic and can miss headings or over-detect short emphasized text
- block quote detection is conservative and currently relies on indentation
- general lists, tables, captions, and code blocks are not modeled as dedicated block types yet
- multi-column and heavily designed PDFs can still produce awkward reading order
- header and footer removal depends on repetition and page-position heuristics
The diagnostic CSV files are part of the workflow: they make it easier to see why a specific line or block was classified the way it was.
Planned next improvements include:
- detection of general lists and preformatted blocks
- merging paragraphs that span more than one page
- correction of font-encoding and ligature-related text extraction errors
- better support for multi-column PDFs
- additional output formats such as HTML, EPUB, ODT, and DOCX
See CONTRIBUTING.md for development setup, style rules, testing commands, and guidance for contributors and coding agents.
Release history and unreleased changes are recorded in CHANGELOG.md.
Development of frompdf has benefited from assistance by ChatGPT and Codex, including support with coding, testing, and documentation.
frompdf is distributed under the MIT license. See LICENSE.txt.