Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,7 @@ Before modifying a subsystem, read its corresponding documentation.
| `ARCHITECTURE.md` | High-level system architecture |
| `ROADMAP.md` | Development roadmap and milestones |
| `MODEL_CONTRACTS.md` | ONNX model interfaces and tensor contracts |
| `KV.md` | Decoder KV cache (host vs CUDA) and GPU decode plan |

Avoid duplicating documentation across multiple files. High-level concepts belong in `ARCHITECTURE.md`, while subsystem-specific implementation details belong in their dedicated documents.

Expand Down
8 changes: 5 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -245,8 +245,10 @@ cargo run --features load-dynamic,cuda --example inference -- \
models/lightonocr examples/SROIE-receipt.jpeg default cuda
```

Optional 5th argument is the CUDA `device_id` (default `0`). On CUDA, decode keeps
KV past/present on the GPU via IoBinding; sampling still runs on the host.
Optional 5th argument is the CUDA `device_id` (default `0`). On CUDA, decode
defaults to device-resident KV via IoBinding (`CudaKVCache`); set
`FAST_LIGHTONOCR_CUDA_HOST_KV` to force the host `KVCache` path. Sampling still
runs on the host. See [`docs/KV.md`](docs/KV.md).

To compare CPU thread settings on your machine:

Expand Down Expand Up @@ -307,7 +309,7 @@ The Python bindings use the same native library and build infrastructure.
- ✅ CPU performance work (KV-cache reuse, top-k/top-p, ORT session tuning, decode host reuse, inference bench)
- 🚧 Generation parity and deterministic seeded generation
- 🚧 Broader processor parity coverage
- ✅ CUDA execution provider + device-resident decoder KV (IoBinding)
- ✅ CUDA execution provider + `KVCacheBackend` (`KVCache` / `CudaKVCache`)
- 🚧 CoreML / DirectML execution providers
- 🚧 Python exposure of runtime / EP options

Expand Down
218 changes: 135 additions & 83 deletions bindings/python/README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# fast-lightonocr

> ⚡ Native Python bindings for the Rust **Fast LightOnOCR** inference engine.
> Native Python bindings for the Rust **Fast LightOnOCR** inference engine.

`fast-lightonocr` provides high-performance OCR for documents and images using
Baidu's **LightOnOCR** model. Model inference runs entirely in native Rust,
Expand All @@ -9,25 +9,25 @@ document parsing.

---

## ✨ Features
## Features

- 🚀 Native Rust inference engine
- 🧠 ONNX Runtime backend
- 📄 OCR for documents and images
- 📝 Structured Markdown output
- 📊 Structured HTML table extraction
- 🎨 Configurable table rendering
- 🎛️ Multiple model presets (`default`, `fp16`, `q4`)
- Native Rust inference engine
- ONNX Runtime backend
- OCR for documents and images
- Structured Markdown output
- Structured HTML table extraction
- Configurable table rendering
- Multiple model presets (`default`, `fp16`, `q4`)

---

## 📦 Installation
## Installation

Install with the matching extra for your backend. Published wheels target
**Linux x86_64** and **macOS arm64** (macOS Intel is not published: ONNX
Runtime 1.28 has no compatible wheel there).

### CPU (default)
### CPU

```bash
pip install "fast-lightonocr[cpu]"
Expand All @@ -37,51 +37,37 @@ CPU wheels bundle ONNX Runtime. No extra environment setup is required.

### CUDA

Published CUDA wheels are a dedicated build profile (default PyPI wheels stay
CPU). Install a CUDA-profile package plus the extra:

```bash
pip install "fast-lightonocr[cuda]"
```

Requires a CUDA-enabled package build and a compatible NVIDIA driver. The
`cuda` extra pulls in `onnxruntime-gpu` (CUDA 13 / cuDNN) and `nvidia-cublas`.

Select CUDA at load time:

```python
model = LightOnOCR.from_pretrained(
"onnx-community/LightOnOCR-2-1B-ONNX",
runtime_kwargs={
"execution_provider": "cuda",
"device_id": 0,
},
)
```

When `execution_provider="cuda"`, `from_pretrained` preloads the pip NVIDIA
CUDA/cuDNN libraries (`onnxruntime.preload_dlls`), so `LD_LIBRARY_PATH` is
usually unnecessary. CPU loads never take that path.
Requires a compatible NVIDIA driver. The `cuda` extra pulls in
`onnxruntime-gpu` (CUDA 13 / cuDNN) and `nvidia-cublas`. Select CUDA at load
time with `runtime_kwargs` — see [Runtime options](#runtime-options).

### Building from source

Source installs use the project build backend. It discovers ONNX Runtime from
Run these from `bindings/python`. The build backend discovers ONNX Runtime from
`ORT_DYLIB_PATH` when set, otherwise from the profile’s Python ORT package,
validates ONNX Runtime 1.28.x (C API level 27), and bundles the native runtime
into the wheel.

#### CPU

```bash
cd bindings/python
pip install -v ".[cpu]"
```

# or explicitly:

```bash
BUILD_PROFILE=cpu pip install -v ".[cpu]"
```

#### CUDA

```bash
cd bindings/python
BUILD_PROFILE=cuda pip install -v ".[cuda]"
```

Expand All @@ -90,9 +76,11 @@ ORT CUDA provider plugins (`libonnxruntime_providers_{shared,cuda}`) into the
wheel. The `[cuda]` extra installs the CUDA 13 / cuDNN / cublas user
libraries used at runtime.

For editable/`maturin develop` workflows, see [Development](#development).

---

## 🚀 Quick Start
## Quick Start

```python
from fast_lightonocr import LightOnOCR
Expand All @@ -109,80 +97,95 @@ them locally.

---

## 📄 OCR Results

The raw model output is available through `result.text`.
## Configuration

```python
print(result.text)
```
`from_pretrained()` accepts a model preset plus two override dicts:
`runtime_kwargs` (ONNX Runtime sessions) and `generation_kwargs` (decode).

The Python bindings also expose a parsed document representation that extracts
embedded HTML tables while preserving the original document structure.
### Model presets

```python
print(result.document)
model = LightOnOCR.from_pretrained(
"onnx-community/LightOnOCR-2-1B-ONNX",
preset="q4",
)
```

Tables can be accessed directly:
Available presets:

```python
for table in result.tables:
print(table.text_rows)
```
- `default`
- `fp16`
- `q4`

---
### Runtime options

## 📋 Table Rendering
Override ONNX Runtime session settings at load time with `runtime_kwargs`.
Unknown keys raise `ValueError`. These options are applied **before** sessions
are created and cannot be changed after load.

By default, tables are rendered using ASCII borders.
Supported keys:

```python
result = model.process(
"receipt.jpg",
table_format="grid",
)
```
| Key | Type | Default | Notes |
| --- | --- | --- | --- |
| `execution_provider` | `"cpu"` \| `"cuda"` | `"cpu"` | `"cuda"` requires a CUDA-enabled build and `[cuda]` extra |
| `device_id` | `int` | `0` | CUDA device index |
| `intra_threads` | `int` | host parallelism | Intra-op threads (no effect if ORT is built with OpenMP; use `OMP_NUM_THREADS`) |
| `inter_threads` | `int` | `1` | Used only when `parallel_execution` is `True` |
| `parallel_execution` | `bool` | `False` | ORT parallel execution mode |

Markdown tables are also supported.
CUDA:

```python
result = model.process(
"receipt.jpg",
table_format="github",
from fast_lightonocr import LightOnOCR

model = LightOnOCR.from_pretrained(
"onnx-community/LightOnOCR-2-1B-ONNX",
preset="q4",
runtime_kwargs={
"execution_provider": "cuda",
"device_id": 0,
},
generation_kwargs={
"max_new_tokens": 1024,
"do_sample": False,
},
)
```

Any table format supported by `tabulate` may be used.
result = model.process("receipt.jpg")
print(result.text)
```

---
When `execution_provider="cuda"`, `from_pretrained` preloads the pip NVIDIA
CUDA/cuDNN libraries (`onnxruntime.preload_dlls`). CPU loads never take that
path. Autoregressive decode keeps KV past/present on the GPU after the first
step (IoBinding); token sampling still runs on the host.

## ⚙️ Model Presets
If CUDA EP registration fails with a missing `libcublasLt` / provider `.so`,
add the pip `nvidia/*/lib` directories and the driver (`libcuda`) to
`LD_LIBRARY_PATH` for that process (common on some notebook runtimes).

`from_pretrained()` supports three ONNX model presets.
CPU thread tuning:

```python
model = LightOnOCR.from_pretrained(
"onnx-community/LightOnOCR-2-1B-ONNX",
preset="q4",
runtime_kwargs={
"execution_provider": "cpu",
"intra_threads": 8,
},
)
```

Available presets:

- `default`
- `fp16`
- `q4`

### Generation overrides

Model defaults come from Hugging Face `generation_config.json` (typically
`do_sample=True`, `temperature=0.2`, `top_k=0`, `top_p=0.9`).
`do_sample=True`, `temperature=0.2`, `top_k=0`, `top_p=0.9`).

Override them at load time with `generation_kwargs` (merged onto the decoder config; unknown keys raise `ValueError`):
Override them at load time with `generation_kwargs` (merged onto the decoder
config; unknown keys raise `ValueError`):

```python
# Faster / deterministic OCR on CPU (greedy decoding)
# Faster / deterministic OCR (greedy decoding)
model = LightOnOCR.from_pretrained(
"onnx-community/LightOnOCR-2-1B-ONNX",
preset="q4",
Expand Down Expand Up @@ -220,13 +223,60 @@ Bare `max_new_tokens=` remains supported as a shorthand:
model = LightOnOCR.from_pretrained("...", max_new_tokens=1024)
```

> On CPU, prefer `do_sample=False` for throughput.
On CPU, prefer `do_sample=False` for throughput. If you need sampling, set a
modest `top_k` (for example `50`) instead of leaving the HF default `top_k=0`.

---

## OCR Results

The raw model output is available through `result.text`.

```python
print(result.text)
```

The Python bindings also expose a parsed document representation that extracts
embedded HTML tables while preserving the original document structure.

```python
print(result.document)
```

Tables can be accessed directly:

```python
for table in result.tables:
print(table.text_rows)
```

---

## Table Rendering

> If you need sampling, set a modest `top_k` (for example `50`) instead of leaving the HF default `top_k=0`.
By default, tables are rendered using ASCII borders.

```python
result = model.process(
"receipt.jpg",
table_format="grid",
)
```

Markdown tables are also supported.

```python
result = model.process(
"receipt.jpg",
table_format="github",
)
```

Any table format supported by `tabulate` may be used.

---

## 🛠 Development
## Development

Install the project and development dependencies:

Expand Down Expand Up @@ -261,6 +311,8 @@ poetry run pip wheel . --wheel-dir dist

# CUDA
BUILD_PROFILE=cuda poetry run pip wheel . --wheel-dir dist
# then install the wheel with the CUDA extra, e.g.
# pip install "dist/fast_lightonocr-<ver>-*.whl[cuda]"
```

> **Note**
Expand All @@ -272,10 +324,10 @@ BUILD_PROFILE=cuda poetry run pip wheel . --wheel-dir dist

---

## 🙏 Acknowledgements
## Acknowledgements

This package wraps the native Rust **Fast LightOnOCR** inference engine and
uses the open-weight **LightOnOCR** model released by Baidu.

- 🤗 https://huggingface.co/onnx-community/LightOnOCR-2-1B-ONNX
- 💻 https://github.com/baidu/LightOnOCR
- https://huggingface.co/onnx-community/LightOnOCR-2-1B-ONNX
- https://github.com/baidu/LightOnOCR
Loading