Skip to content

Commit cdad22f

Browse files
committed
Add --skip_errors to skip documents that failed in a previous run
By default a re-run only skips documents whose TEI output already exists, so documents that failed are sent to GROBID again even though, unless something changed, they will fail again. --skip_errors (skip_errors=... in the library) also skips the documents for which a previous run left an error file <stem>_<status>.txt next to the expected TEI output. --force still reprocesses everything. Since the error file now drives a decision and not only reporting, it is kept in sync: it is removed once the document is processed successfully, and a failure with a different status replaces the previous marker instead of accumulating next to it. Closes #119
1 parent bdee903 commit cdad22f

3 files changed

Lines changed: 292 additions & 31 deletions

File tree

‎Readme.md‎

Lines changed: 30 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -126,15 +126,16 @@ grobid_client [OPTIONS] SERVICE
126126

127127
#### Common Options
128128

129-
| Option | Description | Default |
130-
|-------------|--------------------------|-------------------------|
131-
| `--input` | Input directory path | Required |
132-
| `--output` | Output directory path | Same as input |
133-
| `--server` | GROBID server URL | `http://localhost:8070` |
134-
| `--n` | Concurrency level | 10 |
135-
| `--config` | Config file path | Optional |
136-
| `--force` | Overwrite existing files | False |
137-
| `--verbose` | Enable verbose logging | False |
129+
| Option | Description | Default |
130+
|-----------------|---------------------------------------------------|-------------------------|
131+
| `--input` | Input directory path | Required |
132+
| `--output` | Output directory path | Same as input |
133+
| `--server` | GROBID server URL | `http://localhost:8070` |
134+
| `--n` | Concurrency level | 10 |
135+
| `--config` | Config file path | Optional |
136+
| `--force` | Overwrite existing files | False |
137+
| `--skip_errors` | Also skip documents that failed in a previous run | False |
138+
| `--verbose` | Enable verbose logging | False |
138139

139140
#### Processing Options
140141

@@ -173,6 +174,9 @@ grobid_client --server https://grobid.example.com --input ~/citations.txt proces
173174
# Force reprocessing with sentence segmentation and JSON output
174175
grobid_client --input ~/docs --force --segment_sentences --json processFulltextDocument
175176

177+
# Resume an interrupted run without retrying the documents that already failed
178+
grobid_client --input ~/docs --output ~/results --skip_errors processFulltextDocument
179+
176180
# Process PDFs directly from a zip or tar.gz archive (streamed, not fully decompressed)
177181
grobid_client --input ~/papers.zip --output ~/results processFulltextDocument
178182
grobid_client --input ~/papers.tar.gz --output ~/results processFulltextDocument
@@ -203,6 +207,14 @@ grobid_client --input "~/data/**/*.pdf" --output ~/results processFulltextDocu
203207
> A **manifest of paths** (local, glob or `s3://`, one per line, `#` comments allowed) can be processed together via
204208
> `--input-list paths.txt` (combinable with `--input`).
205209
210+
> [!NOTE]
211+
> **Skipping already handled documents.** By default a re-run skips a document only when its TEI output already exists,
212+
> so documents that failed are sent to GROBID again. Since a failed document generally fails again unless something
213+
> changed, `--skip_errors` also skips the documents for which a previous run left an error file
214+
> (`<name>_<status>.txt`, e.g. `paper_500.txt`) next to the expected TEI output. Drop the flag (or use `--force`) to
215+
> retry them. Error files are kept in sync automatically: the marker is deleted once the document is processed
216+
> successfully, and replaced when the same document fails again with a different status code.
217+
206218
### Python Library
207219
208220
#### Basic Usage
@@ -259,6 +271,15 @@ client.process(
259271
markdown_output=True
260272
)
261273

274+
# Re-run without retrying the documents that failed before
275+
client.process(
276+
service="processFulltextDocument",
277+
input_path="/path/to/pdfs",
278+
output_path="/path/to/output",
279+
force=False,
280+
skip_errors=True
281+
)
282+
262283
```python
263284
# Process citation lists
264285
client.process(

0 commit comments

Comments
 (0)