Skip to content

fix(ocr_bench_v2): write text spotting files as UTF-8 - #1804

Open
RizgarOzan wants to merge 1 commit into
modelscope:mainfrom
RizgarOzan:fix/utf8-loc
Open

RizgarOzan wants to merge 1 commit into
modelscope:mainfrom
RizgarOzan:fix/utf8-loc

Conversation

@RizgarOzan

Copy link
Copy Markdown

Follow-up to the note in #1522 that a few encoding issues are left under benchmarks/.

spotting_evaluation in the OCRBench v2 scorer writes the gt and submission files with the locale encoding, but the scorer reads them back with decode_utf8. On Windows the locale encoding is the ANSI code page, so as soon as a prediction has a character outside it, for example Café → Exit, the write raises UnicodeEncodeError and the sample never gets scored.

I pass encoding='utf-8' to both writes. The new test in tests/benchmark/test_ocr_bench_v2_spotting.py gives open() a cp1252 default and reads the zipped files back with the real decode_utf8, so it fails on Linux too without the fix:

without the fix: 1 failed  (UnicodeEncodeError in cp1252.py)
with the fix:    1 passed

pre-commit run --files (ruff check, ruff format and the other hooks) passes on both files. I couldn't run the full spotting metric on Windows because Polygon3 has no wheel there, so the test stubs main_evaluation and checks what the scorer would read.

spotting_evaluation wrote the gt and submission files with the locale
encoding, and the scorer reads them back with decode_utf8. On Windows the
locale encoding is the ANSI code page, so a prediction like "Café → Exit"
raised UnicodeEncodeError before scoring started.

I pass encoding='utf-8' to both writes and added a test that gives open()
a cp1252 default, so it fails the same way on Linux without the fix.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant