Skip to content

Non-UTF-8 cache file raises an uncaught UnicodeDecodeError instead of a cache miss #160

Description

@MohammedAlkindi

What happens

If models.json or benchmarks.json on disk is not valid UTF-8, whichllm exits with a traceback instead of refreshing the cache:

UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 81: invalid continuation byte

Both loaders document the opposite. load_cache() says "Returns None if expired or missing", and both carry a handler that logs Cache corrupted and returns None. A corrupt cache is meant to be a cache miss.

Why it escapes

src/whichllm/models/cache.py:38 and src/whichllm/models/benchmark_cache.py:29 catch:

except (json.JSONDecodeError, KeyError) as e:

The read one line above is read_text(encoding="utf-8"), which raises UnicodeDecodeError. That is a ValueError, but it is not a json.JSONDecodeError, so the tuple does not catch it. cli.py calls both loaders bare — cached_data = None if refresh else load_cache() at lines 623, 782, 868 — so it propagates all the way out.

How a file gets there

Two ways, and the first is an upgrade path rather than a hypothetical:

  1. Caches written before 07140d5. That commit added encoding="utf-8"; before it, write_text(json.dumps(data, ensure_ascii=False)) used the locale codepage. On a Windows machine a cached model id containing non-ASCII is on disk now as cp1252 bytes. Upgrading to ≥0.5.12 fixes the writer but does not remove the file already written, so the first run after upgrading reads it back and crashes.
  2. A truncated or interrupted write, which needs no version history at all.

Reproduction

Windows 11, Python 3.13, locale.getpreferredencoding(False) == 'cp1252', on ea32ed2:

payload = {"schema_version": CACHE_SCHEMA_VERSION, "cached_at": time.time(),
           "models": [{"id": "test/Café-Münster"}]}
CACHE_FILE.write_text(json.dumps(payload, ensure_ascii=False), encoding="cp1252")
load_cache()
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 81: invalid continuation byte

The same shape reproduces for load_benchmark_cache().

Suggested fix

Add UnicodeDecodeError to both handlers, so an undecodable cache degrades to a miss like every other corruption does. I have that plus regression tests ready and will open a PR against this issue.

Not the same as #158

#158 reports a charmap decode error from the reader at the old cache.py:28, which 07140d5 fixed in v0.5.12; that reporter is on 0.5.8 and upgrading resolves what they filed. This is the adjacent gap that upgrading does not resolve, because the bad file survives the upgrade. I am not proposing to close #158 with this.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions