Conversation
网站移除了缩略图元素上的 title="Page" 属性,导致 flt_pageurl 和 flt_metadata(MPV检测) 中的正则匹配失败,报错 ERR_NO_PAGEURL_FOUND (1008)。 修复方案:不再依赖脆弱的 title 属性,仅通过 URL 模式 (domain/单字符/10位hex/数字-数字) 来识别图片页链接和 MPV 链接。 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
There was a problem hiding this comment.
Code Review
This pull request simplifies the regular expressions used to extract multi-page viewer and gallery page URLs in xeHentai/filters.py. Feedback was provided to improve the robustness of the multi-page viewer regex by ensuring it matches the full href attribute, which helps accommodate various URL suffixes.
| # TODO: remove <img alt="\d+" once e-hentai is updated to align with exhentai | ||
| mpv_urls = re.findall( | ||
| '<a href="(%s/mpv/(\d+)/[a-f0-9]{10})/#page\d+">(?:<div|<img alt="\d+") title="Page' % RESTR_SITE, | ||
| '<a href="(%s/mpv/(\d+)/[a-f0-9]{10})' % RESTR_SITE, |
There was a problem hiding this comment.
For consistency with the change in flt_pageurl and for improved robustness, it would be better to make this regular expression match the entire href attribute. The current regex is a bit too broad as it only matches a prefix of the attribute's value.
By adding [^"*" at the end, you'd match any characters until the closing quote, accommodating potential URL suffixes like /#page1 which was present in the original regex.
| '<a href="(%s/mpv/(\d+)/[a-f0-9]{10})' % RESTR_SITE, | |
| '<a href="(%s/mpv/(\d+)/[a-f0-9]{10})[^"]*"' % RESTR_SITE, |
网站移除了缩略图元素上的 title="Page" 属性,导致 flt_pageurl
和 flt_metadata(MPV检测) 中的正则匹配失败,报错
ERR_NO_PAGEURL_FOUND (1008)。
修复方案:不再依赖脆弱的 title 属性,仅通过 URL 模式
(domain/单字符/10位hex/数字-数字) 来识别图片页链接和 MPV 链接。