Skip to content

fix: remove title="Page" requirement in pageurl and mpv regex - #111

Open
FireHope wants to merge 1 commit into
fffonion:masterfrom
FireHope:fix/remove-title-page-dependency
Open

FireHope wants to merge 1 commit into
fffonion:masterfrom
FireHope:fix/remove-title-page-dependency

Conversation

@FireHope

Copy link
Copy Markdown

网站移除了缩略图元素上的 title="Page" 属性,导致 flt_pageurl
和 flt_metadata(MPV检测) 中的正则匹配失败,报错
ERR_NO_PAGEURL_FOUND (1008)。

修复方案:不再依赖脆弱的 title 属性,仅通过 URL 模式
(domain/单字符/10位hex/数字-数字) 来识别图片页链接和 MPV 链接。

网站移除了缩略图元素上的 title="Page" 属性,导致 flt_pageurl
和 flt_metadata(MPV检测) 中的正则匹配失败,报错
ERR_NO_PAGEURL_FOUND (1008)。

修复方案:不再依赖脆弱的 title 属性,仅通过 URL 模式
(domain/单字符/10位hex/数字-数字) 来识别图片页链接和 MPV 链接。

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request simplifies the regular expressions used to extract multi-page viewer and gallery page URLs in xeHentai/filters.py. Feedback was provided to improve the robustness of the multi-page viewer regex by ensuring it matches the full href attribute, which helps accommodate various URL suffixes.

Comment thread xeHentai/filters.py
# TODO: remove <img alt="\d+" once e-hentai is updated to align with exhentai
mpv_urls = re.findall(
'<a href="(%s/mpv/(\d+)/[a-f0-9]{10})/#page\d+">(?:<div|<img alt="\d+") title="Page' % RESTR_SITE,
'<a href="(%s/mpv/(\d+)/[a-f0-9]{10})' % RESTR_SITE,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

For consistency with the change in flt_pageurl and for improved robustness, it would be better to make this regular expression match the entire href attribute. The current regex is a bit too broad as it only matches a prefix of the attribute's value.

By adding [^"*" at the end, you'd match any characters until the closing quote, accommodating potential URL suffixes like /#page1 which was present in the original regex.

Suggested change
'<a href="(%s/mpv/(\d+)/[a-f0-9]{10})' % RESTR_SITE,
'<a href="(%s/mpv/(\d+)/[a-f0-9]{10})[^"]*"' % RESTR_SITE,

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant