Task.17】DeepEyes processor 动态 patch - #188
Closed
yuki-younai wants to merge 3 commits into
Closed
yuki-younai wants to merge 3 commits into
yuki-younai wants to merge 3 commits into
Conversation
…service - add examples/deepeyes/run_deepeyes_baseline.sh: local 4-GPU Qwen3-VL-4B GRPO baseline derived from run_deepeyes_r3.sh (no cp overlay, no patch), using remote OpenAI-compatible LLM judge. - add examples/deepeyes/openai_judge_service.sh: sources DEEPEYES_JUDGE_* env vars + injects into Ray runtime_env, replaces the local sglang judge. For local baseline testing only; not part of the task-17 patch. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Replace the `cp examples/deepeyes/qwen_vl.py /sgl-workspace/...` whole-file
overlay with a runtime monkey-patch applied at SGLang engine start-up,
injecting only the real DeepEyes-specific differences.
Why: DeepEyes multi-turn rollout sends pre-tokenized `input_ids` (list[int])
to SGLang where each image is already expanded into N `<|image_pad|>` tokens.
The stock `QwenVLImageProcessor.process_mm_data_async` mishandles this:
(A) decode->retokenize drifts the token count, misaligning mrope; (B) the N
image-pad tokens get re-expanded into N×M. The old fix copied a 454-line
legacy `qwen_vl.py` over the install dir, hardcoding `/sgl-workspace` and
breaking on SGLang upgrades (and silently reverting the docker sglang.patch).
Changes:
- add `relax/backends/sglang/patches/qwen_vl_patch.py`: patch
`QwenVLImageProcessor.process_mm_data_async` with two surgical changes —
`_strip_image_token` collapses consecutive `<|image_pad|>` before
`load_mm_data`, and the original `input_ids` are restored + mrope
recomputed from them. Delegates to upstream primitives; idempotent via
`_PATCH_FLAG`; `_verify_upstream_api` fails fast on incompatible sglang.
- add `--deepeyes-qwen-vl-patch` CLI flag, translated to
`RELAX_DEEPEYES_QWEN_VL_PATCH=1` in `sglang_engine._init_normal` and
applied in `_launch_server_with_patches` (same pattern as the OPD patch).
- remove `examples/deepeyes/qwen_vl.py` and the `cp`/`/sgl-workspace`
hardcode in `run_deepeyes_r3.sh`; gate the patch via `DEEPEYES_PATCH` env
in both r3 and baseline scripts. Default off = stock SGLang behavior.
- add unit tests (stubbed sglang, no GPU) covering strip logic, idempotency,
fail-fast, and the text/list input paths.
- add bilingual `docs/{en,zh}/guide/sglang-patches.md` upgrade guide +
register in VitePress sidebar.
Verified: 15 unit tests pass; `_verify_upstream_api` + `apply_qwen_vl_patches`
succeed against real sglang 0.5.12.post1; end-to-end DeepEyes GRPO smoke
(--deepeyes-qwen-vl-patch) completes multiple rollout->train steps with no
token-drift/mrope errors.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Standalone verification for task-17 acceptance redai-studio#4 (same input, key outputs consistent before/after the patch). Compares a faithful re-implementation of the stock upstream `process_mm_data_async` against the patched one on a stubbed `self` (no GPU / real model needed): - Case A (raw-text prompt): patch is a no-op — stock==patch on every key. - Case B (pre-tokenized list[int], N image_pad per image): patch restores the original input_ids and recomputes mrope from them; stock drifts to the retokenized sequence. - Case C: `_strip_image_token` collapse unit check. Run inside the relaxrl image: python3 examples/deepeyes/verify_patch_consistency.py Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
yuki-younai
requested review from
NINGBENZHE,
Yangruipis and
yxyOo
as code owners
July 30, 2026 12:44
Member
|
验收通过 |
Member
|
#164 merged,本 PR close |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
【Task.17】DeepEyes processor 动态 patch
背景
examples/deepeyes/run_deepeyes_r3.sh通过cp把整份examples/deepeyes/qwen_vl.py(454 行的旧版 SGLangqwen_vl.py)覆盖到 SGLang 安装目录:问题:
/sgl-workspace,SGLang 升级后路径/版本不匹配即失效。docker/patch/latest/sglang.patch的改动(该 patch 在 v0.5.12 把load_mm_data切到legacy_load_mm_data),引入隐蔽不一致。目标
启动时用 monkey-patch 注入真实差异,删除
cp覆盖与/sgl-workspace硬编码。真实差异分析
DeepEyes 多轮 rollout 向 SGLang
/generate发送预分词的input_ids: list[int](见examples/deepeyes/rollout.py::_run_inference_step),每张图已被前一轮 processor 展开成 N 个<|image_pad|>。SGLang 原生QwenVLImageProcessor.process_mm_data_async处理不了这种输入:load_mm_data把 int 列表 decode 回文本再重新分词,非无损,token 数可能变化,错位 mrope 位置,破坏 loss_mask / log-prob。<|image_pad|>被当成 N 个图片占位符,每个再展开成 M 个 patch token。真正的定制只有两处,其余 ~100 行都是 SGLang 原生样板。
方案
沿用 Relax 已有的「SGLang 启动时 monkey-patch」机制(与
relax/utils/opd/opd_sglang_patch.py模式一致),通过 env flag 在 SGLang 子进程启动时应用,默认关闭,不影响其他场景。传播链
patch 设计要点
QwenVLImageProcessor.process_mm_data_async,新增_strip_image_token静态函数。其余 helper(smart_resize/preprocess_video/get_mm_data/__init__)一律从上游导入,不复制。_PATCH_FLAG = "_relax_deepeyes_patched"标记类,二次 apply 返回 False。apply_qwen_vl_patches先调_verify_upstream_api,缺符号即RuntimeError并提示「Relax requires sglang 0.5.12.post1」;apply_deepeyes_qwen_vl_patch再包 try/except,patch 失败只 warning,不阻断引擎启动。input_text是str时,strip 直通、original_input_ids保持 None,走原生路径。MultimodalProcessorOutput.from_dict包装;旧版回落 dict。input_ids重算,而原方法内部用 retokenize 后的序列。纯 wrapper 只改返回值会让input_ids与 mrope 错位——正是本 patch 要修的 bug。因此重写process_mm_data_async,但委托上游原语,只内联那两处真实差异。改动文件
核心改动:
relax/backends/sglang/patches/__init__.pyrelax/backends/sglang/patches/qwen_vl_patch.pyrelax/utils/arguments.py--deepeyes-qwen-vl-patch(store_true, 默认 False)relax/backends/sglang/sglang_engine.pyexamples/deepeyes/run_deepeyes_r3.shcp,加DEEPEYES_PATCH开关examples/deepeyes/qwen_vl.pytests/backends/sglang/test_qwen_vl_patch.pydocs/{en,zh}/guide/sglang-patches.mddocs/.vitepress/config.mts辅助测试脚本(本地验证用,非核心):
examples/deepeyes/run_deepeyes_baseline.shDEEPEYES_PATCH=1验证 patchexamples/deepeyes/openai_judge_service.shexamples/deepeyes/verify_patch_consistency.py验收标准对照
cp与/sgl-workspace硬编码qwen_vl.py;run_deepeyes_r3.sh删cp;全仓 grep 无真实硬编码(仅注释保留历史说明)process_mm_data_async+_strip_image_token;_PATCH_FLAG幂等;单测test_apply_is_idempotent_verify_upstream_api+RuntimeError(... Relax requires sglang 0.5.12.post1);单测覆盖 fail-fastverify_patch_consistency.py:Case A 文本输入 no-op,Case B 预分词输入只改 input_ids+mrope,其余不变DEEPEYES_PATCH=1跑 Qwen3-VL-4B GRPO,完成多轮 rollout→train step(step 0/1/2),无 token 漂移/mrope 错误docs/{en,zh}/guide/sglang-patches.md(含 SGLang 升级 checklist)测试
单测(stub sglang,无需 GPU):
pytest tests/backends/sglang/test_qwen_vl_patch.py -v # 15 passed真实 sglang API 校验(relaxrl 镜像,sglang 0.5.12.post1):
前后一致性:
python3 examples/deepeyes/verify_patch_consistency.py # ALL PASS端到端 smoke:
升级说明
SGLang 升级后的 checklist 见
docs/{en,zh}/guide/sglang-patches.md,关键依赖:QwenVLImageProcessor.process_mm_data_async/load_mm_data/process_and_combine_mm_datapreprocess_videoMRotaryEmbedding.get_rope_index(关键字参数签名)MultimodalProcessorOutput.from_dict(sglang ≥0.5.12;缺失则回落 dict)