Skip to content

Task.17】DeepEyes processor 动态 patch - #188

Closed
yuki-younai wants to merge 3 commits into
redai-studio:mainfrom
yuki-younai:task17-deepeyes
Closed

yuki-younai wants to merge 3 commits into
redai-studio:mainfrom
yuki-younai:task17-deepeyes

Conversation

@yuki-younai

Copy link
Copy Markdown

【Task.17】DeepEyes processor 动态 patch

背景

examples/deepeyes/run_deepeyes_r3.sh 通过 cp 把整份 examples/deepeyes/qwen_vl.py(454 行的旧版 SGLang qwen_vl.py)覆盖到 SGLang 安装目录:

cp examples/deepeyes/qwen_vl.py /sgl-workspace/sglang/python/sglang/srt/multimodal/processors/qwen_vl.py

问题:

  • 硬编码 /sgl-workspace,SGLang 升级后路径/版本不匹配即失效。
  • 整文件覆盖会回退 Docker 构建时 docker/patch/latest/sglang.patch 的改动(该 patch 在 v0.5.12 把 load_mm_data 切到 legacy_load_mm_data),引入隐蔽不一致。
  • 维护成本高,真实差异淹没在 400+ 行样板代码里,易随上游升级失效。

目标

启动时用 monkey-patch 注入真实差异,删除 cp 覆盖与 /sgl-workspace 硬编码。

真实差异分析

DeepEyes 多轮 rollout 向 SGLang /generate 发送预分词的 input_ids: list[int](见 examples/deepeyes/rollout.py::_run_inference_step),每张图已被前一轮 processor 展开成 N 个 <|image_pad|>。SGLang 原生 QwenVLImageProcessor.process_mm_data_async 处理不了这种输入:

  • 问题 A(token 漂移):load_mm_data 把 int 列表 decode 回文本再重新分词,非无损,token 数可能变化,错位 mrope 位置,破坏 loss_mask / log-prob。
  • 问题 B(N×M 爆炸):累积的 N 个 <|image_pad|> 被当成 N 个图片占位符,每个再展开成 M 个 patch token。

真正的定制只有两处,其余 ~100 行都是 SGLang 原生样板。

方案

沿用 Relax 已有的「SGLang 启动时 monkey-patch」机制(与 relax/utils/opd/opd_sglang_patch.py 模式一致),通过 env flag 在 SGLang 子进程启动时应用,默认关闭,不影响其他场景。

传播链

--deepeyes-qwen-vl-patch  (CLI, 默认 False)
   →  sglang_engine._init_normal 设 RELAX_DEEPEYES_QWEN_VL_PATCH=1
   →  spawn 的 SGLang 子进程继承
   →  _launch_server_with_patches 读取 env flag → apply_deepeyes_qwen_vl_patch()

patch 设计要点

  • 仅覆盖必要方法:只替换 QwenVLImageProcessor.process_mm_data_async,新增 _strip_image_token 静态函数。其余 helper(smart_resize/preprocess_video/get_mm_data/__init__)一律从上游导入,不复制。
  • 幂等:_PATCH_FLAG = "_relax_deepeyes_patched" 标记类,二次 apply 返回 False。
  • fail fast:apply_qwen_vl_patches 先调 _verify_upstream_api,缺符号即 RuntimeError 并提示「Relax requires sglang 0.5.12.post1」;apply_deepeyes_qwen_vl_patch 再包 try/except,patch 失败只 warning,不阻断引擎启动。
  • 文本输入无副作用:input_text 是 str 时,strip 直通、original_input_ids 保持 None,走原生路径。
  • 返回值版本兼容:sglang ≥0.5.12 用 MultimodalProcessorOutput.from_dict 包装;旧版回落 dict。
  • 为什么重写方法体而非纯 wrapper:mrope 必须用恢复后的原始 input_ids 重算,而原方法内部用 retokenize 后的序列。纯 wrapper 只改返回值会让 input_ids 与 mrope 错位——正是本 patch 要修的 bug。因此重写 process_mm_data_async,但委托上游原语,只内联那两处真实差异。

改动文件

核心改动:

文件 动作 说明
relax/backends/sglang/patches/__init__.py 新增 patches 包
relax/backends/sglang/patches/qwen_vl_patch.py 新增 核心 patch 模块
relax/utils/arguments.py 改 --deepeyes-qwen-vl-patch (store_true, 默认 False)
relax/backends/sglang/sglang_engine.py 改 3 处接线:apply / 日志 / env 翻译
examples/deepeyes/run_deepeyes_r3.sh 改 删 cp,加 DEEPEYES_PATCH 开关
examples/deepeyes/qwen_vl.py 删 旧整文件覆盖物
tests/backends/sglang/test_qwen_vl_patch.py 新增 15 个单测
docs/{en,zh}/guide/sglang-patches.md 新增 升级说明(双语)
docs/.vitepress/config.mts 改 sidebar 注册新文档页

辅助测试脚本(本地验证用,非核心):

文件 说明
examples/deepeyes/run_deepeyes_baseline.sh 本地 4-GPU Qwen3-VL-4B baseline,通过 DEEPEYES_PATCH=1 验证 patch
examples/deepeyes/openai_judge_service.sh 远程 OpenAI 兼容 judge 服务脚本
examples/deepeyes/verify_patch_consistency.py patch 前后关键输出一致性验证脚本

验收标准对照

验收标准 状态 证据
删除所有覆盖安装目录的 cp 与 /sgl-workspace 硬编码 ✅ 删 qwen_vl.py;run_deepeyes_r3.sh 删 cp;全仓 grep 无真实硬编码(仅注释保留历史说明)
patch 仅覆盖必要方法、重复 import/调用幂等 ✅ 只替换 process_mm_data_async + _strip_image_token;_PATCH_FLAG 幂等;单测 test_apply_is_idempotent
不兼容 SGLang 版本时 fail fast 并提示 ✅ _verify_upstream_api + RuntimeError(... Relax requires sglang 0.5.12.post1);单测覆盖 fail-fast
同一输入下 patch 前后关键输出一致 ✅ verify_patch_consistency.py:Case A 文本输入 no-op,Case B 预分词输入只改 input_ids+mrope,其余不变
DeepEyes 端到端 smoke 通过 ✅ DEEPEYES_PATCH=1 跑 Qwen3-VL-4B GRPO,完成多轮 rollout→train step(step 0/1/2),无 token 漂移/mrope 错误
有单测和升级说明 ✅ 15 单测 + 双语 docs/{en,zh}/guide/sglang-patches.md(含 SGLang 升级 checklist)

测试

单测(stub sglang,无需 GPU):

pytest tests/backends/sglang/test_qwen_vl_patch.py -v   # 15 passed

真实 sglang API 校验(relaxrl 镜像,sglang 0.5.12.post1):

from relax.backends.sglang.patches.qwen_vl_patch import _verify_upstream_api, apply_qwen_vl_patches
_verify_upstream_api()        # PASS
apply_qwen_vl_patches()       # True

前后一致性:

python3 examples/deepeyes/verify_patch_consistency.py   # ALL PASS

端到端 smoke:

DEEPEYES_PATCH=1 bash examples/deepeyes/run_deepeyes_baseline.sh
# 日志: deepeyes_qwen_vl_patch=True, deepeyes_qwen_vl=True
# 完成 step 0/1/2, 无 Traceback / ERROR

升级说明

SGLang 升级后的 checklist 见 docs/{en,zh}/guide/sglang-patches.md,关键依赖:

  • QwenVLImageProcessor.process_mm_data_async / load_mm_data / process_and_combine_mm_data
  • 模块级 preprocess_video
  • MRotaryEmbedding.get_rope_index(关键字参数签名)
  • MultimodalProcessorOutput.from_dict(sglang ≥0.5.12;缺失则回落 dict)

yuki-younai and others added 3 commits July 30, 2026 19:21
…service

- add examples/deepeyes/run_deepeyes_baseline.sh: local 4-GPU Qwen3-VL-4B
  GRPO baseline derived from run_deepeyes_r3.sh (no cp overlay, no patch),
  using remote OpenAI-compatible LLM judge.
- add examples/deepeyes/openai_judge_service.sh: sources DEEPEYES_JUDGE_*
  env vars + injects into Ray runtime_env, replaces the local sglang judge.

For local baseline testing only; not part of the task-17 patch.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Replace the `cp examples/deepeyes/qwen_vl.py /sgl-workspace/...` whole-file
overlay with a runtime monkey-patch applied at SGLang engine start-up,
injecting only the real DeepEyes-specific differences.

Why: DeepEyes multi-turn rollout sends pre-tokenized `input_ids` (list[int])
to SGLang where each image is already expanded into N `<|image_pad|>` tokens.
The stock `QwenVLImageProcessor.process_mm_data_async` mishandles this:
(A) decode->retokenize drifts the token count, misaligning mrope; (B) the N
image-pad tokens get re-expanded into N×M. The old fix copied a 454-line
legacy `qwen_vl.py` over the install dir, hardcoding `/sgl-workspace` and
breaking on SGLang upgrades (and silently reverting the docker sglang.patch).

Changes:
- add `relax/backends/sglang/patches/qwen_vl_patch.py`: patch
  `QwenVLImageProcessor.process_mm_data_async` with two surgical changes —
  `_strip_image_token` collapses consecutive `<|image_pad|>` before
  `load_mm_data`, and the original `input_ids` are restored + mrope
  recomputed from them. Delegates to upstream primitives; idempotent via
  `_PATCH_FLAG`; `_verify_upstream_api` fails fast on incompatible sglang.
- add `--deepeyes-qwen-vl-patch` CLI flag, translated to
  `RELAX_DEEPEYES_QWEN_VL_PATCH=1` in `sglang_engine._init_normal` and
  applied in `_launch_server_with_patches` (same pattern as the OPD patch).
- remove `examples/deepeyes/qwen_vl.py` and the `cp`/`/sgl-workspace`
  hardcode in `run_deepeyes_r3.sh`; gate the patch via `DEEPEYES_PATCH` env
  in both r3 and baseline scripts. Default off = stock SGLang behavior.
- add unit tests (stubbed sglang, no GPU) covering strip logic, idempotency,
  fail-fast, and the text/list input paths.
- add bilingual `docs/{en,zh}/guide/sglang-patches.md` upgrade guide +
  register in VitePress sidebar.

Verified: 15 unit tests pass; `_verify_upstream_api` + `apply_qwen_vl_patches`
succeed against real sglang 0.5.12.post1; end-to-end DeepEyes GRPO smoke
(--deepeyes-qwen-vl-patch) completes multiple rollout->train steps with no
token-drift/mrope errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Standalone verification for task-17 acceptance redai-studio#4 (same input, key outputs
consistent before/after the patch). Compares a faithful re-implementation of
the stock upstream `process_mm_data_async` against the patched one on a
stubbed `self` (no GPU / real model needed):

- Case A (raw-text prompt): patch is a no-op — stock==patch on every key.
- Case B (pre-tokenized list[int], N image_pad per image): patch restores
  the original input_ids and recomputes mrope from them; stock drifts to the
  retokenized sequence.
- Case C: `_strip_image_token` collapse unit check.

Run inside the relaxrl image:
    python3 examples/deepeyes/verify_patch_consistency.py

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@yxyOo

yxyOo commented Aug 3, 2026

Copy link
Copy Markdown
Member

验收通过
评价:问题分析和验证材料充分,但异常被外层捕获,启动流程并未真正 fail-fast

@SigureMo

Copy link
Copy Markdown
Member

#164 merged,本 PR close

@SigureMo SigureMo closed this Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants