Skip to content

Fix webui.py validation: return after the warning instead of falling through to inference - #1948

Open
Linxiushen wants to merge 1 commit into
QwenAudio:mainfrom
Linxiushen:fix-webui-validation-return
Open

Linxiushen wants to merge 1 commit into
QwenAudio:mainfrom
Linxiushen:fix-webui-validation-return

Conversation

@Linxiushen

@Linxiushen Linxiushen commented Sep 21, 2026 •

Copy link
Copy Markdown

Problem

generate_audio() in webui.py validates the required inputs for each mode, shows a gr.Warning(...) and yields one frame of silence when something is missing, but never returns. The generator keeps going after the warning and calls the model with the invalid inputs anyway:

  • 预训练音色 with no available voice (sft_dropdown == '') continues into inference_sft('') and crashes with KeyError: ''.
  • 3s极速复刻 / 跨语种复刻 without a prompt wav continues into torchaudio.info(prompt_wav) with prompt_wav=None (webui.py:78) and crashes with TypeError.
  • 3s极速复刻 with an empty prompt text runs zero-shot inference with an empty prompt instead of stopping.
  • 自然语言控制 with an empty instruct text runs instruct inference with an empty instruction (on CosyVoice2/3 that hits the assert self.__class__.__name__ == 'CosyVoice').

So the user sees the warning and then a traceback, or a result that is not what the warning said would happen.

Fix

Add a bare return after each of the six warning yields, so the warning is the outcome. The return is deliberately bare: in a generator, return (sample_rate, data) would not deliver anything to the UI, whereas the existing yield of the silent frame followed by return keeps the current placeholder behaviour and simply stops.

Verification

Local differential with the model and Gradio calls stubbed, running the pristine and patched generate_audio side by side over all 4 modes x sft_dropdown x prompt_text x prompt_wav x instruct_text (64 input combinations), comparing the sequence of gr.* calls, inference calls, yielded frames and exceptions:

  • 28 well-formed combinations: identical before and after (0 changes).
  • 36 malformed combinations: all changed, from "crash or run inference with invalid arguments" to "warning plus one silent frame, then stop". None still reaches inference.

Three of those are worth stating plainly because the old code did not crash there but produced a result the warning text says should not happen: 3s极速复刻 with empty prompt text (previously ran zero-shot with an empty prompt), 自然语言控制 with empty instruct text on CosyVoice v1 (previously ran with an empty instruction), and a prompt wav with a sample rate below 16 kHz (previously warned and synthesized anyway; the upload component's label says the rate must not be below 16 kHz, and this branch did return before 90433f5). All three now stop, matching the warning.

No test file is included: the repo has no test infrastructure and CI runs only flake8, which passes on the changed file.


This fix was developed with AI assistance (Claude); the change was reviewed and verified locally before submission.

…through to inference

generate_audio() checks the required inputs for each mode, shows a
gr.Warning and yields one frame of silence when something is missing, but
never returns, so the generator keeps going and calls the model anyway:

- '预训练音色' with no available voice continues into
  inference_sft('') and crashes with KeyError: ''
- '3s极速复刻' / '跨语种复刻' without a prompt wav continues into
  load_wav(None) and crashes with TypeError
- '3s极速复刻' with an empty prompt text runs zero-shot inference with an
  empty prompt instead of stopping
- '自然语言控制' with an empty instruct text runs instruct inference with
  an empty instruction (and asserts on CosyVoice2/3)

Add a bare `return` after each of the six warning yields so the warning
is the outcome. Well-formed inputs are unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant