Skip to content

[NVBUG-6327718][test] Unwaive test_disaggregated_videomme[nemotron_na… - #17031

Draft
aswinvisva wants to merge 1 commit into
NVIDIA:mainfrom
aswinvisva:avisva/unwaive-nvbug-6327718
Draft

[NVBUG-6327718][test] Unwaive test_disaggregated_videomme[nemotron_na…#17031
aswinvisva wants to merge 1 commit into
NVIDIA:mainfrom
aswinvisva:avisva/unwaive-nvbug-6327718

Conversation

@aswinvisva

@aswinvisva aswinvisva commented Jul 29, 2026

Copy link
Copy Markdown
Collaborator

…no_v3_omni_fp8] on B200 and H20

The root cause of the pre-fix failure — TypeError: isinstance() arg 2 must be a type, a tuple of types, or a union at tensorrt_llm/executor/proxy.py when the test infrastructure's session-reuse cache monkey-patches MpiPoolSession in the proxy module namespace with a factory function — is already fixed on main. The proxy code was refactored to identify pool-backed sessions by excluding (MpiCommSession, RemoteMpiCommSessionClient) instead of isinstance(x, MpiPoolSession).

Remove the two waivers linked to NVBug 6327718 so CI can verify the fix on both platforms.

Dev Engineer Review

  • Removed the two nvbug: 6327718 waivers for the targeted test on B200 and H20.
  • Updated the FP8 test to use the default worker count.
  • Added repetitions rep20 through rep39 and periodic thread-stack dumps.
  • Updated the related QA and test-db entries consistently.
  • Waiver formatting and test-list entries appear correctly scoped.

QA Engineer Review

  • Modified tests/integration/test_lists/qa/llm_function_core.txt, tests/integration/test_lists/test-db/l0_b200.yml, and tests/integration/test_lists/test-db/l0_h100.yml.
  • Updated test_disaggregated_videomme to accept the repetition parameter.
  • Added rep20 through rep39 coverage in the B200 and H100 test lists.
  • The modified test is covered in test-db/ for CI and qa/ for manual QA.
  • Verdict: needs follow-up pending CBTS coverage data.

Description

Test Coverage

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

@aswinvisva
aswinvisva requested review from a team as code owners July 29, 2026 22:47
@aswinvisva
aswinvisva marked this pull request as draft July 29, 2026 22:48
@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_B200-PyTorch-4"

@coderabbitai

coderabbitai Bot commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

The VideoMME E/PD test now runs repetitions 20 through 39 and records periodic thread stacks during execution. B200 and H100 test lists select the repetitions, QA entries use rep20, and related waivers are updated.

Changes

VideoMME repetition coverage

Layer / File(s) Summary
Parameterize repeated VideoMME execution
tests/integration/defs/accuracy/test_epd_disagg_multimodal.py
The test accepts repetitions rep20 through rep39, uses the default worker count for Nano Omni FP8, and cancels periodic thread-stack dump timers in a finally block.
Update test selections and waivers
tests/integration/test_lists/qa/llm_function_core.txt, tests/integration/test_lists/test-db/l0_b200.yml, tests/integration/test_lists/test-db/l0_h100.yml, tests/integration/test_lists/waives.txt
Test lists select the new repetitions. B200 and H20 VideoMME waivers are removed, and two other B200 waivers are added.

Estimated code review effort: 2 (Simple) | ~10 minutes

Suggested reviewers: qijune, tburt-nv, xinhe-nv

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly identifies the NVBug, test change, and affected VideoMME variant.
Description check ✅ Passed The description clearly explains the prior failure and waiver removal, but it does not provide explicit test coverage details or complete the checklist.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #62622 [ run ] triggered by Bot. Commit: 08015b1 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #62622 [ run ] completed with state SUCCESS. Commit: 08015b1
/LLM/main/L0_MergeRequest_PR pipeline #50762 (Partly Tested) completed with status: 'SUCCESS'

CI Report

Link to invocation

@BowenFu

BowenFu commented Aug 1, 2026

Copy link
Copy Markdown

The two lines removed are full:-scoped, i.e. QA/post-merge lists. The green /bot run --stage-list "DGX_B200-PyTorch-4" here executes l0_b200.yml, which lists only the qwen3vl_2b_instruct and nemotron_nano_v3_omni_nvfp4 variants — not the fp8 one being unwaived — so this PR's CI is zero coverage for the change.

Two other gaps against the description: nvbugs/6327718 is still P0 Dev - Open - To fix, and its recorded signature is Fatal Python error: Aborted in executor/ipc.py / inputs/multimodal_data.py, not the isinstance() TypeError; the repair-bot cannot-reproduce comment is for the nvfp4 variant. And the fix you cite (e15883702b, #16444) merged 2026-07-16, while these two waivers were added on 07-19/20 by #16593 / #16602 — after it.

Could you attach a passing run of the fp8 variant on B200 and H20 (host + commit + pytest command), the way repair-bot did for nvfp4? For what it's worth the fp8 variant already runs unwaived pre-merge via l0_h100.yml:166, so the removal is probably right — it just isn't demonstrated on the two platforms being unwaived.

@brnguyen2 brnguyen2 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks like a no-op. [tests/integration/test_lists/waives.txt:16](https://github.com/NVIDIA/TensorRT-LLM/pull/17031/files#diff-621bd2af82a3b97c7a5948368d36c14582ffdf73361deb7f42d2ff395b22167eR16) still waives the same node ID unconditionally (accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_fp8], nvbugs/6478692), so the test stays skipped on every platform and CI can't confirm the fix. Also, the fp8 param only appears in test-db/l0_h100.yml:164 and qa/llm_function_core.txt:86l0_b200.yml lists only the qwen3vl/nvfp4 variants — so the B200 entry was already vestigial. Worth confirming the 6327718 signature matches what you believe is fixed, and updating that bug alongside the removal.

@aswinvisva
aswinvisva force-pushed the avisva/unwaive-nvbug-6327718 branch from 08015b1 to eefdfc4 Compare August 10, 2026 16:25
@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_H100-PyTorch-1, DGX_B200-PyTorch-4" --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65073 [ run ] triggered by Bot. Commit: eefdfc4 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65073 [ run ] completed with state FAILURE. Commit: eefdfc4
/LLM/main/L0_MergeRequest_PR pipeline #52878 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_H100-PyTorch-1, DGX_B200-PyTorch-4" --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65111 [ run ] triggered by Bot. Commit: eefdfc4 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65111 [ run ] completed with state SUCCESS. Commit: eefdfc4
/LLM/main/L0_MergeRequest_PR pipeline #52909 (Partly Tested) completed with status: 'SUCCESS'

CI Report

Link to invocation

@aswinvisva
aswinvisva force-pushed the avisva/unwaive-nvbug-6327718 branch from eefdfc4 to 9264ea2 Compare August 11, 2026 16:18
@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_H100-PyTorch-1, DGX_B200-PyTorch-4" --disable-fail-fast

1 similar comment
@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_H100-PyTorch-1, DGX_B200-PyTorch-4" --disable-fail-fast

@aswinvisva
aswinvisva force-pushed the avisva/unwaive-nvbug-6327718 branch from 9264ea2 to 6dc8867 Compare August 11, 2026 16:47
@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_H100-PyTorch-1, DGX_B200-PyTorch-4" --disable-fail-fast

@aswinvisva
aswinvisva marked this pull request as ready for review August 11, 2026 16:53
@aswinvisva
aswinvisva requested review from a team as code owners August 11, 2026 16:53
@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_H100-PyTorch-1, DGX_B200-PyTorch-4" --disable-fail-fast

@coderabbitai

coderabbitai Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Note

GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer.

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65340 [ run ] triggered by Bot. Commit: 6dc8867 Link to invocation

@aswinvisva
aswinvisva marked this pull request as draft August 11, 2026 16:56

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/integration/test_lists/test-db/l0_b200.yml`:
- Around line 62-82: Add the omitted TestVideoMMEEPD.test_disaggregated_videomme
FP8 entries for rep20–rep39 to tests/integration/test_lists/test-db/l0_b200.yml
to provide B200 waiver coverage. The corresponding site in
tests/integration/test_lists/test-db/l0_h100.yml requires no direct change
because its *h100* targeting does not cover H20; alternatively, create and use
an H20-specific test-db list for H20 waiver coverage.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: a2add11e-57b7-4626-9e05-9ceb25051585

📥 Commits

Reviewing files that changed from the base of the PR and between 0c1b0a8 and 6dc8867.

📒 Files selected for processing (5)
  • tests/integration/defs/accuracy/test_epd_disagg_multimodal.py
  • tests/integration/test_lists/qa/llm_function_core.txt
  • tests/integration/test_lists/test-db/l0_b200.yml
  • tests/integration/test_lists/test-db/l0_h100.yml
  • tests/integration/test_lists/waives.txt
🚧 Files skipped from review as they are similar to previous changes (1)
  • tests/integration/test_lists/waives.txt

Comment on lines +62 to +82
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[qwen3vl_2b_instruct-rep20]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep20]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep21]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep22]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep23]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep24]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep25]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep26]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep27]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep28]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep29]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep30]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep31]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep32]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep33]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep34]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep35]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep36]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep37]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep38]
- accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_nvfp4-rep39]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -eu

printf '%s\n' '--- changed files ---'
git status --short
git diff --stat

printf '%s\n' '--- relevant diff: B200 ---'
git diff -- tests/integration/test_lists/test-db/l0_b200.yml

printf '%s\n' '--- relevant diff: H100 ---'
git diff -- tests/integration/test_lists/test-db/l0_h100.yml

printf '%s\n' '--- matching test files and list files ---'
git ls-files | rg 'test_epd_disagg_multimodal\.py|tests/integration/test_lists/(test-db|qa)/.*(b200|h100|h20|h20|video|multimodal)'
printf '%s\n' '--- GPU/model/variant references ---'
rg -n -i 'H20|H100|B200|fp8|nemotron_nano_v3_omni_fp8|qwen3vl_2b_instruct|rep20|rep39' \
  tests/integration/test_lists tests 2>/dev/null | head -n 500

Repository: NVIDIA/TensorRT-LLM

Length of output: 50375


🏁 Script executed:

#!/bin/bash
set -eu

printf '%s\n' '--- l0_b200 metadata and relevant entries ---'
sed -n '1,35p' tests/integration/test_lists/test-db/l0_b200.yml
sed -n '50,95p' tests/integration/test_lists/test-db/l0_b200.yml

printf '%s\n' '--- l0_h100 metadata and relevant entries ---'
sed -n '1,35p' tests/integration/test_lists/test-db/l0_h100.yml
sed -n '140,185p' tests/integration/test_lists/test-db/l0_h100.yml

printf '%s\n' '--- test source outline and relevant symbols ---'
ast-grep outline tests/integration/defs/accuracy/test_epd_disagg_multimodal.py
rg -n -C 8 'TestVideoMMEEPD|test_disaggregated_videomme|qwen3vl_2b_instruct|nemotron_nano_v3_omni_(fp8|nvfp4)' \
  tests/integration/defs/accuracy/test_epd_disagg_multimodal.py

printf '%s\n' '--- exact matching entries by test-db file ---'
python3 - <<'PY'
from pathlib import Path
needle = "accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme"
for p in sorted(Path("tests/integration/test_lists/test-db").glob("*.yml")):
    hits = [line.strip() for line in p.read_text().splitlines()
            if needle in line]
    if hits:
        print(f"{p}: {len(hits)} entries")
        for line in hits:
            print("  " + line)
PY

Repository: NVIDIA/TensorRT-LLM

Length of output: 22475


🏁 Script executed:

#!/bin/bash
set -eu

printf '%s\n' '--- marker definitions ---'
rg -n -C 5 'skip_pre_hopper|skip_pre_blackwell' tests/integration | head -n 160

printf '%s\n' '--- exact VideoMME waiver records ---'
rg -n -C 2 'test_epd_disagg_multimodal\.py::TestVideoMMEEPD::test_disaggregated_videomme' \
  tests/integration/test_lists/waives.txt || true

printf '%s\n' '--- test-db GPU selectors ---'
python3 - <<'PY'
from pathlib import Path
import re

for p in sorted(Path("tests/integration/test_lists/test-db").glob("*.yml")):
    text = p.read_text()
    if re.search(r'(?i)\bh20\b|\*h20\*|\*b200\*|\bb200\b', text):
        selectors = []
        for m in re.finditer(r'(?im)^\s*-\s*[\'"]?(\*?[a-z0-9_-]*h20[a-z0-9_*-]*|\*?[a-z0-9_-]*b200[a-z0-9_*-]*)[\'"]?\s*$', text):
            selectors.append(m.group(1))
        exact_fp8 = [
            line.strip() for line in text.splitlines()
            if "test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_fp8-" in line
        ]
        print(f"{p}: selectors={selectors or 'inspect'} exact_fp8_entries={len(exact_fp8)}")
PY

printf '%s\n' '--- all exact FP8 entries and their files ---'
rg -l 'accuracy/test_epd_disagg_multimodal\.py::TestVideoMMEEPD::test_disaggregated_videomme\[nemotron_nano_v3_omni_fp8-' \
  tests/integration/test_lists || true

Repository: NVIDIA/TensorRT-LLM

Length of output: 17381


🏁 Script executed:

#!/bin/bash
set -eu

printf '%s\n' '--- marker definitions and hardware predicates ---'
rg -n -C 8 'skip_pre_hopper|skip_pre_blackwell|def is_hopper|def is_blackwell|sm_90|sm_100|sm_120' \
  tests/integration --glob '*.py' | head -n 240

printf '%s\n' '--- QA coverage for the exact test ---'
python3 - <<'PY'
from pathlib import Path

needle = "accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme"
for p in sorted(Path("tests/integration/test_lists/qa").glob("*")):
    if not p.is_file():
        continue
    hits = [line.strip() for line in p.read_text(errors="replace").splitlines()
            if needle in line]
    if hits:
        print(f"{p}: {len(hits)} entries")
        for line in hits:
            print("  " + line)
PY

printf '%s\n' '--- QA file context for the exact test ---'
rg -n -C 4 'test_epd_disagg_multimodal\.py|VideoMMEEPD|nemotron_nano_v3_omni_fp8' \
  tests/integration/test_lists/qa/llm_function_core.txt || true

printf '%s\n' '--- candidate H20 test-db filenames and selectors ---'
find tests/integration/test_lists/test-db -maxdepth 1 -type f -printf '%f\n' | sort | rg -i 'h20|h100|b200'
rg -n -i -C 2 'gpu:|h20|h100|b200' tests/integration/test_lists/test-db/l0_dgx_h100.yml \
  tests/integration/test_lists/test-db/l0_h100.yml | head -n 180

Repository: NVIDIA/TensorRT-LLM

Length of output: 37798


Add FP8 coverage for the intended GPU tier.

TestVideoMMEEPD.test_disaggregated_videomme runs the FP8 variant for rep20rep39 on Hopper and newer GPUs. l0_b200.yml omits these entries, while l0_h100.yml targets only *h100* and cannot cover H20. Add the entries to the B200 list for a B200 waiver, or create/use an H20 test-db list for an H20 waiver.

Coverage summary: no test functions changed; coverage is insufficient for B200 FP8 or H20.

📍 Affects 2 files
  • tests/integration/test_lists/test-db/l0_b200.yml#L62-L82 (this comment)
  • tests/integration/test_lists/test-db/l0_h100.yml#L153-L172
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/integration/test_lists/test-db/l0_b200.yml` around lines 62 - 82, Add
the omitted TestVideoMMEEPD.test_disaggregated_videomme FP8 entries for
rep20–rep39 to tests/integration/test_lists/test-db/l0_b200.yml to provide B200
waiver coverage. The corresponding site in
tests/integration/test_lists/test-db/l0_h100.yml requires no direct change
because its *h100* targeting does not cover H20; alternatively, create and use
an H20-specific test-db list for H20 waiver coverage.

Source: Path instructions

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65343 [ run ] triggered by Bot. Commit: 6dc8867 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65340 [ run ] completed with state ABORTED. Commit: 6dc8867

Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65343 [ run ] completed with state SUCCESS. Commit: 6dc8867
/LLM/main/L0_MergeRequest_PR pipeline #53110 (Partly Tested) completed with status: 'SUCCESS'

CI Report

Link to invocation

@aswinvisva
aswinvisva force-pushed the avisva/unwaive-nvbug-6327718 branch from 6dc8867 to caf2375 Compare August 11, 2026 19:58
@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_H100-PyTorch-1, DGX_B200-PyTorch-4" --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65380 [ run ] triggered by Bot. Commit: caf2375 Link to invocation

Throwaway diagnostic PR — not for merge. Iteration 4.

* Expand to 100 reps (range(60,160)) so each CI run exercises 100
  unique test IDs. With 9 SLURM groups, ~11 reps per group run in
  parallel — all 100 complete in same wallclock time as the 7 we
  were getting before.
* This gives ~63% probability of catching the 1% SIGABRT flake in
  a single run vs ~7% previously.
* faulthandler.dump_traceback_later(60) still active.
* max_workers=128 (workaround removed).

Previous results: 21/21 videomme reps PASSED across 3 CI runs
(rep20-59). Cumulative 0 failures observed. Continuing to accumulate.

Signed-off-by: Aswin Visva <31215515+aswinvisva@users.noreply.github.com>
@aswinvisva
aswinvisva force-pushed the avisva/unwaive-nvbug-6327718 branch from caf2375 to 5f8be05 Compare August 11, 2026 21:26
@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_H100-PyTorch-1, DGX_B200-PyTorch-4" --disable-fail-fast

1 similar comment
@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_H100-PyTorch-1, DGX_B200-PyTorch-4" --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65403 [ run ] triggered by Bot. Commit: 5f8be05 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65404 [ run ] triggered by Bot. Commit: 5f8be05 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65403 [ run ] completed with state ABORTED. Commit: 5f8be05

Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65380 [ run ] completed with state ABORTED. Commit: caf2375

Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65404 [ run ] completed with state SUCCESS. Commit: 5f8be05
/LLM/main/L0_MergeRequest_PR pipeline #53162 (Partly Tested) completed with status: 'SUCCESS'

CI Report

Link to invocation

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants