Skip to content

perf(gemm): update mm_fp4 cute-dsl tactic heuristic and autotune over a top-N ranked list#3948

Merged
bkryu merged 5 commits into
flashinfer-ai:mainfrom
bkryu:mm_fp4_top_n_autotune
Jul 15, 2026
Merged

perf(gemm): update mm_fp4 cute-dsl tactic heuristic and autotune over a top-N ranked list#3948
bkryu merged 5 commits into
flashinfer-ai:mainfrom
bkryu:mm_fp4_top_n_autotune

Conversation

@bkryu

@bkryu bkryu commented Jul 13, 2026

Copy link
Copy Markdown
Collaborator

📌 Description

Autotuning mm_fp4(backend='cute-dsl') currently profiles all ~120 valid tactics per shape. Each kernel compilation requires ~2.5 sec, which substantially deteriorates experience when all tactics are compiled.

Current PR:

  1. Updates the tactic heuristic to provide slightly improved top-1 (non-autotuned) performance, and a suitable top-32 (autotuned) performance.
  2. Makes autotuned mm_fp4 profile top-N ranked candidates instead of the entire list. N is currently set to 32.

A PR for parallelizing and disk-caching ranked tactics for further reduction will follow after #3874

Performance Data measured on B200

Median kernel times are in µs. Top-1 means heuristic pick without autotune. Measured via flashinfer_benchmark.py

Speedup >1.0 means after is faster than before.

Tuning/training testlist for deriving heuristic (147 shapes) — geomean speedup: top-1 1.01x, autotuned 1.00x
M N K Top-1 before (us) Top-1 after (us) Top-1 speedup Autotuned before (us) Autotuned after (us) Autotuned speedup
1 512 7168 6.54 6.54 1.00 6.82 6.85 1.00
4 512 7168 6.61 6.61 1.00 6.85 6.91 0.99
16 512 7168 6.62 6.62 1.00 7.06 7.07 1.00
64 512 7168 7.87 7.78 1.01 8.19 8.22 1.00
256 512 7168 7.62 7.59 1.00 8.08 8.26 0.98
1024 512 7168 7.61 7.82 0.97 8.85 8.61 1.03
1 896 1024 3.58 3.58 1.00 3.90 3.98 0.98
4 896 1024 3.49 3.49 1.00 3.89 4.00 0.97
16 896 1024 3.65 3.65 1.00 3.95 3.90 1.01
64 896 1024 4.13 4.14 1.00 4.27 4.19 1.02
256 896 1024 3.85 4.30 0.90 4.26 4.29 0.99
1024 896 1024 4.13 4.68 0.88 4.45 4.42 1.01
1 896 5120 5.66 5.66 1.00 6.21 6.24 0.99
8 896 5120 5.38 5.38 1.00 6.14 6.19 0.99
64 896 5120 6.66 6.37 1.05 7.12 7.01 1.02
512 896 5120 6.38 8.12 0.79 7.17 7.25 0.99
1 1024 7168 6.30 6.30 1.00 7.42 7.62 0.97
4 1024 7168 6.28 6.28 1.00 7.22 7.22 1.00
16 1024 7168 6.32 6.32 1.00 7.36 7.71 0.95
256 1024 7168 7.60 7.58 1.00 8.45 8.45 1.00
1024 1024 7168 7.94 8.09 0.98 9.15 8.99 1.02
1 1280 8192 7.02 7.02 1.00 7.78 7.90 0.98
8 1280 8192 7.33 7.33 1.00 7.81 7.84 1.00
64 1280 8192 8.41 8.24 1.02 9.06 9.06 1.00
512 1280 8192 8.48 8.50 1.00 9.55 9.58 1.00
1 1792 5120 5.47 5.47 1.00 6.37 6.69 0.95
8 1792 5120 5.64 5.64 1.00 6.45 6.59 0.98
64 1792 5120 6.57 6.57 1.00 7.50 7.44 1.01
512 1792 5120 6.68 6.92 0.97 8.03 7.66 1.05
1 2560 8192 7.08 7.08 1.00 8.42 8.66 0.97
8 2560 8192 7.20 7.20 1.00 8.45 8.64 0.98
64 2560 8192 8.52 8.26 1.03 9.57 9.71 0.99
512 2560 8192 10.07 9.48 1.06 10.98 10.98 1.00
1 3584 5120 5.60 5.60 1.00 7.36 7.31 1.01
8 3584 5120 5.62 5.62 1.00 7.36 7.20 1.02
64 3584 5120 6.72 6.76 0.99 8.48 8.26 1.03
512 3584 5120 8.01 7.83 1.02 9.28 9.10 1.02
1 4608 7168 6.67 6.67 1.00 9.10 8.72 1.04
4 4608 7168 6.53 6.53 1.00 8.75 8.80 0.99
16 4608 7168 6.73 6.73 1.00 9.09 9.01 1.01
64 4608 7168 7.98 7.80 1.02 10.18 10.18 1.00
256 4608 7168 8.20 8.32 0.99 10.48 10.66 0.98
1024 4608 7168 15.42 14.66 1.05 16.13 16.06 1.00
1 5120 640 3.54 3.54 1.00 4.18 4.19 1.00
8 5120 640 3.52 3.52 1.00 4.35 4.18 1.04
64 5120 640 4.10 4.29 0.96 4.54 4.58 0.99
512 5120 640 5.60 5.58 1.00 5.63 5.82 0.97
1 5120 1024 3.65 3.65 1.00 4.54 4.50 1.01
8 5120 1024 3.76 3.76 1.00 4.51 4.54 0.99
64 5120 1024 4.12 4.32 0.95 4.99 4.80 1.04
512 5120 1024 5.74 5.93 0.97 5.92 6.18 0.96
1 5120 1280 3.68 3.68 1.00 4.86 4.75 1.02
8 5120 1280 3.71 3.71 1.00 4.74 4.82 0.98
64 5120 1280 4.57 4.48 1.02 5.41 5.12 1.06
512 5120 1280 5.70 6.38 0.89 7.31 6.53 1.12
1 5120 2048 4.18 4.18 1.00 5.62 5.86 0.96
8 5120 2048 4.17 4.17 1.00 5.57 5.54 1.01
64 5120 2048 4.85 5.08 0.95 5.98 6.14 0.97
512 5120 2048 6.52 6.94 0.94 7.97 7.68 1.04
1 5120 2560 4.55 4.55 1.00 6.02 6.02 1.00
8 5120 2560 4.32 4.32 1.00 6.02 6.02 1.00
64 5120 2560 5.61 5.17 1.08 6.66 6.56 1.01
512 5120 2560 7.19 7.56 0.95 8.54 8.51 1.00
1 5120 4096 6.20 6.20 1.00 6.78 7.07 0.96
8 5120 4096 5.12 5.12 1.00 6.78 7.10 0.96
64 5120 4096 6.17 6.16 1.00 7.55 7.68 0.98
512 5120 4096 8.60 8.86 0.97 10.54 10.24 1.03
1 5120 5120 5.67 5.67 1.00 7.71 7.58 1.02
8 5120 5120 5.63 5.63 1.00 7.68 7.62 1.01
64 5120 5120 8.04 6.58 1.22 8.22 8.32 0.99
512 5120 5120 9.62 9.83 0.98 11.42 11.31 1.01
1 5120 8192 7.29 7.29 1.00 9.65 9.66 1.00
8 5120 8192 7.20 7.20 1.00 9.60 9.60 1.00
64 5120 8192 8.63 8.46 1.02 10.13 10.34 0.98
512 5120 8192 12.31 12.61 0.98 14.26 14.27 1.00
1 5120 16384 11.10 11.10 1.00 15.95 15.94 1.00
8 5120 16384 11.20 11.20 1.00 15.94 15.92 1.00
64 5120 16384 13.58 12.94 1.05 15.63 15.92 0.98
512 5120 16384 20.81 19.69 1.06 22.16 22.34 0.99
1 7168 256 3.07 3.07 1.00 3.74 3.74 1.00
4 7168 256 3.13 3.13 1.00 3.68 3.71 0.99
16 7168 256 3.11 3.11 1.00 3.71 3.65 1.02
64 7168 256 3.79 3.88 0.98 4.26 4.00 1.06
256 7168 256 3.89 4.32 0.90 4.42 4.45 0.99
1024 7168 256 6.09 6.49 0.94 6.27 6.30 0.99
1 7168 512 3.21 3.21 1.00 4.19 4.21 1.00
4 7168 512 3.44 3.44 1.00 4.00 4.22 0.95
16 7168 512 3.38 3.38 1.00 4.29 4.16 1.03
64 7168 512 4.01 4.02 1.00 4.32 4.32 1.00
256 7168 512 4.20 4.51 0.93 4.98 4.82 1.03
1024 7168 512 6.72 6.92 0.97 7.06 7.04 1.00
4 7168 2304 4.28 4.28 1.00 6.18 5.89 1.05
16 7168 2304 4.32 4.32 1.00 6.18 5.95 1.04
64 7168 2304 5.36 5.35 1.00 6.62 6.34 1.05
256 7168 2304 5.85 5.86 1.00 7.55 7.36 1.03
1 7168 4608 5.32 5.32 1.00 7.89 7.86 1.00
4 7168 4608 5.44 5.44 1.00 8.06 7.84 1.03
16 7168 4608 5.33 5.33 1.00 8.03 8.00 1.00
64 7168 4608 6.55 6.47 1.01 8.96 8.80 1.02
256 7168 4608 7.81 7.79 1.00 9.73 9.73 1.00
1024 7168 4608 18.66 17.44 1.07 19.01 19.14 0.99
1 7168 5120 5.63 5.63 1.00 8.19 8.22 1.00
8 7168 5120 5.77 5.77 1.00 8.34 8.14 1.02
64 7168 5120 6.92 6.98 0.99 9.25 9.23 1.00
512 7168 5120 11.75 11.66 1.01 13.47 13.50 1.00
1 8192 1024 6.62 6.62 1.00 4.93 4.82 1.02
8 8192 1024 3.64 3.64 1.00 4.83 4.90 0.99
64 8192 1024 4.44 4.49 0.99 5.23 5.62 0.93
512 8192 1024 5.83 6.01 0.97 6.69 7.01 0.95
1 8192 2048 4.28 4.28 1.00 6.16 6.11 1.01
8 8192 2048 4.07 4.07 1.00 6.16 6.11 1.01
64 8192 2048 7.36 5.16 1.43 6.48 6.48 1.00
512 8192 2048 7.80 8.24 0.95 9.52 9.47 1.01
1 8192 3584 4.91 4.91 1.00 7.42 7.39 1.00
8 8192 3584 5.13 5.13 1.00 7.33 7.34 1.00
64 8192 3584 6.32 5.94 1.06 8.18 8.13 1.01
512 8192 3584 10.21 10.20 1.00 11.52 11.62 0.99
1 8192 4096 5.33 5.33 1.00 7.79 7.74 1.01
8 8192 4096 5.32 5.32 1.00 8.06 7.71 1.05
64 8192 4096 6.60 6.28 1.05 8.58 8.54 1.00
512 8192 4096 10.89 10.70 1.02 12.72 12.43 1.02
1 8192 7168 6.78 6.78 1.00 10.67 10.88 0.98
8 8192 7168 6.94 6.94 1.00 10.82 10.78 1.00
64 8192 7168 8.39 8.02 1.05 11.57 11.68 0.99
512 8192 7168 14.41 14.58 0.99 16.30 16.34 1.00
1 8192 8192 7.30 7.30 1.00 11.62 11.84 0.98
8 8192 8192 7.19 7.19 1.00 11.81 11.68 1.01
64 8192 8192 9.01 8.42 1.07 12.59 12.56 1.00
512 8192 8192 15.77 15.70 1.00 17.78 17.73 1.00
1 8192 14336 10.44 10.44 1.00 18.03 18.37 0.98
8 8192 14336 10.27 10.27 1.00 18.05 18.05 1.00
64 8192 14336 12.81 11.91 1.08 19.23 19.39 0.99
512 8192 14336 33.96 25.90 1.31 25.73 25.90 0.99
1 8192 28672 24.75 24.75 1.00 33.09 32.80 1.01
8 8192 28672 25.10 25.10 1.00 32.24 33.04 0.98
64 8192 28672 29.48 28.28 1.04 33.57 33.39 1.01
512 8192 28672 65.80 54.38 1.21 42.75 43.23 0.99
1 9216 7168 7.17 7.17 1.00 11.49 11.62 0.99
4 9216 7168 6.90 6.90 1.00 11.84 11.60 1.02
16 9216 7168 6.90 6.90 1.00 11.68 11.60 1.01
64 9216 7168 12.33 8.06 1.53 12.40 12.29 1.01
256 9216 7168 9.92 9.47 1.05 13.38 13.31 1.00
1024 9216 7168 29.22 25.99 1.12 28.02 27.97 1.00
1 10240 8192 7.61 7.61 1.00 13.58 13.26 1.02
8 10240 8192 7.18 7.18 1.00 13.50 13.30 1.02
64 10240 8192 10.28 8.77 1.17 14.14 14.05 1.01
512 10240 8192 23.40 20.41 1.15 22.77 22.98 0.99
Holdout testlist for based on realistic transformer shapes (70 shapes) — geomean speedup: top-1 1.05x, autotuned 1.00x
M N K Top-1 before (us) Top-1 after (us) Top-1 speedup Autotuned before (us) Autotuned after (us) Autotuned speedup
2 4096 4096 5.24 5.24 1.00 6.78 6.85 0.99
8 4096 4096 5.10 5.10 1.00 6.85 6.83 1.00
32 4096 4096 5.31 5.31 1.00 7.26 7.31 0.99
48 4096 4096 6.25 6.02 1.04 7.73 7.73 1.00
384 4096 4096 7.60 7.17 1.06 8.94 8.78 1.02
2048 4096 4096 17.34 16.61 1.04 18.03 18.14 0.99
8192 4096 4096 61.67 55.94 1.10 47.71 47.97 0.99
2 4096 7168 15.03 15.03 1.00 8.90 8.96 0.99
8 4096 7168 14.83 14.83 1.00 8.91 8.83 1.01
32 4096 7168 15.31 15.31 1.00 9.23 9.55 0.97
48 4096 4096 6.25 6.02 1.04 7.73 7.73 1.00
384 4096 4096 7.60 7.17 1.06 8.94 8.78 1.02
2048 4096 4096 17.34 16.61 1.04 18.03 18.14 0.99
8192 4096 4096 61.67 55.94 1.10 47.71 47.97 0.99
2 4096 7168 15.03 15.03 1.00 8.90 8.96 0.99
8 4096 7168 14.83 14.83 1.00 8.91 8.83 1.01
32 4096 7168 15.31 15.31 1.00 9.23 9.55 0.97
48 4096 7168 18.44 17.69 1.04 10.00 9.89 1.01
384 4096 7168 21.14 20.31 1.04 11.42 11.38 1.00
2048 4096 7168 58.20 56.75 1.03 25.54 25.94 0.98
8192 4096 7168 174.73 97.13 1.80 74.85 74.93 1.00
2 4096 14336 10.24 10.24 1.00 13.15 13.10 1.00
8 4096 14336 10.31 10.31 1.00 13.14 13.18 1.00
32 4096 14336 10.44 10.44 1.00 14.43 14.40 1.00
48 4096 14336 12.52 11.90 1.05 13.52 13.36 1.01
384 4096 14336 15.02 13.93 1.08 18.02 18.34 0.98
2048 4096 14336 54.14 46.69 1.16 44.93 45.14 1.00
8192 4096 14336 220.24 202.81 1.09 151.43 143.87 1.05
2 6144 4096 5.15 5.15 1.00 7.26 7.23 1.00
8 6144 4096 5.08 5.08 1.00 7.33 7.10 1.03
32 6144 4096 5.34 5.34 1.00 7.62 7.60 1.00
48 6144 4096 6.39 6.11 1.05 8.03 8.05 1.00
384 6144 4096 8.41 9.42 0.89 9.22 9.22 1.00
2048 6144 4096 23.54 23.46 1.00 25.14 24.62 1.02
8192 6144 4096 94.99 84.78 1.12 67.74 67.97 1.00
2 7168 2048 5.29 5.29 1.00 5.84 5.82 1.00
8 7168 2048 5.20 5.20 1.00 5.90 5.92 1.00
32 7168 2048 5.56 5.56 1.00 5.95 5.97 1.00
48 7168 2048 6.76 9.50 0.71 6.30 6.24 1.01
384 7168 2048 9.95 10.27 0.97 7.82 7.81 1.00
2048 7168 2048 24.88 23.07 1.08 18.26 18.27 1.00
8192 7168 2048 82.64 63.81 1.30 47.26 47.30 1.00
2 8192 29568 36.18 36.18 1.00 33.44 33.47 1.00
8 8192 29568 36.76 36.76 1.00 33.22 33.31 1.00
32 8192 29568 41.31 41.31 1.00 36.96 36.24 1.02
48 8192 29568 41.93 41.67 1.01 37.98 37.98 1.00
384 8192 29568 99.73 76.98 1.30 47.54 47.63 1.00
2048 8192 29568 306.94 227.93 1.35 180.37 168.75 1.07
8192 8192 29568 989.21 1206.40 0.82 840.55 817.59 1.03
2 14336 4096 5.75 5.75 1.00 10.24 10.21 1.00
8 14336 4096 5.50 5.50 1.00 10.26 10.19 1.01
32 14336 4096 5.59 5.59 1.00 10.26 10.18 1.01
48 14336 4096 7.62 6.60 1.15 10.53 10.43 1.01
384 14336 4096 17.11 17.73 0.96 16.13 16.10 1.00
2048 14336 4096 57.60 53.07 1.09 46.37 46.69 0.99
8192 14336 4096 215.07 195.82 1.10 163.07 163.49 1.00
2 16384 16384 28.46 28.46 1.00 35.09 35.34 0.99
8 16384 16384 28.65 28.65 1.00 35.63 35.22 1.01
32 16384 16384 27.83 27.83 1.00 36.08 35.65 1.01
48 16384 16384 31.52 28.63 1.10 36.16 36.58 0.99
384 16384 16384 72.79 60.15 1.21 43.41 43.36 1.00
2048 16384 16384 240.52 212.62 1.13 179.57 181.02 0.99
8192 16384 16384 1282.50 1315.63 0.97 899.52 884.71 1.02
2 29568 8192 24.08 24.08 1.00 31.54 31.50 1.00
8 29568 8192 24.42 24.42 1.00 31.15 31.26 1.00
32 29568 8192 25.10 25.10 1.00 31.82 31.71 1.00
48 29568 8192 28.85 25.95 1.11 32.24 32.27 1.00
384 29568 8192 57.70 58.29 0.99 42.91 44.24 0.97
2048 29568 8192 221.47 200.48 1.10 172.13 170.82 1.01
8192 29568 8192 895.00 996.03 0.90 721.30 710.07 1.02
2 36864 4608 18.72 18.72 1.00 23.18 23.15 1.00
8 36864 4608 18.10 18.10 1.00 23.52 23.47 1.00
32 36864 4608 19.32 19.32 1.00 23.07 23.20 0.99
48 36864 4608 44.75 21.42 2.09 23.46 23.71 0.99
384 36864 4608 60.63 67.51 0.90 38.34 37.74 1.02
2048 36864 4608 272.54 218.58 1.25 119.06 138.53 0.86
8192 36864 4608 651.23 574.51 1.13 518.82 540.08 0.96

🔍 Related Issues

🚀 Pull Request Checklist

Thank you for contributing to FlashInfer! Before we review your pull request, please make sure the following items are complete.

✅ Pre-commit Checks

  • I have installed pre-commit by running pip install pre-commit (or used your preferred method).
  • I have installed the hooks with pre-commit install.
  • I have run the hooks manually with pre-commit run --all-files and fixed any reported issues.

If you are unsure about how to set up pre-commit, see the pre-commit documentation.

🧪 Tests

  • Tests have been added or updated as needed.
  • All tests are passing (unittest, etc.).

Reviewer Notes

Summary by CodeRabbit

  • Performance Improvements
    • Enhanced FP4 GEMM tactic selection with a more consistent, workload-aware ranking model to improve efficiency and throughput.
    • Reduced the number of tuning configurations evaluated to help autotuning complete faster.
    • Improved decision-making for memory/compute tradeoffs (including vector-size–dependent behavior).
  • Maintenance
    • Centralized shared tuning and scoring logic for more consistent behavior across GEMM runs.

@bkryu bkryu self-assigned this Jul 13, 2026
@coderabbitai

coderabbitai Bot commented Jul 13, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

📝 Walkthrough

Walkthrough

CuTe DSL SM100 FP4 GEMM tactic selection now uses centralized scoring, shared candidate definitions, and scale-factor-aware caching and feasibility checks. The runner groups tactics by configuration, ranks them, and limits tuning to 32 configurations while preserving prefetch variants.

Changes

SM100 FP4 tactic selection

Layer / File(s) Summary
Centralized tactic scoring and enumeration
flashinfer/gemm/kernels/utils.py
Adds cluster candidates, scale-factor-aware caching and feasibility checks, exhaustively evaluates tile/cluster/swap combinations, and scores them using efficiency, wave quantization, throughput, and tie-breaking penalties.
Bounded CuTe DSL tuning selection
flashinfer/gemm/gemm_base.py
Uses shared SM100 utilities, groups tactics excluding use_prefetch, ranks groups, and returns up to 32 tuning configurations with both prefetch variants.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant CuteDSLFp4GemmRunner
  participant _select_sm100_mm_fp4_cute_dsl_tactic
  participant _compute_tactic_for_m
  participant _score_sm100_mm_fp4_tactic
  CuteDSLFp4GemmRunner->>_select_sm100_mm_fp4_cute_dsl_tactic: request tactic with sf_vec_size
  _select_sm100_mm_fp4_cute_dsl_tactic->>_compute_tactic_for_m: compute tactic for M bucket
  _compute_tactic_for_m->>_compute_tactic_for_m: check candidate feasibility with sf_dtype
  _compute_tactic_for_m->>_score_sm100_mm_fp4_tactic: score feasible tile, cluster, and swap combinations
  _score_sm100_mm_fp4_tactic-->>_compute_tactic_for_m: return candidate score
  _compute_tactic_for_m-->>_select_sm100_mm_fp4_cute_dsl_tactic: return best tactic
  _select_sm100_mm_fp4_cute_dsl_tactic-->>CuteDSLFp4GemmRunner: return selected tactic
Loading

Possibly related PRs

Suggested reviewers: b8zhong, aleozlx, yzh119, dhiraj113, jdebache

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title is concise and accurately summarizes the heuristic update and top-N autotuning change.
Description check ✅ Passed The description follows the template and provides a clear summary, benchmarks, checklist items, and reviewer notes.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@bkryu bkryu added the run-ci label Jul 13, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the SM100 FP4 GEMM tactic selection by introducing a scoring function, _score_sm100_mm_fp4_tactic, to rank and limit the number of autotuning configurations. The feedback suggests adding a guard for zero-dimension inputs to prevent division-by-zero errors, simplifying a redundant tuple lookup in the scoring calculation, and cleaning up an imprecise type annotation.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread flashinfer/gemm/kernels/utils.py
Comment thread flashinfer/gemm/kernels/utils.py Outdated
Comment thread flashinfer/gemm/gemm_base.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@flashinfer/gemm/gemm_base.py`:
- Around line 54-58: Import _SM100_MMA_TILER_MN_CANDIDATES from kernels.utils
alongside _SM100_CLUSTER_SHAPE_MN_CANDIDATES, then remove the duplicate local
definition near the SM100 tactic logic. Keep all existing references using the
shared imported candidate list.

In `@flashinfer/gemm/kernels/utils.py`:
- Around line 55-77: Add a feasibility filter in _compute_tactic_for_m before
calling _score_sm100_mm_fp4_tactic, validating each tile/cluster/swap_ab
combination against the supported Sm100BlockScaledPersistentDenseGemmKernel
contract, including tile_n constraints. Skip unsupported candidates before
ranking while preserving the existing 256/M cluster guard and tactic selection
behavior for feasible candidates.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 74b97ecf-f0ca-4c5a-a9b3-9873d5477508

📥 Commits

Reviewing files that changed from the base of the PR and between e179800 and 95fb4df.

📒 Files selected for processing (2)
  • flashinfer/gemm/gemm_base.py
  • flashinfer/gemm/kernels/utils.py

Comment thread flashinfer/gemm/gemm_base.py
Comment thread flashinfer/gemm/kernels/utils.py Outdated
@bkryu

bkryu commented Jul 13, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run tests/gemm

@flashinfer-bot

Copy link
Copy Markdown
Collaborator

GitLab MR !950 has been created, and the CI pipeline #57846035 is currently running. I'll report back once the pipeline job completes.

@qiching qiching left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM.

one thing i want to highlight: top-N is ranked by the same heuristic as top-1, so anything below the cut never gets profiled. Mostly fine, but i noticed that 2048×36864×4608 is 0.86x autotuned, ~16% slower with tuning on. Since after's set is a subset of the full ~120, tuning should be ≤ before modulo noise, so the best config for that shape is likely falling outside the top-16. Can you check whether it's pruning or just variance? If pruning: only 16 distinct configs are explored, so ranking prefetch or bumping N could recover it.

Worth nailing before we pitch re enabling autotuning to vLLM.

@flashinfer-bot

Copy link
Copy Markdown
Collaborator

[FAILED] Pipeline #57846035: 13/20 passed

@Vinnie6167

Copy link
Copy Markdown
Contributor

My only concern is the Top-1 heuristic runtime on a cache miss would be 10x longer. Previously it was ~55-86 usec (deleted comment on utils.py:144-145) when we only swept the tile shapes. Now sweeping the cross product of tile shapes, cluster shapes, and swap_ab options will be a ~20x increase in runtime (9 cluster shapes and 2 swap_ab options). I believe this is fine assuming in real inference workloads we are serving 1000+ passes where the number of cache misses is upper bound by the number of M buckets (<14), and thus amortized out.

So if 1-2 ms heuristic runtime is okay on the first call/subsequent call with floor(log2(M)) unique, then LGTM.

@bkryu

bkryu commented Jul 15, 2026

Copy link
Copy Markdown
Collaborator Author

LGTM.

one thing i want to highlight: top-N is ranked by the same heuristic as top-1, so anything below the cut never gets profiled. Mostly fine, but i noticed that 2048×36864×4608 is 0.86x autotuned, ~16% slower with tuning on. Since after's set is a subset of the full ~120, tuning should be ≤ before modulo noise, so the best config for that shape is likely falling outside the top-16. Can you check whether it's pruning or just variance? If pruning: only 16 distinct configs are explored, so ranking prefetch or bumping N could recover it.

Worth nailing before we pitch re enabling autotuning to vLLM.

Thanks @qiching, upon remeasuring, the regression is not as severe.

$ python3 flashinfer_benchmark.py --routine mm_fp4 --m 2048 --n 36864 --k 4608 --out_dtype bfloat16 --backends cute-dsl  --use_128x4_sf_layout --use_nvfp4 --refcheck --autotune
2026-07-14 18:57:40,985 - INFO - autotuner.py:757 - flashinfer.jit: [Autotuner]: Autotuning process starts ...
[AutoTuner]: Tuning fp4_gemm: 100%|███████████████████████████████████████| 16/16 [05:35<00:00, 20.98s/profile]
2026-07-14 19:03:16,899 - INFO - autotuner.py:780 - flashinfer.jit: [Autotuner]: Autotuning process ends
[PERF] cute-dsl_autotune:: median time 0.124 ms; std 0.000 ms; achieved tflops 5626.367 TFLOPs/sec; achieved tb_per_sec 1.946 TB/sec

so 0.86x is likely from an unlucky measurement.

In general however, pruning an autotune list means we are losing something; if we do not, it means the configs were not helpful to begin with. We could consider consulting vLLM about it, but we could claim that a max perf loss of <20% with geomean less than 0.5% would be a good tradeoff for substantially reduced autotuning time.

@bkryu
bkryu merged commit 517cca9 into flashinfer-ai:main Jul 15, 2026
31 of 39 checks passed
@bkryu
bkryu deleted the mm_fp4_top_n_autotune branch July 15, 2026 16:36
@qiching

qiching commented Jul 15, 2026

Copy link
Copy Markdown
Collaborator

LGTM.
one thing i want to highlight: top-N is ranked by the same heuristic as top-1, so anything below the cut never gets profiled. Mostly fine, but i noticed that 2048×36864×4608 is 0.86x autotuned, ~16% slower with tuning on. Since after's set is a subset of the full ~120, tuning should be ≤ before modulo noise, so the best config for that shape is likely falling outside the top-16. Can you check whether it's pruning or just variance? If pruning: only 16 distinct configs are explored, so ranking prefetch or bumping N could recover it.
Worth nailing before we pitch re enabling autotuning to vLLM.

Thanks @qiching, upon remeasuring, the regression is not as severe.

$ python3 flashinfer_benchmark.py --routine mm_fp4 --m 2048 --n 36864 --k 4608 --out_dtype bfloat16 --backends cute-dsl  --use_128x4_sf_layout --use_nvfp4 --refcheck --autotune
2026-07-14 18:57:40,985 - INFO - autotuner.py:757 - flashinfer.jit: [Autotuner]: Autotuning process starts ...
[AutoTuner]: Tuning fp4_gemm: 100%|███████████████████████████████████████| 16/16 [05:35<00:00, 20.98s/profile]
2026-07-14 19:03:16,899 - INFO - autotuner.py:780 - flashinfer.jit: [Autotuner]: Autotuning process ends
[PERF] cute-dsl_autotune:: median time 0.124 ms; std 0.000 ms; achieved tflops 5626.367 TFLOPs/sec; achieved tb_per_sec 1.946 TB/sec

so 0.86x is likely from an unlucky measurement.

In general however, pruning an autotune list means we are losing something; if we do not, it means the configs were not helpful to begin with. We could consider consulting vLLM about it, but we could claim that a max perf loss of <20% with geomean less than 0.5% would be a good tradeoff for substantially reduced autotuning time.

make sense to me. thanks for confirming. @bkryu

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants