Description
Hello,
I was wondering why Qwen3.8-27B wasn't showing up for me. Claude was able to find some reasons why it gets deprioritized in the current code, which I believe should not happen on a default run, it hides models I'd like to see:
Qwen/Qwen3.8-27B is tagged pipeline_tag image-text-to-text on HF (natively multimodal) and doesn't get fetched by default unless you specify include_vision=True
# cli.py:408
def _include_vision_candidates(profile: str) -> bool:
return profile.lower() in {"vision", "any"}
Minor versions are collapsed into one family. qwen3.8-27b, qwen3.6-27b, qwen3.5-27b all normalize to qwen3-27b. Inside engine/ranker.py:837 it picks qwen3.6 over qwen3.8 for several reasons:
- Benchmark data is frozen pre-release:
Qwen/Qwen3.6-27B quality=88.95 status=direct selkey=93.95
Qwen/Qwen3.8-27B quality=47.80 status=estimated selkey=47.80
- Lineage table penalizes it since it's a new model.
qwen3.8-27b falls through to the bare qwen3 pattern:
# data/lineage.py
(r"qwen3\.6", 7), (r"qwen3\.5", 6), ..., (r"qwen3", 5),
Steps to Reproduce
uvx whichllm@latest
Hardware Info
➜ ~ uvx whichllm@latest
╭────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── Hardware Info ───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ GPU 0: Apple M2 Max — 96.0 GB shared (budget 94.0 GB) — BW: 400 GB/s │
│ CPU: Apple M2 Max — 12 cores │
│ RAM: 96.0 GB │
│ Disk free: 33.6 GB │
│ OS: darwin │
│ VRAM headroom: 2.0 GB reserved per GPU │
╰────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
Recommended Models
┏━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━┓
┃ ┃ ┃ ┃ Fit / ┃ ┃ ┃ ┃
┃ # ┃ Model ┃ Quant ┃ VRAM ┃ Speed ┃ Published ┃ Score ┃
┡━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━┩
│ 1 │ Qwen/Qwen3.6-27B │ Q5_K_M │ Full GPU │ 8.9 tok/s ~ │ 2026-04-21 │ 88.9 │
│ │ 27.8B │ │ 21.2 GB │ │ │ │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│ 2 │ google/gemma-4-31B-it │ Q4_K_M │ Full GPU │ 10.3 tok/s ~ │ 2026-03-11 │ 85.1 │
│ │ 31.3B │ │ 20.1 GB │ │ │ │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│ 3 │ google/gemma-4-26B-A4B-it │ Q8_0 │ Full GPU │ 36.6 tok/s ? │ 2026-03-11 │ 85.0 │
│ │ 25.8B (3.8Ba) │ │ 27.0 GB │ │ │ │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│ 4 │ Qwen/Qwen3-30B-A3B │ Q8_0 │ Full GPU │ 46.3 tok/s ? │ 2025-04-27 │ 84.3 │
│ │ 30.5B (3.0Ba) │ │ 31.6 GB │ │ │ │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│ 5 │ Qwen/Qwen3-Next-80B-A3B-Instruct │ Q3_K_M │ Full GPU │ 59.0 tok/s ? │ 2025-09-09 │ 79.0 │
│ │ 81.3B (3.0Ba) │ │ 34.5 GB │ │ │ │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│ 6 │ openai/gpt-oss-20b │ Q8_0 │ Full GPU │ 38.6 tok/s ? │ 2025-08-04 │ 78.7 │
│ │ 20.9B (3.6Ba) │ │ 22.0 GB │ │ │ │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│ 7 │ zai-org/GLM-4.7-Flash │ Q6_K │ Full GPU │ 16.8 tok/s ? │ 2026-01-19 │ 76.3 │
│ │ 31.2B (12.0Ba) │ │ 26.2 GB │ │ │ │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│ 8 │ Qwen/QwQ-32B │ Q4_K_M │ Full GPU │ 9.8 tok/s ~ │ 2025-03-05 │ 74.8 │
│ │ 32.8B │ │ 21.0 GB │ │ │ │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│ 9 │ deepseek-ai/DeepSeek-R1-Distill-Qwen-32B │ Q4_K_M │ Full GPU │ 9.8 tok/s ~ │ 2025-01-20 │ 73.8 │
│ │ 32.8B │ │ 21.0 GB │ │ │ │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│ 10 │ Qwen/Qwen3-14B │ Q6_K │ Full GPU │ 13.7 tok/s ~ │ 2025-04-27 │ 73.7 │
│ │ 14.8B │ │ 13.5 GB │ │ │ │
└─────┴──────────────────────────────────────────┴────────┴──────────┴──────────────┴────────────┴───────┘
Speed: ~ = estimated tok/s range, ? = low-confidence/backend-sensitive tok/s
Top pick confidence: High (direct benchmark, gap +3.9)
Benchmark reference: 2026-05 curated snapshot; live AA / LiveBench / Aider merged when reachable.
Speed caution: Low-confidence speed estimates in top ranks: #3
Python Version
No response
Operating System
macos 27
whichllm Version
0.5.16
Description
Hello,
I was wondering why
Qwen3.8-27Bwasn't showing up for me. Claude was able to find some reasons why it gets deprioritized in the current code, which I believe should not happen on a default run, it hides models I'd like to see:Qwen/Qwen3.8-27Bis tagged pipeline_tag image-text-to-text on HF (natively multimodal) and doesn't get fetched by default unless you specifyinclude_vision=TrueMinor versions are collapsed into one family.
qwen3.8-27b,qwen3.6-27b,qwen3.5-27ball normalize toqwen3-27b. Insideengine/ranker.py:837it picks qwen3.6 over qwen3.8 for several reasons:qwen3.8-27bfalls through to the bare qwen3 pattern:Steps to Reproduce
uvx whichllm@latestHardware Info
Python Version
No response
Operating System
macos 27
whichllm Version
0.5.16