Skip to content

Qwen3.8-27B not showing up for multiple reasons #171

Description

@matusfaro

Description

Hello,

I was wondering why Qwen3.8-27B wasn't showing up for me. Claude was able to find some reasons why it gets deprioritized in the current code, which I believe should not happen on a default run, it hides models I'd like to see:

  1. Qwen/Qwen3.8-27B is tagged pipeline_tag image-text-to-text on HF (natively multimodal) and doesn't get fetched by default unless you specify include_vision=True
  # cli.py:408
  def _include_vision_candidates(profile: str) -> bool:
      return profile.lower() in {"vision", "any"}

Minor versions are collapsed into one family. qwen3.8-27b, qwen3.6-27b, qwen3.5-27b all normalize to qwen3-27b. Inside engine/ranker.py:837 it picks qwen3.6 over qwen3.8 for several reasons:

  1. Benchmark data is frozen pre-release:
  Qwen/Qwen3.6-27B   quality=88.95  status=direct     selkey=93.95
  Qwen/Qwen3.8-27B   quality=47.80  status=estimated  selkey=47.80
  1. Lineage table penalizes it since it's a new model. qwen3.8-27b falls through to the bare qwen3 pattern:
 # data/lineage.py
  (r"qwen3\.6", 7), (r"qwen3\.5", 6), ..., (r"qwen3", 5),

Steps to Reproduce

uvx whichllm@latest

Hardware Info

➜  ~ uvx whichllm@latest

╭────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── Hardware Info ───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ GPU 0: Apple M2 Max — 96.0 GB shared (budget 94.0 GB) — BW: 400 GB/s                                                                                                                                                                                                                       │
│ CPU: Apple M2 Max — 12 cores                                                                                                                                                                                                                                                               │
│ RAM: 96.0 GB                                                                                                                                                                                                                                                                               │
│ Disk free: 33.6 GB                                                                                                                                                                                                                                                                         │
│ OS: darwin                                                                                                                                                                                                                                                                                 │
│ VRAM headroom: 2.0 GB reserved per GPU                                                                                                                                                                                                                                                     │
╰────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

                                            Recommended Models                                            
┏━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━┓
┃     ┃                                          ┃        ┃  Fit /   ┃              ┃            ┃       ┃
┃   # ┃ Model                                    ┃ Quant  ┃   VRAM   ┃        Speed ┃ Published  ┃ Score ┃
┡━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━┩
│   1 │ Qwen/Qwen3.6-27B                         │ Q5_K_M │ Full GPU │  8.9 tok/s ~ │ 2026-04-21 │  88.9 │
│     │ 27.8B                                    │        │ 21.2 GB  │              │            │       │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│   2 │ google/gemma-4-31B-it                    │ Q4_K_M │ Full GPU │ 10.3 tok/s ~ │ 2026-03-11 │  85.1 │
│     │ 31.3B                                    │        │ 20.1 GB  │              │            │       │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│   3 │ google/gemma-4-26B-A4B-it                │  Q8_0  │ Full GPU │ 36.6 tok/s ? │ 2026-03-11 │  85.0 │
│     │ 25.8B (3.8Ba)                            │        │ 27.0 GB  │              │            │       │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│   4 │ Qwen/Qwen3-30B-A3B                       │  Q8_0  │ Full GPU │ 46.3 tok/s ? │ 2025-04-27 │  84.3 │
│     │ 30.5B (3.0Ba)                            │        │ 31.6 GB  │              │            │       │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│   5 │ Qwen/Qwen3-Next-80B-A3B-Instruct         │ Q3_K_M │ Full GPU │ 59.0 tok/s ? │ 2025-09-09 │  79.0 │
│     │ 81.3B (3.0Ba)                            │        │ 34.5 GB  │              │            │       │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│   6 │ openai/gpt-oss-20b                       │  Q8_0  │ Full GPU │ 38.6 tok/s ? │ 2025-08-04 │  78.7 │
│     │ 20.9B (3.6Ba)                            │        │ 22.0 GB  │              │            │       │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│   7 │ zai-org/GLM-4.7-Flash                    │  Q6_K  │ Full GPU │ 16.8 tok/s ? │ 2026-01-19 │  76.3 │
│     │ 31.2B (12.0Ba)                           │        │ 26.2 GB  │              │            │       │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│   8 │ Qwen/QwQ-32B                             │ Q4_K_M │ Full GPU │  9.8 tok/s ~ │ 2025-03-05 │  74.8 │
│     │ 32.8B                                    │        │ 21.0 GB  │              │            │       │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│   9 │ deepseek-ai/DeepSeek-R1-Distill-Qwen-32B │ Q4_K_M │ Full GPU │  9.8 tok/s ~ │ 2025-01-20 │  73.8 │
│     │ 32.8B                                    │        │ 21.0 GB  │              │            │       │
├─────┼──────────────────────────────────────────┼────────┼──────────┼──────────────┼────────────┼───────┤
│  10 │ Qwen/Qwen3-14B                           │  Q6_K  │ Full GPU │ 13.7 tok/s ~ │ 2025-04-27 │  73.7 │
│     │ 14.8B                                    │        │ 13.5 GB  │              │            │       │
└─────┴──────────────────────────────────────────┴────────┴──────────┴──────────────┴────────────┴───────┘
  Speed:  ~ = estimated tok/s range,  ? = low-confidence/backend-sensitive tok/s
  Top pick confidence: High (direct benchmark, gap +3.9)
  Benchmark reference: 2026-05 curated snapshot; live AA / LiveBench / Aider merged when reachable.
  Speed caution: Low-confidence speed estimates in top ranks: #3

Python Version

No response

Operating System

macos 27

whichllm Version

0.5.16

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions