Summary: on an RTX 3070 (Ampere, sm_86), mesh generation works perfectly but
the texture stage always produces garbage. I ran a fairly exhaustive elimination
matrix and the result is stable across every axis except the weight format. This
looks like the same problem reported in #81 and #147.
Environment
Two independent installs, both built from scratch:
|
Install A |
Install B |
| OS |
Windows 10 Pro 22H2 |
same |
| GPU |
RTX 3070 8 GB, driver 581.42 |
same |
| Python |
3.13.14 (ComfyUI portable) |
3.12.13 (standalone) |
| torch |
2.10.0+cu130 |
2.8.0+cu128, then 2.7.0+cu128 |
| CUDA wheels |
Torch2100/CUDA 13.1 cp313 |
Torch280 then Torch270, cp312 |
| flash-attn |
not available |
2.7.4.post1+cu128torch2.7.0 installed and active |
| Nodes |
ComfyUI-Trellis2, ComfyUI-Trellis2-GGUF, ComfyUI-GGUF |
same |
Install B was built specifically to match the README's tested configuration
(Python 3.11 + torch 2.7.0+cu128 + flash-attn); 3.12 was the closest Python with
published wheels.
Symptom
Mesh: correct, watertight, good detail. Texture: per-texel noise.
GGUF Q4_K_M → colored noise, mean RGB [124.6 113.1 93.1], std [30 28.3 31.6]
GGUF Q8_0 → different garbage, contour-like patterns with saturated values
FP8 safetensors (visualbruno/TRELLIS.2-4B-FP8) → near-black, mean RGB [0.5 2.0 7.1]
All runs report status: success. No OOM, no exception.
Elimination matrix
| Variable |
Values tried |
Changed the output? |
| Weight format |
Q4_K_M, Q8_0, FP8 safetensors |
yes |
| Node implementation |
ComfyUI-Trellis2, ComfyUI-Trellis2-GGUF |
no |
| Python |
3.13.14, 3.12.13 |
no |
| torch |
2.10.0+cu130, 2.8.0+cu128, 2.7.0+cu128 |
no |
| Attention |
SDPA fallback vs real flash-attn |
no |
FLEX_GEMM_ALGO |
masked_implicit_gemm_splitk, explicit_gemm |
no |
| Tiled encoder / decoder |
on, off |
no |
| Model |
Pixal3D, TRELLIS.2-4B, and cross-combinations |
no |
resolution |
512, 1024 |
no |
texture_size |
2048, 4096 |
no |
The most informative result
Switching FLEX_GEMM_ALGO from masked_implicit_gemm_splitk to explicit_gemm
changed peak VRAM from 3297 MiB to 7791 MiB — so the execution path really
did change — and produced a byte-identical texture.
If a sparse-conv kernel were miscomputing, a different algorithm should produce
different garbage. It produced the same. That argues the corruption is upstream
of the conv kernels.
The rasterizer is not the cause
Two checks:
bake_on_vertices=True, which skips the UV bake entirely, gives the same
corrupted colors ([0, 2, 7] per vertex on FP8) — so the color data is already
wrong before baking.
Trellis2RenderMultiViewNvdiffrast on a textured mesh renders a clean
silhouette with correct shading. nvdiffrast itself works.
What I did not try
- Building the CUDA wheels from source on this machine.
- An RTX 40xx/50xx GPU for comparison.
Possible pattern
#147 reports a 3090 (also Ampere) with self-compiled wheels on Linux, torch 2.11 /
CUDA 13: "on 3090 better geometry, but textures with artifacts", and the inverse
on a 4060. #81 describes "a fine mosaic / pixelated mess / multicolored squares",
which matches what I see exactly.
That is three reports on Ampere, across Windows and Linux, across torch versions,
with self-compiled and prebuilt wheels. If the maintainer's GPU is a 40xx or newer,
that would explain the difficulty reproducing.
Happy to run any diagnostic build or dump intermediate tensors — I can reproduce
this in about 60 seconds, every time.
Summary: on an RTX 3070 (Ampere, sm_86), mesh generation works perfectly but
the texture stage always produces garbage. I ran a fairly exhaustive elimination
matrix and the result is stable across every axis except the weight format. This
looks like the same problem reported in #81 and #147.
Environment
Two independent installs, both built from scratch:
Torch2100/CUDA 13.1cp313Torch280thenTorch270, cp312Install B was built specifically to match the README's tested configuration
(Python 3.11 + torch 2.7.0+cu128 + flash-attn); 3.12 was the closest Python with
published wheels.
Symptom
Mesh: correct, watertight, good detail. Texture: per-texel noise.
GGUF Q4_K_M→ colored noise, mean RGB[124.6 113.1 93.1], std[30 28.3 31.6]GGUF Q8_0→ different garbage, contour-like patterns with saturated valuesFP8 safetensors(visualbruno/TRELLIS.2-4B-FP8) → near-black, mean RGB[0.5 2.0 7.1]All runs report
status: success. No OOM, no exception.Elimination matrix
ComfyUI-Trellis2,ComfyUI-Trellis2-GGUFFLEX_GEMM_ALGOmasked_implicit_gemm_splitk,explicit_gemmresolutiontexture_sizeThe most informative result
Switching
FLEX_GEMM_ALGOfrommasked_implicit_gemm_splitktoexplicit_gemmchanged peak VRAM from 3297 MiB to 7791 MiB — so the execution path really
did change — and produced a byte-identical texture.
If a sparse-conv kernel were miscomputing, a different algorithm should produce
different garbage. It produced the same. That argues the corruption is upstream
of the conv kernels.
The rasterizer is not the cause
Two checks:
bake_on_vertices=True, which skips the UV bake entirely, gives the samecorrupted colors (
[0, 2, 7]per vertex on FP8) — so the color data is alreadywrong before baking.
Trellis2RenderMultiViewNvdiffraston a textured mesh renders a cleansilhouette with correct shading. nvdiffrast itself works.
What I did not try
Possible pattern
#147 reports a 3090 (also Ampere) with self-compiled wheels on Linux, torch 2.11 /
CUDA 13: "on 3090 better geometry, but textures with artifacts", and the inverse
on a 4060. #81 describes "a fine mosaic / pixelated mess / multicolored squares",
which matches what I see exactly.
That is three reports on Ampere, across Windows and Linux, across torch versions,
with self-compiled and prebuilt wheels. If the maintainer's GPU is a 40xx or newer,
that would explain the difficulty reproducing.
Happy to run any diagnostic build or dump intermediate tensors — I can reproduce
this in about 60 seconds, every time.