Skip to content

Texture stage outputs noise on RTX 3070 (Ampere) — identical across weight formats, torch versions and sparse-conv algorithms; geometry is fine #195

Description

@tiagofoil

Summary: on an RTX 3070 (Ampere, sm_86), mesh generation works perfectly but
the texture stage always produces garbage. I ran a fairly exhaustive elimination
matrix and the result is stable across every axis except the weight format. This
looks like the same problem reported in #81 and #147.

Environment

Two independent installs, both built from scratch:

Install A Install B
OS Windows 10 Pro 22H2 same
GPU RTX 3070 8 GB, driver 581.42 same
Python 3.13.14 (ComfyUI portable) 3.12.13 (standalone)
torch 2.10.0+cu130 2.8.0+cu128, then 2.7.0+cu128
CUDA wheels Torch2100/CUDA 13.1 cp313 Torch280 then Torch270, cp312
flash-attn not available 2.7.4.post1+cu128torch2.7.0 installed and active
Nodes ComfyUI-Trellis2, ComfyUI-Trellis2-GGUF, ComfyUI-GGUF same

Install B was built specifically to match the README's tested configuration
(Python 3.11 + torch 2.7.0+cu128 + flash-attn); 3.12 was the closest Python with
published wheels.

Symptom

Mesh: correct, watertight, good detail. Texture: per-texel noise.

  • GGUF Q4_K_M → colored noise, mean RGB [124.6 113.1 93.1], std [30 28.3 31.6]
  • GGUF Q8_0 → different garbage, contour-like patterns with saturated values
  • FP8 safetensors (visualbruno/TRELLIS.2-4B-FP8) → near-black, mean RGB [0.5 2.0 7.1]

All runs report status: success. No OOM, no exception.

Elimination matrix

Variable Values tried Changed the output?
Weight format Q4_K_M, Q8_0, FP8 safetensors yes
Node implementation ComfyUI-Trellis2, ComfyUI-Trellis2-GGUF no
Python 3.13.14, 3.12.13 no
torch 2.10.0+cu130, 2.8.0+cu128, 2.7.0+cu128 no
Attention SDPA fallback vs real flash-attn no
FLEX_GEMM_ALGO masked_implicit_gemm_splitk, explicit_gemm no
Tiled encoder / decoder on, off no
Model Pixal3D, TRELLIS.2-4B, and cross-combinations no
resolution 512, 1024 no
texture_size 2048, 4096 no

The most informative result

Switching FLEX_GEMM_ALGO from masked_implicit_gemm_splitk to explicit_gemm
changed peak VRAM from 3297 MiB to 7791 MiB — so the execution path really
did change — and produced a byte-identical texture.

If a sparse-conv kernel were miscomputing, a different algorithm should produce
different garbage. It produced the same. That argues the corruption is upstream
of the conv kernels.

The rasterizer is not the cause

Two checks:

  1. bake_on_vertices=True, which skips the UV bake entirely, gives the same
    corrupted colors ([0, 2, 7] per vertex on FP8) — so the color data is already
    wrong before baking.
  2. Trellis2RenderMultiViewNvdiffrast on a textured mesh renders a clean
    silhouette with correct shading. nvdiffrast itself works.

What I did not try

  • Building the CUDA wheels from source on this machine.
  • An RTX 40xx/50xx GPU for comparison.

Possible pattern

#147 reports a 3090 (also Ampere) with self-compiled wheels on Linux, torch 2.11 /
CUDA 13: "on 3090 better geometry, but textures with artifacts", and the inverse
on a 4060. #81 describes "a fine mosaic / pixelated mess / multicolored squares",
which matches what I see exactly.

That is three reports on Ampere, across Windows and Linux, across torch versions,
with self-compiled and prebuilt wheels. If the maintainer's GPU is a 40xx or newer,
that would explain the difficulty reproducing.

Happy to run any diagnostic build or dump intermediate tensors — I can reproduce
this in about 60 seconds, every time.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions