Skip to content

[Bug] GGUF fork: sparse_structure_resolution > 32 always OOMs — missing ratio correction in sample_shape_slat_cascade (with fix) #193

Description

@carroyoaesa

Filed here because the fork has issues disabled. This is not a bug in visualbruno/ComfyUI-Trellis2 — the code here is correct. The bug is in the GGUF fork Aero-Ex/ComfyUI-Trellis2-GGUF (has_issues: false), which dropped one line when porting sample_shape_slat_cascade. Filing upstream so it is findable by anyone hitting the same OOM, and in case you want to relay it.

Symptom

On the GGUF fork with the Pixal3D-GGUF model, sparse_structure_resolution = 64 OOMs on a 12 GB card in all five quantizations (Q4_K_M, Q5_K_M, Q6_K, Q8_0, BF16), even with every other lever at minimum (pipeline_type=1024_cascade, steps 8/8/8, max_num_tokens=16384, low_vram=True, keep_models_loaded=False) and the card otherwise empty (1,119 MiB in use). sparse_structure_resolution = 32 always works.

Traceback is always the same:

trellis2_gguf/pipelines/trellis2_image_to_3d.py:1769, in run
trellis2_gguf/pipelines/trellis2_image_to_3d.py:1313, in sample_shape_slat_cascade
trellis2_gguf/pipelines/trellis2_image_to_3d.py:635,  in get_proj_cond_shape
trellis2_gguf/pipelines/trellis2_image_to_3d.py:610,  in optimized_forward_instance
trellis2/trainers/flow_matching/mixins/image_conditioned_proj.py:710, in _proj_naf_tiled
    z_proj_hr = torch.zeros(
torch.OutOfMemoryError: Allocation on device

Root cause

Two halves; fixing either alone does nothing.

1. The ratio correction is missing. This repo computes it in sample_shape_slat_cascade and divides by lr_resolution * ratio:

# visualbruno/ComfyUI-Trellis2, trellis2/pipelines/trellis2_image_to_3d.py:1104-1109
ratio = (sparse_structure_resolution / 32)
while True:
    quant_coords = torch.cat([
        hr_coords[:, :1],
        ((hr_coords[:, 1:] + 0.5) / (lr_resolution * ratio) * (hr_resolution // 16)).int(),
    ], dim=1)

The fork keeps sparse_structure_resolution in the signature but never computes ratio, and divides by lr_resolution alone (fork, trellis2_gguf/pipelines/trellis2_image_to_3d.py:1289).

The span of upsample(slat, upsample_times=4) is ssres * 16. At ssres=32 that is 512, which happens to equal lr_resolution, so the fork is correct by coincidence. At ssres=64 the span is 1024 and every coordinate comes out doubled.

2. The parameter never arrives anyway. This repo passes it as the 10th positional (:1545); the fork's four call sites in run() stop at max_num_tokens and continue with keyword arguments, so sparse_structure_resolution stays at its default of 32 inside the function.

Why that OOMs

The fork has a block this repo does not (:1311), which derives the projection grid from the data:

hr_grid_res = int(coords[:, 1:].max().item()) + 1
cond = self.get_proj_cond_shape(..., grid_resolution_override=hr_grid_res)

ProjGrid is dense in R³, so doubled coordinates mean an 8× allocation. Measured with torch.cuda.max_memory_allocated() wrapped around each stage:

ssres 32 ssres 64 (before fix)
active voxels after sparse structure 977–1,024 4,042–4,223
R passed to get_proj_cond_shape 61–63 124–125
dense [1, R³, 1024] fp32 tensor 887–977 MiB 7,448–7,629 MiB
peak max_memory_reserved 4.2–6.3 GB 11,584 MiB → OOM

Worth noting for anyone debugging this: the sparse-structure stage itself is bit-identical between the two (peak 1,112 / 1,441 / 2,979 MiB for Q4/Q6/BF16 in both). sparse_structure_resolution costs nothing in the stage it is named after; the entire cost lands downstream in the shape stage's projected conditioning.

Also: even on a card big enough to hold it, the output would be wrong — the 1024 shape flow model expects coords in [0, 1024/16 = 64) and would receive them in [0, 128).

Fix

--- a/trellis2_gguf/pipelines/trellis2_image_to_3d.py
+++ b/trellis2_gguf/pipelines/trellis2_image_to_3d.py
@@ -1285,6 +1285,7 @@ def sample_shape_slat_cascade(...)
+        ratio = (sparse_structure_resolution / 32)
         while True:
             quant_coords = torch.cat([
                 hr_coords[:, :1],
-                ((hr_coords[:, 1:] + 0.5) / lr_resolution * (hr_resolution // 16)).int(),
+                ((hr_coords[:, 1:] + 0.5) / (lr_resolution * ratio) * (hr_resolution // 16)).int(),
             ], dim=1)
@@ -1774 (and :1816, :1852, :1888 — all four branches of pipeline_type)
                 coords, shape_slat_sampler_params,
                 max_num_tokens,
+                sparse_structure_resolution,
                 sampler=shape_sampler or sampler,

After the fix

All five quantizations complete at ssres=64 on the same 12 GB card. ssres=32 is bit-identical (ratio == 1), verified: same 88 s and same 5,408 MiB peak before and after.

ssres 32 ssres 64 (fixed)
voxels from sparse structure 1,011 in 32³ 4,214 in 64³
R for that stage's conditioning 32 63
final shape coords 4,310 4,208
peak allocated 4,592 MiB 4,692 MiB
wall clock 66 s 73 s

So the real cost of ssres=64 is +100 MiB and +7 s — the OOM was entirely the bug inflating the grid to 124³.

On quality I do not want to overclaim: paired silhouette-IoU against 6 reference views, same seed, gave Q4 −0.007, Q5 −0.058, Q6 +0.005, Q8 +0.008, BF16 +0.021. With n=1 per cell that is all inside the seed-to-seed noise I measured separately on this pipeline (std 0.0235, range 0.0658 over 6 seeds), so I can only claim it now runs and is geometrically correct, not that it is better.

Environment

ComfyUI-Trellis2-GGUF @ 6bd11ea, ComfyUI-Trellis2 @ 438fe4e, model Pixal3D-GGUF, torch 2.9.1+cu129, PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, RTX 4070 12 GB, Linux.

One caveat on the measurements: sampling VRAM externally with nvidia-smi at 1 Hz reported a "peak" of 3,913 MiB for the failing runs, which is badly misleading — the real peak lives for well under a second. The numbers above come from torch.cuda.max_memory_allocated() / max_memory_reserved() wrapped around each pipeline stage inside the process.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions