Filed here because the fork has issues disabled. This is not a bug in visualbruno/ComfyUI-Trellis2 — the code here is correct. The bug is in the GGUF fork Aero-Ex/ComfyUI-Trellis2-GGUF (has_issues: false), which dropped one line when porting sample_shape_slat_cascade. Filing upstream so it is findable by anyone hitting the same OOM, and in case you want to relay it.
Symptom
On the GGUF fork with the Pixal3D-GGUF model, sparse_structure_resolution = 64 OOMs on a 12 GB card in all five quantizations (Q4_K_M, Q5_K_M, Q6_K, Q8_0, BF16), even with every other lever at minimum (pipeline_type=1024_cascade, steps 8/8/8, max_num_tokens=16384, low_vram=True, keep_models_loaded=False) and the card otherwise empty (1,119 MiB in use). sparse_structure_resolution = 32 always works.
Traceback is always the same:
trellis2_gguf/pipelines/trellis2_image_to_3d.py:1769, in run
trellis2_gguf/pipelines/trellis2_image_to_3d.py:1313, in sample_shape_slat_cascade
trellis2_gguf/pipelines/trellis2_image_to_3d.py:635, in get_proj_cond_shape
trellis2_gguf/pipelines/trellis2_image_to_3d.py:610, in optimized_forward_instance
trellis2/trainers/flow_matching/mixins/image_conditioned_proj.py:710, in _proj_naf_tiled
z_proj_hr = torch.zeros(
torch.OutOfMemoryError: Allocation on device
Root cause
Two halves; fixing either alone does nothing.
1. The ratio correction is missing. This repo computes it in sample_shape_slat_cascade and divides by lr_resolution * ratio:
# visualbruno/ComfyUI-Trellis2, trellis2/pipelines/trellis2_image_to_3d.py:1104-1109
ratio = (sparse_structure_resolution / 32)
while True:
quant_coords = torch.cat([
hr_coords[:, :1],
((hr_coords[:, 1:] + 0.5) / (lr_resolution * ratio) * (hr_resolution // 16)).int(),
], dim=1)
The fork keeps sparse_structure_resolution in the signature but never computes ratio, and divides by lr_resolution alone (fork, trellis2_gguf/pipelines/trellis2_image_to_3d.py:1289).
The span of upsample(slat, upsample_times=4) is ssres * 16. At ssres=32 that is 512, which happens to equal lr_resolution, so the fork is correct by coincidence. At ssres=64 the span is 1024 and every coordinate comes out doubled.
2. The parameter never arrives anyway. This repo passes it as the 10th positional (:1545); the fork's four call sites in run() stop at max_num_tokens and continue with keyword arguments, so sparse_structure_resolution stays at its default of 32 inside the function.
Why that OOMs
The fork has a block this repo does not (:1311), which derives the projection grid from the data:
hr_grid_res = int(coords[:, 1:].max().item()) + 1
cond = self.get_proj_cond_shape(..., grid_resolution_override=hr_grid_res)
ProjGrid is dense in R³, so doubled coordinates mean an 8× allocation. Measured with torch.cuda.max_memory_allocated() wrapped around each stage:
|
ssres 32 |
ssres 64 (before fix) |
| active voxels after sparse structure |
977–1,024 |
4,042–4,223 |
R passed to get_proj_cond_shape |
61–63 |
124–125 |
dense [1, R³, 1024] fp32 tensor |
887–977 MiB |
7,448–7,629 MiB |
peak max_memory_reserved |
4.2–6.3 GB |
11,584 MiB → OOM |
Worth noting for anyone debugging this: the sparse-structure stage itself is bit-identical between the two (peak 1,112 / 1,441 / 2,979 MiB for Q4/Q6/BF16 in both). sparse_structure_resolution costs nothing in the stage it is named after; the entire cost lands downstream in the shape stage's projected conditioning.
Also: even on a card big enough to hold it, the output would be wrong — the 1024 shape flow model expects coords in [0, 1024/16 = 64) and would receive them in [0, 128).
Fix
--- a/trellis2_gguf/pipelines/trellis2_image_to_3d.py
+++ b/trellis2_gguf/pipelines/trellis2_image_to_3d.py
@@ -1285,6 +1285,7 @@ def sample_shape_slat_cascade(...)
+ ratio = (sparse_structure_resolution / 32)
while True:
quant_coords = torch.cat([
hr_coords[:, :1],
- ((hr_coords[:, 1:] + 0.5) / lr_resolution * (hr_resolution // 16)).int(),
+ ((hr_coords[:, 1:] + 0.5) / (lr_resolution * ratio) * (hr_resolution // 16)).int(),
], dim=1)
@@ -1774 (and :1816, :1852, :1888 — all four branches of pipeline_type)
coords, shape_slat_sampler_params,
max_num_tokens,
+ sparse_structure_resolution,
sampler=shape_sampler or sampler,
After the fix
All five quantizations complete at ssres=64 on the same 12 GB card. ssres=32 is bit-identical (ratio == 1), verified: same 88 s and same 5,408 MiB peak before and after.
|
ssres 32 |
ssres 64 (fixed) |
| voxels from sparse structure |
1,011 in 32³ |
4,214 in 64³ |
| R for that stage's conditioning |
32 |
63 |
| final shape coords |
4,310 |
4,208 |
| peak allocated |
4,592 MiB |
4,692 MiB |
| wall clock |
66 s |
73 s |
So the real cost of ssres=64 is +100 MiB and +7 s — the OOM was entirely the bug inflating the grid to 124³.
On quality I do not want to overclaim: paired silhouette-IoU against 6 reference views, same seed, gave Q4 −0.007, Q5 −0.058, Q6 +0.005, Q8 +0.008, BF16 +0.021. With n=1 per cell that is all inside the seed-to-seed noise I measured separately on this pipeline (std 0.0235, range 0.0658 over 6 seeds), so I can only claim it now runs and is geometrically correct, not that it is better.
Environment
ComfyUI-Trellis2-GGUF @ 6bd11ea, ComfyUI-Trellis2 @ 438fe4e, model Pixal3D-GGUF, torch 2.9.1+cu129, PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, RTX 4070 12 GB, Linux.
One caveat on the measurements: sampling VRAM externally with nvidia-smi at 1 Hz reported a "peak" of 3,913 MiB for the failing runs, which is badly misleading — the real peak lives for well under a second. The numbers above come from torch.cuda.max_memory_allocated() / max_memory_reserved() wrapped around each pipeline stage inside the process.
Filed here because the fork has issues disabled. This is not a bug in
visualbruno/ComfyUI-Trellis2— the code here is correct. The bug is in the GGUF forkAero-Ex/ComfyUI-Trellis2-GGUF(has_issues: false), which dropped one line when portingsample_shape_slat_cascade. Filing upstream so it is findable by anyone hitting the same OOM, and in case you want to relay it.Symptom
On the GGUF fork with the Pixal3D-GGUF model,
sparse_structure_resolution = 64OOMs on a 12 GB card in all five quantizations (Q4_K_M, Q5_K_M, Q6_K, Q8_0, BF16), even with every other lever at minimum (pipeline_type=1024_cascade, steps 8/8/8,max_num_tokens=16384,low_vram=True,keep_models_loaded=False) and the card otherwise empty (1,119 MiB in use).sparse_structure_resolution = 32always works.Traceback is always the same:
Root cause
Two halves; fixing either alone does nothing.
1. The
ratiocorrection is missing. This repo computes it insample_shape_slat_cascadeand divides bylr_resolution * ratio:The fork keeps
sparse_structure_resolutionin the signature but never computesratio, and divides bylr_resolutionalone (fork,trellis2_gguf/pipelines/trellis2_image_to_3d.py:1289).The span of
upsample(slat, upsample_times=4)isssres * 16. Atssres=32that is 512, which happens to equallr_resolution, so the fork is correct by coincidence. Atssres=64the span is 1024 and every coordinate comes out doubled.2. The parameter never arrives anyway. This repo passes it as the 10th positional (
:1545); the fork's four call sites inrun()stop atmax_num_tokensand continue with keyword arguments, sosparse_structure_resolutionstays at its default of 32 inside the function.Why that OOMs
The fork has a block this repo does not (
:1311), which derives the projection grid from the data:ProjGridis dense in R³, so doubled coordinates mean an 8× allocation. Measured withtorch.cuda.max_memory_allocated()wrapped around each stage:get_proj_cond_shape[1, R³, 1024]fp32 tensormax_memory_reservedWorth noting for anyone debugging this: the sparse-structure stage itself is bit-identical between the two (peak 1,112 / 1,441 / 2,979 MiB for Q4/Q6/BF16 in both).
sparse_structure_resolutioncosts nothing in the stage it is named after; the entire cost lands downstream in the shape stage's projected conditioning.Also: even on a card big enough to hold it, the output would be wrong — the 1024 shape flow model expects coords in
[0, 1024/16 = 64)and would receive them in[0, 128).Fix
After the fix
All five quantizations complete at
ssres=64on the same 12 GB card.ssres=32is bit-identical (ratio == 1), verified: same 88 s and same 5,408 MiB peak before and after.So the real cost of
ssres=64is +100 MiB and +7 s — the OOM was entirely the bug inflating the grid to 124³.On quality I do not want to overclaim: paired silhouette-IoU against 6 reference views, same seed, gave Q4 −0.007, Q5 −0.058, Q6 +0.005, Q8 +0.008, BF16 +0.021. With n=1 per cell that is all inside the seed-to-seed noise I measured separately on this pipeline (std 0.0235, range 0.0658 over 6 seeds), so I can only claim it now runs and is geometrically correct, not that it is better.
Environment
ComfyUI-Trellis2-GGUF@6bd11ea,ComfyUI-Trellis2@438fe4e, model Pixal3D-GGUF, torch 2.9.1+cu129,PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, RTX 4070 12 GB, Linux.One caveat on the measurements: sampling VRAM externally with
nvidia-smiat 1 Hz reported a "peak" of 3,913 MiB for the failing runs, which is badly misleading — the real peak lives for well under a second. The numbers above come fromtorch.cuda.max_memory_allocated()/max_memory_reserved()wrapped around each pipeline stage inside the process.