Skip to content
Merged

merge #2809

Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
125 commits
Select commit Hold shift + click to select a range
50f1a99
LTX2: unified conditioning mechanism
Jun 19, 2026
701403d
LTX-2 docs clarified
Jun 19, 2026
c2329eb
ltx2 should not expect clean latents unless the conditioning is enabled
Jun 19, 2026
22263f6
LTX2: rewrite docs for ref dataset knobs, add examples for the new modes
Jun 19, 2026
76355bd
LTX2.3 fixes for validation and training with reference inputs
Jun 19, 2026
a34d5d6
LTX2.3: fix for bf16 mask compatibility with varlen fa3 backend
Jun 19, 2026
e868848
LTX2 multi-gpu training fixes for DDP unwrap and FA3 varlen bug in Di…
Jun 19, 2026
e304042
DDP fixes for LTX2 validation
Jun 19, 2026
94455fd
LTX2 DDP validation fix (another)
Jun 19, 2026
d630941
clean-up for unwrap_model overrides
Jun 19, 2026
2ef5ffd
FSDP2 LoRA training (LTX 2.3 verified)
Jun 19, 2026
eef772d
FSDP2 context-parallel fixes for LTX2.3 and data broadcasting
Jun 19, 2026
ff3a88e
fixes for tests
Jun 19, 2026
c7dbf79
Merge pull request #2774 from bghira/feature/ltx2-conditioning-changes
bghira Jun 19, 2026
5965058
tests
Jun 19, 2026
a0ff224
fix stalling test
Jun 19, 2026
eda6231
tests: fix MagicMock infinite recursion
Jun 20, 2026
89842b5
Merge pull request #2775 from bghira/feature/fsdp2-lora-training
bghira Jun 20, 2026
e04081c
update FSDP2 documentation
Jun 20, 2026
acdc0b3
Merge pull request #2776 from bghira/docs/fsdp2
bghira Jun 20, 2026
90a0f87
support fused projections across more models
Jun 20, 2026
1a24c10
handle 4D case
Jun 20, 2026
5afd442
Merge pull request #2777 from bghira/feature/more-qkv-fusion
bghira Jun 20, 2026
1aecc80
add context parallel plans for more model families
Jun 20, 2026
d29ae6a
Merge pull request #2779 from bghira/feature/more-context-parallel
bghira Jun 20, 2026
5aeae88
add more config examples centred around 8xH100 for LTX 2.3
Jun 20, 2026
dedcf8a
Merge pull request #2780 from bghira/docs/multigpu-h100-examples
bghira Jun 20, 2026
7e61dc1
LTX 2.3 context-parallel fixes for uneven sequence lengths
Jun 20, 2026
7bcaeea
Merge pull request #2781 from bghira/bugfix/ltx2-context-parallel
bghira Jun 20, 2026
f7a55d9
diffusers: monkeypatch loader for flash-attn 2 from kernels library
Jun 20, 2026
653a8b1
Merge pull request #2782 from bghira/bugfix/diffusers-flash-attn-2-at…
bghira Jun 20, 2026
4ec49ac
qwen-image: fix 4D attn case in FA
Jun 20, 2026
3918645
Merge pull request #2783 from bghira/bugfix/diffusers-flash-attn-2-at…
bghira Jun 20, 2026
e0742a8
standalone context-parallel training without FSDP2
Jun 20, 2026
56af88d
add packed attn processors tests
Jun 20, 2026
1a1bf30
validation: allow a single video to generate across multiple GPUs wit…
Jun 20, 2026
9c7d375
Merge pull request #2784 from bghira/feature/cp-without-fsdp
bghira Jun 20, 2026
81e465c
Merge pull request #2785 from bghira/tests/more-qkv-fusion-tests
bghira Jun 20, 2026
7a6ab73
Merge pull request #2786 from bghira/feature/context-parallel-validation
bghira Jun 20, 2026
574a356
(#2778) fix for torchao quantize_via=pipeline
Jun 20, 2026
44848c4
Merge pull request #2787 from bghira/bugfix/torchao-contract-diffuser…
bghira Jun 20, 2026
5a61503
refactor torchao fp8 backend to default to dynamic quantisation; add …
Jun 20, 2026
7f39b55
quant docs update
Jun 20, 2026
09c2f0e
Bump version to 4.4.0
bghira Jun 20, 2026
69c4989
fix test
Jun 20, 2026
b3adfd8
run tests
Jun 20, 2026
66a7b6a
Merge pull request #2788 from bghira/refactor/quant-backends
bghira Jun 20, 2026
f1717ec
TransformerEngine FP8 quant backend
Jun 21, 2026
54d06c1
flux.1: fixes for qkv fused training
Jun 22, 2026
f853137
LTX 2.3 fixes for TransformerEngine with activation checkpointing
Jun 22, 2026
c659191
attention backend fixes for kernels library function selection
Jun 22, 2026
7808a8f
PEFT: fix reduce-overhead torch compile mode by cloning result tensor…
Jun 22, 2026
c17bff4
TransformerEngine fixes for DDP, checkpointing, module filtering
Jun 22, 2026
ccd0b9c
torchao: add custom autograd backward for scaled_mm to differentiate …
Jun 22, 2026
fb29bf5
TE: DDP workarounds and cudagraph breaks
Jun 22, 2026
a2fc27c
attn backend tests
Jun 22, 2026
c758ad5
flux.1: fixes for qkv fused training (tests)
Jun 22, 2026
6cd00cc
torchao validation for PEFT gradient passing in fp8 dynamic quant mode
Jun 22, 2026
838b2b3
add proj_out to torchao test coverage expectations
Jun 22, 2026
7b48e83
tests: transformerengine DDP workaround validation
Jun 22, 2026
2765944
sdnq fp8 behavioural/doc updates
Jun 22, 2026
6b6d7d6
lycoris: add support for fp8-native
Jun 22, 2026
b747b4d
fixes for tests
Jun 22, 2026
91884fa
add some missing dependencies to Docker container
Jun 23, 2026
bf1a2ae
add custom transformer path to docs
Jun 23, 2026
ca4e3ee
Merge pull request #2793 from bghira/docs/options-additions
bghira Jun 24, 2026
f3c2e48
Merge pull request #2792 from bghira/chore/dockerfile-sync
bghira Jun 24, 2026
0bc6bac
fix
Jun 24, 2026
1e648a5
add SDR data generator that can downsample native HDR set for convers…
Jun 24, 2026
b4a223e
TransformerEngine CUDA13 install docs update
Jun 24, 2026
f34dc49
CUDA 13 dependencies update for TE
Jun 24, 2026
cd50592
lycoris: enable bypass_mode for all algo
Jun 24, 2026
1837065
lycoris: enable bypass_mode for all algo (tests)
Jun 24, 2026
7a419a8
test for setup dependencies changing
Jun 24, 2026
3ab1750
clean up SDR generator addition
Jun 24, 2026
cdff86a
fix
Jun 24, 2026
9556827
add ref input support to feature matrix in main readme
Jun 24, 2026
babe5ca
test docker dependencies
Jun 24, 2026
1149828
Merge pull request #2791 from bghira/refactor/quants
bghira Jun 24, 2026
16f6924
Merge pull request #2794 from bghira/feature/hdr-sdr-data-generator
bghira Jun 24, 2026
51cbad2
Merge pull request #2795 from bghira/chore/test-docker-deps
bghira Jun 24, 2026
8615992
docker: add SIMPLETUNER_WORKSPACE and ensure fallback discovers corre…
Jun 24, 2026
2d83ab5
(#2790) fix UI bug in prompt creation modal
Jun 24, 2026
59a4d41
Merge pull request #2796 from bghira/chore/docker-path-fixes
bghira Jun 24, 2026
96eb666
IC-LoRA support for Z-Image Turbo
Jun 24, 2026
64effb2
Z-Image IC-LoRA training examples and metadata fix for empty instance…
Jun 24, 2026
76bb691
Merge pull request #2797 from bghira/bugfix/prompt-library-creation-ui
bghira Jun 24, 2026
450f19b
update example duration to match validation run
Jun 25, 2026
0fdefb7
Merge pull request #2798 from bghira/feature/more-ic-lora-support
bghira Jun 25, 2026
08b05a9
krea2 with qwen embed based reference image training and opt-in clean…
Jun 25, 2026
54d5ade
krea2 documentation for clean ref input option
Jun 25, 2026
134dc30
krea2: fix for 4D vs 5D VAE handling
Jun 25, 2026
af3c259
krea2: docs, training vram vs config matrix
Jun 25, 2026
f44caf2
fix test
Jun 26, 2026
650ada8
Merge pull request #2799 from bghira/feature/krea2-training
bghira Jun 26, 2026
ed67009
krea2: CREPA self-flow, LayerSync, TREAD, block swap, cfg-zero star, …
Jun 26, 2026
f885f41
cog: add a script to deploy multiple versions in sequence for multigp…
Jun 27, 2026
b679095
Merge pull request #2800 from bghira/feature/krea2-training
bghira Jun 27, 2026
49c4ac3
Merge pull request #2801 from bghira/util/cog-deployer-script
bghira Jun 27, 2026
176b7cc
cloud: multigpu training options via Replicate
Jun 27, 2026
993040a
docs: update for multigpu cloud training profiles
Jun 27, 2026
d9bde5e
cloud: fixes for profile defaulting to obsolete container path and up…
Jun 27, 2026
e64a139
cloud: update replicate deploy script to push using cog to ensure ver…
Jun 27, 2026
605cbfb
Merge pull request #2803 from bghira/bugfix/cog-push-script2
bghira Jun 27, 2026
4201383
Merge pull request #2802 from bghira/feature/multigpu-cloud-training
bghira Jun 27, 2026
b9b5986
cloud: update configuration wizard to show correct hardware type sele…
Jun 27, 2026
15cb982
cloud: fix UI settings default to actually show the tab by default
Jun 27, 2026
22757d5
fix webui test ambiguity
Jun 27, 2026
5029faa
cloud: fix rounding by doing so at display-time; other fixes for revi…
Jun 27, 2026
68513a2
Merge pull request #2805 from bghira/bugfix/cloud-tab-enablement
bghira Jun 27, 2026
294dd22
Merge pull request #2804 from bghira/bugfix/replicate-training-hw-types
bghira Jun 27, 2026
1706f0c
fix --disable_tf32 in newer torch api checks
Jun 28, 2026
cf90f96
Merge pull request #2806 from bghira/bugfix/tf32-disable-torch-api-ch…
bghira Jun 29, 2026
2b38fa0
feat: add per-dataset timestep_sampling_offset for semantic-aware flo…
okdsf Jul 7, 2026
f00003c
docs: add comparison images for timestep_sampling_offset feature
okdsf Jul 7, 2026
ce828e5
metal-flash-attention integration
Jul 7, 2026
1823443
rename timestep_bias -> timestep_sampling_offset
okdsf Jul 7, 2026
5424211
docs: add histogram and comparison images for timestep_sampling_offset
okdsf Jul 8, 2026
24bddfd
Merge pull request #2807 from okdsf/feat/timestep-sampling-offset
bghira Jul 8, 2026
b5cfbf7
metal-flash-attention autograd support and other fixes
Jul 11, 2026
96c0115
metal-flash-attention int4 and int8 quantised attn backends
Jul 12, 2026
805b8ba
fix warning text
bghira Jul 12, 2026
6499d6c
update error text to include error
bghira Jul 12, 2026
1e7d737
fix error text to include error string
bghira Jul 12, 2026
2a3f2d4
Merge pull request #2808 from bghira/feature/apple-flash-attn
bghira Jul 12, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -1,5 +1,9 @@
*.code-workspace
multidatabackend*json
!simpletuner/examples/multidatabackend-small-video-480p+81f.json
!simpletuner/examples/multidatabackend-small-video-720p+81f.json
!simpletuner/examples/multidatabackend-small-video-720p+49f.json
!simpletuner/examples/multidatabackend-small-video-1080p+49f.json
# Python and virtual environment files
output/
temp/
Expand Down
11 changes: 9 additions & 2 deletions Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -37,21 +37,28 @@ RUN apt-get update -y && apt-get install -y --no-install-recommends \
&& git lfs install

# 2. Python Environment & Core Deps
# Keep container CUDA deps explicit: the full cuda extra installs NVIDIA
# library wheels that conflict with this CUDA base image.
RUN python${PYTHON_VERSION} -m venv /opt/venv \
&& pip install --upgrade pip setuptools wheel \
&& pip install --no-cache-dir \
"huggingface_hub[cli,hf_transfer]" \
wandb \
mpi4py \
ninja \
"torchao>=0.17.0,<0.18.0"
"bitsandbytes>=0.45.0" \
"deepspeed>=0.17.2" \
"torchao>=0.17.0,<0.18.0" \
"nvidia-ml-py>=12.555" \
"lm-eval>=0.4.4" \
ramtorch

# 3. Install SimpleTuner
# Use main for current model integrations such as Boogu-Image.
ARG SIMPLETUNER_BRANCH=main
RUN git clone https://github.com/bghira/SimpleTuner --branch $SIMPLETUNER_BRANCH \
&& cd SimpleTuner \
&& pip install --no-cache-dir -e .[jxl] \
&& pip install --no-cache-dir -e ".[jxl]" \
&& pip install --no-build-isolation --no-cache-dir sageattention==1.0.6

# 4. Setup Runtime
Expand Down
54 changes: 29 additions & 25 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,31 +81,31 @@ For deployment details, see the [Enterprise Guide](/documentation/experimental/s

### Model Architecture Support

| Model | Parameters | PEFT LoRA | Lycoris | Full-Rank | ControlNet | Quantization | Flow Matching | Text Encoders |
|-------|------------|-----------|---------|-----------|------------|--------------|---------------|---------------|
| **Stable Diffusion XL** | 3.5B | ✓ | ✓ | ✓ | ✓ | int8/nf4 | ✗ | CLIP-L/G |
| **Stable Diffusion 3** | 2B-8B | ✓ | ✓ | ✓* | ✓ | int8/fp8/nf4 | ✓ | CLIP-L/G + T5-XXL |
| **Flux.1** | 12B | ✓ | ✓ | ✓* | ✓ | int8/fp8/nf4 | ✓ | CLIP-L + T5-XXL |
| **Flux.2** | 32B | ✓ | ✓ | ✓* | ✗ | int8/fp8/nf4 | ✓ | Mistral-3 Small |
| **Ideogram 4** | 9B | ✓ | ✓ | ✓* | ✗ | fp8/nf4 | ✓ | Qwen3-VL |
| **ACE-Step** | 3.5B | ✓ | ✓ | ✓* | ✗ | int8 | ✓ | UMT5 |
| **HeartMuLa** | 3B | ✓ | ✓ | ✓* | ✗ | int8 | ✗ | None |
| **Chroma 1** | 8.9B | ✓ | ✓ | ✓* | ✗ | int8/fp8/nf4 | ✓ | T5-XXL |
| **Auraflow** | 6.8B | ✓ | ✓ | ✓* | ✓ | int8/fp8/nf4 | ✓ | UMT5-XXL |
| **PixArt Sigma** | 0.6B-0.9B | ✗ | ✓ | ✓ | ✓ | int8 | ✗ | T5-XXL |
| **Sana** | 0.6B-4.8B | ✗ | ✓ | ✓ | ✗ | int8 | ✓ | Gemma2-2B |
| **Lumina2** | 2B | ✓ | ✓ | ✓ | ✗ | int8 | ✓ | Gemma2 |
| **Kwai Kolors** | 5B | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ChatGLM-6B |
| **LTX Video** | 5B | ✓ | ✓ | ✓ | ✗ | int8/fp8 | ✓ | T5-XXL |
| **LTX Video 2** | 19B | ✓ | ✓ | ✓* | ✗ | int8/fp8 | ✓ | Gemma3 |
| **Wan Video** | 1.3B-14B | ✓ | ✓ | ✓* | ✗ | int8 | ✓ | UMT5 |
| **HiDream** | 17B (8.5B MoE) | ✓ | ✓ | ✓* | ✓ | int8/fp8/nf4 | ✓ | CLIP-L + T5-XXL + Llama |
| **Cosmos2** | 2B-14B | ✗ | ✓ | ✓ | ✗ | int8 | ✓ | T5-XXL |
| **OmniGen** | 3.8B | ✓ | ✓ | ✓ | ✗ | int8/fp8 | ✓ | T5-XXL |
| **Qwen Image** | 20B | ✓ | ✓ | ✓* | ✗ | int8/nf4 (req.) | ✓ | T5-XXL |
| **SD 1.x/2.x (Legacy)** | 0.9B | ✓ | ✓ | ✓ | ✓ | int8/nf4 | ✗ | CLIP-L |

*✓ = Supported, ✗ = Not supported, * = Requires DeepSpeed for full-rank training*
| Model | Parameters | PEFT LoRA | Lycoris | Full-Rank | ControlNet | Ref Inputs | Quantization | Flow Matching | Text Encoders |
|-------|------------|-----------|---------|-----------|------------|------------|--------------|---------------|---------------|
| **Stable Diffusion XL** | 3.5B | ✓ | ✓ | ✓ | ✓ | ✗ | int8/nf4 | ✗ | CLIP-L/G |
| **Stable Diffusion 3** | 2B-8B | ✓ | ✓ | ✓* | ✓ | ✗ | int8/fp8/nf4 | ✓ | CLIP-L/G + T5-XXL |
| **Flux.1** | 12B | ✓ | ✓ | ✓* | ✓ | ✓ (Kontext) | int8/fp8/nf4 | ✓ | CLIP-L + T5-XXL |
| **Flux.2** | 32B | ✓ | ✓ | ✓* | ✗ | ✓ opt | int8/fp8/nf4 | ✓ | Mistral-3 Small |
| **Ideogram 4** | 9B | ✓ | ✓ | ✓* | ✗ | ✗ | fp8/nf4 | ✓ | Qwen3-VL |
| **ACE-Step** | 3.5B | ✓ | ✓ | ✓* | ✗ | ✗ | int8 | ✓ | UMT5 |
| **HeartMuLa** | 3B | ✓ | ✓ | ✓* | ✗ | ✗ | int8 | ✗ | None |
| **Chroma 1** | 8.9B | ✓ | ✓ | ✓* | ✗ | ✗ | int8/fp8/nf4 | ✓ | T5-XXL |
| **Auraflow** | 6.8B | ✓ | ✓ | ✓* | ✓ | ✗ | int8/fp8/nf4 | ✓ | UMT5-XXL |
| **PixArt Sigma** | 0.6B-0.9B | ✗ | ✓ | ✓ | ✓ | ✗ | int8 | ✗ | T5-XXL |
| **Sana** | 0.6B-4.8B | ✗ | ✓ | ✓ | ✗ | ✗ | int8 | ✓ | Gemma2-2B |
| **Lumina2** | 2B | ✓ | ✓ | ✓ | ✗ | ✗ | int8 | ✓ | Gemma2 |
| **Kwai Kolors** | 5B | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ChatGLM-6B |
| **LTX Video** | 5B | ✓ | ✓ | ✓ | ✗ | ✓ I2V | int8/fp8 | ✓ | T5-XXL |
| **LTX Video 2** | 19B | ✓ | ✓ | ✓* | ✗ | ✓ opt | int8/fp8 | ✓ | Gemma3 |
| **Wan Video** | 1.3B-14B | ✓ | ✓ | ✓* | ✗ | ✗ | int8 | ✓ | UMT5 |
| **HiDream** | 17B (8.5B MoE) | ✓ | ✓ | ✓* | ✓ | ✗ | int8/fp8/nf4 | ✓ | CLIP-L + T5-XXL + Llama |
| **Cosmos2** | 2B-14B | ✗ | ✓ | ✓ | ✗ | ✗ | int8 | ✓ | T5-XXL |
| **OmniGen** | 3.8B | ✓ | ✓ | ✓ | ✗ | ✗ | int8/fp8 | ✓ | T5-XXL |
| **Qwen Image** | 20B | ✓ | ✓ | ✓* | ✗ | ✓ req (Edit) | int8/nf4 (req.) | ✓ | T5-XXL |
| **SD 1.x/2.x (Legacy)** | 0.9B | ✓ | ✓ | ✓ | ✓ | ✗ | int8/nf4 | ✗ | CLIP-L |

*✓ = Supported, ✗ = Not supported, * = Requires DeepSpeed for full-rank training, Ref Inputs marks existing reference/edit/I2V conditioning paths only*

### Advanced Training Techniques

Expand All @@ -125,6 +125,7 @@ For deployment details, see the [Enterprise Guide](/documentation/experimental/s
### Model-Specific Features

- **Flux Kontext** - Edit conditioning and image-to-image training for Flux models
- **Reference-input training** - Existing paired reference/edit/I2V paths for Flux Kontext, Flux.2, LTX Video 2, Qwen Edit, LongCat edit/I2V, Boogu edit, Hunyuan I2V, and Kandinsky I2I/I2V
- **PixArt two-stage** - eDiff training pipeline support for PixArt Sigma
- **Flow matching models** - Advanced scheduling with beta/uniform distributions
- **HiDream MoE** - Mixture of Experts gate loss augmentation
Expand Down Expand Up @@ -200,6 +201,9 @@ pip install 'simpletuner[cuda]'
# CUDA 13 / Blackwell users (NVIDIA B-series GPUs)
pip install 'simpletuner[cuda13]' --extra-index-url https://download.pytorch.org/whl/cu130

# CUDA 13 with TransformerEngine FP8 support
pip install 'simpletuner[cuda13-transformerengine]' --extra-index-url https://download.pytorch.org/whl/cu130

# ROCm users (AMD GPUs)
pip install 'simpletuner[rocm]' --extra-index-url https://download.pytorch.org/whl/rocm7.1

Expand Down
2 changes: 1 addition & 1 deletion docker-start.sh
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@
# This file can then later be sourced in a login shell
echo "Exporting environment variables..."
printenv |
grep -E '^RUNPOD_|^PATH=|^HF_HOME=|^HF_TOKEN=|^HUGGING_FACE_HUB_TOKEN=|^WANDB_API_KEY=|^WANDB_TOKEN=|^_=' |
grep -E '^RUNPOD_|^PATH=|^HF_HOME=|^HF_TOKEN=|^HUGGING_FACE_HUB_TOKEN=|^SIMPLETUNER_WORKSPACE=|^WANDB_API_KEY=|^WANDB_TOKEN=|^_=' |
sed 's/^\(.*\)=\(.*\)$/export \1="\2"/' >>/etc/rp_environment

# Add it to Bash login script
Expand Down
Binary file added docs/images/timestep_sampling_offset/A1.jpg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/images/timestep_sampling_offset/A2.jpg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/images/timestep_sampling_offset/A3.jpg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/images/timestep_sampling_offset/Q1.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/images/timestep_sampling_offset/Q2.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/images/timestep_sampling_offset/Q3.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
21 changes: 21 additions & 0 deletions documentation/DATALOADER.es.md
Original file line number Diff line number Diff line change
Expand Up @@ -352,6 +352,20 @@ Genera versiones de baja calidad de las imágenes para entrenamiento de super-re
}
```

##### `sdr` / `logc3_sdr`
Genera imágenes SDR/de referencia de condicionamiento para datasets de condicionamiento emparejados. El transform `rec709` predeterminado normaliza y recorta entradas Rec.709 SDR ya existentes, igual que la ruta de acondicionamiento de referencia HDR IC-LoRA de LTX-2:
```json
{
"type": "sdr",
"conditioning_type": "reference_strict",
"transform": "rec709",
"input_scale": 1.0,
"exposure": 0.0,
"captions": false
}
```
Usa `transform: "srgb"` cuando los valores decodificados sean lineales y quieras un proxy SDR de pantalla. Usa `transform: "logc3"` (o el alias `logc3_sdr`) solo cuando quieras muestras codificadas en LogC3. `input_scale` se aplica antes de exposure y sirve cuando los valores decodificados necesitan normalización. Este generador actualmente opera sobre muestras de imagen que el backend de imágenes de SimpleTuner puede decodificar.

##### `jpeg_artifacts`
Crea artefactos de compresión JPEG para entrenamiento de eliminación de artefactos:
```json
Expand Down Expand Up @@ -590,6 +604,13 @@ Por defecto, SimpleTuner escalará hacia arriba imágenes pequeñas para cumplir

- Además de `prepend_instance_prompt`, reemplaza todos los captions del dataset con una sola frase o palabra trigger.

### `timestep_sampling_offset`

- Desplaza el muestreo de timesteps de flow-matching de este dataset hacia niveles de ruido más altos o más bajos. Se aplica al programa por defecto logit-normal (sigmoid).
- Un valor **negativo** sesga el muestreo hacia timesteps de menor ruido, centrando el entrenamiento en el detalle fino (por ejemplo, primeros planos o datos con mucha textura). Un valor **positivo** sesga hacia timesteps de mayor ruido, centrándose en la estructura general (por ejemplo, cuerpo entero o datos centrados en la composición).
- El valor se suma a la muestra normal previa al sigmoid; las magnitudes útiles suelen estar entre `-1.0` y `1.0`. Por defecto es `0.0` (sin sesgo, idéntico al comportamiento original).
- Como cada lote (batch) proviene de un único dataset, agrupa las imágenes por granularidad semántica en datasets separados y define un `timestep_sampling_offset` para cada uno.

### `repeats`

- Especifica cuántas veces se ven todas las muestras del dataset durante una época. Útil para dar más impacto a datasets pequeños o maximizar el uso de objetos de caché VAE.
Expand Down
21 changes: 21 additions & 0 deletions documentation/DATALOADER.hi.md
Original file line number Diff line number Diff line change
Expand Up @@ -352,6 +352,20 @@ Super‑resolution training के लिए images के low‑quality सं
}
```

##### `sdr` / `logc3_sdr`
Paired conditioning datasets के लिए SDR/reference conditioning images बनाता है। Default `rec709` transform पहले से SDR Rec.709 inputs को normalize और clamp करता है, जो LTX-2 HDR IC-LoRA reference-conditioning path से मेल खाता है:
```json
{
"type": "sdr",
"conditioning_type": "reference_strict",
"transform": "rec709",
"input_scale": 1.0,
"exposure": 0.0,
"captions": false
}
```
Decoded source values linear हों और display SDR proxy चाहिए तो `transform: "srgb"` उपयोग करें। LogC3-encoded samples जानबूझकर चाहिए तभी `transform: "logc3"` (या `logc3_sdr` alias) उपयोग करें। `input_scale` exposure से पहले apply होता है और decoded values को normalize करने की जरूरत होने पर उपयोगी है। यह generator अभी उन image samples पर काम करता है जिन्हें SimpleTuner image backend decode कर सकता है।

##### `jpeg_artifacts`
Artifact removal training के लिए JPEG compression artifacts बनाता है:
```json
Expand Down Expand Up @@ -590,6 +604,13 @@ Images cropping से पहले resize नहीं होतीं **जब

- `prepend_instance_prompt` के अतिरिक्त, पूरे dataset के सभी captions को एक ही phrase या trigger word से बदल देता है।

### `timestep_sampling_offset`

- इस dataset के flow-matching timestep sampling को अधिक या कम noise स्तरों की ओर स्थानांतरित करता है। यह डिफ़ॉल्ट logit-normal (sigmoid) schedule पर लागू होता है।
- **ऋणात्मक** मान sampling को कम-noise timesteps की ओर झुकाता है, जिससे प्रशिक्षण बारीक विवरण (जैसे क्लोज़-अप या texture-प्रधान data) पर केंद्रित होता है। **धनात्मक** मान अधिक-noise timesteps की ओर झुकाता है, जिससे समग्र संरचना (जैसे फ़ुल-बॉडी या रचना-प्रधान data) पर ध्यान केंद्रित होता है।
- यह मान sigmoid से पहले के normal sample में जोड़ा जाता है; उपयोगी मात्राएँ आमतौर पर `-1.0` से `1.0` के बीच होती हैं। डिफ़ॉल्ट `0.0` है (कोई bias नहीं, मूल व्यवहार के समान)।
- चूँकि प्रत्येक batch एकल dataset से लिया जाता है, images को semantic granularity के अनुसार अलग-अलग datasets में समूहित करें और प्रत्येक पर `timestep_sampling_offset` सेट करें।

### `repeats`

- यह बताता है कि epoch के दौरान dataset के सभी samples कितनी बार देखे जाते हैं। छोटे datasets का प्रभाव बढ़ाने या VAE cache objects का अधिकतम उपयोग करने के लिए उपयोगी।
Expand Down
21 changes: 21 additions & 0 deletions documentation/DATALOADER.ja.md
Original file line number Diff line number Diff line change
Expand Up @@ -352,6 +352,20 @@ Hugging Face の音声データセットでは、キャプション(プロン
}
```

##### `sdr` / `logc3_sdr`
ペア条件データセット向けに SDR/参照条件画像を生成します。既定の `rec709` transform は、すでに Rec.709 SDR である入力を正規化してクリップし、LTX-2 HDR IC-LoRA の参照条件パスと一致します:
```json
{
"type": "sdr",
"conditioning_type": "reference_strict",
"transform": "rec709",
"input_scale": 1.0,
"exposure": 0.0,
"captions": false
}
```
デコード済みソース値がリニアで、表示用 SDR プロキシが必要な場合は `transform: "srgb"` を使用します。LogC3 エンコードサンプルが明示的に必要な場合だけ `transform: "logc3"`(または `logc3_sdr` エイリアス)を使用してください。`input_scale` は exposure の前に適用され、デコード済み値の正規化が必要な場合に使えます。このジェネレータは現在、SimpleTuner の画像バックエンドがデコードできる画像サンプルを対象とします。

##### `jpeg_artifacts`
アーティファクト除去学習向けに JPEG 圧縮アーティファクトを作成します:
```json
Expand Down Expand Up @@ -590,6 +604,13 @@ Canny エッジ検出マップを生成します:

- `prepend_instance_prompt` に加えて、データセット内のすべてのキャプションを単一のフレーズまたはトリガーワードで置き換えます。

### `timestep_sampling_offset`

- このデータセットのフローマッチングのタイムステップサンプリングを、より高いまたは低いノイズレベルへシフトします。デフォルトの logit-normal(sigmoid)スケジュールに適用されます。
- **負の値**はサンプリングを低ノイズのタイムステップへ偏らせ、細部(クローズアップやテクスチャの多いデータなど)の学習に重点を置きます。**正の値**は高ノイズのタイムステップへ偏らせ、全体の構造(全身や構図中心のデータなど)に重点を置きます。
- 値は sigmoid 適用前の正規サンプルに加算されます。有効な大きさは通常 `-1.0` ~ `1.0` 程度です。デフォルトは `0.0`(バイアスなし、既存の動作と同一)です。
- 各バッチは単一のデータセットから取得されるため、画像を意味的な粒度でデータセットに分け、それぞれに `timestep_sampling_offset` を設定してください。

### `repeats`

- エポック中にデータセット内のすべてのサンプルが表示される回数を指定します。小さなデータセットの影響を強めたり、VAE キャッシュオブジェクトの使用率を最大化したりするのに有用です。
Expand Down
21 changes: 21 additions & 0 deletions documentation/DATALOADER.md
Original file line number Diff line number Diff line change
Expand Up @@ -353,6 +353,20 @@ Generates low-quality versions of images for super-resolution training:
}
```

##### `sdr` / `logc3_sdr`
Generates SDR/reference conditioning images for paired conditioning datasets. The default `rec709` transform normalizes and clamps already-SDR Rec.709 inputs, matching the LTX-2 HDR IC-LoRA reference-conditioning path:
```json
{
"type": "sdr",
"conditioning_type": "reference_strict",
"transform": "rec709",
"input_scale": 1.0,
"exposure": 0.0,
"captions": false
}
```
Use `transform: "srgb"` when your decoded source values are linear and you want a display SDR proxy. Use `transform: "logc3"` (or the `logc3_sdr` alias) only when you intentionally want LogC3-encoded samples. `input_scale` is applied before exposure and is useful when decoded values need normalization. This generator currently operates on image samples that the SimpleTuner image backend can decode.

##### `jpeg_artifacts`
Creates JPEG compression artifacts for artifact removal training:
```json
Expand Down Expand Up @@ -591,6 +605,13 @@ By default, SimpleTuner will upscale small images to meet the target resolution,

- In addition to `prepend_instance_prompt`, replaces all captions in the dataset with a single phrase or trigger word.

### `timestep_sampling_offset`

- Shifts this dataset's flow-matching timestep sampling toward higher or lower noise levels. Applies to the default logit-normal (sigmoid) schedule.
- A **negative** value biases sampling toward lower-noise timesteps, focusing training on fine detail (e.g. close-up or texture-heavy data). A **positive** value biases toward higher-noise timesteps, focusing on overall structure (e.g. full-body or composition-heavy data).
- The value is added to the pre-sigmoid normal sample; useful magnitudes are typically around `-1.0` to `1.0`. Defaults to `0.0` (no bias, identical to stock behaviour).
- Because each batch is drawn from a single dataset, group images by semantic granularity into separate datasets and set a per-dataset `timestep_sampling_offset` on each.

### `repeats`

- Specifies the number of times all samples in the dataset are seen during an epoch. Useful for giving more impact to smaller datasets or maximizing the usage of VAE cache objects.
Expand Down
Loading
Loading