Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 16 additions & 3 deletions docs/src/content/docs/configuration/Optimization/low-vram-mode.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,7 @@ Low-VRAM mode and related workload-specific optimizations include:
- Partial model loading (`enable_partial_loading`)
- PyTorch CUDA allocator config (`pytorch_cuda_alloc_conf`)
- Dynamic RAM and VRAM cache sizes (`max_cache_ram_gb`, `max_cache_vram_gb`)
- Extra VRAM kept outside the model cache (`reserve_vram_gb`)
- Working memory (`device_working_mem_gb`)
- Keeping a RAM weight copy (`keep_ram_copy_of_weights`)
- Tiled VAE decode and encode (`auto_tiled_decode`, `force_tiled_decode`)
Expand Down Expand Up @@ -125,15 +126,27 @@ max_cache_vram_gb: 16
:::caution[Max safe value for `max_cache_vram_gb`]
Most users should not manually configure the `max_cache_vram_gb`. This configuration value caps model-cache residency; `device_working_mem_gb` and operation-specific reservations (e.g. VAE decode) are still subtracted from that cap for every model-cache operation, not only when Wan memory optimization is enabled. A cap below the active working-memory reservation can force aggressive model offloading.

For users who wish to configure `max_cache_vram_gb`, the max safe value can be determined by subtracting `device_working_mem_gb` from your GPU's VRAM. As described below, the default for `device_working_mem_gb` is 3GB.
For users who wish to configure `max_cache_vram_gb`, estimate the max safe value by subtracting both `device_working_mem_gb` and `reserve_vram_gb` from your GPU's VRAM. The default for `device_working_mem_gb` is 3GB, and `reserve_vram_gb` defaults to 0GB.

For example, if you have a 12GB GPU, the max safe value for `max_cache_vram_gb` is `12GB - 3GB = 9GB`.
For example, with a 12GB GPU and default settings, the max safe value is `12GB - 3GB - 0GB = 9GB`.

If you had increased `device_working_mem_gb` to 4GB, then the max safe value for `max_cache_vram_gb` is `12GB - 4GB = 8GB`.
If you increase `device_working_mem_gb` to 4GB and set `reserve_vram_gb` to 1GB, the max safe value is `12GB - 4GB - 1GB = 7GB`.

Most users who override `max_cache_vram_gb` are doing so because they wish to use significantly less VRAM, and should be setting `max_cache_vram_gb` to a value significantly less than the 'max safe value'.
:::

### Reserving VRAM outside the model cache

`reserve_vram_gb` keeps the configured amount of VRAM unavailable to the model cache. This can leave room for other GPU workloads that are not managed by the cache. The default is `0`, so the option has no effect unless configured.

For example, add this to `invokeai.yaml` to reserve 1GB:

```yaml
reserve_vram_gb: 1
```

This reserve is in addition to `device_working_mem_gb` and any operation-specific working-memory requests. A larger reserve can cause models to be offloaded or partially loaded sooner.

### Working memory

Invoke cannot use _all_ of your VRAM for model caching and loading. It requires some VRAM to use as working memory for various operations.
Expand Down
11 changes: 11 additions & 0 deletions docs/src/generated/settings.json
Original file line number Diff line number Diff line change
Expand Up @@ -493,6 +493,17 @@
"type": "typing.Optional[float]",
"validation": {}
},
{
"category": "CACHE",
"default": 0.0,
"description": "The amount of VRAM (in GB) to subtract from the model cache's available memory budget. Defaults to 0.",
"env_var": "INVOKEAI_RESERVE_VRAM_GB",
"literal_values": [],
"name": "reserve_vram_gb",
"required": false,
"type": "<class 'float'>",
"validation": {}
},
{
"category": "CACHE",
"default": false,
Expand Down
2 changes: 2 additions & 0 deletions invokeai/app/services/config/config_default.py
Original file line number Diff line number Diff line change
Expand Up @@ -111,6 +111,7 @@ class InvokeAIAppConfig(BaseSettings):
profiles_dir: Path to profiles output directory.
max_cache_ram_gb: The maximum amount of CPU RAM to use for model caching in GB. If unset, the limit will be configured based on the available RAM. In most cases, it is recommended to leave this unset.
max_cache_vram_gb: The amount of VRAM to use for model caching in GB. If unset, the limit will be configured based on the available VRAM and the device_working_mem_gb. In most cases, it is recommended to leave this unset.
reserve_vram_gb: The amount of VRAM (in GB) to subtract from the model cache's available memory budget. Defaults to 0.
log_memory_usage: If True, a memory snapshot will be captured before and after every model cache operation, and the result will be logged (at debug level). There is a time cost to capturing the memory snapshots, so it is recommended to only enable this feature if you are actively inspecting the model cache's behaviour.
model_cache_keep_alive_min: How long to keep models in cache after last use, in minutes. A value of 0 (the default) means models are kept in cache indefinitely. If no model generations occur within the timeout period, the model cache is cleared using the same logic as the 'Clear Model Cache' button.
device_working_mem_gb: The amount of working memory to keep available on the compute device (in GB). Has no effect if running on CPU. If you are experiencing OOM errors, try increasing this value.
Expand Down Expand Up @@ -234,6 +235,7 @@ class InvokeAIAppConfig(BaseSettings):
# CACHE
max_cache_ram_gb: Optional[float] = Field(default=None, gt=0, description="The maximum amount of CPU RAM to use for model caching in GB. If unset, the limit will be configured based on the available RAM. In most cases, it is recommended to leave this unset.")
max_cache_vram_gb: Optional[float] = Field(default=None, ge=0, description="The amount of VRAM to use for model caching in GB. If unset, the limit will be configured based on the available VRAM and the device_working_mem_gb. In most cases, it is recommended to leave this unset.")
reserve_vram_gb: float = Field(default=0.0, ge=0, description="The amount of VRAM (in GB) to subtract from the model cache's available memory budget. Defaults to 0.")
log_memory_usage: bool = Field(default=False, description="If True, a memory snapshot will be captured before and after every model cache operation, and the result will be logged (at debug level). There is a time cost to capturing the memory snapshots, so it is recommended to only enable this feature if you are actively inspecting the model cache's behaviour.")
model_cache_keep_alive_min: float = Field(default=0, ge=0, description="How long to keep models in cache after last use, in minutes. A value of 0 (the default) means models are kept in cache indefinitely. If no model generations occur within the timeout period, the model cache is cleared using the same logic as the 'Clear Model Cache' button.")
device_working_mem_gb: float = Field(default=3, description="The amount of working memory to keep available on the compute device (in GB). Has no effect if running on CPU. If you are experiencing OOM errors, try increasing this value.")
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -100,6 +100,7 @@ def build_cache(device: torch.device) -> ModelCache:
keep_ram_copy_of_weights=app_config.keep_ram_copy_of_weights,
max_ram_cache_size_gb=app_config.max_cache_ram_gb,
max_vram_cache_size_gb=app_config.max_cache_vram_gb,
reserve_vram_gb=app_config.reserve_vram_gb,
execution_device=device,
storage_device="cpu",
log_memory_usage=app_config.log_memory_usage,
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -578,6 +578,7 @@ def __init__(
keep_alive_minutes: float = 0,
shared_cpu_weights: SharedCpuWeightsStore | None = SHARED_CPU_WEIGHTS,
ram_budget: RamBudget | None = None,
reserve_vram_gb: float = 0.0,
):
"""Initialize the model RAM cache.

Expand All @@ -604,6 +605,7 @@ def __init__(
:param ram_budget: Optional shared RamBudget used as the single global RAM authority across all per-device
caches. When provided, eviction decisions are made against the deduplicated, system-wide RAM total rather
than this cache's local (double-counted) sum. When None, the cache uses its own local RAM accounting.
:param reserve_vram_gb: VRAM (in GB) to subtract from the available memory budget.
"""
self._shared_cpu_weights = shared_cpu_weights
self._ram_budget = ram_budget
Expand All @@ -618,6 +620,7 @@ def __init__(

self._max_ram_cache_size_gb = max_ram_cache_size_gb
self._max_vram_cache_size_gb = max_vram_cache_size_gb
self._reserve_vram_gb = reserve_vram_gb

self._logger = PrefixedLoggerAdapter(
logger or InvokeAILogger.get_logger(self.__class__.__name__), "MODEL CACHE"
Expand Down Expand Up @@ -2343,12 +2346,13 @@ def _get_vram_available(self, working_mem_bytes: Optional[int], honor_cap: bool
`_get_physical_vram_available`.
"""
working_mem_bytes = self._working_mem_reserve(working_mem_bytes)
reserve_vram_bytes = int(self._reserve_vram_gb * GB)

# An explicit cache cap limits model residency, but operation-specific working
# memory still must remain free for activations and temporary tensors.
if honor_cap and self._max_vram_cache_size_gb is not None:
vram_total_available_to_cache = int(self._max_vram_cache_size_gb * GB) - working_mem_bytes
return vram_total_available_to_cache - self._get_vram_in_use()
return vram_total_available_to_cache - self._get_vram_in_use() - reserve_vram_bytes

if self._execution_device.type == "cuda":
vram_allocated = torch.cuda.memory_allocated(self._execution_device)
Expand Down Expand Up @@ -2384,7 +2388,7 @@ def _get_vram_available(self, working_mem_bytes: Optional[int], honor_cap: bool

vram_total_available_to_cache = vram_available_to_process - working_mem_bytes
vram_cur_available_to_cache = vram_total_available_to_cache - self._get_vram_in_use()
return vram_cur_available_to_cache
return vram_cur_available_to_cache - reserve_vram_bytes

def _get_reclaimable_allocator_bytes(self) -> int:
"""Bytes the torch caching allocator holds for this device but is not using, excluding
Expand Down
9 changes: 8 additions & 1 deletion invokeai/frontend/api/openapi.json

Large diffs are not rendered by default.

7 changes: 7 additions & 0 deletions invokeai/frontend/api/schema.ts
Original file line number Diff line number Diff line change
Expand Up @@ -22583,6 +22583,7 @@ export type components = {
* profiles_dir: Path to profiles output directory.
* max_cache_ram_gb: The maximum amount of CPU RAM to use for model caching in GB. If unset, the limit will be configured based on the available RAM. In most cases, it is recommended to leave this unset.
* max_cache_vram_gb: The amount of VRAM to use for model caching in GB. If unset, the limit will be configured based on the available VRAM and the device_working_mem_gb. In most cases, it is recommended to leave this unset.
* reserve_vram_gb: The amount of VRAM (in GB) to subtract from the model cache's available memory budget. Defaults to 0.
* log_memory_usage: If True, a memory snapshot will be captured before and after every model cache operation, and the result will be logged (at debug level). There is a time cost to capturing the memory snapshots, so it is recommended to only enable this feature if you are actively inspecting the model cache's behaviour.
* model_cache_keep_alive_min: How long to keep models in cache after last use, in minutes. A value of 0 (the default) means models are kept in cache indefinitely. If no model generations occur within the timeout period, the model cache is cleared using the same logic as the 'Clear Model Cache' button.
* device_working_mem_gb: The amount of working memory to keep available on the compute device (in GB). Has no effect if running on CPU. If you are experiencing OOM errors, try increasing this value.
Expand Down Expand Up @@ -22921,6 +22922,12 @@ export type components = {
* @description The amount of VRAM to use for model caching in GB. If unset, the limit will be configured based on the available VRAM and the device_working_mem_gb. In most cases, it is recommended to leave this unset.
*/
max_cache_vram_gb?: number | null;
/**
* Reserve Vram Gb
* @description The amount of VRAM (in GB) to subtract from the model cache's available memory budget. Defaults to 0.
* @default 0
*/
reserve_vram_gb?: number;
/**
* Log Memory Usage
* @description If True, a memory snapshot will be captured before and after every model cache operation, and the result will be logged (at debug level). There is a time cost to capturing the memory snapshots, so it is recommended to only enable this feature if you are actively inspecting the model cache's behaviour.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -216,6 +216,33 @@ def test_the_windows_video_memory_budget_caps_the_measured_free_vram(monkeypatch
assert cache._get_vram_available(None) == 6 * GB


@pytest.mark.parametrize("cache_cap_gb", [None, 12.0])
@pytest.mark.parametrize("reserve_vram_gb", [0.0, 1.5])
def test_reserve_vram_reduces_available_memory(
monkeypatch: pytest.MonkeyPatch, cache_cap_gb: float | None, reserve_vram_gb: float
):
cache = ModelCache(
execution_device_working_mem_gb=1.0,
enable_partial_loading=True,
keep_ram_copy_of_weights=True,
max_vram_cache_size_gb=cache_cap_gb,
reserve_vram_gb=reserve_vram_gb,
execution_device="cpu",
storage_device="cpu",
logger=MagicMock(),
shared_cpu_weights=None,
)
cache._execution_device = torch.device("cuda") # Avoid requiring a physical GPU.
monkeypatch.setattr(TorchDevice, "cuda_mem_get_info", classmethod(lambda cls, device: (8 * GB, 16 * GB)))
monkeypatch.setattr(torch.cuda, "memory_allocated", lambda device: 2 * GB)
monkeypatch.setattr(cache, "_get_reclaimable_allocator_bytes", lambda: 0)

reserved_bytes = int(reserve_vram_gb * GB)
expected_cache = (9 if cache_cap_gb is not None else 7) * GB - reserved_bytes
assert cache._get_vram_available(None) == expected_cache
assert cache._get_physical_vram_available() == 7 * GB - reserved_bytes


@pytest.mark.parametrize("peer_busy", [False, True], ids=["alone", "peer-device-busy"])
def test_offloading_under_expandable_segments_stops_once_enough_is_free(monkeypatch: pytest.MonkeyPatch, peer_busy):
"""Under expandable segments the driver sees an offloaded model's pages only after empty_cache(), and no
Expand Down
Loading