Skip to content

Fix skinned MV stalls from reading mapped upload memory - #15

Merged
Bazouz660 merged 1 commit into
Bazouz660:mainfrom
tkosub:fix/skinned-mv-map-proxy
Sep 5, 2026
Merged

Bazouz660 merged 1 commit into
Bazouz660:mainfrom
tkosub:fix/skinned-mv-map-proxy

Conversation

@tkosub

@tkosub tkosub commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Avoid reading from mapped WRITE_DISCARD bone constant buffers.
  • Preserve Dust's exact previous-bone character motion vectors.
  • Replace the unused full geometry-replay snapshot with the few bindings the injected skinned-MV path actually needs.
  • Preserve the default CaptureGeometry=1 behavior.

Root cause

Diagnostic builds isolated the character-dependent performance collapse to the shadow copy performed between Map and Unmap.

The previous implementation saved the pointer returned by:

Map(..., D3D11_MAP_WRITE_DISCARD, ...)

At Unmap, it copied from that mapped pointer into a CPU shadow buffer. On the affected system, reading this upload-oriented memory was extremely slow. Such mapped memory is commonly write-combined: CPU writes are efficient, but CPU reads can be pathological.

Classification, the required D3D state queries, pose matching, constant-buffer upload, and the skinned shader path retained normal performance when tested separately. Enabling the old mapped-memory read alone reproduced the collapse.

Fix

For known skinned bone buffers, the new path:

  1. Saves the real mapped upload pointer.
  2. Gives the game an ordinary cached-RAM shadow buffer as pData.
  3. At Unmap, copies the cached RAM to the real mapped pointer.
  4. Keeps the same cached bytes for previous-pose matching.

The game still uploads exactly the data it wrote, but Dust never reads from the WRITE_DISCARD mapping.

The injected path also identifies skinned draws using only the current vertex shader, bone constant buffer, and index buffer instead of constructing the legacy full CapturedDraw replay state.

Test results

Tested in the same character-heavy scene with approximately 15 visible characters and FSR3/4 temporal AA enabled, using a 9060 XT at 2560x1440:

  • Previous Map/Unmap shadow-copy path: approximately 16–18 FPS
  • Fixed proxy path: approximately 47 FPS
  • A comparison with CaptureGeometry=0 in INI shows near identical performance
  • Final surgical build with CaptureGeometry absent from the INI, therefore using its default value of 1: approximately 47 FPS
  • Upscaler-off reference: approximately 57 FPS

Rendering and exact character motion vectors remained correct.

Avoid the legacy full per-draw geometry snapshot in the injected motion-vector path and classify skinned draws from only the current shader, constant buffer, and index buffer.

Proxy known WRITE_DISCARD bone buffers through cached CPU memory, then copy the cached bytes into the real mapped upload pointer at Unmap. This preserves exact previous-bone motion vectors without reading from slow write-combined mapped memory.
@Bazouz660
Bazouz660 merged commit 5008d03 into Bazouz660:main Sep 5, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants