Problem
Now that #165 has landed, I would like to continue helping with DeepSeek V4 performance work without duplicating work already underway.
The current V4 performance branch appears to include verified drafting / MTP work. I also have other experimentally viable optimization directions, such as W4A8 execution paths, that may improve speed or memory use but can slightly change numerical results and may not remain token-exact at every near tie.
Before implementing another large change, I would like to confirm which direction the maintainers actually want.
Proposed solution
Could the maintainers clarify whether contributions are wanted in either of these areas?
- MTP / DSpark: Is there any remaining part of the MTP or DSpark implementation, validation, or optimization that you would like me to own, or is the current performance work intended to cover it fully?
- Opt-in approximate optimization: Would an experimental W4A8 or similar path be acceptable if the existing reference path remains unchanged and is still the default?
If an approximate path is welcome, I would keep it isolated behind an explicit opt-in flag and provide:
- the unchanged exact/reference fallback;
- wall-clock, tok/s, memory, and I/O comparisons;
- numerical-error and token-divergence measurements against the reference path;
- results across multiple representative prompts rather than a single favorable case;
- clear documentation that the mode trades a small amount of numerical fidelity for performance;
- portable fallback behavior on platforms without the optimized kernel.
I would first publish the experiment and measurements, then submit production integration only if the results justify it.
Alternatives considered
- Restrict future work to token-exact optimizations only. This is the safest policy, but excludes potentially useful speedups on constrained consumer hardware.
- Keep W4A8 as an external experiment until the accuracy/performance tradeoff is better understood.
- Wait for the current V4 performance work to land, then contribute only any missing validation or platform-specific kernels.
Scope and compatibility
The exact DeepSeek V4 target path would remain available and unchanged. Approximate kernels would be disabled by default and clearly identified in logs and documentation. Target-only / --no-dspark operation would continue to work independently.
Please also let me know the preferred base branch and validation bar for this kind of work?for example, whether numerical error and token-divergence measurements are sufficient, or whether you would also want perplexity or downstream task evaluation before considering a PR.
Problem
Now that #165 has landed, I would like to continue helping with DeepSeek V4 performance work without duplicating work already underway.
The current V4 performance branch appears to include verified drafting / MTP work. I also have other experimentally viable optimization directions, such as W4A8 execution paths, that may improve speed or memory use but can slightly change numerical results and may not remain token-exact at every near tie.
Before implementing another large change, I would like to confirm which direction the maintainers actually want.
Proposed solution
Could the maintainers clarify whether contributions are wanted in either of these areas?
If an approximate path is welcome, I would keep it isolated behind an explicit opt-in flag and provide:
I would first publish the experiment and measurements, then submit production integration only if the results justify it.
Alternatives considered
Scope and compatibility
The exact DeepSeek V4 target path would remain available and unchanged. Approximate kernels would be disabled by default and clearly identified in logs and documentation. Target-only /
--no-dsparkoperation would continue to work independently.Please also let me know the preferred base branch and validation bar for this kind of work?for example, whether numerical error and token-divergence measurements are sufficient, or whether you would also want perplexity or downstream task evaluation before considering a PR.