Skip to content

[Feature]: Contribution scope for DeepSeek V4 MTP and opt-in approximate kernels #840

Description

@DrewZt

Problem

Now that #165 has landed, I would like to continue helping with DeepSeek V4 performance work without duplicating work already underway.

The current V4 performance branch appears to include verified drafting / MTP work. I also have other experimentally viable optimization directions, such as W4A8 execution paths, that may improve speed or memory use but can slightly change numerical results and may not remain token-exact at every near tie.

Before implementing another large change, I would like to confirm which direction the maintainers actually want.

Proposed solution

Could the maintainers clarify whether contributions are wanted in either of these areas?

  1. MTP / DSpark: Is there any remaining part of the MTP or DSpark implementation, validation, or optimization that you would like me to own, or is the current performance work intended to cover it fully?
  2. Opt-in approximate optimization: Would an experimental W4A8 or similar path be acceptable if the existing reference path remains unchanged and is still the default?

If an approximate path is welcome, I would keep it isolated behind an explicit opt-in flag and provide:

  • the unchanged exact/reference fallback;
  • wall-clock, tok/s, memory, and I/O comparisons;
  • numerical-error and token-divergence measurements against the reference path;
  • results across multiple representative prompts rather than a single favorable case;
  • clear documentation that the mode trades a small amount of numerical fidelity for performance;
  • portable fallback behavior on platforms without the optimized kernel.

I would first publish the experiment and measurements, then submit production integration only if the results justify it.

Alternatives considered

  • Restrict future work to token-exact optimizations only. This is the safest policy, but excludes potentially useful speedups on constrained consumer hardware.
  • Keep W4A8 as an external experiment until the accuracy/performance tradeoff is better understood.
  • Wait for the current V4 performance work to land, then contribute only any missing validation or platform-specific kernels.

Scope and compatibility

The exact DeepSeek V4 target path would remain available and unchanged. Approximate kernels would be disabled by default and clearly identified in logs and documentation. Target-only / --no-dspark operation would continue to work independently.

Please also let me know the preferred base branch and validation bar for this kind of work?for example, whether numerical error and token-divergence measurements are sufficient, or whether you would also want perplexity or downstream task evaluation before considering a PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    discussionProposta / discussione aperta, non un taskmodel-supportSupporto a nuovi modelliperformanceVelocità / tok-s / ottimizzazioni

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions