Articles
DeepSeek Releases DSpark: A Speculative Decoding Framework That Accelerates DeepSeek-V4 Per-User Generation by 60–85% Over MTP-1⭐7
DeepSeek releases DSpark, a speculative decoding framework boosting DeepSeek-V4 per-user generation speed by 60–85% over MTP-1, with open-source availability.
DFlash Speculative Decoding Drafts Whole Token Blocks in Parallel for Up to 15x Higher Throughput on NVIDIA Blackwell⭐8
DFlash speculative decoding drafts entire token blocks in parallel, achieving up to 15x higher throughput on NVIDIA Blackwell GPUs.
