#speculative decoding
Speculative Decoding: 9 AI articles covering speculative decoding news, analysis, and research
Articles
X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative DecodingNEWβ8
X-CoSD enables lossless, communication-efficient collaborative speculative decoding across heterogeneous SLMβLLM vocabularies, improving token generation speed ...
TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decodingβ10
TreeGraft enables adaptive multi-drafter speculative decoding, boosting LLM inference speed by 15.1% across benchmarks.
Speculative Decoding on CPUs: Nearly 4x Faster Token Generation with DFlashβ10
Speculative decoding with DFlash accelerates CPU token generation up to 3.92x on Qwen3.5-9B, cutting costs by 74%. Learn how vLLM makes it possible.
Liquid AI Unveils LFM2.5-DSpark Draft Models: Up to 3.18x Faster Decoding with Identical Outputsβ7
Liquid AI's DSpark draft models speed up LFM2.5 inference by up to 3.18x on H100, with identical outputs. Self-hosted deployment, open license for small firms.
Meta AI Unveils Muse Glimmer: A 30B Open-Weight Agentic Modelβ7
Metaβs Muse Glimmer is a 30B open-weight agentic model running on a single consumer GPU via 4-bit compression and speculative decoding.
DeepSeek Unveils DeepSeek-V4-Flash-0731: Major Agentic andβ8
DeepSeek launches V4-Flash-0731 with DSpark speculative decoding and Responses API, boosting agentic workflows and coding speed.
Tencent Open-Sources AngelSpec: A Unified Training Framework forβ8
Tencent open-sources AngelSpec, a unified training framework for speculative decoding draft models, supporting MTP and block-parallel DFlash architectures
DeepSeek Releases DSpark: A Speculative Decoding Framework Thatβ7
DeepSeek releases DSpark, a speculative decoding framework boosting DeepSeek-V4 per-user generation speed by 60β85% over MTP-1, with open-source availability.
DFlash Speculative Decoding Drafts Whole Token Blocks inβ8
DFlash speculative decoding drafts entire token blocks in parallel, achieving up to 15x higher throughput on NVIDIA Blackwell GPUs.
