Inco AI released DFlash 2, a speculative-decoding drafter that generates token blocks in parallel and is designed to raise large-language-model inference throughput without changing verified output. The company released drafters for Qwen3.8-27B and Meta’s Muse Glimmer on Hugging Face. Inco AI says DFlash 2 delivers 16% to 25% more accepted output tokens per verification pass than DFlash, with about 1.3% added draft-verify cycle latency for its combined changes. It reports 2.7× to 3.4× autoregressive-decoding throughput for Qwen3.8-27B and 3.1× to 4.6× for Muse Glimmer. DFlash 2 adds a 2 million-parameter path selector that retains 16 candidate tokens at each draft position, scores adjacent candidate pairs in parallel, and selects a more coherent sequence. It also adds 16.5 million parameters, or 3%, through dynamic two-tap depthwise convolutions that model short-range within-block dependencies and reduce declining accuracy at later token positions. Across GSM8K, MATH-500, HumanEval, MBPP, and MT-Bench, Inco AI reports a mean acceptance length of 5.97 tokens for DFlash 2, versus 4.92 for DFlash and 5.49 for DSpark under matched training setups. DFlash, released in January, is supported by SGLang, vLLM, TensorRT-LLM, and llama.cpp; Inco AI says its models had exceeded 3.5 million Hugging Face downloads as of August 2026.
inco.ai
14 min
6h ago
Inco AI released DFlash 2, a speculative-decoding drafter that generates token blocks in parallel and is designed to raise large-language-model inference throughput without changing verified output. The company released drafters for Qwen3.8-27B and Meta’s Muse Glimmer on Hugging Face. Inco AI says DFlash 2 delivers 16% to 25% more accepted output tokens per verification pass than DFlash, with about 1.3% added draft-verify cycle latency for its combined changes. It reports 2.7× to 3.4× autoregressive-decoding throughput for Qwen3.8-27B and 3.1× to 4.6× for Muse Glimmer. DFlash 2 adds a 2 million-parameter path selector that retains 16 candidate tokens at each draft position, scores adjacent candidate pairs in parallel, and selects a more coherent sequence. It also adds 16.5 million parameters, or 3%, through dynamic two-tap depthwise convolutions that model short-range within-block dependencies and reduce declining accuracy at later token positions. Across GSM8K, MATH-500, HumanEval, MBPP, and MT-Bench, Inco AI reports a mean acceptance length of 5.97 tokens for DFlash 2, versus 4.92 for DFlash and 5.49 for DSpark under matched training setups. DFlash, released in January, is supported by SGLang, vLLM, TensorRT-LLM, and llama.cpp; Inco AI says its models had exceeded 3.5 million Hugging Face downloads as of August 2026.
inco.ai
14 min
6h ago
Inco AI released DFlash 2, a speculative-decoding drafter that generates token blocks in parallel and is designed to raise large-language-model inference throughput without changing verified output. The company released drafters for Qwen3.8-27B and Meta’s Muse Glimmer on Hugging Face. Inco AI says DFlash 2 delivers 16% to 25% more accepted output tokens per verification pass than DFlash, with about 1.3% added draft-verify cycle latency for its combined changes. It reports 2.7× to 3.4× autoregressive-decoding throughput for Qwen3.8-27B and 3.1× to 4.6× for Muse Glimmer. DFlash 2 adds a 2 million-parameter path selector that retains 16 candidate tokens at each draft position, scores adjacent candidate pairs in parallel, and selects a more coherent sequence. It also adds 16.5 million parameters, or 3%, through dynamic two-tap depthwise convolutions that model short-range within-block dependencies and reduce declining accuracy at later token positions. Across GSM8K, MATH-500, HumanEval, MBPP, and MT-Bench, Inco AI reports a mean acceptance length of 5.97 tokens for DFlash 2, versus 4.92 for DFlash and 5.49 for DSpark under matched training setups. DFlash, released in January, is supported by SGLang, vLLM, TensorRT-LLM, and llama.cpp; Inco AI says its models had exceeded 3.5 million Hugging Face downloads as of August 2026.
inco.ai
14 min
6h ago
No more articles to load