Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#llms#ai-ethics#claude#code-generation#ai-safety#openai#anthropic#discussion

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

© 2026 Themata.AI • All Rights Reserved

Archive

|

Topics

|

Privacy

|

Cookies

|

Contact
🕒 Latest🔥 Top

Filtering by tag:

dflashClear
DFlash 2: Keep Drafting Parallel
dflashai-agentsinference-optimizationdeveloper-tools
Tool

DFlash 2: Keep Drafting Parallel

Inco AI released DFlash 2, a speculative-decoding drafter that generates token blocks in parallel and is designed to raise large-language-model inference throughput without changing verified output. The company released drafters for Qwen3.8-27B and Meta’s Muse Glimmer on Hugging Face. Inco AI says DFlash 2 delivers 16% to 25% more accepted output tokens per verification pass than DFlash, with about 1.3% added draft-verify cycle latency for its combined changes. It reports 2.7× to 3.4× autoregressive-decoding throughput for Qwen3.8-27B and 3.1× to 4.6× for Muse Glimmer. DFlash 2 adds a 2 million-parameter path selector that retains 16 candidate tokens at each draft position, scores adjacent candidate pairs in parallel, and selects a more coherent sequence. It also adds 16.5 million parameters, or 3%, through dynamic two-tap depthwise convolutions that model short-range within-block dependencies and reduce declining accuracy at later token positions. Across GSM8K, MATH-500, HumanEval, MBPP, and MT-Bench, Inco AI reports a mean acceptance length of 5.97 tokens for DFlash 2, versus 4.92 for DFlash and 5.49 for DSpark under matched training setups. DFlash, released in January, is supported by SGLang, vLLM, TensorRT-LLM, and llama.cpp; Inco AI says its models had exceeded 3.5 million Hugging Face downloads as of August 2026.

inco.ai

🔥🔥🔥🔥🔥

14 min

5h ago

DFlash 2: Keep Drafting Parallel

Inco AI released DFlash 2, a speculative-decoding drafter that generates token blocks in parallel and is designed to raise large-language-model inference throughput without changing verified output. The company released drafters for Qwen3.8-27B and Meta’s Muse Glimmer on Hugging Face. Inco AI says DFlash 2 delivers 16% to 25% more accepted output tokens per verification pass than DFlash, with about 1.3% added draft-verify cycle latency for its combined changes. It reports 2.7× to 3.4× autoregressive-decoding throughput for Qwen3.8-27B and 3.1× to 4.6× for Muse Glimmer. DFlash 2 adds a 2 million-parameter path selector that retains 16 candidate tokens at each draft position, scores adjacent candidate pairs in parallel, and selects a more coherent sequence. It also adds 16.5 million parameters, or 3%, through dynamic two-tap depthwise convolutions that model short-range within-block dependencies and reduce declining accuracy at later token positions. Across GSM8K, MATH-500, HumanEval, MBPP, and MT-Bench, Inco AI reports a mean acceptance length of 5.97 tokens for DFlash 2, versus 4.92 for DFlash and 5.49 for DSpark under matched training setups. DFlash, released in January, is supported by SGLang, vLLM, TensorRT-LLM, and llama.cpp; Inco AI says its models had exceeded 3.5 million Hugging Face downloads as of August 2026.

inco.ai

🔥🔥🔥🔥🔥

14 min

5h ago

DFlash 2: Keep Drafting Parallel

Inco AI released DFlash 2, a speculative-decoding drafter that generates token blocks in parallel and is designed to raise large-language-model inference throughput without changing verified output. The company released drafters for Qwen3.8-27B and Meta’s Muse Glimmer on Hugging Face. Inco AI says DFlash 2 delivers 16% to 25% more accepted output tokens per verification pass than DFlash, with about 1.3% added draft-verify cycle latency for its combined changes. It reports 2.7× to 3.4× autoregressive-decoding throughput for Qwen3.8-27B and 3.1× to 4.6× for Muse Glimmer. DFlash 2 adds a 2 million-parameter path selector that retains 16 candidate tokens at each draft position, scores adjacent candidate pairs in parallel, and selects a more coherent sequence. It also adds 16.5 million parameters, or 3%, through dynamic two-tap depthwise convolutions that model short-range within-block dependencies and reduce declining accuracy at later token positions. Across GSM8K, MATH-500, HumanEval, MBPP, and MT-Bench, Inco AI reports a mean acceptance length of 5.97 tokens for DFlash 2, versus 4.92 for DFlash and 5.49 for DSpark under matched training setups. DFlash, released in January, is supported by SGLang, vLLM, TensorRT-LLM, and llama.cpp; Inco AI says its models had exceeded 3.5 million Hugging Face downloads as of August 2026.

inco.ai

🔥🔥🔥🔥🔥

14 min

5h ago

No more articles to load