Atomic Chat has integrated DFlash, a block-diffusion speculative decoding mode, natively into its llama.cpp backend, promising up to 2.2x faster local inference on Qwen models with output the developers describe as byte-for-byte identical to standard generation. The feature is available across the app's macOS, Windows, and Linux builds.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.