This edition of The Weekly Kaitchup reviews several NVFP4 quantization versions of the Qwen3.6 27B model, comparing NVIDIA's mixed-precision approach with community alternatives like Unsloth and PrismaQuant. It also details DSpark, DeepSeek's new speculative decoding method that uses a parallel draft backbone and a confidence head to significantly accelerate large language model generation speeds.
* Comparison of Qwen3.6 27B NVFP4 quantization variants
* Guidance on selecting models based on accuracy versus memory footprint
* Technical overview of DSpark architecture and suffix decay mitigation
* Performance improvements and vLLM support for DSpark