This page provides GGUF quantized versions of DiffusionGemma 26B A4B-it, a multimodal model from Google DeepMind based on the Gemma 4 architecture. The model employs discrete text diffusion through block-autoregressive multi-canvas sampling to achieve significantly faster decoding speeds than standard autoregressive models. It is capable of processing interleaved inputs consisting of text, images with variable resolutions and aspect ratios, and video content for generating textual outputs.
Key topics:
- Mixture-of-Experts architecture with 3.8 billion active parameters.
- High-speed generation through parallel denoising of token blocks.
- Multimodal input support including image and video understanding.
- Extensive context window capability up to 256K tokens.
- Integrated reasoning modes for step-by-step thought processes.