The user bartowski provides GGUF quantizations of the MiMo-V2.6-Distill-Qwen-9B model, which is a 9 billion parameter multimodal model based on Qwen3.5 designed for image-text tasks. These files are optimized for use with llama.cpp and various local applications like LM Studio, Ollama, and Jan AI using imatrix quantization techniques to preserve quality at lower bitrates.
- Supports both text and image inputs via a separate mmproj file
- Uses per-tensor layout computations to optimize precision for sensitive weights
- Includes calibration datasets that combine prose with tool-calling and reasoning conversations
- Compatible with various hardware architectures, including ARM and AVX through online repacking