Yohei Nakajima writes about glance, a tool designed to allow users to ask an open vision-language model (VLM) typed questions about images and receive probability data directly on their own machine. Rather than generating new text or training models, it acts as a measurement harness that reads logits from frozen models—such as Qwen3-VL-4B by default—to provide yes/no answers, single-choice selections, and qualitative ratings without any image data leaving the user's device.
- The tool provides three response types: "noul" (yes/no), choice (pick one from a list), and score (a rating on a specified scale).
- It includes an experimental MLX backend to provide faster runtimes specifically for Apple Silicon users.
- Glance can perform self-calibration using unlabeled data or precise calibration through labeled datasets to improve rating accuracy.
- The software is designed with privacy in mind, ensuring all inference and logging stay local on the user's hardware.