TokenWatt is a transparent, OpenAI-compatible proxy designed to measure the actual electricity cost of running local Large Language Model (LLM) inference on Apple Silicon hardware. By sitting in front of local inference servers and utilizing Apple's IOReport via SoC rail energy measurements, it provides real-time pricing for requests based on user-defined utility rates without requiring sudo privileges. The tool allows users to compare the cost-efficiency of local execution versus cloud API providers, particularly highlighting the economic advantages of running high-context agentic loops locally where context re-processing is essentially free (limited only by power).
- Uses Apple's IOReport for sudoless energy measurement on macOS/Apple Silicon.
- Provides an OpenAI-compatible interface that forwards requests byte-for-byte to backends like LM Studio or MLX.
- Supports dynamic model discovery so routing updates automatically when models are loaded into memory.
- Offers a calibration feature to replace estimates (±15–30%) with highly accurate measurements via smart plugs.
>"I Measured Every Watt on Apple Silicon Five models, sustained generation, real wall-socket energy at $0.31/kWh — and the surprise the RTX-3090 numbers predicted, only bigger."
Justin Stewart writes about how the energy cost of running local Large Language Models (LLMs) on Apple Silicon depends more on throughput than parameter count. Using an M3 Ultra Mac Studio, he demonstrates that large Mixture-of-Experts (MoE) models can be significantly cheaper to operate per token than smaller dense models because they only activate a fraction of their parameters during generation. Ultimately, the study reveals that efficiency is driven by how much data must be moved from memory for every token produced.
* The measurements were calibrated against actual wall power using a Shelly Plug US Gen4 meter.
* A custom tool called TokenWatt was used to measure marginal energy consumption via Apple’s IOReport interface.
* In real-world "lumpy" traffic scenarios, the cost of dense models compared to MoE models actually widens even further.