klotz: power*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. Arsen Apostolov writes about the actual electrical cost of running local Large Language Models on a single NVIDIA RTX 3090 compared to hosted cloud APIs.

    >"I measured the actual GPU electricity for eight local models on one RTX 3090 — and the cheapest wasn't the smallest, nor the priciest the biggest"

    Cost of Generating 1 Million Tokens Locally

    | MODEL | PARAMS (Billions) | MEAN SPEED (tok/s) | AVG GPU DRAW (W) | € / 1M OUTPUT TOKENS |
    | :--- | :---: | :---: | :---: | :---: |
    | **gemma3:1b** | 1B | 136 tok/s | 154 W | €0.060 |
    | **Qwen3-Coder** | 30.5B | 130 tok/s | 233 W | €0.112 |
    | **gemma4:26b** | 26B | 85 tok/s | 246 W | €0.139 |
    | **Devstral** | 24B | 49 tok/s | 320 W | €0.321 |
    | **gemma3:27b** | 27B | 36 tok/s | 283 W | €0.361 |
    | **Seed-OSS** | 36B | 4.5 tok/s | 186 W | €0.946 |
    | **GLM-4.5-Air** | 106B | 5.7 tok/s | 141 W | €1.040 |
    | **DeepSeek-R1-Distill** | 32.8B | 6.9 tok/s | 155 W | €1.526 |

    By measuring real-time GPU power consumption through a custom dashboard, he discovered that token costs are driven by effective wall-clock throughput rather than model parameter size or raw generation speed alone. The results show that while small and fast models can be more economical than cloud services, reasoning-heavy models may actually become the most expensive to run locally due to the time spent "deliberating" between tokens.

    * Measurements were performed using HomeLab Monitor, an open-source dashboard that integrates live power data from `nvidia-smi`.
    * DeepSeek-R1-Distill emerged as the most expensive model per million tokens because its effective throughput is slowed by reasoning delays.
    * The findings focus on marginal electricity costs and exclude total cost of ownership factors like hardware amortization or idle draw.
  2. >"I Measured Every Watt on Apple Silicon Five models, sustained generation, real wall-socket energy at $0.31/kWh — and the surprise the RTX-3090 numbers predicted, only bigger."

    Justin Stewart writes about how the energy cost of running local Large Language Models (LLMs) on Apple Silicon depends more on throughput than parameter count. Using an M3 Ultra Mac Studio, he demonstrates that large Mixture-of-Experts (MoE) models can be significantly cheaper to operate per token than smaller dense models because they only activate a fraction of their parameters during generation. Ultimately, the study reveals that efficiency is driven by how much data must be moved from memory for every token produced.

    * The measurements were calibrated against actual wall power using a Shelly Plug US Gen4 meter.
    * A custom tool called TokenWatt was used to measure marginal energy consumption via Apple’s IOReport interface.
    * In real-world "lumpy" traffic scenarios, the cost of dense models compared to MoE models actually widens even further.
  3. This Decoder podcast episode features host Jon Fortt interviewing author Gil Duran about the growing influence of ultra-wealthy tech billionaires and their potentially anti-democratic ideologies. Duran latest book, *The Nerd Reich*, argues that tech billionaires are increasingly acting like dictators, and illustrates the dangers that poses to democracy.

    Duran argues a group of billionaires (Thiel, Musk, Andreessen, Altman, etc.) are pushing a "Dark Enlightenment" or "neo-reactionary" philosophy – essentially advocating for a shift away from democracy towards a system resembling tech-led feudalism. This ideology, inspired by figures like Curtis Yarvin, envisions a future where tech corporations wield significant power, potentially even governing territories with limited individual freedoms.

    Duran connects this movement to the rise of Donald Trump and the current political climate, suggesting a dangerous alliance between the far-right and tech elites. He believes these billionaires are not simply ideologues, but actively seeking to reshape government to suit their vision.

    The conversation also touches on the historical parallels to Ayn Rand's philosophies, the potential for emergency powers to be abused, and the role of Silicon Valley’s culture of self-importance. Duran criticizes the lack of critical coverage in mainstream media and calls for a broader public discussion about these issues.

    Ultimately, Duran expresses concern about the concentration of power and wealth, and argues that a robust political movement is needed to counter this trend and preserve democratic values. He believes a future where technology serves the majority, rather than a select few, is still possible, but requires active engagement and a rejection of this emerging "corporate dictatorship."
  4. Four decades after his death, the French philosopher's ideas continue to resonate in contemporary discussions around social media, power, and identity.

    The philosophy of Michel Foucault, who died in 1984, is still relevant in today's world saturated with social media. His work emphasizes the importance of understanding the self and the power dynamics at play in knowledge acquisition. In the current digital age, social media platforms have become central in shaping the self, as individuals compete to be knowledgeable and powerful through their online personas. A key aspect of Foucault's work is the need to recognize the power dynamics at play, as well as the importance of self-awareness and resistance to prefabricated identities. His viral philosophy has made him an influential figure in the contemporary world, even if his work sometimes appears tedious or infuriating.
  5. 2024-01-30 Tags: , , by klotz
  6. 2022-01-29 Tags: , , , , , by klotz
  7. 2019-06-05 Tags: , , , , by klotz
  8. 2019-06-05 Tags: , , , , , by klotz
  9. 2019-06-05 Tags: , , , , , by klotz

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: power

About - Propulsed by SemanticScuttle