klotz: llm benchmarks*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. A collection of 20 copy-paste prompts designed to test local coding LLMs by asking them to generate complete, playable browser games in a single HTML file. Each challenge enforces the same constraints: no external assets, libraries, or network requests, offline playable, and a priority on polish and game feel over explanations. The prompts span genres from arcade racing and survival horror to stealth and colony management, and the document suggests specific prompts as stronger benchmark candidates for evaluating complex coding ability.
    - The prompts intentionally stress-test gameplay state, input handling, rendering, collision, enemy AI, progression systems, UI, persistence, and visual polish all in one shot.
    - Prompts like OUTBREAK, DUNGEON ZERO, and CYBER SURVIVOR are flagged as particularly good for benchmarking because they demand more than basic rendering and movement.
    - All graphics must be drawn procedurally using Canvas and CSS; Web Audio is used for synthesized sound effects.
    - The single-file constraint makes it easy to compare models because there is no build system, dependency installation, or asset pipeline hiding behind the output.
  2. This research provides a forensic analysis and benchmark comparison of three abliteration techniques—Heretic, HauhauCS Aggressive, and Huihui—applied to the Qwen 3 and 3.5 model families. The study assesses these methods across various parameter sizes ranging from 2B to 27B using safety evaluations, capability benchmarks including MMLU and GSM8K, KL divergence, and tensor weight analysis.

    - Large models suffer more significant collateral damage during abliteration than smaller ones.
    - The Heretic technique shows the most consistency in preserving model capabilities across different scales.
    - Huihui performs inconsistently, showing catastrophic failures on certain hybrid Mamba2 and Transformer architectures at specific sizes.
    - Claims of "lossless" modifications are largely contradicted by performance drops in larger models.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: llm benchmarks

About - Propulsed by SemanticScuttle