klotz: model safety*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. This research provides a forensic analysis and benchmark comparison of three abliteration techniques—Heretic, HauhauCS Aggressive, and Huihui—applied to the Qwen 3 and 3.5 model families. The study assesses these methods across various parameter sizes ranging from 2B to 27B using safety evaluations, capability benchmarks including MMLU and GSM8K, KL divergence, and tensor weight analysis.

    - Large models suffer more significant collateral damage during abliteration than smaller ones.
    - The Heretic technique shows the most consistency in preserving model capabilities across different scales.
    - Huihui performs inconsistently, showing catastrophic failures on certain hybrid Mamba2 and Transformer architectures at specific sizes.
    - Claims of "lossless" modifications are largely contradicted by performance drops in larger models.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: model safety

About - Propulsed by SemanticScuttle