Geiping writes about a language model architecture that scales test-time compute by iterating a recurrent block to reason implicitly in latent space, unrolling to arbitrary depth at inference time rather than generating more tokens. Unlike chain-of-thought methods, this approach needs no specialized training data, operates with small context windows, and can represent reasoning that is difficult to express linguistically. The authors scale a proof-of-concept to 3.5 billion parameters trained on 800 billion tokens, showing dramatic benchmark improvements at computation loads equivalent to 50 billion parameters.
- Hugging Face "huginn-0125"
- GitHub com/seal-rg/huginn-0125