klotz: 24gb* + llama.cpp*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. Pedro Cuenca writes Meta released Muse Glimmer-30B, a local, open-source multimodal model distilled from its larger Muse architecture. Designed for agentic workflows, it combines a 28B text decoder with a 2B vision encoder, supporting image, video, and multimodal tool calling out of the box. The release includes immediate compatibility with major inference frameworks like transformers, llama.cpp, and vLLM, alongside built-in speculative decoding for faster generation.

    - Features a hybrid attention pattern alternating between three sliding window layers and one full attention layer.
    - Incorporates a DFlash block-diffusion drafter to accelerate structured text generation like coding.
    - Supports fine-tuning via TRL with practical minimums ranging from one to eight H100 GPUs depending on the method.
    - Demonstrates autonomous agent capabilities such as self-quantization, self-deployment, and hardware-specific optimization.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: 24gb + llama.cpp

About - Propulsed by SemanticScuttle