klotz: mixture-of-experts*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. Mark Marosi writes about decider, a family of models fine-tuned from Qwen3.5 that produce typed decisions (choice, score, boolean) in a single forward pass without text generation, returning calibrated probability distributions over user-defined options. The project is an open reproduction of TypeSafe AI's "System One" model class (Jev), released in sizes from 0.8B to 35B mixture-of-experts with 3B active parameters.
    - The schema cache stores K/V states for repeated question prefixes, achieving up to 19x speedup on large option sets by running only the state per request
    - v10 adds calibration-aware RL on live MiniWoB++ browser tasks and exact games, lifting browser accuracy from 83% to 93% and halving the belief gap
    - TypeSafe's SDKs work unchanged by pointing TYPESAFE_BASE_URL at the decider server
    - The 35B model outperforms the 2B on 93 of 95 regression tasks but costs 3-4x per decision and lacks the RL stage
    - Co-developed with Claude (Anthropic) as a listed co-author on commits
  2. Carl Franzen writes that DeepSeek has launched V4.1-Flash, a model featuring a 552-billion-parameter mixture-of-experts backbone designed to drastically reduce costs for long-context workflows through specialized caching and architecture. The model offers extremely low off-peak rates of $0.003 per million cached input tokens, making it highly competitive against frontier models like GPT-5.6 Sol and Claude Opus 5 when used in repetitive agentic loops. While its total parameter count has increased significantly compared to previous versions, its Causal Encoder-Decoder architecture aims to minimize compute requirements during the prefill stage of inference.

    - V4.1-Flash utilizes a Causal Encoder-Decoder architecture that activates 8 billion parameters during prefill and 16 billion during decoding.
    - The model features an extremely high context window of up to 1 million tokens.
    - DeepSeek's technical report highlights the use of FP4 KV caching, which reduces global KV cache size by approximately one-quarter compared to its predecessor.
    - Off-peak hours for lower pricing are scheduled from Monday through Friday, specifically between 01:00–04:00 UTC and 06:00–10:00 UTC.
  3. Asif Razzaq writes that Alibaba's Qwen team has released Qwen3.8-Flash-Next, an open-weight multimodal Mixture-of-Experts model pairing a 125B backbone with a 51B N-gram embedding table and a 4B multi-token prediction module, totaling 180B on disk but activating only 6B parameters per token. The architecture blends Gated DeltaNet linear-attention layers (three of every four) with Qwen Sparse Attention at micro-block granularity, and is positioned as the architectural preview of Qwen4 in the same role Qwen3-Next played for Qwen3.5. Training cost is reported at roughly one-ninth that of Qwen3.7-Plus.

    - Licensed under qwen-community-1.0, not Apache-2.0; verify terms before commercial use.
    - FP8 checkpoint is 172.78 GiB; minimum validated config is TP2 on GB300, so self-hosting requires a multi-GPU node.
    - The 20M-entry bigram/trigram table at layer 2 can be offloaded to host memory with asynchronous prefetch (NVIDIA only).
    - Claude Opus 4.6 (Max) still leads on HLE (40.0 vs. Qwen's 35.9); DeepSeek-V4-Flash-0731 leads NL2Repo-Bench (54.2 vs. 48.1).
    - Native context is 262,144 tokens, extensible to 1,000,000 via YaRN.
  4. >"I Measured Every Watt on Apple Silicon Five models, sustained generation, real wall-socket energy at $0.31/kWh — and the surprise the RTX-3090 numbers predicted, only bigger."

    Justin Stewart writes about how the energy cost of running local Large Language Models (LLMs) on Apple Silicon depends more on throughput than parameter count. Using an M3 Ultra Mac Studio, he demonstrates that large Mixture-of-Experts (MoE) models can be significantly cheaper to operate per token than smaller dense models because they only activate a fraction of their parameters during generation. Ultimately, the study reveals that efficiency is driven by how much data must be moved from memory for every token produced.

    * The measurements were calibrated against actual wall power using a Shelly Plug US Gen4 meter.
    * A custom tool called TokenWatt was used to measure marginal energy consumption via Apple’s IOReport interface.
    * In real-world "lumpy" traffic scenarios, the cost of dense models compared to MoE models actually widens even further.
  5. Thinking Machines has released Inkling, an open-weights Mixture-of-Experts transformer model featuring 975B total parameters and a context window of up to 1M tokens. The model was trained on 45 trillion tokens across text, images, audio, and video to enable native multimodal reasoning. It is designed with controllable thinking effort to optimize the balance between performance and cost/latency, alongside strong capabilities for agentic coding and tool use.
    - Native multimodality in vision, audio, and text
    - Controllable computational effort settings
    - High proficiency in agentic workflows and design tasks
    - Available on Tinker for custom fine-tuning
  6. OpenMythos is an open-source PyTorch project by Kye Gomez that proposes a theoretical reconstruction of Anthropic's Claude Mythos architecture. Instead of standard transformer layers, it suggests a Recurrent-Depth Transformer (RDT) design where weights loop through multiple iterations to increase reasoning depth during inference. By combining Mixture-of-Experts with Multi-Latent Attention and stability constraints, the model achieves performance parity between 770M parameters and a 1.3B parameter standard transformer.

    * open-source PyTorch reconstruction of claude mythos
    * proposes recurrent-depth transformer architecture
    * reasoning depth scales via inference-time loops rather than parameter count
    * uses mixture-of-experts for domain breadth
    * implements multi-latent attention to reduce memory usage
    * employs lti injection and adaptive computation time for stability
    * achieves 1.3b parameter performance with only 770m parameters
  7. OpenAI's release of GPT-OSS marks their first major open source LLM since GPT-2, featuring improvements in reasoning, tool usage, and problem-solving capabilities. The article explores its architecture, message formatting, reasoning modes, and tokenizer details.
  8. This article discusses Time-MOE, an open-source time-series foundation model using Mixture-of-Experts (MOE) to improve forecasting accuracy while reducing computational costs. Key contributions include the Time-300B dataset, scaling laws for time series, and the Time-MOE architecture.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: mixture-of-experts

About - Propulsed by SemanticScuttle