Tags: z.ai*

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. Unsloth provides documentation and support for running the GLM-5.3-Flash (ox-alpha) multimodal model locally using Unsloth Desktop or llama.cpp. Developed by Z.ai, this 320B parameter model features a hybrid sparse and linear attention architecture designed to improve scaling through Manifold-Constrained Hyper Connections. Users can utilize various quantization levels'' from 1-bit for low RAM requirements (approx. 93GB) up to higher bitrates for improved accuracy'' to run the model on hardware ranging from Mac systems to NVIDIA DGX Spark setups.

    - The model features three thinking modes: Low, High, and Max reasoning effort.
    - It is designed to rival Claude Opus 4.8 in coding and agentic benchmarks.
    - Unsloth's dynamic 1-bit quantization retains 71% of top-1% accuracy while being 85% smaller than the BF16 version.
    - The model can be run via a local API using `unsloth run` with llama-server runtime flags.
  2. Maria Deutscher writes that Z.ai has open-sourced GLM-5.3-Flash, a 320-billion-parameter mixture-of-experts LLM that debuted the prior week under the anonymous codename "Ox Alpha" on OpenRouter, sparking industry speculation about its origin. The model is 10x more cost-efficient than Z.ai's predecessor, using sparse and linear attention mechanisms to dramatically reduce memory and processing overhead while supporting a 1-million-token context window.
    - Linear attention replaces the softmax function, cutting memory scaling from quadratic to linear as prompt size grows
    - Trained on 30 trillion tokens using an mHC technique that prevents gradient distortion through neuron layers
    - Scored highest among Claude Opus 4.8, GPT-5.6 Terra, and Gemini 3.7 Flash on GDPval-AA v2, a knowledge-work evaluation

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: tagged with "z.ai"

About - Propulsed by SemanticScuttle