llama.cpp now supports decision models via a `/v1/systemone` endpoint, which accepts a state and typed questions to return probabilities for each provided option in a single forward pass. The API utilizes the System One format introduced with TypeSafe's Jev model, meaning existing clients only need a new base URL. Five initial models are available ranging from 144M to 27B parameters, supporting use cases like request routing, content moderation, and agent action selection with median response times as low as 3 ms.
- The `state` field accepts text, JSON, screenshots, or a list of chat messages with `image_url` parts.
- Three question types are supported: `choice` for categorical selection, `score` for a continuous level between 2 and 10 options, and `noul` for yes/no probabilities.
- Router mode allows loading multiple models on a single server and selecting one per request.
- Adding descriptions to option labels can significantly improve accuracy, with Julia-1's routing confidence for "charged twice" jumping from a misroute to 0.99 when descriptions were provided.
- Cloudflare's Clef is the next model planned for integration.
The Laya AI model is a family of open-weight, non-autoregressive decision models from Convai Innovations designed to return structured answers and probabilities rather than free-form text. By mapping input states to specific question types like choices, scores, or propositions (noul), it provides deterministic outputs suitable for application policies in workflows such as support ticket routing. The guide details its architecture—utilizing bidirectional encoders like ModernBERT—and offers practical advice on running the model locally via Python and evaluating performance through metrics like calibration and accuracy.
- Laya uses a "state + typed questions" pattern to ensure output validity without needing complex parsing of generative prose.
- It offers three specific checkpoints: an English version, a multilingual version (mmBERT), and one optimized for typed decisions.
- The model's architecture relies on bidirectional encoders with decision heads rather than token-by-token generation.
- Users are encouraged to implement "abstention policies" where uncertain predictions (based on low confidence/calibration) are routed to humans.
Simon Willison writes about Jev, a new category of models from TypeSafe AI called "System One models" or decision models. Unlike standard large language models that output text, Jev accepts unstructured input and returns structured probabilistic decisions such as floating-point numbers for yes/no questions (Noul), choices between options, or numeric scores. These models are designed to be extremely fast and inexpensive, charging only for input tokens while providing free output.
- Jev is optimized for classification tasks like spam detection, ranking, and labeling.
- The model's "Noul" question type refers to the Bernoulli distribution.
- Using such black-box decision models raises concerns about hidden biases that are difficult to audit without explanations.
- There is an emerging trend of open-weight recreations of Jev-class models, including projects like Kev and benchmarks like JevBench.
Alvaro Bartolome provides a Rust-based implementation of the System One compatible API, designed specifically for open decision models such as Laya. The project features dynamic token-based batching and supports hardware acceleration via CPU, CUDA, and Metal (MPS). It is built using modern asynchronous frameworks like tokio and axum to provide high performance for model queries.
- Achieves approximately 14ms latency per query on an NVIDIA RTX Pro 6000.
- Includes support for ModernBert with custom decision heads for Laya models.
- Utilizes the Candle machine learning framework by Hugging Face.
- Supports multiple installation features via cargo, including specific flags for metal or cuda.