The Laya AI model is a family of open-weight, non-autoregressive decision models from Convai Innovations designed to return structured answers and probabilities rather than free-form text. By mapping input states to specific question types like choices, scores, or propositions (noul), it provides deterministic outputs suitable for application policies in workflows such as support ticket routing. The guide details its architecture—utilizing bidirectional encoders like ModernBERT—and offers practical advice on running the model locally via Python and evaluating performance through metrics like calibration and accuracy.
- Laya uses a "state + typed questions" pattern to ensure output validity without needing complex parsing of generative prose.
- It offers three specific checkpoints: an English version, a multilingual version (mmBERT), and one optimized for typed decisions.
- The model's architecture relies on bidirectional encoders with decision heads rather than token-by-token generation.
- Users are encouraged to implement "abstention policies" where uncertain predictions (based on low confidence/calibration) are routed to humans.
Simon Willison writes about Jev, a new category of models from TypeSafe AI called "System One models" or decision models. Unlike standard large language models that output text, Jev accepts unstructured input and returns structured probabilistic decisions such as floating-point numbers for yes/no questions (Noul), choices between options, or numeric scores. These models are designed to be extremely fast and inexpensive, charging only for input tokens while providing free output.
- Jev is optimized for classification tasks like spam detection, ranking, and labeling.
- The model's "Noul" question type refers to the Bernoulli distribution.
- Using such black-box decision models raises concerns about hidden biases that are difficult to audit without explanations.
- There is an emerging trend of open-weight recreations of Jev-class models, including projects like Kev and benchmarks like JevBench.
Alvaro Bartolome provides a Rust-based implementation of the System One compatible API, designed specifically for open decision models such as Laya. The project features dynamic token-based batching and supports hardware acceleration via CPU, CUDA, and Metal (MPS). It is built using modern asynchronous frameworks like tokio and axum to provide high performance for model queries.
- Achieves approximately 14ms latency per query on an NVIDIA RTX Pro 6000.
- Includes support for ModernBert with custom decision heads for Laya models.
- Utilizes the Candle machine learning framework by Hugging Face.
- Supports multiple installation features via cargo, including specific flags for metal or cuda.