Clef is a 27B multimodal model developed by Cloudflare that transforms state data and schemas of typed questions into structured decisions. By accepting text, JSON, images, or video as input, it returns a probability for every allowed option of each question in a single forward pass, entirely eliminating free-form text generation and output parsing. The model is post-trained from Qwen/Qwen3.8-27B and features a joint schema head that routes evidence to each question and scores all options simultaneously. It is fully compatible with Jev and SystemOne APIs, making it well-suited for deterministic, structured decision-making tasks.
- Supports `choice`, `score`, and `noul` (true/false) question types
- A smaller, faster variant called Clef-Flash is also available
- Achieves strong benchmark results, particularly on BFCL (98.5%), ToolRet (69.2), and API-Bank (91.9%)
- Significantly faster than comparable models, with a median latency of 209.3 ms
- Released under the Apache-2.0 license
The Laya AI model is a family of open-weight, non-autoregressive decision models from Convai Innovations designed to return structured answers and probabilities rather than free-form text. By mapping input states to specific question types like choices, scores, or propositions (noul), it provides deterministic outputs suitable for application policies in workflows such as support ticket routing. The guide details its architecture—utilizing bidirectional encoders like ModernBERT—and offers practical advice on running the model locally via Python and evaluating performance through metrics like calibration and accuracy.
- Laya uses a "state + typed questions" pattern to ensure output validity without needing complex parsing of generative prose.
- It offers three specific checkpoints: an English version, a multilingual version (mmBERT), and one optimized for typed decisions.
- The model's architecture relies on bidirectional encoders with decision heads rather than token-by-token generation.
- Users are encouraged to implement "abstention policies" where uncertain predictions (based on low confidence/calibration) are routed to humans.
Convai Innovations presents Laya, a multilingual, non-autoregressive system 1 decision model designed to provide typed answers with mathematically calibrated probabilities in a single forward pass. Unlike generative models, it does not generate text, thereby eliminating hallucinations and the need for parsing. The framework includes an automated Router that detects language and script to dispatch tasks to the most efficient checkpoint (English or Multilingual) within approximately 35ms on GPU.
- It is trained using Reinforcement Learning with Calibrated Decisions (RLCD) to ensure honest probability reporting.
- Laya can support context lengths of up to 8,192 tokens in its multilingual version.
- The model family includes specialized checkpoints like `laya-typed-decisions` which achieves significantly higher accuracy through fine-tuning on specific workflows.
- Performance benchmarks show it is roughly 6–8× faster than TypeSafe Jev for single question latency on a T4 GPU.
NandhaKishorM writes about Laya, a non-autoregressive System 1 decision engine designed to perform typed decisions—such as choices, scores, and yes/no answers—over text in over 100 languages. Using a single forward pass with an integrated router that selects the appropriate checkpoint per request, it achieves high speed (e.g., ~33 ms on a T4 GPU) without the need for text generation or parsing.
- Features three specific checkpoints: English, Multilingual (for 100+ languages), and Typed Decisions.
- Supports multiple decision primitives: `choice` (labels/probabilities), `score` (ordinal rubrics), and `noul` (binary probability).
- Offers a "Fast Path" using TileLang GPU kernels for significant latency reduction on NVIDIA hardware.
- Includes an HTTP server implementation (`laya-serve`) that is Jev-compatible via the `/v1/systemone` protocol.
- Provides TypeScript support through `laya-ts`, allowing inference in Node.js and browser environments via ONNX Runtime.
yibie maintains a curated awesome list of over 300 public projects built on Jev, TypeSafe AI's System One model that takes unstructured state plus a typed question and returns a typed decision (Choice, Score, or Noul) with a confidence value in a single forward pass. The list spans 14 categories from classification and routing to agent safety guardrails, database extensions, and open-source replicas, enforcing inclusion rules that require public, citable sources demonstrating a genuine typed-decision use of Jev.
- The repo explicitly warns that same-day bulk submissions from one author sharing a scaffold can satisfy every inclusion rule while remaining unproven, and treats volume as not evidence of quality
- Open replicas range from 547k parameters (minojev) to 395M (von); one 706K-parameter model beats Jev at form-filling (99.7% vs 83.6%)
- Jev performs no token generation; its speed comes from parallel constrained decoding, so an inference engine can expose a Jev-like API over any open-weight model
- A CAPTCHA arbitrage example illustrates per-decision pricing: solving at $0.0068 per hundred against a marketplace paying a cent each
- LangChain published both a Jev harness wiring guide and an independent evaluation concluding Jev is the cheaper and more consistent judge for online evals versus LLM judges
TianyuCodings writes about JevHarness, a system where an LLM authors a task-specific harness for Jev (a lightweight judgment model from TypeSafe). The harness converts task observations into features, constructs Jev questions and criteria, and combines structured answers into actions. Once the harness is frozen, execution runs only the harness code and Jev calls without the authoring LLM on every decision. An optional GEPA evolution loop refines the harness using rewards and complete execution traces.
- Pokemon eval: 5 reflection rounds improved win rate from 25% (3/12) to 75% (9/12); the round-3 candidate was selected as best
- Selected harness median full decision: 568 ms (P95 657 ms); individual Jev request median: 269 ms (P95 348 ms)
- Installable as a Claude Code plugin or Codex skill; the skill guides the agent to clarify task inputs, legal actions, and success criteria before building
- Requires Python 3.11+; functional Python nodes need a macOS native sandbox and fail closed if unavailable
- The archived website and recorded-call inspection require no model credentials
jev-social is an open-source tool that pairs the Jev model (from Typesafe AI) with the socai browser CLI to perform grounded social media research on Instagram, TikTok, and LinkedIn. Jev makes sequential decisions about which operation to perform next (search, open a profile, read comments, download media), while socai executes each chosen command in a real Chrome instance and returns the result before Jev makes the next decision.
- Packaged as a Codex plugin and installable via `npx` or the cross-agent Skills CLI
- Each step's choice, confidence, command, and elapsed time are stored for auditability
- Enforces read-only and login-gate boundaries; unsupported or low-confidence decisions are rejected rather than executed
- A recorded Instagram run produced four source-linked records in ~64 seconds
- Runs loopback-only on localhost:8766 and requires an OpenRouter API key with Jev access
Mark Marosi writes about decider, a family of models fine-tuned from Qwen3.5 that produce typed decisions (choice, score, boolean) in a single forward pass without text generation, returning calibrated probability distributions over user-defined options. The project is an open reproduction of TypeSafe AI's "System One" model class (Jev), released in sizes from 0.8B to 35B mixture-of-experts with 3B active parameters.
- The schema cache stores K/V states for repeated question prefixes, achieving up to 19x speedup on large option sets by running only the state per request
- v10 adds calibration-aware RL on live MiniWoB++ browser tasks and exact games, lifting browser accuracy from 83% to 93% and halving the belief gap
- TypeSafe's SDKs work unchanged by pointing TYPESAFE_BASE_URL at the decider server
- The 35B model outperforms the 2B on 93 of 95 regression tasks but costs 3-4x per decision and lacks the RL stage
- Co-developed with Claude (Anthropic) as a listed co-author on commits
w3cj writes about jev-chat, a tool-calling chat bot that routes user requests to real tools using Jev, a non-generative classifier from TypeSafe, with no LLM writing any output. Every value on screen was either typed by the user or returned by a tool, so the assistant cannot invent a fact. The system supports weather, unit conversion, Wikipedia lookups, recipes, web search, Todoist, and Home Assistant via MCP servers, with an inspector pane exposing every decision and probability for each turn.
- Jev answers only two question types — Choice and Noul — and never produces text; all reply wording is templated in code
- Pre-processing handles spell-check (cspell + compromise) and resolves short follow-ups like "what about Boston?" by swapping in the new value
- Multi-step tools (Wikipedia, web search) chain multiple Jev requests: pick topic, then article, then the exact line that answers
- Confidence-gated routing shows two buttons when the top tools are close rather than guessing
- The repo is a proof of concept; the author will not accept PRs for new features
- English only; no compound requests or multi-step reasoning supported
jeff is a self-hosted drop-in replacement for TypeSafe's jev System One API, powered by the 400M-parameter GLiFormer-large-v1 model. It serves `choice`, `score`, and `noul` classification questions over a compatible wire format, so existing applications using the official `typesafe-sdk` can point at it by changing a single base URL. Deployment targets include GPU (L4, A10G) and CPU (ONNX Runtime with int8 quantization) via Modal, with ~50 req/s per container throughput on L4. Benchmarks on 1,600 labeled items show it costs roughly a quarter of jev per million requests but trails significantly on reasoning-heavy tasks like irony and reading comprehension.
- `JEFF_ISOLATE=nouls` (default) gives separate encoder passes per noul question to reduce cross-question interference; choice and score questions share a pass unless set to `all`
- Default temperature of 3.2 calibrates noul probabilities; setting it to 1 makes `score` output match the weighted average of displayed probabilities
- A smaller `gliformer-base-v1` variant with `JEFF_NOUL_MODE=single` is recommended for faster local iteration