Mark Marosi writes about decider, a family of models fine-tuned from Qwen3.5 that produce typed decisions (choice, score, boolean) in a single forward pass without text generation, returning calibrated probability distributions over user-defined options. The project is an open reproduction of TypeSafe AI's "System One" model class (Jev), released in sizes from 0.8B to 35B mixture-of-experts with 3B active parameters.
- The schema cache stores K/V states for repeated question prefixes, achieving up to 19x speedup on large option sets by running only the state per request
- v10 adds calibration-aware RL on live MiniWoB++ browser tasks and exact games, lifting browser accuracy from 83% to 93% and halving the belief gap
- TypeSafe's SDKs work unchanged by pointing TYPESAFE_BASE_URL at the decider server
- The 35B model outperforms the 2B on 93 of 95 regression tasks but costs 3-4x per decision and lacks the RL stage
- Co-developed with Claude (Anthropic) as a listed co-author on commits
Diogo Almeida writes that TypeSafe AI is releasing Jev, its first System One Model—a new class of frontier model built for fast, structured decisions that software can consume directly. Unlike autoregressive language models that generate strings token by token, Jev outputs type-safe structured values with calibrated probabilities in a single parallel query, achieving frontier-level intelligence on decision tasks at roughly 40–200× lower latency and cost. The company's new training method, Reinforcement Learning for Calibrated Decisions (RLCD), optimizes for epistemically honest probability estimates rather than human preference or verifiable rewards, and the architecture is mathematically incapable of producing type errors or hallucinations.
- Named after William Stanley Jevons, whose paradox predicted that efficiency gains would increase (not decrease) total demand; TypeSafe expects each order-of-magnitude cost drop to unlock orders of magnitude more use cases.
- Workflow evals benchmark Jev against the average of GPT-6 Astra and Fable 5.1 as reference probabilities, claiming 193.6× speed and 444.6× cost advantages on production-shaped tasks.
- The team demonstrated real-time intelligence with a Doom bot making 10 structured queries per second (~$7/hour) and a Wikiracing bot that outperforms LLMs at high-cardinality link selection.
- Jev supports output cardinality up to 255; for higher-cardinality choices it falls back to a two-stage scoring system that scores independently then makes an explicit selection.