yibie maintains a curated awesome list of over 300 public projects built on Jev, TypeSafe AI's System One model that takes unstructured state plus a typed question and returns a typed decision (Choice, Score, or Noul) with a confidence value in a single forward pass. The list spans 14 categories from classification and routing to agent safety guardrails, database extensions, and open-source replicas, enforcing inclusion rules that require public, citable sources demonstrating a genuine typed-decision use of Jev.
- The repo explicitly warns that same-day bulk submissions from one author sharing a scaffold can satisfy every inclusion rule while remaining unproven, and treats volume as not evidence of quality
- Open replicas range from 547k parameters (minojev) to 395M (von); one 706K-parameter model beats Jev at form-filling (99.7% vs 83.6%)
- Jev performs no token generation; its speed comes from parallel constrained decoding, so an inference engine can expose a Jev-like API over any open-weight model
- A CAPTCHA arbitrage example illustrates per-decision pricing: solving at $0.0068 per hundred against a marketplace paying a cent each
- LangChain published both a Jev harness wiring guide and an independent evaluation concluding Jev is the cheaper and more consistent judge for online evals versus LLM judges
Mark Marosi writes about decider, a family of models fine-tuned from Qwen3.5 that produce typed decisions (choice, score, boolean) in a single forward pass without text generation, returning calibrated probability distributions over user-defined options. The project is an open reproduction of TypeSafe AI's "System One" model class (Jev), released in sizes from 0.8B to 35B mixture-of-experts with 3B active parameters.
- The schema cache stores K/V states for repeated question prefixes, achieving up to 19x speedup on large option sets by running only the state per request
- v10 adds calibration-aware RL on live MiniWoB++ browser tasks and exact games, lifting browser accuracy from 83% to 93% and halving the belief gap
- TypeSafe's SDKs work unchanged by pointing TYPESAFE_BASE_URL at the decider server
- The 35B model outperforms the 2B on 93 of 95 regression tasks but costs 3-4x per decision and lacks the RL stage
- Co-developed with Claude (Anthropic) as a listed co-author on commits