klotz: typed decisions*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. yibie maintains a curated awesome list of over 300 public projects built on Jev, TypeSafe AI's System One model that takes unstructured state plus a typed question and returns a typed decision (Choice, Score, or Noul) with a confidence value in a single forward pass. The list spans 14 categories from classification and routing to agent safety guardrails, database extensions, and open-source replicas, enforcing inclusion rules that require public, citable sources demonstrating a genuine typed-decision use of Jev.
    - The repo explicitly warns that same-day bulk submissions from one author sharing a scaffold can satisfy every inclusion rule while remaining unproven, and treats volume as not evidence of quality
    - Open replicas range from 547k parameters (minojev) to 395M (von); one 706K-parameter model beats Jev at form-filling (99.7% vs 83.6%)
    - Jev performs no token generation; its speed comes from parallel constrained decoding, so an inference engine can expose a Jev-like API over any open-weight model
    - A CAPTCHA arbitrage example illustrates per-decision pricing: solving at $0.0068 per hundred against a marketplace paying a cent each
    - LangChain published both a Jev harness wiring guide and an independent evaluation concluding Jev is the cheaper and more consistent judge for online evals versus LLM judges
  2. Mark Marosi writes about decider, a family of models fine-tuned from Qwen3.5 that produce typed decisions (choice, score, boolean) in a single forward pass without text generation, returning calibrated probability distributions over user-defined options. The project is an open reproduction of TypeSafe AI's "System One" model class (Jev), released in sizes from 0.8B to 35B mixture-of-experts with 3B active parameters.
    - The schema cache stores K/V states for repeated question prefixes, achieving up to 19x speedup on large option sets by running only the state per request
    - v10 adds calibration-aware RL on live MiniWoB++ browser tasks and exact games, lifting browser accuracy from 83% to 93% and halving the belief gap
    - TypeSafe's SDKs work unchanged by pointing TYPESAFE_BASE_URL at the decider server
    - The 35B model outperforms the 2B on 93 of 95 regression tasks but costs 3-4x per decision and lacks the RL stage
    - Co-developed with Claude (Anthropic) as a listed co-author on commits

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: typed decisions

About - Propulsed by SemanticScuttle