Zhening Li and colleagues introduce JAZ, an LLM agent framework designed to minimize the complexity of agent loops by treating them as a programming language primitive called `invoke`. Instead of relying on external specialized systems for memory or self-improvement, JAZ enables agents to achieve these capabilities through code execution where all interactions are treated as variables within the environment. This minimalist approach allows highly expressive workflows, such as long-horizon recall and continual self-improvement, using only prompting rather than manually designed tools or complex architectures.
- The `invoke` primitive allows for recursive calls, enabling LLMs to write arbitrary executable code that includes further iterations of itself.
- In testing on the StuLife dataset, JAZ outperformed MemGPT (Letta) by 8% in recall performance while costing half as much.
- On self-improvement tasks using AppWorld, JAZ demonstrated a 4% improvement over ACE at a lower computational cost.
Convai Innovations presents Laya, a multilingual, non-autoregressive system 1 decision model designed to provide typed answers with mathematically calibrated probabilities in a single forward pass. Unlike generative models, it does not generate text, thereby eliminating hallucinations and the need for parsing. The framework includes an automated Router that detects language and script to dispatch tasks to the most efficient checkpoint (English or Multilingual) within approximately 35ms on GPU.
- It is trained using Reinforcement Learning with Calibrated Decisions (RLCD) to ensure honest probability reporting.
- Laya can support context lengths of up to 8,192 tokens in its multilingual version.
- The model family includes specialized checkpoints like `laya-typed-decisions` which achieves significantly higher accuracy through fine-tuning on specific workflows.
- Performance benchmarks show it is roughly 6–8× faster than TypeSafe Jev for single question latency on a T4 GPU.
Mark Marosi writes about decider, a family of models fine-tuned from Qwen3.5 that produce typed decisions (choice, score, boolean) in a single forward pass without text generation, returning calibrated probability distributions over user-defined options. The project is an open reproduction of TypeSafe AI's "System One" model class (Jev), released in sizes from 0.8B to 35B mixture-of-experts with 3B active parameters.
- The schema cache stores K/V states for repeated question prefixes, achieving up to 19x speedup on large option sets by running only the state per request
- v10 adds calibration-aware RL on live MiniWoB++ browser tasks and exact games, lifting browser accuracy from 83% to 93% and halving the belief gap
- TypeSafe's SDKs work unchanged by pointing TYPESAFE_BASE_URL at the decider server
- The 35B model outperforms the 2B on 93 of 95 regression tasks but costs 3-4x per decision and lacks the RL stage
- Co-developed with Claude (Anthropic) as a listed co-author on commits
Diogo Almeida writes that TypeSafe AI is releasing Jev, its first System One Model—a new class of frontier model built for fast, structured decisions that software can consume directly. Unlike autoregressive language models that generate strings token by token, Jev outputs type-safe structured values with calibrated probabilities in a single parallel query, achieving frontier-level intelligence on decision tasks at roughly 40–200× lower latency and cost. The company's new training method, Reinforcement Learning for Calibrated Decisions (RLCD), optimizes for epistemically honest probability estimates rather than human preference or verifiable rewards, and the architecture is mathematically incapable of producing type errors or hallucinations.
- Named after William Stanley Jevons, whose paradox predicted that efficiency gains would increase (not decrease) total demand; TypeSafe expects each order-of-magnitude cost drop to unlock orders of magnitude more use cases.
- Workflow evals benchmark Jev against the average of GPT-6 Astra and Fable 5.1 as reference probabilities, claiming 193.6× speed and 444.6× cost advantages on production-shaped tasks.
- The team demonstrated real-time intelligence with a Doom bot making 10 structured queries per second (~$7/hour) and a Wikiracing bot that outperforms LLMs at high-cardinality link selection.
- Jev supports output cardinality up to 255; for higher-cardinality choices it falls back to a two-stage scoring system that scores independently then makes an explicit selection.
Milan Minsky writes that Leela AI transforms standard factory and warehouse cameras into smart sensors, offering an alternative to traditional IoT sensors by leveraging existing video feeds instead of physical hardware. The platform provides contextual visibility into operations, identifies bottlenecks, and tracks interactions between machines, operators, and materials without requiring retrofitting. It complements IoT systems by integrating with platforms like Velotic ThingWorx and AVEVA to create a comprehensive digital twin of manufacturing floors. The core technology utilizes MIT research-based AI, combining causal and neural networks for efficient data processing.
Dan Russell writes about the power of AI-augmented search to retrieve hard-to-find information, using an example of finding a study on how the gender of lab assistants affects experimental outcomes on lab mice. He demonstrates how a simple query with AI can yield relevant results, leading to original source papers. The study highlights the impact of experimenter gender on reproducibility in scientific research.
Michal Sutter writes about Pollen Robotics, a Bordeaux-based team at Hugging Face, which has opened pre-orders for Microduck, a 25 cm bipedal robot priced at $399. Unlike most robotics launches that rely on demo videos, Microduck ships with its full training loop — every movement (walking, sitting, kicking, roller-skating, self-recovery) is a neural policy trained in a physics simulator and exported to hardware. The robot carries 15 motors, a camera, LiDAR, two IMUs, and a Rockchip RK3566, with policies trained via PPO in MuJoCo Warp in roughly one to two hours on a CUDA GPU.
- Sim-to-real hinges on a BAM actuator model (voltage control law, back-EMF, Coulomb/Stribeck/load-dependent friction) plus randomization of battery voltage, command delay, and ±1° backlash per joint
- Every policy shares a 61-dimensional actor observation (48 proprioception + twist, head pose, body pose commands), enabling hot-swap between walk, recover, and trick policies mid-run
- Software is Apache-2.0, but mechanical and electronic design files are not open
- The robot generates a unique audio identity on first wake that persists permanently; it does not speak in a linguistic sense
- Pre-orders opened August 27, 2026, with deliveries targeted before Christmas
Chris Patrick writes that SLAC researchers have built a neural-network method for compressing large scientific datasets while preserving fine details that conventional compression erases. The approach uses wavelet analysis to separate data features by scale, then encodes each scale separately through a neural network, enabling 10- to 100-fold file size reductions and selective decompression of only the regions a researcher needs.
- Published in Nature Machine Intelligence (August 24, 2026)
- Motivated by upcoming LCLS upgrades that will generate nearly one terabyte of data per second
- Tested successfully on X-ray diffraction data, solar magnetic field measurements, and photographs
- Neural networks trained on Perlmutter at NERSC (Lawrence Berkeley National Lab)
- Co-developers include researchers from UC Davis and Carnegie Mellon University
El Assadi et al. compare ten LLMs (six families) and 26 embedding models (118M - 14B parameters) on 37 tasks, considering cost. In aggregate, the two paradigms are effectively tied (best LLM scores 77.6 versus best embedding model 77.2), yet their strengths diverge by task: LLMs lead on reasoning-heavy retrieval while embedding models lead on classification, and the two match on clustering, STS, and pair classification.
LLMs are significantly more expensive (up to 1,431x) and slower (2.5-736x) than embedding models for certain tasks. The authors suggest using embedding models for similarity, classification, and clustering, and LLMs for reasoning in retrieval.
Reasoning tokens are 28-81% of LLM inference cost; lower budgets maintain or boost retrieval quality for most tested models.
- Only Gemini 3.1 Pro breaks into the Pareto frontier alongside the leading embedding models.
- Accepted to COLM 2026; code, datasets, and results are publicly released on GitHub.
Alibaba has open-sourced Qwen-UI-Agent, a GUI agent foundation model that operates across mobile, desktop, web, and deep-search environments on real hardware rather than relying on simulation. It achieves top benchmark results: 82.1% on MobileWorld, 79.5% on OSWorld-Verified, and first on WebArena. It also introduces MobileWorld-Real, a 400+ task benchmark on 100+ phones and 150+ apps, with a 92.2% success rate.
- Supports command-line execution alongside standard GUI operations and batches multiple actions into a single decision step to shorten trajectories.
- Built-in safety layer refuses illegal or high-risk requests outright and pauses at sensitive operations (payments, data deletion, privacy grants) for explicit user confirmation.
- Trained via online reinforcement learning on trajectories exceeding 100 steps, paired with adaptive curriculum learning to progressively tackle longer tasks.