Donald Papp writes about Jev, a new class of model that takes text input but outputs only floating-point numbers, making it fast and cheap for classification tasks. Rather than generating sentences, it returns direct answers to yes/no questions, multiple-choice lists, and scoring requests. The concept has quickly gained traction, with developers already building their own decision-type models like Kev and Nimble.
- Jev outputs a confidence score for every answer, derived from the relative token probabilities
- Nimble is small enough to run locally and was recently added as a supported model in Ollama
- Simon Willison provided a concise summary of what Jev does
- The comment section sparked debate over whether this is truly novel, with some noting it's essentially an LLM with constrained outputs
Codewhale is a Rust-based terminal agent designed to read codebases, edit files, run commands, and verify results in a continuous loop, distinguishing it from tools that only generate code. It supports both hosted and local model providers and offers a two-mode system (`plan` for read-only exploration and `work` for active modification) that increases developer confidence in its autonomy.
- Written in Rust and distributed via crates.io, npm, Docker, Nix, and Android/Termux.
- Supports headless execution via `codewhale exec` for use in CI pipelines or shell scripts.
- Allows delegation of different subtasks to different models or roles within a single project.
- Maintainers note that the changelog currently describes an unreleased candidate not yet in published downloads.
Nolen Jonker writes that consolidating multiple LLMs into a single open-source client, Cherry Studio, allows users to compare outputs, manage privacy, and control costs more effectively than using separate vendor subscriptions. The tool acts as a unified workspace where API-based cloud models and local instances can be queried simultaneously, letting users route sensitive data to private local systems while utilizing specialized cloud models for complex tasks.
- Cherry Studio supports simultaneous multi-model responses, enabling side-by-side comparisons and acting as a basic hallucination check.
- For most users, pay-per-token APIs are cheaper than flat-rate premium subscriptions unless they use high-end models for extended periods daily.
- Providers like Anthropic and OpenAI do not train on API inputs or outputs by default, offering better privacy than their respective standard consumer apps.
- Alternatives include self-hosted options like LibreChat and Open WebUI, or simpler desktop clients like AnythingLLM and Jan.
llama.cpp now supports decision models via a `/v1/systemone` endpoint, which accepts a state and typed questions to return probabilities for each provided option in a single forward pass. The API utilizes the System One format introduced with TypeSafe's Jev model, meaning existing clients only need a new base URL. Five initial models are available ranging from 144M to 27B parameters, supporting use cases like request routing, content moderation, and agent action selection with median response times as low as 3 ms.
- The `state` field accepts text, JSON, screenshots, or a list of chat messages with `image_url` parts.
- Three question types are supported: `choice` for categorical selection, `score` for a continuous level between 2 and 10 options, and `noul` for yes/no probabilities.
- Router mode allows loading multiple models on a single server and selecting one per request.
- Adding descriptions to option labels can significantly improve accuracy, with Julia-1's routing confidence for "charged twice" jumping from a misroute to 0.99 when descriptions were provided.
- Cloudflare's Clef is the next model planned for integration.
autoharness is a self-learning skill layer for Claude Code that distills reusable skills from a user's real sessions, merges near-duplicates, updates them in use, and prunes those that stop getting used — all without a daemon or an external benchmark. It fires on tool-call count rather than turns, keeps only the skills it authored, and validates a skill's worth by adherence in later turns rather than a held-out score.
- Skills are stored as plain native SKILL.md files in `.claude/skills/`; the plugin's own recall index is injected on top of the host's native mechanism
- Three distinct lifecycle signals are tracked: load (model invoked the skill), view (session read into the skill's directory), and patch (promoter landed an improvement)
- The `/learn` command allows on-demand distillation of the current session through the same proposal-and-validation chain
GitButler is a version control tool designed to optimize workflows for developers and AI agents by introducing features like stacked branches, parallel branching, and unlimited undo capabilities. It layers seamlessly onto existing Git repositories without requiring new configuration, aiming to reduce the friction often associated with complex Git operations through structured output and idempotent commands.
- Agents are reported to be 60% faster using GitButler compared to vanilla Git
- The software includes a CLI that offers JSON output mode for better AI parsing
- Features include "Smartlog" and simplified history editing/rebasing
- It is free and open source software
Gregory Gibson writes about applying the Linux kernel's tool-generated content guidance to his vibe-coded Morse decoder app, CW Inspector. He used OpenAI Codex to build the app, then subjected it to the kernel's five transparency rules: name the tool, preserve inputs, keep the prompt trail, record exactly what the tool changed, and make the generated code prove itself. The process exposed a critical flaw — the decoder confidently produced a wrong answer (reading TTT E as SE) with zero errors, demonstrating that a clean run doesn't guarantee correctness.
- The original algorithm treated the shortest 55% of keyed pulses as dots, which broke on dash-heavy messages; the fix looks for a large ratio between short and long pulse clusters before recording dot duration.
- The author deliberately withheld the actual WAV file from Codex, providing only metadata and the expected message, to create an independent acceptance test rather than letting the model optimize against the exact sample it would later decode.
- The kernel's guidance applies only when a tool generates something substantial (functions, files, fixes, translations), not for spelling corrections, autocomplete, or variable renames.
@omarsar0 writes on X that the fastest path to genuinely understanding agent harnesses is to build one from scratch in TypeScript or Python, starting with a minimal ReAct implementation prompted from Google's original paper, targeting three clean components—an LLM inference module (multi-model, OpenRouter-backed, with separable system prompt), an MCP tools module for interoperability, and a simple agent loop that ties them together—then logging every input/output at each boundary and iterating against a small set of diverse test tasks so each change is inspectable. The punchline: skip the framework first, because only once you've felt the loop, the tokens, and the tool calls in your own code do the "next steps"—memory, skills, subagents—stop being black boxes you configure and become modules you actually know how to tune.
- LLM module: wraps inference across multiple frontier models via OpenRouter; system prompt either embedded or isolated for context-engineering experiments
- Tools module: implement as MCP (Model Context Protocol) tools for cross-harness interoperability, or as bespoke functions if experienced
- Agent loop: ReAct pattern (alternating reasoning traces and action calls) encapsulating both LLM and tools; exit conditions handled via system-prompt instructions (non-deterministic), code-level checks (deterministic), or both
- Logging strategy: capture loop in/out, every LLM call in/out, and every tool-call in/out; run a fixed diverse task suite after each modification
- Scaling path: keep architecture modular so memory, skills, and subagent orchestration can be bolted on once the core loop is understood
- Shortcut alternatives (if not building from scratch): Pi SDK or LangChain harness tooling
mini-swe-agent is a radically simple Python-based agent from the Princeton and Stanford team behind SWE-bench that uses only bash as its tool, maintains a completely linear message history, and executes each action via independent subprocess.run calls. Despite being roughly 100 lines of core agent code, it scores over 74% on SWE-bench verified and is used by organizations including Meta, NVIDIA, IBM, and Anyscale.
- The core design argument is that as language models grow more capable, elaborate tool scaffolds become unnecessary and the LM itself should drive the shell
- Supports sandboxed deployment via docker, podman, singularity, bwrap, and others; installable from PyPI via uvx, pipx, or pip
Anurag Singh writes that providing Claude Code with read-only access to a SaaS application's server logs allowed the coding agent to identify and propose fixes for real performance issues. By observing error patterns, traces, and metrics directly within the environment rather than relying on manual bug reports, the agent was able to autonomously trace bugs back to specific lines of code across various files.
- The experiment highlights a shift toward AI agents joining the "on-call" workflow by inspecting live operational telemetry.
- To mitigate security risks, it is recommended using Model Context Protocol (MCP) servers to restrict an agent's tools to read-only actions.
- Major observability companies like Sentry and Datadog are already implementing similar features to automate root cause analysis and pull request generation.