Nolen Jonker writes that consolidating multiple LLMs into a single open-source client, Cherry Studio, allows users to compare outputs, manage privacy, and control costs more effectively than using separate vendor subscriptions. The tool acts as a unified workspace where API-based cloud models and local instances can be queried simultaneously, letting users route sensitive data to private local systems while utilizing specialized cloud models for complex tasks.
- Cherry Studio supports simultaneous multi-model responses, enabling side-by-side comparisons and acting as a basic hallucination check.
- For most users, pay-per-token APIs are cheaper than flat-rate premium subscriptions unless they use high-end models for extended periods daily.
- Providers like Anthropic and OpenAI do not train on API inputs or outputs by default, offering better privacy than their respective standard consumer apps.
- Alternatives include self-hosted options like LibreChat and Open WebUI, or simpler desktop clients like AnythingLLM and Jan.
llama.cpp now supports decision models via a `/v1/systemone` endpoint, which accepts a state and typed questions to return probabilities for each provided option in a single forward pass. The API utilizes the System One format introduced with TypeSafe's Jev model, meaning existing clients only need a new base URL. Five initial models are available ranging from 144M to 27B parameters, supporting use cases like request routing, content moderation, and agent action selection with median response times as low as 3 ms.
- The `state` field accepts text, JSON, screenshots, or a list of chat messages with `image_url` parts.
- Three question types are supported: `choice` for categorical selection, `score` for a continuous level between 2 and 10 options, and `noul` for yes/no probabilities.
- Router mode allows loading multiple models on a single server and selecting one per request.
- Adding descriptions to option labels can significantly improve accuracy, with Julia-1's routing confidence for "charged twice" jumping from a misroute to 0.99 when descriptions were provided.
- Cloudflare's Clef is the next model planned for integration.
autoharness is a self-learning skill layer for Claude Code that distills reusable skills from a user's real sessions, merges near-duplicates, updates them in use, and prunes those that stop getting used — all without a daemon or an external benchmark. It fires on tool-call count rather than turns, keeps only the skills it authored, and validates a skill's worth by adherence in later turns rather than a held-out score.
- Skills are stored as plain native SKILL.md files in `.claude/skills/`; the plugin's own recall index is injected on top of the host's native mechanism
- Three distinct lifecycle signals are tracked: load (model invoked the skill), view (session read into the skill's directory), and patch (promoter landed an improvement)
- The `/learn` command allows on-demand distillation of the current session through the same proposal-and-validation chain
GitButler is a version control tool designed to optimize workflows for developers and AI agents by introducing features like stacked branches, parallel branching, and unlimited undo capabilities. It layers seamlessly onto existing Git repositories without requiring new configuration, aiming to reduce the friction often associated with complex Git operations through structured output and idempotent commands.
- Agents are reported to be 60% faster using GitButler compared to vanilla Git
- The software includes a CLI that offers JSON output mode for better AI parsing
- Features include "Smartlog" and simplified history editing/rebasing
- It is free and open source software
Gregory Gibson writes about applying the Linux kernel's tool-generated content guidance to his vibe-coded Morse decoder app, CW Inspector. He used OpenAI Codex to build the app, then subjected it to the kernel's five transparency rules: name the tool, preserve inputs, keep the prompt trail, record exactly what the tool changed, and make the generated code prove itself. The process exposed a critical flaw — the decoder confidently produced a wrong answer (reading TTT E as SE) with zero errors, demonstrating that a clean run doesn't guarantee correctness.
- The original algorithm treated the shortest 55% of keyed pulses as dots, which broke on dash-heavy messages; the fix looks for a large ratio between short and long pulse clusters before recording dot duration.
- The author deliberately withheld the actual WAV file from Codex, providing only metadata and the expected message, to create an independent acceptance test rather than letting the model optimize against the exact sample it would later decode.
- The kernel's guidance applies only when a tool generates something substantial (functions, files, fixes, translations), not for spelling corrections, autocomplete, or variable renames.
@omarsar0 writes on X that the fastest path to genuinely understanding agent harnesses is to build one from scratch in TypeScript or Python, starting with a minimal ReAct implementation prompted from Google's original paper, targeting three clean components—an LLM inference module (multi-model, OpenRouter-backed, with separable system prompt), an MCP tools module for interoperability, and a simple agent loop that ties them together—then logging every input/output at each boundary and iterating against a small set of diverse test tasks so each change is inspectable. The punchline: skip the framework first, because only once you've felt the loop, the tokens, and the tool calls in your own code do the "next steps"—memory, skills, subagents—stop being black boxes you configure and become modules you actually know how to tune.
- LLM module: wraps inference across multiple frontier models via OpenRouter; system prompt either embedded or isolated for context-engineering experiments
- Tools module: implement as MCP (Model Context Protocol) tools for cross-harness interoperability, or as bespoke functions if experienced
- Agent loop: ReAct pattern (alternating reasoning traces and action calls) encapsulating both LLM and tools; exit conditions handled via system-prompt instructions (non-deterministic), code-level checks (deterministic), or both
- Logging strategy: capture loop in/out, every LLM call in/out, and every tool-call in/out; run a fixed diverse task suite after each modification
- Scaling path: keep architecture modular so memory, skills, and subagent orchestration can be bolted on once the core loop is understood
- Shortcut alternatives (if not building from scratch): Pi SDK or LangChain harness tooling
mini-swe-agent is a radically simple Python-based agent from the Princeton and Stanford team behind SWE-bench that uses only bash as its tool, maintains a completely linear message history, and executes each action via independent subprocess.run calls. Despite being roughly 100 lines of core agent code, it scores over 74% on SWE-bench verified and is used by organizations including Meta, NVIDIA, IBM, and Anyscale.
- The core design argument is that as language models grow more capable, elaborate tool scaffolds become unnecessary and the LM itself should drive the shell
- Supports sandboxed deployment via docker, podman, singularity, bwrap, and others; installable from PyPI via uvx, pipx, or pip
Anurag Singh writes that providing Claude Code with read-only access to a SaaS application's server logs allowed the coding agent to identify and propose fixes for real performance issues. By observing error patterns, traces, and metrics directly within the environment rather than relying on manual bug reports, the agent was able to autonomously trace bugs back to specific lines of code across various files.
- The experiment highlights a shift toward AI agents joining the "on-call" workflow by inspecting live operational telemetry.
- To mitigate security risks, it is recommended using Model Context Protocol (MCP) servers to restrict an agent's tools to read-only actions.
- Major observability companies like Sentry and Datadog are already implementing similar features to automate root cause analysis and pull request generation.
Yuhao Wu writes about HarnessDev, a benchmark that evaluates LLMs' ability to build and iteratively improve their own agent harness—the model-external execution infrastructure that wraps a model and shapes its task performance. The benchmark has two stages: Creation, where the agent builds a complete execution system from a minimal seed and a few cases, and Evolution, where it revises its own harness using downstream execution feedback. Generated harnesses substantially lag behind mature human-engineered references on code and search/research, while matching or exceeding them on writing and machine-learning experimentation, with large variation in execution cost.
- Covers six creator LLMs across four domains and five downstream benchmarks (2,207 unique instances).
- Hidden evaluation tasks are withheld from development to prevent overfitting.
- Evolution gains are unstable and transfer only partially to held-out tasks.
- Performance gains depend strongly on which model executes the harness, indicating limited cross-model transfer.
NPC-Worldwide (primary contributor cagostino) presents npcsh, a composable multi-agent shell that interprets both bash commands and natural language within a single interactive interface. Built primarily in Rust with a Python backend (npcpy) for the LLM inference loop, it lets users delegate tasks to named agents, define custom "Jinxes" (Jinja Execution templates) for tool-use and skills, and works with any model provider LiteLLM supports. A 100-task benchmark suite scores how well various models can drive the shell, with results ranging from 23% (Qwen3.5 0.8b) to 97% (Qwen3.5 35b, Ornith 35b, Kimi K2.7-Code 1t).
- Agent definitions support three interchangeable formats: .npc YAML files, agents.md markdown, and agents/ directories with per-agent .md files
- The Python backend (npcpy) is explicitly temporary and slated for replacement by a Rust-native runner (npcrs)
- The project references an arxiv paper on "ALARA for Agents: Least-Privilege Context Engineering Through Portable Composable Multi-Agent Teams"
- Supports local model runtimes including Ollama, LM Studio, and MLX (Apple Silicon)
- Currently at v2.1.16 with 128 releases, 473 stars, and 854 commits