CodeAF is an open-source software factory designed for open models, aiming to optimize cost and efficiency in agentic coding workflows. Unlike traditional AI copilots that assist with line-by-line typing, CodeAF operates as a single terminal interface where users can delegate complex tasks, manage multiple projects simultaneously via subharnesses, and supervise autonomous "crews" of specialized agents (worker, planner, and checker). It is built in Go to be a lightweight, high-performance binary that supports various providers like OpenRouter, DeepSeek, Ollama, and Codex.
- Ranked #1 on the DeepSWE benchmark for cost efficiency per solved issue.
- Uses "Pareto Crewing" to automatically select different models for planning, working, and checking tasks based on performance/cost profiles.
- Supports a headless mode (`codeaf do`) designed specifically for CI/CD pipelines and automated workflows.
- Features a remote execution capability that allows users to drive the engine from any terminal or mobile device via SSH without network latency in UI rendering.
Daniel Furman and colleagues argue that traditional model routers are limited because a router is inherently less capable than the LLM it selects; instead, Replit Agent empowers the core model to act as its own orchestrator. By providing composable primitives—such as domain-aware subagents with varying effort levels and reusable context—the agent can decide when to delegate tasks, how much computational effort to apply, and which specialist models to invoke. This approach moves away from rigid human-designed scaffolding toward a system that leverages the emergent reasoning capabilities of frontier models like GPT-6 Astra and Claude Fable 5.1.
- Replit Agent outperformed sidekick architectures by up to 16 points on benchmarks like DeepSWE and Terminal-Bench while maintaining better cost efficiency.
- "Claudish" refers to a distinct, jargon-heavy prose register used by certain coding agents that allows for model identification from text alone.
- Frontier models have shown an emergent tendency in production to naturally delegate work to specialized subagents without explicit prompting.
TianyuCodings writes about JevHarness, a system where an LLM authors a task-specific harness for Jev (a lightweight judgment model from TypeSafe). The harness converts task observations into features, constructs Jev questions and criteria, and combines structured answers into actions. Once the harness is frozen, execution runs only the harness code and Jev calls without the authoring LLM on every decision. An optional GEPA evolution loop refines the harness using rewards and complete execution traces.
- Pokemon eval: 5 reflection rounds improved win rate from 25% (3/12) to 75% (9/12); the round-3 candidate was selected as best
- Selected harness median full decision: 568 ms (P95 657 ms); individual Jev request median: 269 ms (P95 348 ms)
- Installable as a Claude Code plugin or Codex skill; the skill guides the agent to clarify task inputs, legal actions, and success criteria before building
- Requires Python 3.11+; functional Python nodes need a macOS native sandbox and fail closed if unavailable
- The archived website and recorded-call inspection require no model credentials
@omarsar0 writes on X that the fastest path to genuinely understanding agent harnesses is to build one from scratch in TypeScript or Python, starting with a minimal ReAct implementation prompted from Google's original paper, targeting three clean components—an LLM inference module (multi-model, OpenRouter-backed, with separable system prompt), an MCP tools module for interoperability, and a simple agent loop that ties them together—then logging every input/output at each boundary and iterating against a small set of diverse test tasks so each change is inspectable. The punchline: skip the framework first, because only once you've felt the loop, the tokens, and the tool calls in your own code do the "next steps"—memory, skills, subagents—stop being black boxes you configure and become modules you actually know how to tune.
- LLM module: wraps inference across multiple frontier models via OpenRouter; system prompt either embedded or isolated for context-engineering experiments
- Tools module: implement as MCP (Model Context Protocol) tools for cross-harness interoperability, or as bespoke functions if experienced
- Agent loop: ReAct pattern (alternating reasoning traces and action calls) encapsulating both LLM and tools; exit conditions handled via system-prompt instructions (non-deterministic), code-level checks (deterministic), or both
- Logging strategy: capture loop in/out, every LLM call in/out, and every tool-call in/out; run a fixed diverse task suite after each modification
- Scaling path: keep architecture modular so memory, skills, and subagent orchestration can be bolted on once the core loop is understood
- Shortcut alternatives (if not building from scratch): Pi SDK or LangChain harness tooling
The OpenAI Agents API architecture consists of three primary components: the harness, which is a hosted Codex instance that manages model loops and sessions; the environment, where compute or file operations occur via sandboxes or local infrastructure; and the application server, which acts as the bridge between the user's product and the agent. Depending on requirements, an environment can be non-existent (using only external tools), OpenAI-hosted in a managed sandbox, or self-hosted on private infrastructure through an executor connection.
- Users can use "none" for `environment.type` if agents only need to call external services via function tools without local compute.
- Self-hosted environments require the developer to manage provisioning, reconnection, and shutdown of the lifecycle.
- Progress can be tracked using streaming for real-time events or webhooks for asynchronous state changes.
- Managed sandboxes allow developers to pre-configure specific packages, files, and network access levels.
Frederic Lardinois writes that Harness field CTO Martin Reynolds is addressing the surge in pull requests caused by coding agents, which can increase new code volume from 1.5x to as much as 50x. To manage this "review bottleneck," Harness has launched a rebuilt Code Repository and an AI Code Review product designed specifically for high-frequency agent traffic rather than just human teams. The company's approach focuses on using a software delivery knowledge graph to provide reviewers with context quickly, helping them distinguish critical code changes from routine dependency updates.
- Coding agents can increase the volume of pull requests by 10x to 50x compared to traditional developer workflows.
- Harness rebuilt its repository service as an "AI-first" platform that is Kubernetes-based and runs across multiple clouds.
- The new AI Code Review tool integrates with existing GitHub repositories, allowing teams to use it without migrating their entire codebase.
/u/locbuilds on r/LocalLLM gives advice for an issue where the Qwen 3.8-27b model enters repetitive loops when making tool calls during debugging sessions. Community members suggest that this is often a bug within the agent harness rather than the model itself, recommending several technical mitigations to manage these failures effectively.
- Implement hard loop breakers in the application harness to detect and stop identical consecutive tool calls.
- Provide explicit "error" or "already tried" feedback in tool observations to signal failure back to the model.
- Lower temperature (0.1–0.3) for tool-heavy turns and apply repetition penalties via the sampler.
- Use specialized chat templates, such as Froggeric's Qwen fixed template, which may alleviate looping issues.
Cobus Greyling provides a practical pattern library, starter templates, and CLI tools for loop engineering using AI coding agents. This repository aims to help developers design systems that orchestrate agents to discover work, execute tasks, verify results, and persist state—moving beyond simple prompting toward automated agentic workflows.
- Includes the `@cobusgreyling/loop` unified CLI with commands like `init`, `doctor`, `status`, `audit`, and `cost`.
- Offers various patterns such as Daily Triage, PR Babysitter, CI Sweeper, and Dependency Sweeper.
- Features a tiered rollout strategy: L1 (report) $rightarrow$ L2 (assisted) $rightarrow$ L3 (unattended).
- Includes tools for observability like `loop-cost` to estimate token spend and ROI.
- **Inference** – Platforms and engines for running models, plus user interfaces.
- **Models** – LLMs (general, coding, multimodal, image, audio), model providers, and specific model highlights.
- **RAG** – Retrieval-Augmented Generation tools.
- **Safeguards** – Safety and content filtering.
- **Agents & Tools** – Agent frameworks, Model Context Protocol, coding agents, computer/browser automation, memory management, and testing/evaluation.
- **Research, Training & Fine-tuning** – Security, sandboxing, and model development.
- **Hardware** – Local hardware options.
- **Tutorials** – Guides covering models, prompt/context engineering, inference, agents, and RAG.
- **Communities** – Places to connect and share knowledge.
Borui Kang presents Harness Continual Learning (HCL), where an agent's harness (prompts, memories, tools, skills, routing rules) evolves around a frozen foundation model, unlike updating model parameters. The paper defines "harness-level forgetting" as losing reliable behavior due to harness updates and proposes a guarded evolution mechanism. A Continual Optimizer generates candidate harnesses from feedback, and a Continual Evaluator commits changes only after verifying improvement, retention, and validity. Experiments in textual reasoning, multimodal perception, and open-world interaction show capability accumulation and failure recovery, with over 10% relative gains versus baselines.
- Four execution-facing harness components: Task Interface, Experience Memory, Capability Map, and Adaptive Router.
- Controlled retention sweeps show the stability''-plasticity trade-off can be explicitly adjusted at the harness level.
- The work reframes continual learning away from parameter updates toward externalized, inspectable agent state.