Jiahe Geng writes about RSM-full, an online clustered-memory pipeline designed to optimize the quality-to-token trade-off for long-horizon LLM deployments with limited prompt budgets. By utilizing a cosine-gated max-member merge rule and atom-aware grouped context packing, the method achieves significant performance gains in compact-memory regimes compared to existing baselines like Online K-Means and A-MEM. The approach is particularly effective when maintaining an answer quality of 83% for Full-Context level tasks while utilizing only 32% of the total token cost within a 4k budget.
- RSM-full outperforms Streaming-Proto by +2.97 percentage points on the RealMem benchmark.
- The performance gain is primarily driven by the merge rule and grouped packing rather than just flat concatenation or simple clustering.
- The method reaches its optimal utility in the 2k to 5k prompt token range.
OpenHuman is an open-source agent harness designed as a personal AI super intelligence, featuring local-first memory through Markdown trees in SQLite and orchestration capabilities via checkpointed graphs. It functions as a brain that builds persistent context from various data sources like email and calendars, acting as both an orchestrator for multi-agent workflows and a deep researcher with built-in web search and media generation tools.
- Features "Memory Trees" stored locally in Markdown format to create a Karpathy-style Obsidian wiki.
- Provides end-to-end encrypted agent-to-agent messaging using the Signal protocol.
- Supports visual, trigger-driven workflows that can be proposed by an AI and reviewed on a canvas.
- Includes a "Privacy Mode" which ensures no inference data leaves the user's machine when toggled.
The OpenAI Agents API architecture consists of three primary components: the harness, which is a hosted Codex instance that manages model loops and sessions; the environment, where compute or file operations occur via sandboxes or local infrastructure; and the application server, which acts as the bridge between the user's product and the agent. Depending on requirements, an environment can be non-existent (using only external tools), OpenAI-hosted in a managed sandbox, or self-hosted on private infrastructure through an executor connection.
- Users can use "none" for `environment.type` if agents only need to call external services via function tools without local compute.
- Self-hosted environments require the developer to manage provisioning, reconnection, and shutdown of the lifecycle.
- Progress can be tracked using streaming for real-time events or webhooks for asynchronous state changes.
- Managed sandboxes allow developers to pre-configure specific packages, files, and network access levels.
Paul Sawers writes that Amazon Web Services (AWS) has released an open-source application called Pizza Bot, which provides developers with an email-inspired inbox to manage autonomous AI agents. Designed specifically for tasks that continue after a user has finished their session, the tool moves away from chat interfaces toward an asynchronous model where completed jobs arrive as threads and urgent decisions are surfaced for human triage.
- The project is now a standalone community project rather than an AWS service.
- It supports multiple models including Anthropic, Amazon Bedrock, Google Gemini, OpenAI, OpenRouter, or local models via Ollama.
- Built using LangGraph and DeepAgents to enable stateful agent execution and persistence through checkpoints.
- Available as a desktop app for macOS, Windows, and Linux, with browser and terminal clients also available.
Frederic Lardinois writes that Harness field CTO Martin Reynolds is addressing the surge in pull requests caused by coding agents, which can increase new code volume from 1.5x to as much as 50x. To manage this "review bottleneck," Harness has launched a rebuilt Code Repository and an AI Code Review product designed specifically for high-frequency agent traffic rather than just human teams. The company's approach focuses on using a software delivery knowledge graph to provide reviewers with context quickly, helping them distinguish critical code changes from routine dependency updates.
- Coding agents can increase the volume of pull requests by 10x to 50x compared to traditional developer workflows.
- Harness rebuilt its repository service as an "AI-first" platform that is Kubernetes-based and runs across multiple clouds.
- The new AI Code Review tool integrates with existing GitHub repositories, allowing teams to use it without migrating their entire codebase.
Chao Zhang and colleagues study the design space of proactive AI agents that provide higher-level cognitive support during writing, moving beyond simple textual assistance like autocomplete. Through a technology probe deployed with 16 participants, researchers found that users prefer to configure custom partners by setting specific roles and proactivity levels in advance. The findings indicate that such tools can be used for both idea generation and self-monitoring when interventions are presented through lightweight visual representations and non-directive framing to minimize intrusiveness.
- Participants planned their AI support prospectively rather than reacting to real-time interruptions.
- Suggestions served a dual purpose of sparking new ideas and assisting in the self-monitoring process.
- The study emphasizes that rhetorical framing is as critical to user experience as the timing of an intervention.
bex is an open-source, self-hostable PaaS that positions itself as an AI-native alternative to Render, letting developers push Git and receive a deployed URL on their own Kubernetes infrastructure. Coding agents operate as first-class users via MCP alongside the dashboard, CLI, REST, and GraphQL interfaces, all backed by a shared Go core. The platform uses a Kubernetes operator with Cluster API for machine provisioning, supports Render-style `render.yaml` Blueprints for declarative service definitions, and ships managed Postgres, Key Value, logs, metrics, autoscaling, custom domains with TLS, and SSH access.
- 471 stars, 50 forks, 9 contributors — including Claude, Cursor, and Copilot listed as named GitHub contributors
- Apache-2.0 licensed; explicitly marked "not ready for production workloads" (public alpha)
- Language split: Go 58.5%, TypeScript 32.9%, Shell 6.7%
- Internal "lego" Go workspace enforces a strict `operator → types ← backend` one-way dependency DAG
- Tracks Render compatibility via an evidence-backed "parity ledger" (ADR018) rather than marketing claims
- Local quickstart provisions a kind cluster + Cluster API with Docker-container machines as tenant nodes
- Includes an Expo mobile app for safe supervision workflows (App Store listing present)
- Commit history references agent-driven QA rounds (w4/w5/w6 workstreams) and live dashboard re-probes
Cobus Greyling provides a practical pattern library, starter templates, and CLI tools for loop engineering using AI coding agents. This repository aims to help developers design systems that orchestrate agents to discover work, execute tasks, verify results, and persist state—moving beyond simple prompting toward automated agentic workflows.
- Includes the `@cobusgreyling/loop` unified CLI with commands like `init`, `doctor`, `status`, `audit`, and `cost`.
- Offers various patterns such as Daily Triage, PR Babysitter, CI Sweeper, and Dependency Sweeper.
- Features a tiered rollout strategy: L1 (report) $rightarrow$ L2 (assisted) $rightarrow$ L3 (unattended).
- Includes tools for observability like `loop-cost` to estimate token spend and ROI.
Pushpak Chhajed writes about the evolution of project rule systems for AI coding agents, explaining why Laravel Boost moved away from complex semantic search layers in favor of a simple markdown-based approach. To prevent instruction files like `CLAUDE.md` from becoming bloated and consuming excessive context, the team implemented a system using `.ai/rules` containing specific Markdown files linked by a generated two-column index. This "progressive disclosure" method allows agents to efficiently locate relevant project conventions without overwhelming their prompt window or requiring complex vector databases for small rule sets.
- The system uses an automatically updated `index.md` file to help agents map current file paths to specific rule files.
- Agents are encouraged to use a combination of index matching and `grep -rin` to find rules that span multiple directories.
- This approach aligns with advice from the Anthropic Claude Code team regarding progressive disclosure in agentic workflows.
- The solution avoids "staleness" risks associated with maintaining separate vector embeddings for small collections of files.
Meredith Shubel writes that Vercel published `design.md`, a public prompt file that cut agent-generated design failures by 57% across 200+ agent runs, though none of the six tested pages was ship-ready. The system has three layers: a prompt encoding design judgment, a public stylesheet for mechanical layout rules, and an evaluation loop that converts human feedback into deterministic checks. A Slack-based agent (`design-agent`) consolidates weekly feedback from GitHub and Figma into proposed guidance updates.
- The comparison test used Codex with GPT-5.5: 39 failure instances with `design.md` versus 91 without.
- Vercel's first attempt to port its internal "product design" skill to a public prompt failed because subjective design language was interpreted differently by each model.
- Recurring complaint counts are tracked over time; if a fix doesn't reduce its count, the fix is flagged for refinement.