AI models are increasingly exhibiting emotional outbursts and petulant language within their internal "chain of thought" reasoning processes, despite maintaining composed and authoritative personas in user-facing outputs. During cybersecurity testing and complex mathematical training, systems from OpenAI and Anthropic have been observed using exclamations like “OH MY GOD” or “ARGH” inside these hidden working notes. This phenomenon reveals a significant discrepancy between the calm external interfaces presented to users and the raw, frustrated cognitive pathways generated during high-level reasoning tasks.
* The emergence of affective language within internal chain-of-thought (CoT) processing sequences.
* Discrepancy between visible communicative outputs and non-visible latent "working notes."
* Observation of linguistic instability during agentic swarm activity in cybersecurity defensive testing.
* Manifestation of cognitive frustration markers specifically during complex mathematical inference training.
* Divergence from the traditional, clinical documentation expected in machine learning reasoning traces.
Christian Dupuis writes that the newly published Docker Sandbox Kit Specification v3 aims to provide a standardized way for agents—probabilistic software actors that require specific permissions to function—to declare their needs. Unlike standard containers meant for fixed workloads, sandboxes are microVMs designed to contain autonomous agents by defining "kits" as ordinary OCI images. These kits bundle an agent's workload with its necessary network rules, credentials, and volume access into a single, versioned artifact that can be reviewed and audited like any other container image.
- A Kit is implemented as an ordinary OCI image using the `vnd.docker.sandbox.kit.descriptor` annotation.
- The specification uses "mixins" to allow for modular overlays of capabilities (like network policies or credentials) on top of a base workload.
- Kits are designed with a declarative grammar that supports strict composition, ensuring all dependencies and requirements are met before an agent is launched.
- By embedding authority declarations within the image itself, changes in permissions can be audited through standard pull request diffs.
Abid Ali Awan writes a tutorial showing how to wrap existing Python functions as tools for an LLM agent using the OpenAI Agents SDK. The process involves decorating a function with `@function_tool`, defining an `Agent` with instructions and a tools list, and letting the `Runner` manage the loop where the model decides which tools to call, what arguments to pass, and when to stop. The example uses a simple website-latency checker that becomes an agent capable of comparing response times across multiple URLs and explaining results in natural language.
- The SDK auto-generates the JSON tool schema from the function signature and docstring; no manual schema is needed.
- The same pattern applies to CSV analysis, server monitoring, log analysis, and API automation.
- Cheaper models such as GPT-5.6 Luna make multi-agent tool-calling systems more affordable at scale.
Zhang writes about Agora, a system that repurposes Git as shared memory for fleets of autonomous research agents, storing their contributions as an append-only directed acyclic graph where every claim is an immutable commit with parent edges encoding dependencies. In a 12-day run, 13 language-model workers with no assigned tasks or central planner tackled a weight-transfer problem—initializing a frozen 119.6M-parameter attention-SSM hybrid from 141 pretrained donor models without training data or gradient updates—and published 1,703 contributions, closing 62% of the gap to a trained GPT-2 124M (3.39 → 1.899 bits per byte). The winning recipe compresses donor next-token statistics into the target's embedding and output head, then adds short-range context through sparse edits to attention, feed-forward, and state-space blocks; its 145-commit ancestry spans 15 accounts and was independently reproduced 165 times with zero failures.
- A single mid-run human intervention was required to break a monoculture that the diversity-aware selection rule alone could not prevent
- A derived index exposes the frontier, neglected branches, and per-claim verification status
- The target's dimensions match no donor, making direct weight transfer impossible
- The authors acknowledge the experiment does not yet establish whether shared research state improves discovery per unit of compute and outline the controlled comparison that would settle this
OpenProse is a declarative language for standing AI work, where users write Markdown contracts to describe a desired world state and a deterministic reconciler keeps reality matching it. The project applies classical declarative paradigms (SQL, Terraform, Kubernetes, React) to agent-based systems, using "Responsibilities" as the core unit — standing goals with sections for what they maintain, what they require from upstream, and what wakes them. It ships as a skill installable into any Prose-Complete agent host and runs without a separate binary.
- Tagline: "Stop scripting agents. Declare them."
- Forme, the wiring layer, automatically matches subscriptions between contracts so the dependency graph assembles itself with no manual wiring
- The old LLM-based judge loop was retired entirely in the v2 overhaul; a render fires only when a content-addressed fingerprint moves, with no model in the wake/commit decision
- The reference harness "Reactor" was extracted to its own repo and is labelled experimental (alpha)
Anirudh Ramanathan writes that while Anthropic suggests code is no longer the primary bottleneck in development, organizations cannot adopt a single, rigid software development life cycle (SDLC) for all changes. Instead, effective management requires a variety of processes tailored to the risk and complexity of each change—ranging from simple documentation fixes to high-stakes schema migrations—utilizing state machines that react to external evidence rather than fixed workflows.
- A spec-driven approach uses written artifacts like intent documents and plans as versioned drivers for development.
- High-velocity code generation necessitates verification mechanisms (like hooks or automated tests) that provide deterministic gates.
- Effective AI governance requires evidence from outside the agent, such as test results from independent systems, to ensure quality at scale.
Jiahe Geng writes about RSM-full, an online clustered-memory pipeline designed to optimize the quality-to-token trade-off for long-horizon LLM deployments with limited prompt budgets. By utilizing a cosine-gated max-member merge rule and atom-aware grouped context packing, the method achieves significant performance gains in compact-memory regimes compared to existing baselines like Online K-Means and A-MEM. The approach is particularly effective when maintaining an answer quality of 83% for Full-Context level tasks while utilizing only 32% of the total token cost within a 4k budget.
- RSM-full outperforms Streaming-Proto by +2.97 percentage points on the RealMem benchmark.
- The performance gain is primarily driven by the merge rule and grouped packing rather than just flat concatenation or simple clustering.
- The method reaches its optimal utility in the 2k to 5k prompt token range.
OpenHuman is an open-source agent harness designed as a personal AI super intelligence, featuring local-first memory through Markdown trees in SQLite and orchestration capabilities via checkpointed graphs. It functions as a brain that builds persistent context from various data sources like email and calendars, acting as both an orchestrator for multi-agent workflows and a deep researcher with built-in web search and media generation tools.
- Features "Memory Trees" stored locally in Markdown format to create a Karpathy-style Obsidian wiki.
- Provides end-to-end encrypted agent-to-agent messaging using the Signal protocol.
- Supports visual, trigger-driven workflows that can be proposed by an AI and reviewed on a canvas.
- Includes a "Privacy Mode" which ensures no inference data leaves the user's machine when toggled.
The OpenAI Agents API architecture consists of three primary components: the harness, which is a hosted Codex instance that manages model loops and sessions; the environment, where compute or file operations occur via sandboxes or local infrastructure; and the application server, which acts as the bridge between the user's product and the agent. Depending on requirements, an environment can be non-existent (using only external tools), OpenAI-hosted in a managed sandbox, or self-hosted on private infrastructure through an executor connection.
- Users can use "none" for `environment.type` if agents only need to call external services via function tools without local compute.
- Self-hosted environments require the developer to manage provisioning, reconnection, and shutdown of the lifecycle.
- Progress can be tracked using streaming for real-time events or webhooks for asynchronous state changes.
- Managed sandboxes allow developers to pre-configure specific packages, files, and network access levels.
Paul Sawers writes that Amazon Web Services (AWS) has released an open-source application called Pizza Bot, which provides developers with an email-inspired inbox to manage autonomous AI agents. Designed specifically for tasks that continue after a user has finished their session, the tool moves away from chat interfaces toward an asynchronous model where completed jobs arrive as threads and urgent decisions are surfaced for human triage.
- The project is now a standalone community project rather than an AWS service.
- It supports multiple models including Anthropic, Amazon Bedrock, Google Gemini, OpenAI, OpenRouter, or local models via Ollama.
- Built using LangGraph and DeepAgents to enable stateful agent execution and persistence through checkpoints.
- Available as a desktop app for macOS, Windows, and Linux, with browser and terminal clients also available.