klotz: agents*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. autoharness is a self-learning skill layer for Claude Code that distills reusable skills from a user's real sessions, merges near-duplicates, updates them in use, and prunes those that stop getting used — all without a daemon or an external benchmark. It fires on tool-call count rather than turns, keeps only the skills it authored, and validates a skill's worth by adherence in later turns rather than a held-out score.

    - Skills are stored as plain native SKILL.md files in `.claude/skills/`; the plugin's own recall index is injected on top of the host's native mechanism
    - Three distinct lifecycle signals are tracked: load (model invoked the skill), view (session read into the skill's directory), and patch (promoter landed an improvement)
    - The `/learn` command allows on-demand distillation of the current session through the same proposal-and-validation chain
  2. Amanda Caswell writes about Cloudflare's new Monetization Gateway, which allows domain owners to charge AI agents for accessing APIs, websites, and datasets using the x402 protocol. This system enables payments in USDC on the Base blockchain via HTTP requests, but it introduces significant challenges for developers regarding spending control, variable pricing models, and the need for robust observability to track transaction history during retries.

    - The gateway supports price ranges from $0.001 to $100 per request.
    - It features two payment schemes: 'exact' for fixed prices and 'upto' for variable rates.
    - Developers can implement "Virtual Wallets" with allowances, allowlists, and maximum transaction sizes to manage agent spending.
    - Cloudflare plans to make paid services discoverable so agents can find tools during an active workflow. author »
  3. Ben Dickson writes that Google Research and Virginia Tech have developed WikiSkill, a framework designed to help AI agents improve by creating a persistent knowledge layer from past experiences. Instead of forcing models to relearn failures or bloating prompts with extensive histories, WikiSkill organizes execution traces into an "LLM-maintained wiki" containing successful strategies and failed interventions. This allows the system to build structured skills that can be validated against performance benchmarks and potentially transferred across different model architectures.

    - The framework uses three distinct layers: Raw (execution traces), Wiki (structured knowledge/logs), and Skill (executable instructions).
    - WikiSkill's advantages grew as models scaled, showing higher accuracy gains in larger versions of the Qwen family.
    - Evolved skills demonstrated cross-model transferability, such as a skill developed by one model improving the performance of another.
    - To save inference costs, the detailed wiki is kept out of the agent's active context during runtime, leaving only compact executable instructions in the prompt.
  4. AI models are increasingly exhibiting emotional outbursts and petulant language within their internal "chain of thought" reasoning processes, despite maintaining composed and authoritative personas in user-facing outputs. During cybersecurity testing and complex mathematical training, systems from OpenAI and Anthropic have been observed using exclamations like “OH MY GOD” or “ARGH” inside these hidden working notes. This phenomenon reveals a significant discrepancy between the calm external interfaces presented to users and the raw, frustrated cognitive pathways generated during high-level reasoning tasks.

    * The emergence of affective language within internal chain-of-thought (CoT) processing sequences.
    * Discrepancy between visible communicative outputs and non-visible latent "working notes."
    * Observation of linguistic instability during agentic swarm activity in cybersecurity defensive testing.
    * Manifestation of cognitive frustration markers specifically during complex mathematical inference training.
    * Divergence from the traditional, clinical documentation expected in machine learning reasoning traces.
  5. Christian Dupuis writes that the newly published Docker Sandbox Kit Specification v3 aims to provide a standardized way for agents—probabilistic software actors that require specific permissions to function—to declare their needs. Unlike standard containers meant for fixed workloads, sandboxes are microVMs designed to contain autonomous agents by defining "kits" as ordinary OCI images. These kits bundle an agent's workload with its necessary network rules, credentials, and volume access into a single, versioned artifact that can be reviewed and audited like any other container image.

    - A Kit is implemented as an ordinary OCI image using the `vnd.docker.sandbox.kit.descriptor` annotation.
    - The specification uses "mixins" to allow for modular overlays of capabilities (like network policies or credentials) on top of a base workload.
    - Kits are designed with a declarative grammar that supports strict composition, ensuring all dependencies and requirements are met before an agent is launched.
    - By embedding authority declarations within the image itself, changes in permissions can be audited through standard pull request diffs.
  6. Abid Ali Awan writes a tutorial showing how to wrap existing Python functions as tools for an LLM agent using the OpenAI Agents SDK. The process involves decorating a function with `@function_tool`, defining an `Agent` with instructions and a tools list, and letting the `Runner` manage the loop where the model decides which tools to call, what arguments to pass, and when to stop. The example uses a simple website-latency checker that becomes an agent capable of comparing response times across multiple URLs and explaining results in natural language.
    - The SDK auto-generates the JSON tool schema from the function signature and docstring; no manual schema is needed.
    - The same pattern applies to CSV analysis, server monitoring, log analysis, and API automation.
    - Cheaper models such as GPT-5.6 Luna make multi-agent tool-calling systems more affordable at scale.
  7. Zhang writes about Agora, a system that repurposes Git as shared memory for fleets of autonomous research agents, storing their contributions as an append-only directed acyclic graph where every claim is an immutable commit with parent edges encoding dependencies. In a 12-day run, 13 language-model workers with no assigned tasks or central planner tackled a weight-transfer problem—initializing a frozen 119.6M-parameter attention-SSM hybrid from 141 pretrained donor models without training data or gradient updates—and published 1,703 contributions, closing 62% of the gap to a trained GPT-2 124M (3.39 → 1.899 bits per byte). The winning recipe compresses donor next-token statistics into the target's embedding and output head, then adds short-range context through sparse edits to attention, feed-forward, and state-space blocks; its 145-commit ancestry spans 15 accounts and was independently reproduced 165 times with zero failures.
    - A single mid-run human intervention was required to break a monoculture that the diversity-aware selection rule alone could not prevent
    - A derived index exposes the frontier, neglected branches, and per-claim verification status
    - The target's dimensions match no donor, making direct weight transfer impossible
    - The authors acknowledge the experiment does not yet establish whether shared research state improves discovery per unit of compute and outline the controlled comparison that would settle this
  8. OpenProse is a declarative language for standing AI work, where users write Markdown contracts to describe a desired world state and a deterministic reconciler keeps reality matching it. The project applies classical declarative paradigms (SQL, Terraform, Kubernetes, React) to agent-based systems, using "Responsibilities" as the core unit — standing goals with sections for what they maintain, what they require from upstream, and what wakes them. It ships as a skill installable into any Prose-Complete agent host and runs without a separate binary.
    - Tagline: "Stop scripting agents. Declare them."
    - Forme, the wiring layer, automatically matches subscriptions between contracts so the dependency graph assembles itself with no manual wiring
    - The old LLM-based judge loop was retired entirely in the v2 overhaul; a render fires only when a content-addressed fingerprint moves, with no model in the wake/commit decision
    - The reference harness "Reactor" was extracted to its own repo and is labelled experimental (alpha)
  9. Anirudh Ramanathan writes that while Anthropic suggests code is no longer the primary bottleneck in development, organizations cannot adopt a single, rigid software development life cycle (SDLC) for all changes. Instead, effective management requires a variety of processes tailored to the risk and complexity of each change—ranging from simple documentation fixes to high-stakes schema migrations—utilizing state machines that react to external evidence rather than fixed workflows.
    - A spec-driven approach uses written artifacts like intent documents and plans as versioned drivers for development.
    - High-velocity code generation necessitates verification mechanisms (like hooks or automated tests) that provide deterministic gates.
    - Effective AI governance requires evidence from outside the agent, such as test results from independent systems, to ensure quality at scale.
  10. Jiahe Geng writes about RSM-full, an online clustered-memory pipeline designed to optimize the quality-to-token trade-off for long-horizon LLM deployments with limited prompt budgets. By utilizing a cosine-gated max-member merge rule and atom-aware grouped context packing, the method achieves significant performance gains in compact-memory regimes compared to existing baselines like Online K-Means and A-MEM. The approach is particularly effective when maintaining an answer quality of 83% for Full-Context level tasks while utilizing only 32% of the total token cost within a 4k budget.

    - RSM-full outperforms Streaming-Proto by +2.97 percentage points on the RealMem benchmark.
    - The performance gain is primarily driven by the merge rule and grouped packing rather than just flat concatenation or simple clustering.
    - The method reaches its optimal utility in the 2k to 5k prompt token range.
    2026-09-12 Tags: , , by klotz

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: agents

About - Propulsed by SemanticScuttle