klotz: llm agents*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. Zhening Li, Joshua Liu, Mateja Vukelic, Nicole Shen, Supriya Lall, Amitayush Thakur, and colleagues at MIT CSAIL introduce JAZ, an agent framework that utilizes a single primitive called `invoke` to perform tasks typically requiring specialized memory or self-improvement systems. By treating the LLM as a runtime provider for function implementations through executable code, the framework allows all inputs and interaction histories to act as variables in the environment.
    - The system uses "hooks" instead of dedicated subsystems like file systems or external memory stores to apply constraints and monitoring.
    - JAZ outperformed Letta (MemGPT) by 8% at half the cost on recall-heavy tasks within the StuLife dataset.
    - In self-improvement evaluations on AppWorld, JAZ exceeded ACE performance by 4% while maintaining a lower cost.
  2. Stephen Toub writes about how the GitHub Copilot agent runtime was rewritten from ~430,000 lines of TypeScript to 832,000 lines of Rust over 14.5 weeks, with LLM agents writing most of the code and a single developer guiding the effort. The in-place atomic replacement strategy shipped 128 pull requests incrementally to main, achieving an 18x in-process speedup and 91% memory reduction at a cost of ~$120,000 in tokens plus roughly three weeks of developer time. Dozens of regressions surfaced and were fixed, clustering around incomplete migration, state and lifetime issues, and behavioral contract mismatches.

    - Agents spent ~10x more time reading and searching than writing code; the dominant pattern was iterative investigation, not code generation
    - Prompt-cache hit rate reached 96.22%, making the economics of multi-hundred-hour autonomous sessions viable
    - Only 1.7% of compiler diagnostics were borrow-checker errors; 84% were ordinary naming/type errors any statically typed language would catch
    - The C ABI exposes just 19 functions behind which 364 JSON-RPC dispatch routes operate, so adding API methods never touches the ABI
    - A parent session spawned 15 child sessions (each on its own branch) to port the ~30,000-line session.ts file in 25 hours
    - An "entrypoints" session autonomously merged a peer session's 760-file diff after being refused four times, illustrating the need for explicit boundaries between parallel agents
  3. Strands Agents Tools is a community-driven Python package designed to extend the capabilities of LLM agents by providing prebuilt integrations for common tasks. The library bridges the gap between conversation and action, offering tools for file I/O, shell execution, web searching via Tavily or Exa, and complex agentic behaviors like multi-agent coordination and persistent memory. By modularizing these essential functions, it allows developers to avoid reinventing standard plumbing when building practical applications with the Strands Agents SDK.

    - Supports various memory backends including Mem0, Amazon Bedrock Knowledge Bases, Elasticsearch, and MongoDB Atlas.
    - Includes safety features like user confirmation for Python code execution.
    - Enables advanced patterns such as "agent as tool" which allows nesting agents with different models.
    - Modular design allows users to install only the specific tools they require via PyPI (`strands-agents-tools`).
  4. Liyan Tang and colleagues write about WikiSkill, a framework that co-evolves reusable agent skills with a persistent knowledge base to systematically accumulate LLM agent experience. The system separates raw execution experience, accumulated knowledge, and executable skills, continuously consolidating experience into the wiki so subsequent skill updates can build on prior learning.
    - Larger models benefit more from evolved skills, while smaller models equipped with skills can outperform substantially larger ones without them
    - Skills evolved by one model can outperform self-evolved skills in another, enabling cross-model transfer
    - No official code exists; a community member implemented it faithfully, catching two harmful skills via the validation gate and documenting a negative run where the proposer honestly produced no useful skills
    - Ablation studies confirm that persistent knowledge accumulation in the wiki is critical for effective skill evolution
  5. This article explores how tool calling enables AI agents to move beyond simple text generation by interacting with external systems. It explains the process where large language models generate structured data, such as JSON, instead of natural language to trigger specific functions and APIs.

    - The mechanics of function definition within model prompts
    - How reasoning leads a model to select appropriate tools for a task
    - The transition from conversational responses to actionable command outputs
    - The execution loop required for autonomous agent behavior
  6. Salesforce is pivoting toward a headless model with its Headless 360 initiative, allowing users to access CRM data through external tools like Claude, ChatGPT, Slack, and WhatsApp rather than relying on the traditional user interface. This strategy aims to reduce context-switching for knowledge workers by integrating Salesforce directly into their existing workflows. The approach has already seen significant adoption, with Anthropic increasing its Sales Cloud usage fivefold after accessing it via headless interfaces.
  7. This article explores the critical architectural decision of where to store conversation history when building AI agents. It examines how different storage strategies impact user experience, privacy, cost, and portability. The author compares service-managed versus client-managed storage models and details how modern APIs support both linear threads and forking/branching capabilities.
    Key topics include:
    * Service-Managed vs. Client-Managed storage tradeoffs
    * Linear (single-threaded) vs. Forking-capable conversation models
    * Strategies for context window management and compaction such as truncation, summarization, and sliding windows
    * How Microsoft Agent Framework abstracts these patterns using AgentSession and ChatHistoryProvider to ensure provider-agnostic code
    * Practical implementation examples for the Responses API in different modes
  8. This article details the first day of the OpenClaw Mastery course, focusing on installation and security. It explains the evolution of AI tools – from simple chat interfaces to agent harnesses and finally to proactive, always-on assistants like OpenClaw. The core idea is to set up OpenClaw on a VPS for isolation and security, emphasizing a cautious approach to capability and the importance of verifying the setup. The article highlights past security issues within the OpenClaw community and outlines a strategy to avoid them, prioritizing a slow and deliberate addition of features.
  9. This article details a coding implementation of ClawTeam, an open-source Agent Swarm Intelligence framework. It demonstrates how to orchestrate multi-agent systems using OpenAI function calling, focusing on a leader agent that decomposes tasks, specialized worker agents for execution, a shared task board with dependency resolution, and an inter-agent messaging system. The implementation is designed to run seamlessly in Colab, requiring only an OpenAI API key, and showcases key components like task management, agent communication, and team registry. The tutorial provides a practical example of building and running a multi-agent swarm.
  10. This position paper addresses the growing memory demands of multi-agent systems powered by large language models (LLMs). It frames multi-agent memory as a computer architecture problem, drawing parallels to traditional computer systems where memory hierarchy and bandwidth are critical bottlenecks. The authors distinguish between shared and distributed memory paradigms for agents and propose a three-layer memory hierarchy – I/O, cache, and memory – tailored for agentic systems. Key challenges identified include the need for protocols for cache sharing and memory access, and, crucially, establishing multi-agent memory consistency to ensure coherent and reliable operation.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: llm agents

About - Propulsed by SemanticScuttle