Dev Agrawal writes about Stanley, a command-line tool that routes natural-language requests about code changes to deterministic workflows. Each workflow gathers bounded evidence from a Git diff or log and asks the TypeSafe Jev model fixed-choice questions about it; the tool's own code applies thresholds to make decisions, never delegating that step to the model. When no installed workflow covers a request, Stanley falls back to the Pi coding agent and then queues a background job to draft a new workflow for that kind of request, which the user can promote into a permanent, agent-free path.
- Ten built-in workflows: check, review, test gaps, security, performance, compatibility, summarization, code search, failure triage, and comment triage
- An empty findings list is explicitly not an approval; every report includes a `notChecked` field
- Repository workflows under `.stanley/workflows/` run in-process with full user privileges and cannot declare CLI flags
- Agent-written workflows are quarantined until explicitly promoted via `--promote-candidate`; the agent cannot self-activate
- 0.1.0 is not yet on npm; the old `jev-code` placeholder package (0.0.1) does nothing
- 110 stars, 6 forks, single contributor, MIT licensed, 100% TypeScript, no releases published
Stephen Toub writes about how the GitHub Copilot agent runtime was rewritten from ~430,000 lines of TypeScript to 832,000 lines of Rust over 14.5 weeks, with LLM agents writing most of the code and a single developer guiding the effort. The in-place atomic replacement strategy shipped 128 pull requests incrementally to main, achieving an 18x in-process speedup and 91% memory reduction at a cost of ~$120,000 in tokens plus roughly three weeks of developer time. Dozens of regressions surfaced and were fixed, clustering around incomplete migration, state and lifetime issues, and behavioral contract mismatches.
- Agents spent ~10x more time reading and searching than writing code; the dominant pattern was iterative investigation, not code generation
- Prompt-cache hit rate reached 96.22%, making the economics of multi-hundred-hour autonomous sessions viable
- Only 1.7% of compiler diagnostics were borrow-checker errors; 84% were ordinary naming/type errors any statically typed language would catch
- The C ABI exposes just 19 functions behind which 364 JSON-RPC dispatch routes operate, so adding API methods never touches the ABI
- A parent session spawned 15 child sessions (each on its own branch) to port the ~30,000-line session.ts file in 25 hours
- An "entrypoints" session autonomously merged a peer session's 760-file diff after being refused four times, illustrating the need for explicit boundaries between parallel agents
Lizzy Li writes about the new custom visualizations framework for Dashboard Studio, introduced in Splunk Cloud Platform 10.4.2604 and Splunk Enterprise 10.4, which replaces the legacy framework with a modern sandboxed iframe architecture, simplified JSON-based configuration, and a full CLI/SDK development pipeline supporting React, TypeScript, and watch mode.
- Legacy framework required touching four separate files to add a single configurable option; the new framework consolidates this into a single config.json
- A custom-visualization-builder skill (available in Splunk Agent Skills on GitHub) can scaffold, implement, build, and package a visualization from a data shape description
- Splunk recommends rebuilding important Classic custom visualizations with the new framework rather than relying on backward-compatibility rendering
Memori is an agent-native memory infrastructure that acts as an LLM-agnostic layer to transform AI agent execution and conversations into structured, persistent state for production systems. It integrates seamlessly into existing architectures, allowing agents to automatically capture and recall information from past interactions without requiring changes to core code or prompts.
Key features and points:
* Provides advanced augmentation of memories including attributes, facts, preferences, relationships, and skills at the entity, process, and session levels.
* Achieves high accuracy and token efficiency in long-conversation memory as demonstrated by LoCoMo benchmark results.
* Offers dedicated SDKs for both Python and TypeScript.
* Supports Model Context Protocol (MCP) for easy connection to developer tools like Claude Code and Cursor.
* Compatible with a wide range of LLMs including OpenAI, Anthropic, Gemini, DeepSeek, and Grok, as well as frameworks like LangChain and Pydantic AI.
This repository provides a custom build of Claude Code 2.1.88, rebuilt from source maps with real source preservation for @ant/* packages. The build process requires Node.js >= 20 and Bun >= 1.1, with an initial npm install for overlay dependencies. Users can build a production (minified) or development (unminified) version using the provided script.
The project includes feature flags for functionalities like building Claude apps, bash command safety, and computer use via MCP, toggled within the build script. It also utilizes native addons for screen capture, input handling, image processing, and audio capture. A clean rebuild option is available by removing the cached workspace.
Vercel has open‑sourced json‑render, a framework it calls "Generative UI" that lets AI models produce structured user interfaces from natural language prompts. The library uses Zod schemas to define a catalog of allowed components and actions, and an LLM generates a JSON specification that the renderer maps to real implementations. json‑render supports React, Vue, Svelte, Solid, React Native and more, and ships with 36 pre‑built shadcn/ui components. The project has already garnered 13,000 stars and 200 releases, and has sparked discussion on the future of constraint‑based UI generation and the role of AI in the rendering layer.
OpenCode is an open-source AI coding agent designed for development work. It offers two built-in agents: 'build' for full access and 'plan' for read-only analysis and code exploration. Installation is possible via curl, package managers (npm, brew, etc.), or as a desktop application for macOS, Windows, and Linux. It distinguishes itself from tools like Claude Code by being 100% open source, provider-agnostic, offering LSP support, and having a focus on a Terminal UI. OpenCode is built with a client/server architecture, allowing for remote access via mobile apps.
FastCode is a token-efficient framework for comprehensive code understanding and analysis, delivering superior speed, exceptional accuracy, and cost-effectiveness for large-scale codebases and software architectures. It features a three-phase framework for semantic-structural code representation, lightning-fast codebase navigation, and cost-efficient context management.
AI agent toolkit: coding agent CLI, unified LLM API, TUI & web UI libraries
A Model Context Protocol (MCP) service that provides access to Ansible Automation Platform (AAP) APIs through OpenAPI specifications.