klotz: model context protocol*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. SandBase Harness is an open-source Node.js runtime that keeps agent sessions, sandboxed tools, memory, credentials, audit trails, and a web Console entirely on the user's own machine. An `init` command scaffolds a workspace with a `config.yaml` pointing to a provider environment variable, and `start` serves both an API and a dashboard on localhost. It supports OpenAI, Anthropic, MiniMax, or any OpenAI-compatible endpoint, and Docker is optional, used only for Docker-backed sandboxes. A CLI `chat` command is available for terminal-based interaction.
    - Tool approval is a first-class concept; by default tool calls park for human approval, though `--tool-approval allow` can preauthorize them.
    - The model ID must match what the provider actually serves (e.g., `deepseek-chat` for DeepSeek), or the turn fails with `model_not_found`.
    - Provider settings saved via the Console only take effect after a restart, a gotcha the documentation calls out explicitly.
    - The project is listed in the Official MCP Registry.
  2. Joseph Dolivo and Igor Fil write about Pizza Bot, an open source application that runs background AI agents and provides an email-like inbox for their outputs. The design pattern moves agent interaction away from live chat sessions toward asynchronous threads, allowing users to delegate tasks that require multiple steps or human approval without having to monitor a terminal in real time. Because it is fully self-hosted and supports local models, it allows users to maintain strict data boundaries while utilizing tools they already own through MCP servers and agent skills.
    - The architecture is built on DeepAgents and LangGraph, using SQLite for local persistence and checkpointing.
    - It features an "Action" queue specifically for tasks paused on human approval or decision-making.
    - Cron schedules run once after a machine wakes from sleep rather than replaying every missed occurrence.
    - The project name references Amazon's two-pizza team methodology.
    - It currently ships with a browser-automation skill and a documentation guide skill.
  3. Thomas Claburn writes that the Pi coding agent hit its 1.0 milestone with a notable reversal on Model Context Protocol (MCP) support, which its creator Mario Zechner previously dismissed as unnecessary. Earendil, the company that acquired Pi in April 2026, attributed the about-face to improvements in MCP and the developers' own changes that made integrating other capabilities easier. The release also introduced Pi Durable, a separate component for orchestrating long-running agentic applications that was split out to preserve Pi's minimalist design.

    - Earendil was formed by Armin Ronacher and Colin Daymond Hanna.
    - The MCP changes also enable easier use of the Jev decision model within Pi.
    - Pi is described as the Flask of agent harnesses (minimal and extensible), contrasted with Django.
    - New features include Codemode (a sandbox for tool calls), virtual model support, deferred tool loading, and cache warming for Anthropic models.
  4. Anmol Baranwal writes a comprehensive guide to AG-UI, an open protocol that standardizes how agents communicate with user-facing applications by replacing custom streaming layers with a unified event stream. The protocol connects agentic backends to frontends across any framework and surface, defining 31 event types grouped into 8 categories covering runs, text, tool calls, reasoning, state, activity, subagents, and custom events. It operates alongside MCP (agent-to-tool) and A2A (agent-to-agent) as the user-facing layer, and its 1.0 specification is stable, backwards compatible, and backed by major tech companies.
    - Google, Microsoft, Amazon, and Oracle have adopted the protocol
    - SDKs for TypeScript, Python, and .NET are generated from a shared JSON Schema, with community SDKs for Go, Ruby, Rust, Dart, Kotlin, Java, and C++
    - Transport-agnostic design supports HTTP SSE, WebSockets, and Protobuf binary format
    - AG-UI 1.0 adds subagent support, multimodal tool results, human-in-the-loop interrupts, and token usage tracking
    - CopilotKit Intelligence uses captured interaction data with Automatic Learning to extract insights and generate Skills that improve agent behavior over time
  5. Chat On Steroids is an open-source desktop workspace that connects your local files and terminal to an existing ChatGPT conversation, allowing the model to read, edit, run tests, and manage terminals within your actual project rather than a sandbox. It operates through MCP and a companion component that observes and automates the ChatGPT browser UI, while offering worker-based task management with persistent context to handle multi-step jobs. The project is explicitly an independent beta that rides on your existing ChatGPT plan, meaning shared usage limits apply, and it ships for Windows x64, macOS Apple silicon, and Linux x64.
    - Workers maintain their context across tasks so you don't re-explain project state on every handoff.
    - Goal, Loop, Compact & Resume, and mid-run corrections address the all-or-nothing nature of long agentic jobs.
    - The README is direct about not bypassing usage limits, account restrictions, or safety controls.
  6. w3cj writes about jev-chat, a tool-calling chat bot that routes user requests to real tools using Jev, a non-generative classifier from TypeSafe, with no LLM writing any output. Every value on screen was either typed by the user or returned by a tool, so the assistant cannot invent a fact. The system supports weather, unit conversion, Wikipedia lookups, recipes, web search, Todoist, and Home Assistant via MCP servers, with an inspector pane exposing every decision and probability for each turn.

    - Jev answers only two question types — Choice and Noul — and never produces text; all reply wording is templated in code
    - Pre-processing handles spell-check (cspell + compromise) and resolves short follow-ups like "what about Boston?" by swapping in the new value
    - Multi-step tools (Wikipedia, web search) chain multiple Jev requests: pick topic, then article, then the exact line that answers
    - Confidence-gated routing shows two buttons when the top tools are close rather than guessing
    - The repo is a proof of concept; the author will not accept PRs for new features
    - English only; no compound requests or multi-step reasoning supported
  7. Leela Kumili writes about DoorDash's multi-agent LLM system that automates stale feature flag cleanup across 623 repositories. In an evaluation of 50 stale flags, the system produced usable pull requests for 45, averaging 13.8 minutes and $4.79 per cleanup versus an estimated one to two hours for manual work. The two-phase workflow uses Claude Sonnet as an orchestrator to retrieve Jira tickets and query experimentation metadata via MCP, then Claude Opus agents in isolated Git worktrees to perform code changes and validation.

    - A single Boolean flag can require changes across 5–20 files due to dependency-injected wrappers
    - Uber's AST-based Piranha couldn't handle DoorDash's DI patterns where flag-to-logic relationships are semantic
    - Outcomes: 31 first-pass merges, 14 revisions, 5 engineer interventions, zero regressions
    - Gradle runs without its daemon to prevent state sharing between concurrent worktrees
    - Work accepted for the ICSME 2026 industry track
  8. @omarsar0 writes on X that the fastest path to genuinely understanding agent harnesses is to build one from scratch in TypeScript or Python, starting with a minimal ReAct implementation prompted from Google's original paper, targeting three clean components—an LLM inference module (multi-model, OpenRouter-backed, with separable system prompt), an MCP tools module for interoperability, and a simple agent loop that ties them together—then logging every input/output at each boundary and iterating against a small set of diverse test tasks so each change is inspectable. The punchline: skip the framework first, because only once you've felt the loop, the tokens, and the tool calls in your own code do the "next steps"—memory, skills, subagents—stop being black boxes you configure and become modules you actually know how to tune.

    - LLM module: wraps inference across multiple frontier models via OpenRouter; system prompt either embedded or isolated for context-engineering experiments
    - Tools module: implement as MCP (Model Context Protocol) tools for cross-harness interoperability, or as bespoke functions if experienced
    - Agent loop: ReAct pattern (alternating reasoning traces and action calls) encapsulating both LLM and tools; exit conditions handled via system-prompt instructions (non-deterministic), code-level checks (deterministic), or both
    - Logging strategy: capture loop in/out, every LLM call in/out, and every tool-call in/out; run a fixed diverse task suite after each modification
    - Scaling path: keep architecture modular so memory, skills, and subagent orchestration can be bolted on once the core loop is understood
    - Shortcut alternatives (if not building from scratch): Pi SDK or LangChain harness tooling
  9. OpenHuman is an open-source agent harness designed as a personal AI super intelligence, featuring local-first memory through Markdown trees in SQLite and orchestration capabilities via checkpointed graphs. It functions as a brain that builds persistent context from various data sources like email and calendars, acting as both an orchestrator for multi-agent workflows and a deep researcher with built-in web search and media generation tools.

    - Features "Memory Trees" stored locally in Markdown format to create a Karpathy-style Obsidian wiki.
    - Provides end-to-end encrypted agent-to-agent messaging using the Signal protocol.
    - Supports visual, trigger-driven workflows that can be proposed by an AI and reviewed on a canvas.
    - Includes a "Privacy Mode" which ensures no inference data leaves the user's machine when toggled.
  10. bex is an open-source, self-hostable PaaS that positions itself as an AI-native alternative to Render, letting developers push Git and receive a deployed URL on their own Kubernetes infrastructure. Coding agents operate as first-class users via MCP alongside the dashboard, CLI, REST, and GraphQL interfaces, all backed by a shared Go core. The platform uses a Kubernetes operator with Cluster API for machine provisioning, supports Render-style `render.yaml` Blueprints for declarative service definitions, and ships managed Postgres, Key Value, logs, metrics, autoscaling, custom domains with TLS, and SSH access.
    - 471 stars, 50 forks, 9 contributors — including Claude, Cursor, and Copilot listed as named GitHub contributors
    - Apache-2.0 licensed; explicitly marked "not ready for production workloads" (public alpha)
    - Language split: Go 58.5%, TypeScript 32.9%, Shell 6.7%
    - Internal "lego" Go workspace enforces a strict `operator → types ← backend` one-way dependency DAG
    - Tracks Render compatibility via an evidence-backed "parity ledger" (ADR018) rather than marketing claims
    - Local quickstart provisions a kind cluster + Cluster API with Docker-container machines as tenant nodes
    - Includes an Expo mobile app for safe supervision workflows (App Store listing present)
    - Commit history references agent-driven QA rounds (w4/w5/w6 workstreams) and live dashboard re-probes

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: model context protocol

About - Propulsed by SemanticScuttle