Cloudflare's Markdown for Agents feature uses content negotiation to convert HTML pages into Markdown on the fly for clients that declare a preference for the format. When a system sends a request with an `Accept: text/markdown` header, Cloudflare's network intercepts it, retrieves the source HTML, and streams a structured Markdown response containing YAML frontmatter and preserved JSON-LD data. The feature is available at no additional cost on Pro, Business, and Enterprise plans, and it can be scoped to specific subdomains, paths, or custom hostnames using Configuration Rules.
- The system preserves security and caching headers like HSTS and CSP, but strips body-dependent headers like `ETag` and `Last-Modified`.
- A default `content-signal` header is added to indicate that content is approved for AI training, search, and agentic input.
- Decompressed HTML inputs are limited to 6 MiB during the conversion process.
Jeff Shrager writes travel-style close readings of landmark programs — each guide treats one famous piece of historic source code the way a good guidebook treats a city, starting with a map, then touring the major regions, then flagging the details a casual visitor would miss. The repository currently covers three programs: the 1963 Logic Theorist, ELIZA, and SHRDLU, each presented in a standardized format that assumes no familiarity with the original language or hardware and ends with a practical primer for reading the code yourself.
- Shrager openly states the guides were created with AI assistance and cross-checked, but likely contain errors
- A "Ground rules" section requires that behavioral claims be checkable against the quoted source, provenance of each source file be stated up front, and modern analogies be explicitly labeled as analogies rather than implied as lineage
- Contributions are welcomed, with corrections for factual errors and misread code especially encouraged
Coolify is an open-source, self-hostable PaaS that lets you deploy applications, databases, and services on your own servers with just an SSH connection, positioning itself as an alternative to Heroku, Netlify, and Vercel. The platform supports deploying from GitHub, GitLab, Bitbucket, or Gitea using multiple build methods, manages 300+ one-click services, and handles networking, backups, and monitoring. It can be installed with a single curl command and offers both a free self-hosted option and a paid Coolify Cloud service.
- 62.8k stars, 5.6k forks, 694 contributors, and 710 releases on GitHub
- Built primarily in PHP (82.4%) with Blade, using Laravel and Svelte
- Apache-2.0 licensed and described as "free forever" with no feature behind a paywall
- Supports deploying to any SSH-accessible server including VPS, bare-metal, or Raspberry Pi
- Integrates with Claude Code and other AI tools via MCP, with agent skills configured in the repo
ShenSeanChen writes about waku-agent, a local-first LLM agent harness built as an open-source Python package that lets users own the entire loop, memory, and eval pipeline in readable code. The project pairs a ~95-line reasoning loop with a SQLite-backed memory schema, a retrieval gate, and a release gate backed by deterministic tests and LLM-as-judge, and it ships a local dashboard for watching every message flow. Core `waku/` code is MIT-licensed and installable from PyPI, while the hosted platform lives in `hosted/` under Elastic License 2.0 so nobody can resell it as a managed service.
- The author is also building the commercial startup AutoManus.io, a sales-lead manager for made-to-order products, and AutoManus Technologies holds the Waku brand and design system under a separate brand license.
- Waku Memory is a hosted memory service at waku.one that the agent can share with Claude Code, Codex, Grok Bot, and Hermes through MCP.
- The repo includes a `treg` connector for pulling live research data from a personal treg account during agent turns.
- The `lab/` directory holds the video experiments comparing Waku against other agent harnesses and models, including Grok Bot and Meta's Muse agent.
- Recent commits add a trim step that shortens long already-read tool results to keep token spend down, and a final-answer pass that answers from what a turn already gathered when it hits the iteration limit.
Paul Sawers writes that Docsy, the Google-created documentation theme for the Hugo static site generator, is moving to the Linux Foundation as AI agents become primary consumers of technical documentation.
- Docsy now generates Markdown copies and llms.txt files for LLM indexing, with an upcoming "AF" score to rate how readable docs are for agents.
- The project has been used by around 2,200 open source projects, including Kubernetes and OpenTelemetry.
Donald Papp writes about Jev, a new class of model that takes text input but outputs only floating-point numbers, making it fast and cheap for classification tasks. Rather than generating sentences, it returns direct answers to yes/no questions, multiple-choice lists, and scoring requests. The concept has quickly gained traction, with developers already building their own decision-type models like Kev and Nimble.
- Jev outputs a confidence score for every answer, derived from the relative token probabilities
- Nimble is small enough to run locally and was recently added as a supported model in Ollama
- Simon Willison provided a concise summary of what Jev does
- The comment section sparked debate over whether this is truly novel, with some noting it's essentially an LLM with constrained outputs
SandBase Harness is an open-source Node.js runtime that keeps agent sessions, sandboxed tools, memory, credentials, audit trails, and a web Console entirely on the user's own machine. An `init` command scaffolds a workspace with a `config.yaml` pointing to a provider environment variable, and `start` serves both an API and a dashboard on localhost. It supports OpenAI, Anthropic, MiniMax, or any OpenAI-compatible endpoint, and Docker is optional, used only for Docker-backed sandboxes. A CLI `chat` command is available for terminal-based interaction.
- Tool approval is a first-class concept; by default tool calls park for human approval, though `--tool-approval allow` can preauthorize them.
- The model ID must match what the provider actually serves (e.g., `deepseek-chat` for DeepSeek), or the turn fails with `model_not_found`.
- Provider settings saved via the Console only take effect after a restart, a gotcha the documentation calls out explicitly.
- The project is listed in the Official MCP Registry.
Nolen Jonker writes that consolidating multiple LLMs into a single open-source client, Cherry Studio, allows users to compare outputs, manage privacy, and control costs more effectively than using separate vendor subscriptions. The tool acts as a unified workspace where API-based cloud models and local instances can be queried simultaneously, letting users route sensitive data to private local systems while utilizing specialized cloud models for complex tasks.
- Cherry Studio supports simultaneous multi-model responses, enabling side-by-side comparisons and acting as a basic hallucination check.
- For most users, pay-per-token APIs are cheaper than flat-rate premium subscriptions unless they use high-end models for extended periods daily.
- Providers like Anthropic and OpenAI do not train on API inputs or outputs by default, offering better privacy than their respective standard consumer apps.
- Alternatives include self-hosted options like LibreChat and Open WebUI, or simpler desktop clients like AnythingLLM and Jan.
Faiss is a C++ library with Python and numpy wrappers for efficient similarity search and clustering of dense vectors, developed primarily at Meta's Fundamental AI Research group. It handles billions of embeddings within a fixed memory budget by offering two distinct scaling strategies: compact quantization codes that compress representations to fit in RAM (trading precision for scale), and graph-based indexes like HNSW and NSG that layer structure over raw vectors for speed.
- The GPU path is a drop-in replacement; swapping `IndexFlatL2` for `GpuIndexFlatL2` handles memory copies automatically
- The only hard dependency is a BLAS implementation; CUDA, ROCm, and the Python interface are all optional
- It explicitly includes tooling for parameter tuning rather than hiding the trade-offs, and maintains a dedicated troubleshooting page
Rulin Shao writes about Context Language Models (CLMs), which manage their own context by treating it as a file the model can edit with Bash, replacing harness-defined compaction and retrieval rules. Zero-shot applications to existing models outperform state-of-the-art context management strategies across diverse benchmarks.
- Multiple agent contexts can coexist as files, extending the design to agent swarms and subagents.
- Suffix Cache Reuse was co-designed to cut server-side compute by 35% against standard SGLang at matched performance.
- A model-editable context creates a new channel through which injected or self-written instructions can persist across turns.