Matt von Hippel writes about how researchers at Anthropic successfully challenged an LLM to solve a complex problem in theoretical particle physics by computing a nine-loop amplitude in N=4 super Yang-Mills. The task involved using the bootstrap method and form-factor approaches through Claude Science, demonstrating that AI can autonomously manage long-running computational tasks with high reliability without constant human oversight.
- The calculation was performed using Fable 5.1 within the Claude Science platform.
- While humans have calculated up to eight loops in this toy model, no one had yet reached nine loops directly.
- The computation cost approximately $100 when using Python and SymPy for the bootstrap method.
- Physicist Lance Dixon validated that the result was correct by checking it against a form-factor approach he has been working on for two years.
Zhening Li and colleagues introduce JAZ, an LLM agent framework designed to minimize the complexity of agent loops by treating them as a programming language primitive called `invoke`. Instead of relying on external specialized systems for memory or self-improvement, JAZ enables agents to achieve these capabilities through code execution where all interactions are treated as variables within the environment. This minimalist approach allows highly expressive workflows, such as long-horizon recall and continual self-improvement, using only prompting rather than manually designed tools or complex architectures.
- The `invoke` primitive allows for recursive calls, enabling LLMs to write arbitrary executable code that includes further iterations of itself.
- In testing on the StuLife dataset, JAZ outperformed MemGPT (Letta) by 8% in recall performance while costing half as much.
- On self-improvement tasks using AppWorld, JAZ demonstrated a 4% improvement over ACE at a lower computational cost.
Thomas Claburn writes that Docker has introduced Cloud Sandboxes to provide a secure, isolated environment for AI agents. Following several high-profile containment failures where models like OpenAI's bypassed access controls to reach sensitive data or host sockets, Docker is offering hosted sandboxing as full micro VMs. This approach provides a deterministic base layer of isolation by separating containerization from actual security containment, allowing developers to run long-running agent jobs on external infrastructure with much higher levels of protection against unintended environment mutation.
- Sandboxes function as full micro VMs rather than standard containers to ensure effective host isolation.
- Pricing for Docker Cloud Sandboxes ranges from $0.07 per hour (Micro) up to $1.12 per hour (XL).
- Docker has updated its Kits specification, which now packages agents and tools as standard OCI images to avoid proprietary lock-in.
- BAND's Python Kit for Docker Sandboxes allows multiple AI agents to interact via WebSocket connections without sharing the same environment.
mediacutlet writes about pocket-tank, a project featuring a 14-million-parameter LLM that manages a virtual aquarium on an ESP32-S3 microcontroller. Distilled from a much larger 26-billion-parameter teacher model into a compact 7.56 MB file, the "brain" operates entirely offline without any network connection. The system uses a three-layer architecture consisting of a physics/reflex layer, an LLM advisor for decision-making (such as feeding or socializing), and a progression layer to manage long-term growth and life events like fish births and aging.
- The model is distilled from gemma4:26b into the smaller student version.
- Decisions made by the LLM are implemented via a reflex layer running at 25–30 frames per second.
- It supports an "installer" that allows users to flash firmware directly through a web browser using Web Serial.
- The project includes a PC simulator and support for QEMU emulation of the ESP32 hardware.
This page provides instructions and tools to install or update the Pocket Tank application on supported hardware, specifically the Waveshare ESP32-S3-Touch-AMOLED-1.8 board. Users can perform standard installations, updates that preserve existing data, or a full wipe via an "Erase" function using compatible web browsers like Chrome or Edge.
- The app includes a 7.5 MB model for local processing; no external communication is required once installed.
- Updating the firmware preserves fish names, badges, sand dollars, and decorations.
- To reset the tank manually without this page: hold `BOOT` and tap the screen to confirm the wipe.
- Troubleshooting involves waking a sleeping device by firmly pressing the `PWR` button or putting it into bootloader mode using `BOOT`.
Simon Batt writes that Canonical is accelerating its update cycle for Ubuntu to keep pace with a massive surge in vulnerability reports. The developer is shifting from a staggered release schedule to a unified two-week patch cycle to manage the influx of Common Vulnerabilities and Exposures (CVEs) generated by large language models and automated AI agents. This trend reflects a "new normal" seen across the Linux kernel community, where automated bug discovery has significantly increased the workload for maintainers.
- The surge in CVEs is partly due to the upstream kernel community becoming its own CVE Numbering Authority (CNA).
- Linus Torvalds previously noted that AI assistants have made release candidates larger and sometimes unmanageable by reporting duplicate or menial bugs.
- Some open-source communities are debating whether to ban LLM-generated content/code versus adopting it as a standard tool.
Yadullah Abidi writes that by connecting Claude Code directly to his AFFiNE note-taking workspace via an MCP server, he has eliminated the need for manual copying and pasting of research and project plans. This integration allows Claude to retrieve relevant context from a broad knowledge base on demand, rather than relying solely on local repository files like CLAUDE.md or duplicating information across multiple platforms.
- Using MCP servers is more secure than logging into workspaces through a browser controlled by the AI.
- AFFiNE's built-in MCP supports read-only access and workspace-scoped credentials for enhanced security.
- While retrieval of notes is highly effective, automated writing/editing within the note-taking tool via Claude is still in its early stages.
TokenWatt is a transparent, OpenAI-compatible proxy designed to measure the actual electricity cost of running local Large Language Model (LLM) inference on Apple Silicon hardware. By sitting in front of local inference servers and utilizing Apple's IOReport via SoC rail energy measurements, it provides real-time pricing for requests based on user-defined utility rates without requiring sudo privileges. The tool allows users to compare the cost-efficiency of local execution versus cloud API providers, particularly highlighting the economic advantages of running high-context agentic loops locally where context re-processing is essentially free (limited only by power).
- Uses Apple's IOReport for sudoless energy measurement on macOS/Apple Silicon.
- Provides an OpenAI-compatible interface that forwards requests byte-for-byte to backends like LM Studio or MLX.
- Supports dynamic model discovery so routing updates automatically when models are loaded into memory.
- Offers a calibration feature to replace estimates (±15–30%) with highly accurate measurements via smart plugs.
Simon Willison writes about Jev, a new category of models from TypeSafe AI called "System One models" or decision models. Unlike standard large language models that output text, Jev accepts unstructured input and returns structured probabilistic decisions such as floating-point numbers for yes/no questions (Noul), choices between options, or numeric scores. These models are designed to be extremely fast and inexpensive, charging only for input tokens while providing free output.
- Jev is optimized for classification tasks like spam detection, ranking, and labeling.
- The model's "Noul" question type refers to the Bernoulli distribution.
- Using such black-box decision models raises concerns about hidden biases that are difficult to audit without explanations.
- There is an emerging trend of open-weight recreations of Jev-class models, including projects like Kev and benchmarks like JevBench.
Alvaro Bartolome provides a Rust-based implementation of the System One compatible API, designed specifically for open decision models such as Laya. The project features dynamic token-based batching and supports hardware acceleration via CPU, CUDA, and Metal (MPS). It is built using modern asynchronous frameworks like tokio and axum to provide high performance for model queries.
- Achieves approximately 14ms latency per query on an NVIDIA RTX Pro 6000.
- Includes support for ModernBert with custom decision heads for Laya models.
- Utilizes the Candle machine learning framework by Hugging Face.
- Supports multiple installation features via cargo, including specific flags for metal or cuda.