llama.cpp now supports decision models via a `/v1/systemone` endpoint, which accepts a state and typed questions to return probabilities for each provided option in a single forward pass. The API utilizes the System One format introduced with TypeSafe's Jev model, meaning existing clients only need a new base URL. Five initial models are available ranging from 144M to 27B parameters, supporting use cases like request routing, content moderation, and agent action selection with median response times as low as 3 ms.
- The `state` field accepts text, JSON, screenshots, or a list of chat messages with `image_url` parts.
- Three question types are supported: `choice` for categorical selection, `score` for a continuous level between 2 and 10 options, and `noul` for yes/no probabilities.
- Router mode allows loading multiple models on a single server and selecting one per request.
- Adding descriptions to option labels can significantly improve accuracy, with Julia-1's routing confidence for "charged twice" jumping from a misroute to 0.99 when descriptions were provided.
- Cloudflare's Clef is the next model planned for integration.
Jeff Shrager's repository introduces "executable archaeology" — the practice of recovering, transcribing, and running the original historical source of early computing programs so that the surviving artifact itself, rather than a modern rewrite, can serve as experimental evidence. The main project here is an IPL-V reimplementation of Victor Yngve's 1959–61 MIT sentence generator, run under Herbert A. Simon's account on 25 June 1962, whose surviving printout preserves the program, its phrase-structure grammar, execution traces, and handwritten corrections by Simon and his daughter Katherine.
- The repository also points to separate repos reconstructing the Logic Theorist in two historically distinct forms (LT1 in IPL-I and LT5 in IPL-V) and a full CTSS/IBM 7094 reanimation of Weizenbaum's original ELIZA
- A 2026 arXiv paper by Shrager and a 2025 IEEE Annals paper by Lane et al. provide the formal background
- Six stated principles govern the work, including distinguishing original behavior from behavior introduced by the reconstruction
autoharness is a self-learning skill layer for Claude Code that distills reusable skills from a user's real sessions, merges near-duplicates, updates them in use, and prunes those that stop getting used — all without a daemon or an external benchmark. It fires on tool-call count rather than turns, keeps only the skills it authored, and validates a skill's worth by adherence in later turns rather than a held-out score.
- Skills are stored as plain native SKILL.md files in `.claude/skills/`; the plugin's own recall index is injected on top of the host's native mechanism
- Three distinct lifecycle signals are tracked: load (model invoked the skill), view (session read into the skill's directory), and patch (promoter landed an improvement)
- The `/learn` command allows on-demand distillation of the current session through the same proposal-and-validation chain
Jeff Shrager writes a comprehensive guide to Herbert Simon's 1961 Heuristic Compiler, a program that uses General Problem Solver (GPS) means-end analysis to generate IPL-V code. After being dormant for roughly 65 years, the archival card deck was transcribed and successfully run on a modern Common Lisp interpreter. The system is composed of three parts: a State Description Compiler that derives code from before-and-after memory states, a Functional Description Compiler that modifies existing code based on imperative phrases, and a General Compiler executive that manages the search process. Notably, the program reproduces the exact machine code printed in Simon's 1963 paper, such as the "INSERT AT END OF VALUE LIST" and "SET SIGNAL MINUS" examples, without any modern language rewrites.
- The guide highlights that the compiler treats routine generation as a problem to be solved via operator search rather than a direct translation task.
- A debugging print statement left in routine U113 happens to output the exact stage-by-stage compilation sequence described in the 1963 paper.
- The original 1961 deck fails to compile the J3 routine; a specific card must be modified to erase a redundant description before the state compiler can execute successfully.
Maya Posch writes about Aaron Christophel disassembling a €15 UGreen USB-C cable rated at 240 watts, which features an integrated IPS LCD screen and a single touch control. The teardown reveals a PCB with an EN32LF056 Cortex-M0+ MCU, voltage regulators, and a MOSFET for backlighting. By probing the exposed SWD interface, Christophel confirmed the chip has 64 kB of Flash and 4 kB of SRAM, then flashed custom firmware to display video frames on the tiny display.
- The MCU is not wired to the USB data interface, so it cannot act as a man-in-the-middle or keylogger
- A commenter noted that such dense integrated cables raise concerns about secure pairing and encryption for USB connections
Chat On Steroids is an open-source desktop workspace that connects your local files and terminal to an existing ChatGPT conversation, allowing the model to read, edit, run tests, and manage terminals within your actual project rather than a sandbox. It operates through MCP and a companion component that observes and automates the ChatGPT browser UI, while offering worker-based task management with persistent context to handle multi-step jobs. The project is explicitly an independent beta that rides on your existing ChatGPT plan, meaning shared usage limits apply, and it ships for Windows x64, macOS Apple silicon, and Linux x64.
- Workers maintain their context across tasks so you don't re-explain project state on every handoff.
- Goal, Loop, Compact & Resume, and mid-run corrections address the all-or-nothing nature of long agentic jobs.
- The README is direct about not bypassing usage limits, account restrictions, or safety controls.
Dazzle aims to transform personal media and interactions into a proactive, predictive intelligence layer by extracting multidimensional context from camera rolls and conversational histories. This approach moves beyond generic large language models to provide hyper-personalized experiences through technologies that understand visual media, locations, people, and conversation topics.
- The company has filed foundational patents for building contextual user stores and forming location/relationship maps from visual media.
- Their technology includes methods for grouping photos to suggest activities and segregating conversations by topic.
- They focus on "intent-based" management of visual media.
Amanda Caswell writes about Cloudflare's new Monetization Gateway, which allows domain owners to charge AI agents for accessing APIs, websites, and datasets using the x402 protocol. This system enables payments in USDC on the Base blockchain via HTTP requests, but it introduces significant challenges for developers regarding spending control, variable pricing models, and the need for robust observability to track transaction history during retries.
- The gateway supports price ranges from $0.001 to $100 per request.
- It features two payment schemes: 'exact' for fixed prices and 'upto' for variable rates.
- Developers can implement "Virtual Wallets" with allowances, allowlists, and maximum transaction sizes to manage agent spending.
- Cloudflare plans to make paid services discoverable so agents can find tools during an active workflow. author »
Ben Dickson writes that Google Research and Virginia Tech have developed WikiSkill, a framework designed to help AI agents improve by creating a persistent knowledge layer from past experiences. Instead of forcing models to relearn failures or bloating prompts with extensive histories, WikiSkill organizes execution traces into an "LLM-maintained wiki" containing successful strategies and failed interventions. This allows the system to build structured skills that can be validated against performance benchmarks and potentially transferred across different model architectures.
- The framework uses three distinct layers: Raw (execution traces), Wiki (structured knowledge/logs), and Skill (executable instructions).
- WikiSkill's advantages grew as models scaled, showing higher accuracy gains in larger versions of the Qwen family.
- Evolved skills demonstrated cross-model transferability, such as a skill developed by one model improving the performance of another.
- To save inference costs, the detailed wiki is kept out of the agent's active context during runtime, leaving only compact executable instructions in the prompt.
This open-source project provides firmware for an ESP32-C3 Super Mini connected to a 1.28″ round GC9A01 display, creating a sonar-style aircraft radar that visualizes live ADS-B data from adsb.fi around the user's location. The device features WiFiManager for easy initial configuration via a captive portal and allows users to cycle through range presets or reset network credentials using hardware buttons.
- Visualizes aircraft with red heading triangles, magenta speed vectors, and callsign tags.
- Supports runway overlays from major airports based on OurAirports data.
- Displays "rim dots" for approaching aircraft located outside the current radar ring but within range of ADS-B data.
- Uses NVS to persist user settings like latitude, longitude, units (mi/km), and WiFi credentials.