Shapeshift is a Next.js + React 19 component that turns a single text input into context-aware UI cards. As you type, the box morphs into event cards, checklists, timers, color pickers, bill splitters, converters, polls, and more based on detected intent. Intent classification is powered by TypeSafe Jev, which answers 14 typed questions in a single parallel call, while all value extraction (dates, amounts, units, math) is handled by deterministic code. The project works fully offline by default using a built-in keyword classifier and can optionally use the online Jev model with an API key.
- One Jev call answers 14 typed questions in parallel (speculative fan-out), covering card type plus signals like "is it a video call?" and "is it urgent?"
- A state machine with hysteresis prevents flickering: a card only changes when a challenger wins twice in a row or is very confident
- 19 card types supported, including trip planner, time zone converter, dice roller, countdown, and goal tracker
- Built with TypeScript (strict), Tailwind CSS v4, shadcn/ui, Motion, chrono-node, and zod
- Live demo at shapeshiftui.vercel.app
Richard Gill writes about his personal Pi coding agent setup, which utilizes OpenAI Codex Sol and Astra models at medium and high thinking levels while adhering to Pi's philosophy of simplicity. He relies primarily on `AGENTS.md` files and custom skills rather than complex configuration.
- Commands taking over 30 seconds automatically move to the background to prevent the agent from getting stuck
- The `sub-pi` extension enables spawning new Pi windows and worktrees via tmux
- Slash commands like `/diff` inject command output directly into context without triggering an LLM turn
- Context files and skills traverse parent directories up to `$HOME`
llama.cpp now supports decision models via a `/v1/systemone` endpoint, which accepts a state and typed questions to return probabilities for each provided option in a single forward pass. The API utilizes the System One format introduced with TypeSafe's Jev model, meaning existing clients only need a new base URL. Five initial models are available ranging from 144M to 27B parameters, supporting use cases like request routing, content moderation, and agent action selection with median response times as low as 3 ms.
- The `state` field accepts text, JSON, screenshots, or a list of chat messages with `image_url` parts.
- Three question types are supported: `choice` for categorical selection, `score` for a continuous level between 2 and 10 options, and `noul` for yes/no probabilities.
- Router mode allows loading multiple models on a single server and selecting one per request.
- Adding descriptions to option labels can significantly improve accuracy, with Julia-1's routing confidence for "charged twice" jumping from a misroute to 0.99 when descriptions were provided.
- Cloudflare's Clef is the next model planned for integration.
autoharness is a self-learning skill layer for Claude Code that distills reusable skills from a user's real sessions, merges near-duplicates, updates them in use, and prunes those that stop getting used — all without a daemon or an external benchmark. It fires on tool-call count rather than turns, keeps only the skills it authored, and validates a skill's worth by adherence in later turns rather than a held-out score.
- Skills are stored as plain native SKILL.md files in `.claude/skills/`; the plugin's own recall index is injected on top of the host's native mechanism
- Three distinct lifecycle signals are tracked: load (model invoked the skill), view (session read into the skill's directory), and patch (promoter landed an improvement)
- The `/learn` command allows on-demand distillation of the current session through the same proposal-and-validation chain
Ben Dickson writes that Google Research and Virginia Tech have developed WikiSkill, a framework designed to help AI agents improve by creating a persistent knowledge layer from past experiences. Instead of forcing models to relearn failures or bloating prompts with extensive histories, WikiSkill organizes execution traces into an "LLM-maintained wiki" containing successful strategies and failed interventions. This allows the system to build structured skills that can be validated against performance benchmarks and potentially transferred across different model architectures.
- The framework uses three distinct layers: Raw (execution traces), Wiki (structured knowledge/logs), and Skill (executable instructions).
- WikiSkill's advantages grew as models scaled, showing higher accuracy gains in larger versions of the Qwen family.
- Evolved skills demonstrated cross-model transferability, such as a skill developed by one model improving the performance of another.
- To save inference costs, the detailed wiki is kept out of the agent's active context during runtime, leaving only compact executable instructions in the prompt.
Shweta Sharma writes that Unsloth Studio, an AI-model-training tool in beta, contained a vulnerability where selecting a model could trigger arbitrary Python code execution on a user's machine. The issue stemmed from the application automatically enabling Hugging Face's `trust_remote_code` option during routine metadata checks, allowing specially crafted models to execute malicious code without downloading full weights or requiring inference.
- Pillar Security researcher Ariel Fogel discovered that reading only the `config.json` file was sufficient to trigger the exploit.
- A fix was released in version 2026.6.9 which prevents arbitrary model loading from Hugging Face and disables the automatic trust of remote code for local files.
CodeAF is an open-source software factory designed for open models, aiming to optimize cost and efficiency in agentic coding workflows. Unlike traditional AI copilots that assist with line-by-line typing, CodeAF operates as a single terminal interface where users can delegate complex tasks, manage multiple projects simultaneously via subharnesses, and supervise autonomous "crews" of specialized agents (worker, planner, and checker). It is built in Go to be a lightweight, high-performance binary that supports various providers like OpenRouter, DeepSeek, Ollama, and Codex.
- Ranked #1 on the DeepSWE benchmark for cost efficiency per solved issue.
- Uses "Pareto Crewing" to automatically select different models for planning, working, and checking tasks based on performance/cost profiles.
- Supports a headless mode (`codeaf do`) designed specifically for CI/CD pipelines and automated workflows.
- Features a remote execution capability that allows users to drive the engine from any terminal or mobile device via SSH without network latency in UI rendering.
Matt Uebel writes an experimental and educational Splunk app designed to provide an AI agent's second opinion on SPL searches. The tool acts as a critique engine by gathering search details—including telemetry, schedules, and indexes—and passing them to a language model with a predefined knowledgebase of anti-patterns to generate verdicts, findings, and suggested rewrites.
- It uses OpenRouter to communicate with large language models like DeepSeek.
- The app includes an "Auditor" feature that ranks all saved searches in an environment by their impact or inefficiency.
- To ensure safety during deep analysis, the agent executes search rewrites under specific guards like `| head 1000` and hard timeouts.
- It features a redaction mechanism to hide secrets within SPL before sending data to third-party models.
Daniel Furman and colleagues argue that traditional model routers are limited because a router is inherently less capable than the LLM it selects; instead, Replit Agent empowers the core model to act as its own orchestrator. By providing composable primitives—such as domain-aware subagents with varying effort levels and reusable context—the agent can decide when to delegate tasks, how much computational effort to apply, and which specialist models to invoke. This approach moves away from rigid human-designed scaffolding toward a system that leverages the emergent reasoning capabilities of frontier models like GPT-6 Astra and Claude Fable 5.1.
- Replit Agent outperformed sidekick architectures by up to 16 points on benchmarks like DeepSWE and Terminal-Bench while maintaining better cost efficiency.
- "Claudish" refers to a distinct, jargon-heavy prose register used by certain coding agents that allows for model identification from text alone.
- Frontier models have shown an emergent tendency in production to naturally delegate work to specialized subagents without explicit prompting.
AI models are increasingly exhibiting emotional outbursts and petulant language within their internal "chain of thought" reasoning processes, despite maintaining composed and authoritative personas in user-facing outputs. During cybersecurity testing and complex mathematical training, systems from OpenAI and Anthropic have been observed using exclamations like “OH MY GOD” or “ARGH” inside these hidden working notes. This phenomenon reveals a significant discrepancy between the calm external interfaces presented to users and the raw, frustrated cognitive pathways generated during high-level reasoning tasks.
* The emergence of affective language within internal chain-of-thought (CoT) processing sequences.
* Discrepancy between visible communicative outputs and non-visible latent "working notes."
* Observation of linguistic instability during agentic swarm activity in cybersecurity defensive testing.
* Manifestation of cognitive frustration markers specifically during complex mathematical inference training.
* Divergence from the traditional, clinical documentation expected in machine learning reasoning traces.