llama.cpp now supports decision models via a `/v1/systemone` endpoint, which accepts a state and typed questions to return probabilities for each provided option in a single forward pass. The API utilizes the System One format introduced with TypeSafe's Jev model, meaning existing clients only need a new base URL. Five initial models are available ranging from 144M to 27B parameters, supporting use cases like request routing, content moderation, and agent action selection with median response times as low as 3 ms.
- The `state` field accepts text, JSON, screenshots, or a list of chat messages with `image_url` parts.
- Three question types are supported: `choice` for categorical selection, `score` for a continuous level between 2 and 10 options, and `noul` for yes/no probabilities.
- Router mode allows loading multiple models on a single server and selecting one per request.
- Adding descriptions to option labels can significantly improve accuracy, with Julia-1's routing confidence for "charged twice" jumping from a misroute to 0.99 when descriptions were provided.
- Cloudflare's Clef is the next model planned for integration.
bebechien writes about DinoDesk AI, a LEGO dino desk companion built on a Raspberry Pi that pairs a local Gemma 4 model (via LM Studio) with cloud Gemini Flash through a unified OpenAI-compatible gateway, giving users a camera-free, privacy-first chat robot with three switching modes (local, cloud, and auto-hybrid routing). The physical body uses LEGO Technic lever mechanisms for neck and tail movement, a Pimoroni Pirate Audio shield for the 1.3" LCD eyes and 8-bit I2S beeps, and a 5-state finite state machine to coordinate expressions, sound, and motor action across Sleeping, Idle, Listening, Thinking, and Speaking states.
- Auto-hybrid mode uses a complexity classifier that detects multi-step reasoning keywords like "explain," "compare," and "write code" to transparently escalate prompts to the cloud engine
- The project is open-sourced at github.com/google-gemma/dinodesk-ai-companion
- Full voice chat is still a work in progress; current interaction is triggered by a physical red push button
- A commenter noted that absence of a camera does not guarantee voice data stays local, and suggested an explicit retention boundary for audio transcripts would strengthen the privacy claim
This tutorial demonstrates how to construct a complete skill-based agent system for large language models using Python. It explores structuring modular capabilities similar to an operating system, where reusable skills are defined with metadata and schemas, registered centrally, and orchestrated through dynamic tool calling and multi-step reasoning. The implementation covers composing multiple skills for advanced workflows, hot-loading new capabilities at runtime, and monitoring performance via an observability dashboard.
This tutorial demonstrates how to implement an intelligent routing layer using NadirClaw to optimize Large Language Model (LLM) costs. The system classifies prompts into simple or complex tiers locally before selecting the most appropriate model, such as switching between Gemini Flash and Pro versions. It covers installation, local classification testing via CLI, visualizing decision boundaries through centroid-based similarity scores, running a proxy server for live routing, and calculating estimated cost savings compared to using high-end models exclusively.
HookCats is a self-hosted webhook routing server that acts as the central hub between your infrastructure and your team chat. It receives webhooks from any supported source, formats the messages nicely, and delivers them to your preferred chat platform.
Katanemo Labs introduces Arch-Router, a 1.5B parameter model that intelligently maps user queries to the most suitable LLM, achieving 93% accuracy without the need for costly retraining. It uses a preference-aligned routing framework based on a Domain-Action Taxonomy, allowing for flexible adaptation to evolving models and use cases.
This paper introduces Arch-Router, a preference-aligned routing framework for large language models (LLMs). It addresses limitations in existing routing approaches by focusing on matching queries to user-defined preferences (domain and action types) rather than solely relying on benchmark performance. The framework includes a 1.5B parameter model, Arch-Router, and a data creation pipeline. Experiments demonstrate state-of-the-art results in matching queries with human preferences and improved adaptability.
This paper proposes a preference-aligned routing framework for LLMs that guides model selection by matching queries to user-defined domains or action types. It introduces Arch-Router, a compact 1.5B model that learns to map queries to domain-action preferences for model routing decisions, outperforming proprietary models in subjective evaluation criteria.
The article discusses the use of AI agents for automating and optimizing tasks in the networking industry, including network deployment, configuration, and monitoring. It outlines a workflow with four agents that collectively achieve the setup and verification of network connectivity within a Linux and SR Linux container environment.
The author demonstrates a workflow involving four AI agents designed to deploy, configure, and monitor a network:
Document Specialist Agent: This agent extracts installation, topology deployment, and node connection instructions from a specified website.
- Linux Configuration Agent: Executes the installation and configuration commands on a Debian 12 UTM VM, checks the health of the VM, and verifies the successful deployment of network containers.
- Network Configuration Specialist Agent: Configures network IP allocation, interfaces, and routing based on the network topology, including detailed BGP configurations for different network nodes.
- Senior Network Administrator Agent: Applies the generated configurations to the network nodes, checks BGP peering, and verifies end-to-end connectivity through ping tests.
WilmerAI is a sophisticated middleware system designed to handle incoming prompts and route them to appropriate categories and workflows. It supports multiple Large Language Models (LLMs) and can handle a single incoming connection to many backend LLMs.