klotz: large language model*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. ShenSeanChen writes about waku-agent, a local-first LLM agent harness built as an open-source Python package that lets users own the entire loop, memory, and eval pipeline in readable code. The project pairs a ~95-line reasoning loop with a SQLite-backed memory schema, a retrieval gate, and a release gate backed by deterministic tests and LLM-as-judge, and it ships a local dashboard for watching every message flow. Core `waku/` code is MIT-licensed and installable from PyPI, while the hosted platform lives in `hosted/` under Elastic License 2.0 so nobody can resell it as a managed service.
    - The author is also building the commercial startup AutoManus.io, a sales-lead manager for made-to-order products, and AutoManus Technologies holds the Waku brand and design system under a separate brand license.
    - Waku Memory is a hosted memory service at waku.one that the agent can share with Claude Code, Codex, Grok Bot, and Hermes through MCP.
    - The repo includes a `treg` connector for pulling live research data from a personal treg account during agent turns.
    - The `lab/` directory holds the video experiments comparing Waku against other agent harnesses and models, including Grok Bot and Meta's Muse agent.
    - Recent commits add a trim step that shortens long already-read tool results to keep token spend down, and a final-answer pass that answers from what a turn already gathered when it hits the iteration limit.
  2. Paul Sawers writes that Docsy, the Google-created documentation theme for the Hugo static site generator, is moving to the Linux Foundation as AI agents become primary consumers of technical documentation.
    - Docsy now generates Markdown copies and llms.txt files for LLM indexing, with an upcoming "AF" score to rate how readable docs are for agents.
    - The project has been used by around 2,200 open source projects, including Kubernetes and OpenTelemetry.
  3. Donald Papp writes about Jev, a new class of model that takes text input but outputs only floating-point numbers, making it fast and cheap for classification tasks. Rather than generating sentences, it returns direct answers to yes/no questions, multiple-choice lists, and scoring requests. The concept has quickly gained traction, with developers already building their own decision-type models like Kev and Nimble.
    - Jev outputs a confidence score for every answer, derived from the relative token probabilities
    - Nimble is small enough to run locally and was recently added as a supported model in Ollama
    - Simon Willison provided a concise summary of what Jev does
    - The comment section sparked debate over whether this is truly novel, with some noting it's essentially an LLM with constrained outputs
  4. SandBase Harness is an open-source Node.js runtime that keeps agent sessions, sandboxed tools, memory, credentials, audit trails, and a web Console entirely on the user's own machine. An `init` command scaffolds a workspace with a `config.yaml` pointing to a provider environment variable, and `start` serves both an API and a dashboard on localhost. It supports OpenAI, Anthropic, MiniMax, or any OpenAI-compatible endpoint, and Docker is optional, used only for Docker-backed sandboxes. A CLI `chat` command is available for terminal-based interaction.
    - Tool approval is a first-class concept; by default tool calls park for human approval, though `--tool-approval allow` can preauthorize them.
    - The model ID must match what the provider actually serves (e.g., `deepseek-chat` for DeepSeek), or the turn fails with `model_not_found`.
    - Provider settings saved via the Console only take effect after a restart, a gotcha the documentation calls out explicitly.
    - The project is listed in the Official MCP Registry.
  5. Nolen Jonker writes that consolidating multiple LLMs into a single open-source client, Cherry Studio, allows users to compare outputs, manage privacy, and control costs more effectively than using separate vendor subscriptions. The tool acts as a unified workspace where API-based cloud models and local instances can be queried simultaneously, letting users route sensitive data to private local systems while utilizing specialized cloud models for complex tasks.
    - Cherry Studio supports simultaneous multi-model responses, enabling side-by-side comparisons and acting as a basic hallucination check.
    - For most users, pay-per-token APIs are cheaper than flat-rate premium subscriptions unless they use high-end models for extended periods daily.
    - Providers like Anthropic and OpenAI do not train on API inputs or outputs by default, offering better privacy than their respective standard consumer apps.
    - Alternatives include self-hosted options like LibreChat and Open WebUI, or simpler desktop clients like AnythingLLM and Jan.
  6. Faiss is a C++ library with Python and numpy wrappers for efficient similarity search and clustering of dense vectors, developed primarily at Meta's Fundamental AI Research group. It handles billions of embeddings within a fixed memory budget by offering two distinct scaling strategies: compact quantization codes that compress representations to fit in RAM (trading precision for scale), and graph-based indexes like HNSW and NSG that layer structure over raw vectors for speed.
    - The GPU path is a drop-in replacement; swapping `IndexFlatL2` for `GpuIndexFlatL2` handles memory copies automatically
    - The only hard dependency is a BLAS implementation; CUDA, ROCm, and the Python interface are all optional
    - It explicitly includes tooling for parameter tuning rather than hiding the trade-offs, and maintains a dedicated troubleshooting page
  7. Rulin Shao writes about Context Language Models (CLMs), which manage their own context by treating it as a file the model can edit with Bash, replacing harness-defined compaction and retrieval rules. Zero-shot applications to existing models outperform state-of-the-art context management strategies across diverse benchmarks.

    - Multiple agent contexts can coexist as files, extending the design to agent swarms and subagents.
    - Suffix Cache Reuse was co-designed to cut server-side compute by 35% against standard SGLang at matched performance.
    - A model-editable context creates a new channel through which injected or self-written instructions can persist across turns.
  8. Shapeshift is a Next.js + React 19 component that turns a single text input into context-aware UI cards. As you type, the box morphs into event cards, checklists, timers, color pickers, bill splitters, converters, polls, and more based on detected intent. Intent classification is powered by TypeSafe Jev, which answers 14 typed questions in a single parallel call, while all value extraction (dates, amounts, units, math) is handled by deterministic code. The project works fully offline by default using a built-in keyword classifier and can optionally use the online Jev model with an API key.
    - One Jev call answers 14 typed questions in parallel (speculative fan-out), covering card type plus signals like "is it a video call?" and "is it urgent?"
    - A state machine with hysteresis prevents flickering: a card only changes when a challenger wins twice in a row or is very confident
    - 19 card types supported, including trip planner, time zone converter, dice roller, countdown, and goal tracker
    - Built with TypeScript (strict), Tailwind CSS v4, shadcn/ui, Motion, chrono-node, and zod
    - Live demo at shapeshiftui.vercel.app
  9. Richard Gill writes about his personal Pi coding agent setup, which utilizes OpenAI Codex Sol and Astra models at medium and high thinking levels while adhering to Pi's philosophy of simplicity. He relies primarily on `AGENTS.md` files and custom skills rather than complex configuration.
    - Commands taking over 30 seconds automatically move to the background to prevent the agent from getting stuck
    - The `sub-pi` extension enables spawning new Pi windows and worktrees via tmux
    - Slash commands like `/diff` inject command output directly into context without triggering an LLM turn
    - Context files and skills traverse parent directories up to `$HOME`
  10. llama.cpp now supports decision models via a `/v1/systemone` endpoint, which accepts a state and typed questions to return probabilities for each provided option in a single forward pass. The API utilizes the System One format introduced with TypeSafe's Jev model, meaning existing clients only need a new base URL. Five initial models are available ranging from 144M to 27B parameters, supporting use cases like request routing, content moderation, and agent action selection with median response times as low as 3 ms.
    - The `state` field accepts text, JSON, screenshots, or a list of chat messages with `image_url` parts.
    - Three question types are supported: `choice` for categorical selection, `score` for a continuous level between 2 and 10 options, and `noul` for yes/no probabilities.
    - Router mode allows loading multiple models on a single server and selecting one per request.
    - Adding descriptions to option labels can significantly improve accuracy, with Julia-1's routing confidence for "charged twice" jumping from a misroute to 0.99 when descriptions were provided.
    - Cloudflare's Clef is the next model planned for integration.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: large language model

About - Propulsed by SemanticScuttle