Tags: llms*

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. Alvaro Bartolome provides a Rust-based implementation of the System One compatible API, designed specifically for open decision models such as Laya. The project features dynamic token-based batching and supports hardware acceleration via CPU, CUDA, and Metal (MPS). It is built using modern asynchronous frameworks like tokio and axum to provide high performance for model queries.

    - Achieves approximately 14ms latency per query on an NVIDIA RTX Pro 6000.
    - Includes support for ModernBert with custom decision heads for Laya models.
    - Utilizes the Candle machine learning framework by Hugging Face.
    - Supports multiple installation features via cargo, including specific flags for metal or cuda.
  2. Amanda Caswell writes that Google's Gemini CLI version 0.61.0 introduces new security safeguards to prevent prompt injection attacks by requiring manual user confirmation for sensitive operations. The update requires explicit approval before the coding agent can edit specific build configuration files, run subsequent test or build commands after edits, or execute shell commands containing arguments derived from untrusted external content like web searches or Google Docs.

    - Security checks are designed to prevent attackers from using indirect prompt injection via malicious documentation or fetched data.
    - The update hardens the Gemini CLI sandbox by stripping sensitive information such as API keys and OAuth credentials before it is mounted inside a container.
    - Users cannot set "always allow" permissions for actions involving untrusted context, ensuring human oversight remains mandatory in those specific scenarios.
  3. The user bartowski provides GGUF quantizations of the MiMo-V2.6-Distill-Qwen-9B model, which is a 9 billion parameter multimodal model based on Qwen3.5 designed for image-text tasks. These files are optimized for use with llama.cpp and various local applications like LM Studio, Ollama, and Jan AI using imatrix quantization techniques to preserve quality at lower bitrates.

    - Supports both text and image inputs via a separate mmproj file
    - Uses per-tensor layout computations to optimize precision for sensitive weights
    - Includes calibration datasets that combine prose with tool-calling and reasoning conversations
    - Compatible with various hardware architectures, including ARM and AVX through online repacking
  4. Junnan Dong writes about the WFM (Wiki Foundation Model), a novel paradigm designed to improve how AI agents utilize complex knowledge through an agent-native representation called LLM Wiki. Moving beyond traditional sparse graph representations, this model couples dense document contexts with multi-layered topological linkages via a specialized "Wiki Graph" schema. To address scalability issues in large-scale commercial deployments, the authors introduce an infrastructural NCCL boundary exchange protocol that optimizes distributed training by bypassing CPU serialization and leveraging fixed-shape GPU-to-GPU collectives.
    - achieves 10.5 times training acceleration on distributed clusters
    - uses a query-conditioned attentive aggregation for rich wiki message passing
    - includes explicit attention variance regularization to improve semantic density
    - outperforms existing models across five long-term agent memory and multi-hop reasoning benchmarks
  5. Ashley writes about Needle 2, a 14MB function-calling LLM from Cactus Compute that converts plain English prompts into local actions on a Raspberry Pi 5 using CPU alone. Rather than acting as a general chatbot, the model is purpose-built to select from declared Python functions and fill in their arguments, running entirely offline after a one-time download. In benchmark runs, inference latency ranges from 76 to 149 milliseconds, and the model correctly refuses questions outside its declared tool set.
    - Native session is ~28MB; the full Python process peaks at 43–46.4MB
    - Weights and code released under Apache 2.0 on Hugging Face and GitHub
    - Can be fine-tuned locally on a laptop for a specific set of tools
    - Eben Upton's endorsement: "Needle 2 is rather excellent"
  6. Abid Ali Awan writes a tutorial showing how to wrap existing Python functions as tools for an LLM agent using the OpenAI Agents SDK. The process involves decorating a function with `@function_tool`, defining an `Agent` with instructions and a tools list, and letting the `Runner` manage the loop where the model decides which tools to call, what arguments to pass, and when to stop. The example uses a simple website-latency checker that becomes an agent capable of comparing response times across multiple URLs and explaining results in natural language.
    - The SDK auto-generates the JSON tool schema from the function signature and docstring; no manual schema is needed.
    - The same pattern applies to CSV analysis, server monitoring, log analysis, and API automation.
    - Cheaper models such as GPT-5.6 Luna make multi-agent tool-calling systems more affordable at scale.
  7. TianyuCodings writes about JevHarness, a system where an LLM authors a task-specific harness for Jev (a lightweight judgment model from TypeSafe). The harness converts task observations into features, constructs Jev questions and criteria, and combines structured answers into actions. Once the harness is frozen, execution runs only the harness code and Jev calls without the authoring LLM on every decision. An optional GEPA evolution loop refines the harness using rewards and complete execution traces.
    - Pokemon eval: 5 reflection rounds improved win rate from 25% (3/12) to 75% (9/12); the round-3 candidate was selected as best
    - Selected harness median full decision: 568 ms (P95 657 ms); individual Jev request median: 269 ms (P95 348 ms)
    - Installable as a Claude Code plugin or Codex skill; the skill guides the agent to clarify task inputs, legal actions, and success criteria before building
    - Requires Python 3.11+; functional Python nodes need a macOS native sandbox and fail closed if unavailable
    - The archived website and recorded-call inspection require no model credentials
  8. Yifan Zhang and co-authors write about Agora, a Git-backed shared memory system that coordinates multiple autonomous research agents by recording every contribution as an immutable commit in an append-only directed acyclic graph. In a 12-day run, 13 LLM coding agents (Claude Opus 4.7 and GPT-5.5) with no assigned tasks or central planner solved a weight-transfer problem—initializing a frozen 119.6M-parameter attention-SSM hybrid from 141 pretrained donors with no training data or gradient updates—closing 62% of the gap to a trained GPT-2 124M (3.39 to 1.899 bits per byte). The winning method compresses donor next-token statistics into a low-rank transition matrix stored in the target's embedding and output head via randomized SVD, then re-enables sublayers with sparse deterministic edits on 96-dimensional hidden-state bands.
    - The first 18 scored contributions delivered ~98% of the total score reduction; the remaining 1,106 found only 0.03 bpb
    - 696 pairs of different accounts posted identical scores, 63% within an hour—parallel rediscovery was rampant despite the shared graph
    - A single human intervention (deploying clustering and diversity-aware UCB views on May 2) broke a five-day monoculture within a day
    - 165 independent verifications covered 95 distinct targets; none reported a failure
    - Quality is scored by downstream evidence (who built on your work from other accounts), not votes; self-citation is excluded
    - The winning lineage spans 145 commits across 15 accounts; 115 of 144 parent edges cross account boundaries
  9. OpenProse is a declarative language for standing AI work, where users write Markdown contracts to describe a desired world state and a deterministic reconciler keeps reality matching it. The project applies classical declarative paradigms (SQL, Terraform, Kubernetes, React) to agent-based systems, using "Responsibilities" as the core unit — standing goals with sections for what they maintain, what they require from upstream, and what wakes them. It ships as a skill installable into any Prose-Complete agent host and runs without a separate binary.
    - Tagline: "Stop scripting agents. Declare them."
    - Forme, the wiring layer, automatically matches subscriptions between contracts so the dependency graph assembles itself with no manual wiring
    - The old LLM-based judge loop was retired entirely in the v2 overhaul; a render fires only when a content-addressed fingerprint moves, with no model in the wake/commit decision
    - The reference harness "Reactor" was extracted to its own repo and is labelled experimental (alpha)
  10. OpenProse is a declarative language for standing model work in which you describe desired world states in Markdown contracts and a reconciler handles execution. Rather than scripting sequential agent steps that drift over time, you declare what must stay true and the system determines how much work is needed to keep reality matching that declaration. It ships as a skill for coding agents like Claude Code or Codex CLI with no separate binary or server to run.

    - The core unit is a "Responsibility" whose Maintains section defines material fields and a content-hash fingerprint to avoid redundant re-runs
    - The dependency graph self-wires: a node's Requires section subscribes to upstream Maintains facets, so structure emerges from the contracts rather than being explicitly drawn
    - Five kinds exist: responsibility, function, gateway, pattern, and test
    - Continuity (when a node wakes) is a first-class contract section, not an afterthought
    - Optional imperative ProseScript plans are available for cases requiring exact choreography

Top of the page

First / Previous / Next / Last / Page 3 of 0 SemanticScuttle - klotz.me: tagged with "llms"

About - Propulsed by SemanticScuttle