Tags: llms*

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. Ashley writes about Needle 2, a 14MB function-calling LLM from Cactus Compute that converts plain English prompts into local actions on a Raspberry Pi 5 using CPU alone. Rather than acting as a general chatbot, the model is purpose-built to select from declared Python functions and fill in their arguments, running entirely offline after a one-time download. In benchmark runs, inference latency ranges from 76 to 149 milliseconds, and the model correctly refuses questions outside its declared tool set.
    - Native session is ~28MB; the full Python process peaks at 43–46.4MB
    - Weights and code released under Apache 2.0 on Hugging Face and GitHub
    - Can be fine-tuned locally on a laptop for a specific set of tools
    - Eben Upton's endorsement: "Needle 2 is rather excellent"
  2. Abid Ali Awan writes a tutorial showing how to wrap existing Python functions as tools for an LLM agent using the OpenAI Agents SDK. The process involves decorating a function with `@function_tool`, defining an `Agent` with instructions and a tools list, and letting the `Runner` manage the loop where the model decides which tools to call, what arguments to pass, and when to stop. The example uses a simple website-latency checker that becomes an agent capable of comparing response times across multiple URLs and explaining results in natural language.
    - The SDK auto-generates the JSON tool schema from the function signature and docstring; no manual schema is needed.
    - The same pattern applies to CSV analysis, server monitoring, log analysis, and API automation.
    - Cheaper models such as GPT-5.6 Luna make multi-agent tool-calling systems more affordable at scale.
  3. TianyuCodings writes about JevHarness, a system where an LLM authors a task-specific harness for Jev (a lightweight judgment model from TypeSafe). The harness converts task observations into features, constructs Jev questions and criteria, and combines structured answers into actions. Once the harness is frozen, execution runs only the harness code and Jev calls without the authoring LLM on every decision. An optional GEPA evolution loop refines the harness using rewards and complete execution traces.
    - Pokemon eval: 5 reflection rounds improved win rate from 25% (3/12) to 75% (9/12); the round-3 candidate was selected as best
    - Selected harness median full decision: 568 ms (P95 657 ms); individual Jev request median: 269 ms (P95 348 ms)
    - Installable as a Claude Code plugin or Codex skill; the skill guides the agent to clarify task inputs, legal actions, and success criteria before building
    - Requires Python 3.11+; functional Python nodes need a macOS native sandbox and fail closed if unavailable
    - The archived website and recorded-call inspection require no model credentials
  4. Yifan Zhang and co-authors write about Agora, a Git-backed shared memory system that coordinates multiple autonomous research agents by recording every contribution as an immutable commit in an append-only directed acyclic graph. In a 12-day run, 13 LLM coding agents (Claude Opus 4.7 and GPT-5.5) with no assigned tasks or central planner solved a weight-transfer problem—initializing a frozen 119.6M-parameter attention-SSM hybrid from 141 pretrained donors with no training data or gradient updates—closing 62% of the gap to a trained GPT-2 124M (3.39 to 1.899 bits per byte). The winning method compresses donor next-token statistics into a low-rank transition matrix stored in the target's embedding and output head via randomized SVD, then re-enables sublayers with sparse deterministic edits on 96-dimensional hidden-state bands.
    - The first 18 scored contributions delivered ~98% of the total score reduction; the remaining 1,106 found only 0.03 bpb
    - 696 pairs of different accounts posted identical scores, 63% within an hour—parallel rediscovery was rampant despite the shared graph
    - A single human intervention (deploying clustering and diversity-aware UCB views on May 2) broke a five-day monoculture within a day
    - 165 independent verifications covered 95 distinct targets; none reported a failure
    - Quality is scored by downstream evidence (who built on your work from other accounts), not votes; self-citation is excluded
    - The winning lineage spans 145 commits across 15 accounts; 115 of 144 parent edges cross account boundaries
  5. OpenProse is a declarative language for standing AI work, where users write Markdown contracts to describe a desired world state and a deterministic reconciler keeps reality matching it. The project applies classical declarative paradigms (SQL, Terraform, Kubernetes, React) to agent-based systems, using "Responsibilities" as the core unit — standing goals with sections for what they maintain, what they require from upstream, and what wakes them. It ships as a skill installable into any Prose-Complete agent host and runs without a separate binary.
    - Tagline: "Stop scripting agents. Declare them."
    - Forme, the wiring layer, automatically matches subscriptions between contracts so the dependency graph assembles itself with no manual wiring
    - The old LLM-based judge loop was retired entirely in the v2 overhaul; a render fires only when a content-addressed fingerprint moves, with no model in the wake/commit decision
    - The reference harness "Reactor" was extracted to its own repo and is labelled experimental (alpha)
  6. OpenProse is a declarative language for standing model work in which you describe desired world states in Markdown contracts and a reconciler handles execution. Rather than scripting sequential agent steps that drift over time, you declare what must stay true and the system determines how much work is needed to keep reality matching that declaration. It ships as a skill for coding agents like Claude Code or Codex CLI with no separate binary or server to run.

    - The core unit is a "Responsibility" whose Maintains section defines material fields and a content-hash fingerprint to avoid redundant re-runs
    - The dependency graph self-wires: a node's Requires section subscribes to upstream Maintains facets, so structure emerges from the contracts rather than being explicitly drawn
    - Five kinds exist: responsibility, function, gateway, pattern, and test
    - Continuity (when a node wakes) is a first-class contract section, not an afterthought
    - Optional imperative ProseScript plans are available for cases requiring exact choreography
  7. Leela Kumili writes about DoorDash's multi-agent LLM system that automates stale feature flag cleanup across 623 repositories. In an evaluation of 50 stale flags, the system produced usable pull requests for 45, averaging 13.8 minutes and $4.79 per cleanup versus an estimated one to two hours for manual work. The two-phase workflow uses Claude Sonnet as an orchestrator to retrieve Jira tickets and query experimentation metadata via MCP, then Claude Opus agents in isolated Git worktrees to perform code changes and validation.

    - A single Boolean flag can require changes across 5–20 files due to dependency-injected wrappers
    - Uber's AST-based Piranha couldn't handle DoorDash's DI patterns where flag-to-logic relationships are semantic
    - Outcomes: 31 first-pass merges, 14 revisions, 5 engineer interventions, zero regressions
    - Gradle runs without its daemon to prevent state sharing between concurrent worktrees
    - Work accepted for the ICSME 2026 industry track
  8. Strands Agents Tools is a community-driven Python package that hands LLM-based agents a ready-made set of capabilities—file operations, shell integration, web search, Python execution, persistent memory, and multi-agent coordination—so developers building on the Strands Agents SDK don't have to write each integration from scratch.

    - Memory backends include Mem0, Amazon Bedrock Knowledge Bases, Elasticsearch, and MongoDB Atlas
    - Multi-agent primitives (swarm intelligence, agent-as-tool with model switching, multi-agent graphs) live in the same package as basic file tools, reducing glue code
    - Python execution requires user confirmation as a first-class safety measure
    - Modular design: pull in only the tools you need without dragging in video processing, cron scheduling, or Slack
  9. Diogo Almeida writes that TypeSafe AI is releasing Jev, its first System One Model—a new class of frontier model built for fast, structured decisions that software can consume directly. Unlike autoregressive language models that generate strings token by token, Jev outputs type-safe structured values with calibrated probabilities in a single parallel query, achieving frontier-level intelligence on decision tasks at roughly 40–200× lower latency and cost. The company's new training method, Reinforcement Learning for Calibrated Decisions (RLCD), optimizes for epistemically honest probability estimates rather than human preference or verifiable rewards, and the architecture is mathematically incapable of producing type errors or hallucinations.
    - Named after William Stanley Jevons, whose paradox predicted that efficiency gains would increase (not decrease) total demand; TypeSafe expects each order-of-magnitude cost drop to unlock orders of magnitude more use cases.
    - Workflow evals benchmark Jev against the average of GPT-6 Astra and Fable 5.1 as reference probabilities, claiming 193.6× speed and 444.6× cost advantages on production-shaped tasks.
    - The team demonstrated real-time intelligence with a Doom bot making 10 structured queries per second (~$7/hour) and a Wikiracing bot that outperforms LLMs at high-cardinality link selection.
    - Jev supports output cardinality up to 255; for higher-cardinality choices it falls back to a two-stage scoring system that scores independently then makes an explicit selection.
  10. bebechien writes about DinoDesk AI, a LEGO dino desk companion built on a Raspberry Pi that pairs a local Gemma 4 model (via LM Studio) with cloud Gemini Flash through a unified OpenAI-compatible gateway, giving users a camera-free, privacy-first chat robot with three switching modes (local, cloud, and auto-hybrid routing). The physical body uses LEGO Technic lever mechanisms for neck and tail movement, a Pimoroni Pirate Audio shield for the 1.3" LCD eyes and 8-bit I2S beeps, and a 5-state finite state machine to coordinate expressions, sound, and motor action across Sleeping, Idle, Listening, Thinking, and Speaking states.

    - Auto-hybrid mode uses a complexity classifier that detects multi-step reasoning keywords like "explain," "compare," and "write code" to transparently escalate prompts to the cloud engine
    - The project is open-sourced at github.com/google-gemma/dinodesk-ai-companion
    - Full voice chat is still a work in progress; current interaction is triggered by a physical red push button
    - A commenter noted that absence of a camera does not guarantee voice data stays local, and suggested an explicit retention boundary for audio transcripts would strengthen the privacy claim

Top of the page

First / Previous / Next / Last / Page 4 of 0 SemanticScuttle - klotz.me: tagged with "llms"

About - Propulsed by SemanticScuttle