Tags: python* + llm*

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. The article clarifies that RAG and fine-tuning are complementary rather than competing techniques for LLM development. RAG works by retrieving external information at inference time, which enables models to access new data and provide citable answers without changing the model weights. In contrast, fine-tuning adjusts a model's internal weights to improve its behavior, such as tone or adherence to specific output formats like JSON.

    - RAG provides dynamic knowledge retrieval for accuracy and traceability.
    - Fine-tuning improves task performance, style, and formatting consistency.
    - Combining both methods allows developers to manage both what a model knows and how it communicates.
  2. An examination of the hype surrounding autonomous AI agent frameworks and why they may add unnecessary complexity to software development. The author argues that for most production use cases, structured workflows using LLM function calling are more reliable than fully autonomous agents.

    - Complexity vs control in agentic systems
    - Limitations of current models regarding long-term autonomy
    - Advantages of explicit programming over unpredictable loops
  3. An open-source command-line tool designed to identify the optimal local Large Language Model specifically suited for a user's existing or planned hardware. It automatically detects GPU, CPU, and RAM capacity to rank HuggingFace models using real performance benchmarks instead of relying on parameter size alone.

    * Hardware auto-detection for NVIDIA, AMD, Apple Silicon, and CPUs
    * Intelligent ranking based on benchmark evidence and recency awareness
    * Capability to simulate different GPUs for hardware upgrade planning
    * Support for GGUF, AWQ, and GPTQ model formats
    * Streamlined workflows including one-command chat sessions and Python code snippet generation
  4. Simon Willison explores his latest approach to running untrusted Python code safely within applications by utilizing MicroPython inside a WebAssembly (WASM) sandbox. The project addresses the security risks of plugin systems where code normally executes with full privileges, potentially leading to data leaks or system compromise. By leveraging wasmtime and an alpha package called micropython-wasm, Willison demonstrates how to enforce memory and CPU limits while providing controlled access to host functions through a custom thread-based request queue for persistent interpreter state.

    Main topics:
    - Security challenges in Python plugin systems
    - Advantages of WebAssembly as a sandboxing technology
    - Building the micropython-wasm alpha package
    - Implementation details for persistent state and host functions
    - Integration with Datasette Agent to execute code via LLMs
  5. This tutorial demonstrates how to evolve a standard chatbot into a truly agentic system using the Gemma 4 model family. Instead of relying solely on remote web APIs, it shows how to provide the model with tools that interact directly with the local environment—specifically a sandboxed filesystem explorer and a restricted Python interpreter. By implementing security measures like path-traversal guards for file access and whitelisted builtins for code execution, users can safely allow small models running locally on laptops to observe their surroundings and perform deterministic calculations.
    Main topics:
    * Transitioning from API retrieval to true agency through local system interaction.
    * Building a secure filesystem explorer with path-traversal protection.
    * Implementing a restricted Python interpreter using exec() and whitelisted builtins.
    * Orchestrating tool calls using Gemma 4 and Ollama for local agentic workflows.
  6. AI agents operate through a ReAct (Reason + Act) pattern implemented as a deterministic Python `while` loop that maintains conversation history within the context window to serve as short-term memory. The core logic involves sending the system prompt and cumulative tool results to an LLM, which returns either a final answer or structured function calls; if tools are requested, their outputs are executed and appended back into the message list for subsequent reasoning iterations. This architecture supports local execution via Ollama's OpenAI-compatible API, mixed-mode orchestration by delegating complex tasks from local models to cloud APIs through specialized tool functions, and scalable tool integration using the Model Context Protocol (MCP) to dynamically discover and invoke external services via JSON-RPC.
    2026-05-18 Tags: , , , by klotz
  7. >"A practical comparison between rule-based PDF extraction using pytesseract and an LLM-based approach with Ollama and LLaMA 3, based on a realistic B2B order scenario."

    - Regex approach excels with stable, standardized layouts.
    - Regex offers speed, low cost, and deterministic results but requires high maintenance as document variety increases.
    - LLM approach (using LLaMA 3 via Ollama) leverages semantic context to handle diverse field names and formats automatically.
    - LLM approach reduces manual rule updates.
    - Trade-offs: LLMs provide superior flexibility for complex layouts but incur higher latency, greater infrastructure costs, and probabilistic uncertainty compared to traditional methods.
    - Selection depends on document stability, required throughput, and need for explainability.
    2026-05-16 Tags: , , , by klotz
  8. Memori is an agent-native memory infrastructure that acts as an LLM-agnostic layer to transform AI agent execution and conversations into structured, persistent state for production systems. It integrates seamlessly into existing architectures, allowing agents to automatically capture and recall information from past interactions without requiring changes to core code or prompts.
    Key features and points:
    * Provides advanced augmentation of memories including attributes, facts, preferences, relationships, and skills at the entity, process, and session levels.
    * Achieves high accuracy and token efficiency in long-conversation memory as demonstrated by LoCoMo benchmark results.
    * Offers dedicated SDKs for both Python and TypeScript.
    * Supports Model Context Protocol (MCP) for easy connection to developer tools like Claude Code and Cursor.
    * Compatible with a wide range of LLMs including OpenAI, Anthropic, Gemini, DeepSeek, and Grok, as well as frameworks like LangChain and Pydantic AI.
  9. This tutorial demonstrates how to construct a complete skill-based agent system for large language models using Python. It explores structuring modular capabilities similar to an operating system, where reusable skills are defined with metadata and schemas, registered centrally, and orchestrated through dynamic tool calling and multi-step reasoning. The implementation covers composing multiple skills for advanced workflows, hot-loading new capabilities at runtime, and monitoring performance via an observability dashboard.
    2026-05-11 Tags: , , , , , by klotz
  10. A post-retrieval temporal layer designed to improve RAG systems by addressing time-blindness in vector searches. This library implements validity filtering, document kind classification, and exponential decay scoring to ensure retrieved information is fresh and accurate. It functions downstream of existing vector search systems without requiring re-indexing or new infrastructure.

Top of the page

First / Previous / Next / Last / Page 2 of 0 SemanticScuttle - klotz.me: tagged with "python+llm"

About - Propulsed by SemanticScuttle