All Bookmarks

Welcome to SemanticScuttle! Social bookmarking for small communities.

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. Ivan Maradzhiyski writes about innomd, a command-line tool designed for Linux and macOS that renders LaTeX math formulas as clean Unicode directly in the terminal. Built on top of the `rich` library, it serves scientists, engineers, and students by providing human-readable mathematical notation (such as Greek letters, operators, and fractions) instead of raw LaTeX source code. The tool also features support for rendering Mermaid and PlantUML diagrams into ASCII/Unicode representations and includes a live reload mode for real-time Markdown previewing.

    - Renders common subsets of Mermaid and PlantUML (including flowchart, sequence, class, Gantt, C4 architecture, and activity diagrams).
    - Supports Jupyter notebook (`.ipynb`) files by rendering cells as Markdown with syntax highlighting.
    - Developed in pure Python using `rich` and `grandalf`, requiring no external binaries like Graphviz or Node.js for diagram layout.
    - Features nine built-in color themes, including Nord, Dracula, and Solarized.
    - Uses Unicode approximation rather than pixel-perfect rendering, making it ideal for terminal use but not as a substitute for PDF/LaTeX compilation.
  2. GitButler is a version control tool designed to optimize workflows for developers and AI agents by introducing features like stacked branches, parallel branching, and unlimited undo capabilities. It layers seamlessly onto existing Git repositories without requiring new configuration, aiming to reduce the friction often associated with complex Git operations through structured output and idempotent commands.

    - Agents are reported to be 60% faster using GitButler compared to vanilla Git
    - The software includes a CLI that offers JSON output mode for better AI parsing
    - Features include "Smartlog" and simplified history editing/rebasing
    - It is free and open source software
    2026-09-28 Tags: , , , , , by klotz
  3. Ruoqi Guo et al. present RLCDAlignBench, a new benchmark designed to evaluate the ability of Jev—a model trained via reinforcement learning for calibrated decisions (RLCD)—to detect various types of alignment failures in language models zero-shot. The study examines ten specific failure modes, including sycophancy, jailbreaks, and hallucination, across 44 benchmarks and five target models. Results indicate that a single generic question applied to Jev achieves a median AUROC of 0.886, outperforming supervised baselines in many cases while being significantly more cost-effective than using LLM-judge scorers.
    - Evaluates ten failure modes: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking.
    - Uses a "relational" detection strategy by varying the question wording separately from input fields to handle failures defined against external references.
    - Jev costs 63x less than traditional LLM-judge scorers.
  4. Maya Posch writes about the reverse-engineering of the Intel 8087 FPU, specifically focusing on how it implements trigonometric functions like `FPTAN`. To achieve high accuracy for 64-bit values efficiently, the hardware uses a hybrid approach that combines the CORDIC algorithm with Padé approximants. The system first performs most calculations using CORDIC and then switches to polynomial approximation once the remaining value is small enough, allowing it to avoid large look-up tables or excessive processing time.

    - For one calculated value of `FPTAN`, CORDIC pseudo-division takes 33% of the time, while pseudo-multiplication takes 47%.
    - The polynomial approximation stage accounts for only about 15% of the execution time with a 5% overhead.
    - Later CPUs like the Pentium series moved away from CORDIC because it is difficult to scale for high bit-accuracy without significant performance penalties.
  5. Zhening Li, Joshua Liu, Mateja Vukelic, Nicole Shen, Supriya Lall, Amitayush Thakur, and colleagues at MIT CSAIL introduce JAZ, an agent framework that utilizes a single primitive called `invoke` to perform tasks typically requiring specialized memory or self-improvement systems. By treating the LLM as a runtime provider for function implementations through executable code, the framework allows all inputs and interaction histories to act as variables in the environment.
    - The system uses "hooks" instead of dedicated subsystems like file systems or external memory stores to apply constraints and monitoring.
    - JAZ outperformed Letta (MemGPT) by 8% at half the cost on recall-heavy tasks within the StuLife dataset.
    - In self-improvement evaluations on AppWorld, JAZ exceeded ACE performance by 4% while maintaining a lower cost.
  6. Yohei Nakajima writes about glance, a tool designed to allow users to ask an open vision-language model (VLM) typed questions about images and receive probability data directly on their own machine. Rather than generating new text or training models, it acts as a measurement harness that reads logits from frozen models—such as Qwen3-VL-4B by default—to provide yes/no answers, single-choice selections, and qualitative ratings without any image data leaving the user's device.

    - The tool provides three response types: "noul" (yes/no), choice (pick one from a list), and score (a rating on a specified scale).
    - It includes an experimental MLX backend to provide faster runtimes specifically for Apple Silicon users.
    - Glance can perform self-calibration using unlabeled data or precise calibration through labeled datasets to improve rating accuracy.
    - The software is designed with privacy in mind, ensuring all inference and logging stay local on the user's hardware.
  7. The Laya AI model is a family of open-weight, non-autoregressive decision models from Convai Innovations designed to return structured answers and probabilities rather than free-form text. By mapping input states to specific question types like choices, scores, or propositions (noul), it provides deterministic outputs suitable for application policies in workflows such as support ticket routing. The guide details its architecture—utilizing bidirectional encoders like ModernBERT—and offers practical advice on running the model locally via Python and evaluating performance through metrics like calibration and accuracy.

    - Laya uses a "state + typed questions" pattern to ensure output validity without needing complex parsing of generative prose.
    - It offers three specific checkpoints: an English version, a multilingual version (mmBERT), and one optimized for typed decisions.
    - The model's architecture relies on bidirectional encoders with decision heads rather than token-by-token generation.
    - Users are encouraged to implement "abstention policies" where uncertain predictions (based on low confidence/calibration) are routed to humans.
  8. Matt von Hippel writes about how researchers at Anthropic successfully challenged an LLM to solve a complex problem in theoretical particle physics by computing a nine-loop amplitude in N=4 super Yang-Mills. The task involved using the bootstrap method and form-factor approaches through Claude Science, demonstrating that AI can autonomously manage long-running computational tasks with high reliability without constant human oversight.
    - The calculation was performed using Fable 5.1 within the Claude Science platform.
    - While humans have calculated up to eight loops in this toy model, no one had yet reached nine loops directly.
    - The computation cost approximately $100 when using Python and SymPy for the bootstrap method.
    - Physicist Lance Dixon validated that the result was correct by checking it against a form-factor approach he has been working on for two years.
  9. Zhening Li and colleagues introduce JAZ, an LLM agent framework designed to minimize the complexity of agent loops by treating them as a programming language primitive called `invoke`. Instead of relying on external specialized systems for memory or self-improvement, JAZ enables agents to achieve these capabilities through code execution where all interactions are treated as variables within the environment. This minimalist approach allows highly expressive workflows, such as long-horizon recall and continual self-improvement, using only prompting rather than manually designed tools or complex architectures.

    - The `invoke` primitive allows for recursive calls, enabling LLMs to write arbitrary executable code that includes further iterations of itself.
    - In testing on the StuLife dataset, JAZ outperformed MemGPT (Letta) by 8% in recall performance while costing half as much.
    - On self-improvement tasks using AppWorld, JAZ demonstrated a 4% improvement over ACE at a lower computational cost.
  10. Tyler August writes that the Cheap Yellow Display (CYD) ESP32 dev boards can be transformed into a Portable Digital Assistant, reminiscent of 1990s Palm devices. Rather than using an emulator, user sau412 developed custom firmware for the CYD that includes features like a Gopher browser, RSS and Wikipedia readers, various games, a BASIC interpreter, a calendar, calculator, contact list, ebook reader, and translator.
    - The project can be set up in approximately five minutes using a web-based flasher.
    - An SD card is required to test all the available applications.
    - This project provides an alternative for those looking for Palm OS-style functionality on modern hardware without needing x86 or ARM emulators like Pumpkin OS.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Recent bookmarks

About - Propulsed by SemanticScuttle