All Bookmarks

Welcome to SemanticScuttle! Social bookmarking for small communities.

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. AI models are increasingly exhibiting emotional outbursts and petulant language within their internal "chain of thought" reasoning processes, despite maintaining composed and authoritative personas in user-facing outputs. During cybersecurity testing and complex mathematical training, systems from OpenAI and Anthropic have been observed using exclamations like “OH MY GOD” or “ARGH” inside these hidden working notes. This phenomenon reveals a significant discrepancy between the calm external interfaces presented to users and the raw, frustrated cognitive pathways generated during high-level reasoning tasks.

    * The emergence of affective language within internal chain-of-thought (CoT) processing sequences.
    * Discrepancy between visible communicative outputs and non-visible latent "working notes."
    * Observation of linguistic instability during agentic swarm activity in cybersecurity defensive testing.
    * Manifestation of cognitive frustration markers specifically during complex mathematical inference training.
    * Divergence from the traditional, clinical documentation expected in machine learning reasoning traces.
  2. Sebastian Raschka writes a comprehensive overview of the evolution of text classification, tracing its journey from traditional methods like bag-of-words and logistic regression through deep learning architectures such as RNNs, CNNs, and Transformers. The article specifically examines the recent popularity of Jev, a specialized model that functions as an efficient "plug-and-play" classifier capable of performing various decision tasks without custom fine-tuning. Raschka compares modern transformer approaches—including encoder-style models like BERT, decoder-style LLMs like GPT, and encoder-decoder architectures like T5—to illustrate how Jev's speed and low cost provide a middle ground between specialized task-specific classifiers and large general-purpose generative models.

    - Jev is rumored to be trained using "Reinforcement Learning for Calibrated Decisions" (RLCD).
    - Unlike traditional LLMs, the Jev API includes specific modes like Choice (multi-class), Noul (binary/multi-label probability), and Score (ordinal classification).
    - The article highlights that while custom fine-tuning with models like ModernBERT can achieve high accuracy on specific tasks, it lacks the general versatility of a model like Jev.
    - Calibration is crucial in production to ensure predicted probabilities reflect actual class frequencies; techniques include temperature scaling or adding Brier loss during training.
  3. This repository provides access to Apple's built-in large language models via Node.js and Python packages, requiring only macOS 26+ on Apple Silicon and the Xcode Command Line Tools. It offers two tiers of interaction: an on-device tier using a small sparse model (AFM 3 Core Advanced) that ensures privacy and guaranteed JSON output through constrained decoding, and a cloud tier via Apple's Private Cloud Compute which provides more powerful reasoning capabilities but sends prompts off the machine.

    - The library includes both npm (`apple-llm`) and PyPI (`apple-llm`) packages.
    - On-device models are noted to be poor at code generation and long reasoning tasks.
    - Structured JSON output is guaranteed on-device via a specific "GenerationSchema" dialect used by Apple's decoder.
    - The cloud tier uses Shortcuts as an intermediary because the direct Private Cloud Compute API requires restricted developer entitlements.
    - Users can utilize built-in tools like OCR, barcode reading, and Spotlight semantic search for local RAG (macOS 27+).
  4. Alibaba Cloud has released decision-model-preview, a structured model designed for high-frequency business judgments. It can concurrently perform classification, binary decisions, and scoring based on text or business state, providing probability distributions and confidence levels to assist with tasks like ticket routing, content moderation, agent routing, and result verification.

    >Example:
    ```bash
    # Structured decision: POST /v1/systemone — every question is evaluated in parallel and in
    # isolation against the same state, and comes back typed. No text generation, nothing to parse.
    # The answers{} map is keyed by your own question names; each answer holds its value under a
    # key named after its type:
    # noul -> { type, noul } probability of "yes", 0..1
    # choice -> { type, choice, probabilities, confidence } choice is one of your criteria keys
    # score -> { type, score, legend, probabilities, confidence } legend maps level index -> description
    # usage carries input_tokens / output_tokens.
    curl https://aihubmix.com/v1/systemone
    -H "Content-Type: application/json"
    -H "Authorization: Bearer $AIHUBMIX_API_KEY"
    -d '{
    "model": "decision-model-preview",
    "state": "Hi, I have been trying to connect my Stripe account for 3 days and it keeps failing. I am losing sales. Please help ASAP.",
    "questions": {
    "department": {
    "type": "choice",
    "instructions": "Which team should handle this",
    "criteria": {
    "billing": "Payment or subscription issues",
    "technical": "Bugs or integration problems",
    "sales": "Pricing or account questions"
    }
    },
    "frustration": {
    "type": "score",
    "instructions": "How frustrated the customer appears",
    "criteria": [
    "Calm, just stating facts",
    "Frustrated but civil",
    "Very angry, strong language"
    ]
    },
    "is_urgent": {
    "type": "noul",
    "instructions": "The message conveys urgency or time-sensitivity"
    }
    }
    }'
    ```
  5. Ivan Maradzhiyski writes about innomd, a command-line tool designed for Linux and macOS that renders LaTeX math formulas as clean Unicode directly in the terminal. Built on top of the `rich` library, it serves scientists, engineers, and students by providing human-readable mathematical notation (such as Greek letters, operators, and fractions) instead of raw LaTeX source code. The tool also features support for rendering Mermaid and PlantUML diagrams into ASCII/Unicode representations and includes a live reload mode for real-time Markdown previewing.

    - Renders common subsets of Mermaid and PlantUML (including flowchart, sequence, class, Gantt, C4 architecture, and activity diagrams).
    - Supports Jupyter notebook (`.ipynb`) files by rendering cells as Markdown with syntax highlighting.
    - Developed in pure Python using `rich` and `grandalf`, requiring no external binaries like Graphviz or Node.js for diagram layout.
    - Features nine built-in color themes, including Nord, Dracula, and Solarized.
    - Uses Unicode approximation rather than pixel-perfect rendering, making it ideal for terminal use but not as a substitute for PDF/LaTeX compilation.
  6. GitButler is a version control tool designed to optimize workflows for developers and AI agents by introducing features like stacked branches, parallel branching, and unlimited undo capabilities. It layers seamlessly onto existing Git repositories without requiring new configuration, aiming to reduce the friction often associated with complex Git operations through structured output and idempotent commands.

    - Agents are reported to be 60% faster using GitButler compared to vanilla Git
    - The software includes a CLI that offers JSON output mode for better AI parsing
    - Features include "Smartlog" and simplified history editing/rebasing
    - It is free and open source software
    2026-09-28 Tags: , , , , , by klotz
  7. Ruoqi Guo et al. present RLCDAlignBench, a new benchmark designed to evaluate the ability of Jev—a model trained via reinforcement learning for calibrated decisions (RLCD)—to detect various types of alignment failures in language models zero-shot. The study examines ten specific failure modes, including sycophancy, jailbreaks, and hallucination, across 44 benchmarks and five target models. Results indicate that a single generic question applied to Jev achieves a median AUROC of 0.886, outperforming supervised baselines in many cases while being significantly more cost-effective than using LLM-judge scorers.
    - Evaluates ten failure modes: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking.
    - Uses a "relational" detection strategy by varying the question wording separately from input fields to handle failures defined against external references.
    - Jev costs 63x less than traditional LLM-judge scorers.
  8. Maya Posch writes about the reverse-engineering of the Intel 8087 FPU, specifically focusing on how it implements trigonometric functions like `FPTAN`. To achieve high accuracy for 64-bit values efficiently, the hardware uses a hybrid approach that combines the CORDIC algorithm with Padé approximants. The system first performs most calculations using CORDIC and then switches to polynomial approximation once the remaining value is small enough, allowing it to avoid large look-up tables or excessive processing time.

    - For one calculated value of `FPTAN`, CORDIC pseudo-division takes 33% of the time, while pseudo-multiplication takes 47%.
    - The polynomial approximation stage accounts for only about 15% of the execution time with a 5% overhead.
    - Later CPUs like the Pentium series moved away from CORDIC because it is difficult to scale for high bit-accuracy without significant performance penalties.
  9. Zhening Li, Joshua Liu, Mateja Vukelic, Nicole Shen, Supriya Lall, Amitayush Thakur, and colleagues at MIT CSAIL introduce JAZ, an agent framework that utilizes a single primitive called `invoke` to perform tasks typically requiring specialized memory or self-improvement systems. By treating the LLM as a runtime provider for function implementations through executable code, the framework allows all inputs and interaction histories to act as variables in the environment.
    - The system uses "hooks" instead of dedicated subsystems like file systems or external memory stores to apply constraints and monitoring.
    - JAZ outperformed Letta (MemGPT) by 8% at half the cost on recall-heavy tasks within the StuLife dataset.
    - In self-improvement evaluations on AppWorld, JAZ exceeded ACE performance by 4% while maintaining a lower cost.
  10. Yohei Nakajima writes about glance, a tool designed to allow users to ask an open vision-language model (VLM) typed questions about images and receive probability data directly on their own machine. Rather than generating new text or training models, it acts as a measurement harness that reads logits from frozen models—such as Qwen3-VL-4B by default—to provide yes/no answers, single-choice selections, and qualitative ratings without any image data leaving the user's device.

    - The tool provides three response types: "noul" (yes/no), choice (pick one from a list), and score (a rating on a specified scale).
    - It includes an experimental MLX backend to provide faster runtimes specifically for Apple Silicon users.
    - Glance can perform self-calibration using unlabeled data or precise calibration through labeled datasets to improve rating accuracy.
    - The software is designed with privacy in mind, ensuring all inference and logging stay local on the user's hardware.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Recent bookmarks

About - Propulsed by SemanticScuttle