Tags: classification*

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. llama.cpp now supports decision models via a `/v1/systemone` endpoint, which accepts a state and typed questions to return probabilities for each provided option in a single forward pass. The API utilizes the System One format introduced with TypeSafe's Jev model, meaning existing clients only need a new base URL. Five initial models are available ranging from 144M to 27B parameters, supporting use cases like request routing, content moderation, and agent action selection with median response times as low as 3 ms.
    - The `state` field accepts text, JSON, screenshots, or a list of chat messages with `image_url` parts.
    - Three question types are supported: `choice` for categorical selection, `score` for a continuous level between 2 and 10 options, and `noul` for yes/no probabilities.
    - Router mode allows loading multiple models on a single server and selecting one per request.
    - Adding descriptions to option labels can significantly improve accuracy, with Julia-1's routing confidence for "charged twice" jumping from a misroute to 0.99 when descriptions were provided.
    - Cloudflare's Clef is the next model planned for integration.
  2. Sebastian Raschka writes a comprehensive overview of the evolution of text classification, tracing its journey from traditional methods like bag-of-words and logistic regression through deep learning architectures such as RNNs, CNNs, and Transformers. The article specifically examines the recent popularity of Jev, a specialized model that functions as an efficient "plug-and-play" classifier capable of performing various decision tasks without custom fine-tuning. Raschka compares modern transformer approaches—including encoder-style models like BERT, decoder-style LLMs like GPT, and encoder-decoder architectures like T5—to illustrate how Jev's speed and low cost provide a middle ground between specialized task-specific classifiers and large general-purpose generative models.

    - Jev is rumored to be trained using "Reinforcement Learning for Calibrated Decisions" (RLCD).
    - Unlike traditional LLMs, the Jev API includes specific modes like Choice (multi-class), Noul (binary/multi-label probability), and Score (ordinal classification).
    - The article highlights that while custom fine-tuning with models like ModernBERT can achieve high accuracy on specific tasks, it lacks the general versatility of a model like Jev.
    - Calibration is crucial in production to ensure predicted probabilities reflect actual class frequencies; techniques include temperature scaling or adding Brier loss during training.
  3. Alibaba Cloud has released decision-model-preview, a structured model designed for high-frequency business judgments. It can concurrently perform classification, binary decisions, and scoring based on text or business state, providing probability distributions and confidence levels to assist with tasks like ticket routing, content moderation, agent routing, and result verification.

    >Example:
    ```bash
    # Structured decision: POST /v1/systemone — every question is evaluated in parallel and in
    # isolation against the same state, and comes back typed. No text generation, nothing to parse.
    # The answers{} map is keyed by your own question names; each answer holds its value under a
    # key named after its type:
    # noul -> { type, noul } probability of "yes", 0..1
    # choice -> { type, choice, probabilities, confidence } choice is one of your criteria keys
    # score -> { type, score, legend, probabilities, confidence } legend maps level index -> description
    # usage carries input_tokens / output_tokens.
    curl https://aihubmix.com/v1/systemone
    -H "Content-Type: application/json"
    -H "Authorization: Bearer $AIHUBMIX_API_KEY"
    -d '{
    "model": "decision-model-preview",
    "state": "Hi, I have been trying to connect my Stripe account for 3 days and it keeps failing. I am losing sales. Please help ASAP.",
    "questions": {
    "department": {
    "type": "choice",
    "instructions": "Which team should handle this",
    "criteria": {
    "billing": "Payment or subscription issues",
    "technical": "Bugs or integration problems",
    "sales": "Pricing or account questions"
    }
    },
    "frustration": {
    "type": "score",
    "instructions": "How frustrated the customer appears",
    "criteria": [
    "Calm, just stating facts",
    "Frustrated but civil",
    "Very angry, strong language"
    ]
    },
    "is_urgent": {
    "type": "noul",
    "instructions": "The message conveys urgency or time-sensitivity"
    }
    }
    }'
    ```
  4. Convai Innovations presents Laya, a multilingual, non-autoregressive system 1 decision model designed to provide typed answers with mathematically calibrated probabilities in a single forward pass. Unlike generative models, it does not generate text, thereby eliminating hallucinations and the need for parsing. The framework includes an automated Router that detects language and script to dispatch tasks to the most efficient checkpoint (English or Multilingual) within approximately 35ms on GPU.
    - It is trained using Reinforcement Learning with Calibrated Decisions (RLCD) to ensure honest probability reporting.
    - Laya can support context lengths of up to 8,192 tokens in its multilingual version.
    - The model family includes specialized checkpoints like `laya-typed-decisions` which achieves significantly higher accuracy through fine-tuning on specific workflows.
    - Performance benchmarks show it is roughly 6–8× faster than TypeSafe Jev for single question latency on a T4 GPU.
  5. jeff is a self-hosted drop-in replacement for TypeSafe's jev System One API, powered by the 400M-parameter GLiFormer-large-v1 model. It serves `choice`, `score`, and `noul` classification questions over a compatible wire format, so existing applications using the official `typesafe-sdk` can point at it by changing a single base URL. Deployment targets include GPU (L4, A10G) and CPU (ONNX Runtime with int8 quantization) via Modal, with ~50 req/s per container throughput on L4. Benchmarks on 1,600 labeled items show it costs roughly a quarter of jev per million requests but trails significantly on reasoning-heavy tasks like irony and reading comprehension.
    - `JEFF_ISOLATE=nouls` (default) gives separate encoder passes per noul question to reduce cross-question interference; choice and score questions share a pass unless set to `all`
    - Default temperature of 3.2 calibrates noul probabilities; setting it to 1 makes `score` output match the weighted average of displayed probabilities
    - A smaller `gliformer-base-v1` variant with `JEFF_NOUL_MODE=single` is recommended for faster local iteration
  6. Iván Palomares Carrascosa writes about methods for interpreting the dense numerical vector representations, or embeddings, generated by large language models (LLMs). By using a combination of probing classifiers like logistic regression, UMAP dimensionality reduction for visualization, and SHAP values to identify influential latent dimensions, one can analyze the quality and semantic structure captured within LLM-generated embedding spaces.

    - Probing classifiers help determine if embeddings are rich enough to distinguish between classes by testing them with simpler models.
    - UMAP is used to project high-dimensional embeddings into 2D space for visual inspection of natural groupings.
    - SHAP values can pinpoint which specific dimensions in an embedding most significantly influence a classifier's decisions.
    - The article demonstrates using Scikit-LLM alongside local Ollama models to generate embeddings cost-effectively.
  7. occlupanid data writes that the Holotypic Occlupanid Research Group hosts several years of research classifying occlupanids, small ubiquitous objects dotting supermarket aisles and sidewalks, as the most common yet puzzling member of phylum Plasticae within a synthetic taxonomy database.

    - The site catalogs dozens of families such as Acutignathidae, Archignathidae, Corrugatidae and Toxodentidae with individual species pages.
    - Navigation includes Identification Guide, Publications and Reports, Cartonalia: The Occlupanopsida, and a Guide to symbols for ecological, geographical and taxonomic classification.
    - The project also covers morphology, growth and development, origins of the Occlupanida, history of occlupanology, and a Pseudo-occlupanids section.
  8. Firecrawl introduces pdf-inspector, a high-performance Rust library designed for rapid PDF classification, text extraction, and Markdown conversion. By sampling content streams to quickly distinguish between text-based and scanned documents, the tool enables intelligent routing that bypasses costly OCR services for standard PDFs. It delivers position-aware text extraction, automated table and column detection, and robust encoding handling while maintaining a lightweight footprint with no external ML dependencies or model training requirements.

    - Provides bindings for Python, Node.js, and browser WebAssembly environments.
    - Achieves sub-200ms processing times on large corpora while outperforming several established local parsers in reading order and table accuracy.
    - Features per-page OCR routing suggestions to optimize mixed-format document workflows.
    - Handles complex layouts including RTL text, multi-column newspapers, and CID-encoded fonts.
    - Released under the MIT license with active community contributions and CI/CD automation.
  9. This tutorial demonstrates how to implement an intelligent routing layer using NadirClaw to optimize Large Language Model (LLM) costs. The system classifies prompts into simple or complex tiers locally before selecting the most appropriate model, such as switching between Gemini Flash and Pro versions. It covers installation, local classification testing via CLI, visualizing decision boundaries through centroid-based similarity scores, running a proxy server for live routing, and calculating estimated cost savings compared to using high-end models exclusively.
    2026-05-11 Tags: , , , by klotz
  10. This article introduces Scikit-LLM, a Python library that integrates large language models like OpenAI's GPT with the Scikit-learn framework to simplify text analysis tasks. It explains and demonstrates two primary classification methods: zero-shot classification, which assigns labels based solely on the model's general knowledge without prior examples, and few-shot classification, which uses a small set of labeled examples within the prompt to improve accuracy. By following a Scikit-learn-style workflow using fit() and predict() methods, users can easily implement these advanced NLP techniques for tasks such as sentiment analysis and topic labeling.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: tagged with "classification"

About - Propulsed by SemanticScuttle