klotz: llm pipelines*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. Iván Palomares Carrascosa writes about using the Scikit-LLM library alongside MLflow to build, track, compare, and register scikit-learn pipelines that incorporate large language models. The article provides a workflow for ensuring model versioning and reproducibility by logging different LLM backends as environment parameters and promoting successful pipeline versions into an MLflow Model Registry.

    - Uses the `scikit-llm gpt4all » ` installation option to ensure compatibility with local execution.
    - Demonstrates how to use `cloudpickle` for serialization when working with scikit-learn models in MLflow.
    - Highlights a two-step workflow of logging experiments first and then registering only "winner" models to keep the registry clean.
  2. This article proposes the DataBook, a design pattern that utilizes Markdown to bridge the gap between large-scale RDF knowledge graphs and small, ephemeral, task-specific semantic content. By combining YAML frontmatter for metadata, inline identifiers for addressability, and typed fenced code blocks for data payloads, DataBooks create self-describing and portable semantic artifacts. The authors argue that this approach allows for a microdatabase model where structured data can exist without the overhead of a full triple store.
    Key points include:
    The use of Markdown as a substrate for semantic infrastructure.
    Defining the microdatabase for small-scale, non-indexed knowledge work.
    Inverting the LLM role to act as a transformation engine within a DataBook pipeline.
    Implementing provenance through process stamps in YAML metadata.
    Managing complex dependencies via manifest DataBooks and build graphs.
    Supporting secure data transfer through designed-in encryption profiles.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: llm pipelines

About - Propulsed by SemanticScuttle