SemanticScuttle - klotz.me » klotz: vllm+openai

klotz: vllm* + openai*

vLLM: Serve LLMs at Scale

High-performance deployment of the vLLM serving engine, optimized for serving large language models at scale.

2024-08-16 Tags: vllm, llm, scalability, openai, api, production engineering by klotz
LLooM: Leverage raw LLM logits to weave threads

This page provides information about LLooM, a tool that uses raw LLM logits to weave threads in a probabilistic way. It includes instructions on how to use LLooM with various environments, such as vLLM, llama.cpp, and OpenAI. The README also explains the parameters and configurations for LLooM.

2024-07-04 Tags: lloom, llm, logits, vllm, llama.cpp, openai, greedy decoding, beamsearch, github by klotz

First / Previous / Next / Last / Page 1 of 0