SemanticScuttle - klotz.me » klotz: llm+nvidia

klotz: llm* + nvidia*

An Introduction to Model Merging for LLMs

This article introduces model merging, a technique that combines the weights of multiple customized large language models to increase resource utilization and add value to successful models.

2024-10-30 Tags: model merging, llm, nvidia by klotz

NVIDIA launches ‘easy button’ for creating gen AI workflows

NVIDIA introduces NIM Agent Blueprints, a collection of pre-trained, customizable AI workflows for common use cases like customer service avatars, PDF extraction, and drug discovery, aiming to simplify generative AI development for businesses.

2024-08-30 Tags: nvidia, workflows, blueprints, llm, agents by klotz

Run:ai - Accelerate AI Development & Innovation

Run:ai offers a platform to accelerate AI development, optimize GPU utilization, and manage AI workloads. It is designed for GPUs, offers CLI & GUI interfaces, and supports various AI tools & frameworks.

2024-08-26 Tags: llm, orchestration, infrastructure, gpu, workload management, k8s, nvidia, production engineering by klotz

Benchmarks show even an old Nvidia RTX 3090 is enough to serve LLMs to thousands

A startup called Backprop has demonstrated that a single Nvidia RTX 3090 GPU, released in 2020, can handle serving a modest large language model (LLM) like Llama 3.1 8B to over 100 concurrent users with acceptable throughput. This suggests that expensive enterprise GPUs may not be necessary for scaling LLMs to a few thousand users.

2024-08-24 Tags: nvidia, rtx 3090, llm, gpu, performance, benchmark, llama 3.1 8b, vllm, production engineering, backprop.co by klotz

RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs

A method that uses instruction tuning to adapt LLMs for knowledge-intensive tasks. RankRAG simultaneously trains the models for context ranking and answer generation, enhancing their retrieval-augmented generation (RAG) capabilities.

2024-07-10 Tags: natural language processing, large language models, instruction tuning, context ranking, retrieval-augmented generation, nvidia, arxiv by klotz

NVIDIA Introduces RankRAG: Enhancing LLMs with Instruction Tuning

NVIDIA and Georgia Tech researchers introduce RankRAG, a novel framework instruction-tuning a single LLM for top-k context ranking and answer generation. Aiming to improve RAG systems, it enhances context relevance assessment and answer generation.

2024-07-10 Tags: rankrag, nvidia, llm, rag, instruction tuning, natural language processing by klotz

Fine Tuning LLM on RTX 3090

- Discusses the use of consumer graphics cards for fine-tuning large language models (LLMs)
- Compares consumer graphics cards, such as NVIDIA GeForce RTX Series GPUs, to data center and cloud computing GPUs
- Highlights the differences in GPU memory and price between consumer and data center GPUs
- Shares the author's experience using a GeForce 3090 RTX card with 24GB of GPU memory for fine-tuning LLMs

2024-02-02 Tags: llm, fine tuning, nvidia, rtx, 3090, self-hosted by klotz

NVIDIA AI Introduces ChatQA: A Family of Conversational Question Answering (QA) Models that Obtain GPT-4 Level Accuracies

ChatQA, a new family of conversational question-answering (QA) models developed by NVIDIA AI. These models employ a unique two-stage instruction tuning method that significantly improves zero-shot conversational QA results from large language models (LLMs). The ChatQA-70B variant has demonstrated superior performance compared to GPT-4 across multiple conversational QA datasets.

2024-01-24 Tags: llm, instruction tuning, nvidia, chatqa, sft by klotz

NVIDIAs new paper introduces ChatQA model that is GPT-4 Level --- ChatQA-70B : r/LocalLLaMA

2024-01-20 Tags: nvidia, chatqa, llm, localllama, reddit, open world by klotz

Nvidia RAG Demo

Windows only

2024-01-11 Tags: github, nvidia, llm, rag by klotz

First / Previous / Next / Last / Page 1 of 0

SemanticScuttle - klotz.me

klotz: llm* + nvidia*

Linked Tags

Related Tags