SemanticScuttle - klotz.me » klotz: localllama

klotz: localllama*

Optimal settings for running gpt-oss-120b on 2x 3090s and 128gb system ram

A user shares their optimal settings for running the gpt-oss-120b model on a system with dual RTX 3090 GPUs and 128GB of RAM, aiming for a balance between performance and quality.

2025-12-04 Tags: gpt-oss-120b, localllama, llm, gpu, rtx 3090, lm studio, optimization, settings by klotz

This is GPT-OSS 120b on Ollama, running on a i7

A user shares their experience running the GPT-OSS 120b model on Ollama with an i7 6700, 64GB DDR4 RAM, RTX 3090, and a 1TB SSD. They note slow initial token generation but acceptable performance overall, highlighting it's possible on a relatively modest setup. The discussion includes comparisons to other hardware configurations, optimization techniques (llama.cpp), and the model's quality.

>I have a 3090 with 64gb ddr4 3200 RAM and am getting around 50 t/s prompt processing speed and 15 t/s generation speed using the following:
>
>`llama-server -m <path to gpt-oss-120b> --ctx-size 32768 --temp 1.0 --top-p 1.0 --jinja -ub 2048 -b 2048 -ngl 99 -fa 'on' --n-cpu-moe 24`
> This about fills up my VRAM and RAM almost entirely. For more wiggle room for other applications use `--n-cpu-moe 26`.

2025-09-01 Tags: gpt-oss, 120b, reddit, localllama, llm, inference, rtx 3090, llama.cpp, hardware by klotz

NotebookLM is perfect for documenting my home lab, and here's how I use it

The article discusses how NotebookLM can be used to document and troubleshoot a home lab setup. It highlights its ability to consolidate documentation, simplify complex tasks, and provide step-by-step instructions. The author shares practical examples of using NotebookLM for learning, troubleshooting, and managing a home lab environment.

2025-08-24 Tags: notebooklm, llm, self-hosting, localllama by klotz

120B runs awesome on just 8GB VRAM!

A user demonstrates how to run a 120B model efficiently on hardware with only 8GB VRAM by offloading MOE layers to CPU and keeping only attention layers on GPU, achieving high performance with minimal VRAM usage.

2025-08-21 Tags: 120b, moe, llama.cpp, gpt-oss, localllama, gpt-oss-120b, openai, llm by klotz

llama-swap

llama-swap is a lightweight, transparent proxy server that provides automatic model swapping to llama.cpp's server. It allows you to easily switch between different language models on a local server, supporting OpenAI API compatible endpoints and offering features like model grouping, automatic unloading, and a web UI for monitoring.

2025-08-08 Tags: golang, openai, llama, openai-api, llamacpp, vllm, localllm, localllama, model swapping, local llm server by klotz

I get a perfect weather report on my Home Assistant dashboard, here's how I do it with a local LLM

This article details how to set up a weather report on a Home Assistant dashboard using a local LLM (Ollama) for more user-friendly summaries and clothing suggestions, avoiding cloud-based services for privacy reasons. It covers the setup process, prompt engineering, and hardware considerations.

2025-08-06 Tags: home assistant, ollama, llm, localllama, smart home, weather, iot, raspberry pi by klotz

Paperless-ngx is already great, but here's how to make it even better with a local LLM

This article details how to enhance the Paperless-ngx document management system by integrating a local Large Language Model (LLM) like Ollama. It covers the setup process, including installing Docker, Ollama, and configuring Paperless AI, to enable AI-powered features such as improved search and document understanding.

2025-07-23 Tags: paperless-ngx, llm, localllama, document management, self-hosted, pagemill by klotz

llm-observe-hub

Real-time observability and analytics platform for local LLMs, with dashboard and API.

2025-07-22 Tags: llm, observability, analytics, dashboard, localllama, production engineering, github by klotz

Building LLM Workflows - - some observations

A post with pithy observations and clear conclusions from building complex LLM workflows, covering topics like prompt chaining, data structuring, model limitations, and fine-tuning strategies.

2025-05-09 Tags: llm, localllama, prompt engineering, fine-tuning, agentic loops, context window, bert, xml, cot, workflow, reddit by klotz

Server approved! 4xH100 (320gb vram). Looking for advice

A user is seeking advice on deploying a new server with 4x H100 GPUs (320GB VRAM) for on-premise AI workloads. They are considering a Kubernetes-based deployment with RKE2, Nvidia GPU Operator, and tools like vLLM, llama.cpp, and Litellm. They are also exploring the option of GPU pass-through with a hypervisor. The post details their current infrastructure and asks for potential gotchas or best practices.

2025-04-28 Tags: h100, kubernetes, vllm, llama.cpp, gpu, ai, deployment, rke2, litellm, quantization, sxm, fp8, awq, gguf, production engineering, inference engineering, scale, reddit, localllama by klotz

First / Previous / Next / Last / Page 1 of 0

SemanticScuttle - klotz.me

klotz: localllama*

Linked Tags

Related Tags