SemanticScuttle - klotz.me » Tags: llm+llama.cpp+gguf

Tags: llm* + llama.cpp* + gguf*

0 bookmark(s) - Sort by: Date ↓ / Title /

guide : running gpt-oss with llama.cpp · Discussion #15396

A detailed guide for running the new gpt-oss models locally with the best performance using `llama.cpp`. The guide covers a wide range of hardware configurations and provides CLI argument explanations and benchmarks for Apple Silicon devices.

2025-10-04 Tags: llama.cpp, gpt-oss, large language model, inference, apple silicon, benchmarks, performance, gguf by klotz
Gemma 3: How to Run & Fine-tune

How to run Gemma 3 effectively with our GGUFs on llama.cpp, Ollama, Open WebUI and how to fine-tune with Unsloth! This page details running Gemma 3 on various platforms, including phones, and fine-tuning it using Unsloth, addressing potential issues with float16 precision and providing optimal configuration settings.

2025-08-16 Tags: gemma 3, llm, fine-tuning, llama.cpp, unsloth, gguf, gpu, colab, vision, audio, oobabooga by klotz
TIL: Building llamafiles from Llama 3.2 GGUFs

A step-by-step guide on building llamafiles from Llama 3.2 GGUFs, including scripting and Dockerization.

2024-09-28 Tags: llamafile, llama.cpp, llm, llama 3.2, gguf, model quantization, docker, mozilla-ocho by klotz
GoogleCloudPlatform/localllm: Run LLMs locally on Cloud Workstations

- create a custom base image for a Cloud Workstation environment using a Dockerfile
. Uses:

Quantized models from

2024-02-08 Tags: llm, google, locallama, github, foss, gguf, huggingface, llama.cpp by klotz
Democratizing LLMs: 4-bit Quantization for Optimal LLM Inference

A deep dive into model quantization with GGUF and llama.cpp and model evaluation with LlamaIndex

2024-01-15 Tags: llm, gguf, georgi gerganov, llama.cpp, llamaindex, huggingface, rag by klotz

Top of the page

First / Previous / Next / Last / Page 1 of 0

About - Propulsed by SemanticScuttle

SemanticScuttle - klotz.me

Tags: llm* + llama.cpp* + gguf*

Linked Tags

Related Tags