SemanticScuttle - klotz.me » Tags: deployment

Tags: deployment*

0 bookmark(s) - Sort by: Date ↓ / Title /

From Flask to vLLM: How Model Inference has evolved (2017-2025)

The article discusses the evolution of model inference techniques from 2017 to a projected 2025, highlighting the progression from simple frameworks like Flask and FastAPI to more advanced solutions like Triton Inference Server and vLLM. It details the increasing demands on inference infrastructure driven by larger and more complex models, and the need for optimization in areas like throughput, latency, and cost.

2025-08-06 Tags: model inference, machine learning, deep learning, llm, vllm, triton, flask, fastapi, deployment by klotz

About billing for GitHub Spark

This article details the billing structure for GitHub Spark, covering costs associated with app creation (based on premium requests) and current limits for deployed apps. It also outlines future billing plans for deployed apps once limits are reached.

2025-07-28 Tags: github, github spark, deployment, cost, github copilot, llm, no code, mini-apps by klotz

Awesome Code Sandboxing for AI

A curated guide to code sandboxing solutions, covering technologies like MicroVMs, application kernels, language runtimes, and containerization. It provides a feature matrix, in-depth platform profiles (e2b, Daytona, microsandbox, WebContainers, Replit, Cloudflare Workers, Fly.io, Kata Containers), and a decision framework for choosing the right sandboxing solution based on security, performance, workload type, and hosting preferences.

2025-07-15 Tags: code sandboxing, llm, security, microvm, gvisor, webassembly, containers, e2b, daytona, microsandbox, cloudflare workers, kata containers, serverless, isolation, virtualization, production engineering, deployment by klotz

El Reg's essential guide to deploying LLMs in production

Running GenAI models is easy. Scaling them to thousands of users, not so much. This guide details avenues for scaling AI workloads from proofs of concept to production-ready deployments, covering API integration, on-prem deployment considerations, hardware requirements, and tools like vLLM and Nvidia NIMs.

2025-04-28 Tags: llm, ai, production engineering, inference engineering, deployment, vllm, nvidia, kubernetes, inference, api, scaling, gpu, machine learning by klotz

production-stack

K8S-native cluster-wide deployment for vLLM. Provides a reference implementation for building an inference stack on top of vLLM, enabling scaling, monitoring, request routing, and KV cache offloading with easy cloud deployment.

2025-04-28 Tags: vllm, kubernetes, inference, deployment, scaling, monitoring, request routing, kv cache, cloud, inference engineering, production engineering, llm by klotz

vLLM Production Stack: reference stack for production vLLM deployment

vLLM Production Stack provides a reference implementation on how to build an inference stack on top of vLLM, allowing for scalable, monitored, and performant LLM deployments using Kubernetes and Helm.

2025-04-28 Tags: vllm, kubernetes, helm, llm, inference, deployment, observability, kv cache, scalability, production engineering, inference engineering by klotz

Server approved! 4xH100 (320gb vram). Looking for advice

A user is seeking advice on deploying a new server with 4x H100 GPUs (320GB VRAM) for on-premise AI workloads. They are considering a Kubernetes-based deployment with RKE2, Nvidia GPU Operator, and tools like vLLM, llama.cpp, and Litellm. They are also exploring the option of GPU pass-through with a hypervisor. The post details their current infrastructure and asks for potential gotchas or best practices.

2025-04-28 Tags: h100, kubernetes, vllm, llama.cpp, gpu, ai, deployment, rke2, litellm, quantization, sxm, fp8, awq, gguf, production engineering, inference engineering, scale, reddit, localllama by klotz

Flagger Documentation - Introduction

Flagger is a progressive delivery Kubernetes operator that automates the release process for applications running on Kubernetes. It reduces the risk of introducing a new software version in production by gradually shifting traffic to the new version while measuring metrics and running conformance tests.

2024-06-17 Tags: flagger, kubernetes operator, deployment, production engineering by klotz

How to Deploy ML Solutions with FastAPI, Docker, and GCP

This is a hands-on guide with Python example code that walks through the deployment of an ML-based search API using a simple 3-step approach. The article provides a deployment strategy applicable to most machine learning solutions, and the example code is available on GitHub.

2024-06-09 Tags: machine learning, fastapi, docker, gcp, deployment, python, llm, tutorial, production engineering by klotz

Running Machine Learning Workloads on Kubernetes with Google AI Platform and TensorFlow Serving

In this article, we explore how to deploy and manage machine learning models using Google Kubernetes Engine (GKE), Google AI Platform, and TensorFlow Serving. We will cover the steps to create a machine learning model and deploy it on a Kubernetes cluster for inference.

2024-05-15 Tags: gcp, kubernetes, machine learning, tensorflow sing, deployment, inference, mlops, production engineering by klotz

First / Previous / Next / Last / Page 1 of 0

SemanticScuttle - klotz.me

Tags: deployment*

Linked Tags

Related Tags