Posts for: #Docker

Basic LLM Monitoring with Grafana and Prometheus

Note, this post assumes you have a basic understanding of containerization and kubernetes works. I highly suggest skimming my introduction to k3s first.

Repo for this project lives here: llm-serving-vllm-k8s

In this blog post, I walk through how to setup Prometheus for monitoring the pods we set up in my previous post. Additionally, we walk through setting up monitoring for the local vLLM instance using Grafana.

Some Context

Prometheus is the thing that we used to monitor vLLM. It works on a pull-based model, which means that it periodically queries some endpoint (in this case /metrics using a GET request) on each target, and stores the numbers it scrapes. This means that “deploying” Prometheus involves the following:

[Read more]

Spinning up a simple k3s to manage a local LLM Docker Container

Link to Github

In this blog post, I walk through how to take a lone Docker image and use k3s to manage it on a single-node machine (i.e. my desktop). The idea behind Docker is an isolated, self-sufficient container that can run anywhere, k3s and its bigger brother k8s (Kubernetes) is used to manage, in a declarative fashinon, these containers.

The core model of Kubernetes, and the main difference from simply calling docker run is the declarative reconciliation. Kubernetes is a control system built around a single loop:

[Read more]

Building a multi-stage Docker image for locally serving an LLM

Serving a small, open-source LLM behind vLLM, containerized on a single desktop GPU.

The repo is here: https://github.com/codecalligrapher/llm-serving-vllm-k8s

The target is to have a single-node Kubernetes, with Grafana to monitor and a throughput benchmark.

Status: containerized OpenAI-compatible endpoint running on an RTX 3060 Ti (8GB). K8s, monitoring, and benchmark are scoped below but not yet built.

Why vLLM

Mainly because of the OpenAI-compatible server, so the endpoint is v1/chat/completions and any OpenAI client works against it.

[Read more]