Note, this post assumes you have a basic understanding of containerization and kubernetes works. I highly suggest skimming my introduction to k3s first.

Repo for this project lives here: llm-serving-vllm-k8s

In this blog post, I walk through how to setup Prometheus for monitoring the pods we set up in my previous post. Additionally, we walk through setting up monitoring for the local vLLM instance using Grafana.

Some Context

Prometheus is the thing that we used to monitor vLLM. It works on a pull-based model, which means that it periodically queries some endpoint (in this case /metrics using a GET request) on each target, and stores the numbers it scrapes. This means that “deploying” Prometheus involves the following: