Running a large language model in production means more than loading weights and calling model.generate(). You need efficient memory management, high throughput under concurrent requests, and an API that integrates with existing tooling. vLLM handles all of this. It is an open-source inference and serving library that consistently delivers higher throughput than naive HuggingFace pipelines, and it exposes an OpenAI-compatible API server so you can swap it into any codebase that already talks to OpenAI's endpoints.
This guide takes you from installation to a production-ready deployment, covering offline batch inference, API serving, GPU memory tuning, quantization, and Docker.
What Is vLLM and Why Is It Fast
vLLM is a high-throughput inference and serving engine for large language models. It was developed at UC Berkeley and introduced PagedAttention, a memory management technique inspired by virtual memory in operating systems.
The core problem it solves is KV cache fragmentation. During autoregressive generation, each token produced requires storing key-value pairs from the attention mechanism. Naive implementations allocate contiguous memory blocks for each sequence's KV cache, sized for the maximum possible sequence length. This wastes GPU memory because most sequences don't reach the maximum length, and the pre-allocated blocks can't be shared or reclaimed.
PagedAttention stores KV cache in non-contiguous blocks (pages) that are allocated on demand. This eliminates internal fragmentation and enables memory sharing across sequences (useful for beam search and parallel sampling). In practice, this means vLLM can serve more concurrent requests with the same GPU memory compared to implementations that allocate fixed-size contiguous buffers.
Beyond PagedAttention, vLLM uses continuous batching (new requests are added to the running batch without waiting for all current requests to finish), optimized CUDA kernels, and support for tensor parallelism across multiple GPUs.
Installation
vLLM requires a CUDA-capable GPU. The primary supported platform is Linux with CUDA 12.1 or later and Python 3.9+. There is experimental support for other platforms, but Linux with NVIDIA GPUs is the production target.
Install with pip:
pip install vllm
This pulls in PyTorch and other dependencies automatically. The installation is large (several GB) because it includes CUDA-compiled kernels. If you're working on a remote GPU server, make sure you have sufficient disk space and that your CUDA drivers are up to date.
Verify the installation:
python -c "import vllm; print(vllm.__version__)"
If you're running on a machine without a GPU (for example, a laptop for development), vLLM will not work for inference. You need an NVIDIA GPU with enough VRAM for your target model.
Offline Inference: Batch Processing
The simplest way to use vLLM is for offline batch inference, where you have a list of prompts and want to generate completions for all of them. vLLM processes them together, taking advantage of continuous batching to maximize GPU utilization.
from vllm import LLM, SamplingParams
sampling_params = SamplingParams(temperature=0.8, top_p=0.95, max_tokens=256)
llm = LLM(model="meta-llama/Llama-3.1-8B-Instruct")
prompts = [
"Explain the difference between TCP and UDP in one paragraph.",
"Write a Python function that checks if a string is a palindrome.",
]
outputs = llm.generate(prompts, sampling_params)
for output in outputs:
print(f"Prompt: {output.prompt!r}")
print(f"Output: {output.outputs[0].text}")
print("---")
The LLM class loads the model and allocates the KV cache. The generate method processes all prompts as a batch, scheduling them through vLLM's engine for maximum throughput. This is significantly faster than generating completions one at a time in a loop.
SamplingParams controls how tokens are selected during generation:
- temperature: Controls randomness. 0.0 is greedy (always pick the most likely token), higher values add more randomness. 0.8 is a reasonable default for varied outputs.
- top_p: Nucleus sampling. Only considers tokens whose cumulative probability mass reaches this threshold. 0.95 means the model samples from the top 95% of the probability distribution.
- top_k: Limits sampling to the top K most likely tokens. Set to -1 (default) to disable.
- max_tokens: Maximum number of tokens to generate per prompt.
- stop: A list of strings that, when generated, cause the output to stop. Useful for structured outputs (e.g.,
stop=["\n\n"]to stop after a double newline).
Note that accessing gated models like Llama 3.1 requires accepting the license on Hugging Face and setting your access token via huggingface-cli login or the HF_TOKEN environment variable.
OpenAI-Compatible API Server
For production deployments, you typically want an HTTP API rather than inline Python. vLLM ships with a built-in server that implements OpenAI's Chat Completions and Completions API format. This means any client library or tool that works with OpenAI's API can point at your vLLM server instead.
Start the server:
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--host 0.0.0.0 \
--port 8000
The server loads the model into GPU memory and begins accepting requests. You can query it with curl:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"messages": [{"role": "user", "content": "What is PagedAttention?"}],
"temperature": 0.7,
"max_tokens": 256
}'
Or use the OpenAI Python client with no code changes beyond the base URL:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
response = client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": "Explain vLLM in two sentences."}],
temperature=0.7,
)
print(response.choices[0].message.content)
The api_key parameter is required by the OpenAI client but vLLM does not enforce authentication by default, so any non-empty string works. If you need API key validation, you can pass --api-key your-secret-key when starting the server.
This drop-in compatibility is the main reason vLLM is popular for production deployments. Existing application code, agent frameworks, and orchestration tools that use OpenAI's API format work with vLLM by changing a single URL. Streaming is also supported via the stream: true parameter, which returns server-sent events in the same format as OpenAI's streaming API.
GPU Memory and Performance Tuning
The default vLLM settings work well for many cases, but tuning a few parameters can make the difference between a deployment that handles your workload and one that runs out of memory or underperforms.
Key flags for the vllm serve command:
--gpu-memory-utilization 0.9 (default: 0.9) -- Controls the fraction of GPU memory that vLLM pre-allocates for the KV cache and model weights. If you're running other processes on the same GPU, lower this value (e.g., 0.7) to leave headroom. If vLLM is the only process, 0.9 is a reasonable default.
--max-model-len 4096 -- Limits the maximum context length (input + output tokens). If your use case doesn't need the model's full context window, setting a lower value reduces KV cache memory requirements and lets vLLM serve more concurrent requests. For example, a model that supports 128K context might only need 4096 tokens for your chatbot use case.
--tensor-parallel-size 2 -- Splits the model across multiple GPUs using tensor parallelism. Use this when a model doesn't fit in a single GPU's memory, or to increase throughput by distributing computation. The value must evenly divide the model's attention heads.
--dtype auto -- Uses the model's default data type. You can explicitly set float16 or bfloat16. Use bfloat16 on Ampere or newer GPUs (A100, H100, RTX 30xx+) for better numerical stability. Use float16 on older architectures like V100.
--enforce-eager -- Disables CUDA graph capture. vLLM normally captures CUDA graphs to reduce kernel launch overhead, which improves throughput. Disabling this is useful for debugging or when running on GPUs that don't support CUDA graphs well. Expect lower performance with this flag.
A practical example for a multi-GPU deployment:
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--tensor-parallel-size 2 \
--gpu-memory-utilization 0.85 \
--max-model-len 8192 \
--dtype bfloat16
This splits the model across 2 GPUs, reserves 85% of each GPU's memory, limits context to 8192 tokens, and uses bfloat16 precision.
Quantization
Quantized models use lower-precision weights (typically 4-bit instead of 16-bit), which dramatically reduces VRAM requirements. A 7B parameter model requires roughly 14GB of GPU memory in float16. The same model quantized to 4-bit AWQ uses approximately 4GB, making it feasible to run on consumer GPUs with 8GB of VRAM.
vLLM supports several quantization formats natively, including AWQ, GPTQ, SqueezeLLM, and FP8. You don't need to quantize the model yourself; pre-quantized models are available on Hugging Face.
To serve a quantized model:
vllm serve TheBloke/Llama-2-7B-Chat-AWQ \
--quantization awq \
--dtype float16
The --quantization flag tells vLLM which dequantization kernels to use. The --dtype float16 sets the compute dtype for the dequantized activations.
The quality trade-off is real but often acceptable. AWQ 4-bit quantization typically causes a small degradation on benchmarks (a few percentage points on perplexity), which is hard to notice in practical chat or code generation tasks. For latency-sensitive applications where VRAM is the bottleneck, quantization is one of the most effective optimizations available.
Docker Deployment
For reproducible production deployments, Docker is the standard approach. The official vLLM Docker image includes all CUDA dependencies and is ready to serve models.
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai:latest \
--model meta-llama/Llama-3.1-8B-Instruct \
--gpu-memory-utilization 0.9
Breaking down the flags:
--runtime nvidia --gpus all: Gives the container access to all NVIDIA GPUs on the host. Requires the NVIDIA Container Toolkit to be installed.-v ~/.cache/huggingface:/root/.cache/huggingface: Mounts the host's Hugging Face cache directory into the container. This avoids re-downloading models on every container restart. If you're using a gated model, make sure your HF token is available in this cache or pass-e HF_TOKEN=your_token.-p 8000:8000: Maps the container's port 8000 to the host's port 8000.--ipc=host: Shares the host's IPC namespace with the container. This is required for PyTorch's shared memory usage, especially when using tensor parallelism across multiple GPUs. Without it, you may see shared memory errors.
Everything after the image name (vllm/vllm-openai:latest) is passed as arguments to the vLLM server, so you can use all the same flags from the previous sections.
For production, you would typically place this behind a reverse proxy (nginx, Caddy, or a cloud load balancer), add health checks, and manage the container with Docker Compose or Kubernetes.
Wrapping Up
vLLM gives you a fast, well-tested path from "I have a model on Hugging Face" to "I have an OpenAI-compatible API serving concurrent requests." The combination of PagedAttention for memory efficiency, continuous batching for throughput, and the built-in OpenAI-compatible server makes it a practical choice for self-hosted LLM inference.
For developers running AI coding agents or building AI-powered applications, self-hosting with vLLM provides full control over latency, cost, and data privacy. You decide which model to run, where the data goes, and how much hardware to allocate. Tools like Agents UI make it straightforward to manage SSH sessions to remote GPU servers where vLLM is running, so you can monitor, restart, and interact with your deployment without juggling terminal windows.
The commands and code in this guide are enough to get a working deployment. From there, the vLLM documentation covers advanced topics like LoRA adapter serving, speculative decoding, and structured output generation for when your use case demands them.