vLLM Profiling
See the Quick Start guide on how to install Graphsignal.
graphsignal-run recognizes vllm serve and configures vLLM’s Prometheus metrics and GPU profiling for you. Everything measured is served at the local /signals endpoint while the server runs.
What’s captured
Section titled “What’s captured”- GPU profiling: cumulative time per CUDA kernel, per CUDA graph, and per memcpy/memset/synchronization kind, plus transfer byte counters. See CUDA Profiling for details on what’s captured GPU-side.
- Engine metrics: vLLM’s
/metricsendpoint is scraped automatically (vllm:*counters, gauges, histograms). Graphsignal keeps metrics enabled — it removes--disable-log-statsif present. The scrape port derives from vLLM’s--port(default8000); override with--metrics-port. - System metrics: CPU, host memory, and GPU telemetry via NVML — collected by the profiler sidecar regardless of engine.
- Errors: warnings, errors, and tracebacks extracted from vLLM’s console output, including crashes.
Run vllm serve with graphsignal-run
Section titled “Run vllm serve with graphsignal-run”graphsignal-run vllm serve Qwen/Qwen2.5-1.5B-Instruct --port 8000While the server runs, read everything at:
curl -s http://127.0.0.1:18259/signalsCUDA graphs
Section titled “CUDA graphs”vLLM decode paths replay CUDA graphs. Kernels launched eagerly appear in the cuda_kernels_nanoseconds profile by symbol; graph replays appear in the cuda_graphs_nanoseconds profile as one frame per unique graph structure, with cumulative replay time. To rank the individual kernels inside those replays instead, relaunch with --cuda-graph-trace node: the decode kernels then appear in cuda_kernels_nanoseconds by symbol and cuda_graphs_nanoseconds stays empty. See CUDA Profiling.
Add Graphsignal to a vLLM Docker image
Section titled “Add Graphsignal to a vLLM Docker image”vLLM images may not include the CUPTI library. Install the matching Graphsignal CUPTI extra (e.g. graphsignal[cu12] for CUDA 12.x).
Example: a modified docker run that installs Graphsignal with CUPTI support and runs vLLM with it:
docker run --gpus all \ -p 8000:8000 \ --ipc=host \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --entrypoint bash \ vllm/vllm-openai:latest \ -lc 'pip install --no-cache-dir graphsignal[cu12] \ && exec graphsignal-run vllm serve \ --model Qwen/Qwen2-VL-7B-Instruct \ --trust-remote-code'To also read /signals from outside the container, publish the profiler’s port (-p 18259:18259) and bind the endpoint beyond loopback with graphsignal-run --listen-host 0.0.0.0 ... — the endpoint is unauthenticated, so expose it only on networks you trust. To upload signals to Graphsignal instead, add -e GRAPHSIGNAL_API_KEY=... — see the production feedback loop.