Skip to content

vLLM Profiling

See the Quick Start guide on how to install Graphsignal.

graphsignal-run recognizes vllm serve and configures vLLM’s Prometheus metrics and GPU profiling for you. Everything measured is served at the local /signals endpoint while the server runs.

  • GPU profiling: cumulative time per CUDA kernel, per CUDA graph, and per memcpy/memset/synchronization kind, plus transfer byte counters. See CUDA Profiling for details on what’s captured GPU-side.
  • Engine metrics: vLLM’s /metrics endpoint is scraped automatically (vllm:* counters, gauges, histograms). Graphsignal keeps metrics enabled — it removes --disable-log-stats if present. The scrape port derives from vLLM’s --port (default 8000); override with --metrics-port.
  • System metrics: CPU, host memory, and GPU telemetry via NVML — collected by the profiler sidecar regardless of engine.
  • Errors: warnings, errors, and tracebacks extracted from vLLM’s console output, including crashes.
Terminal window
graphsignal-run vllm serve Qwen/Qwen2.5-1.5B-Instruct --port 8000

While the server runs, read everything at:

Terminal window
curl -s http://127.0.0.1:18259/signals

vLLM decode paths replay CUDA graphs. Kernels launched eagerly appear in the cuda_kernels_nanoseconds profile by symbol; graph replays appear in the cuda_graphs_nanoseconds profile as one frame per unique graph structure, with cumulative replay time. To rank the individual kernels inside those replays instead, relaunch with --cuda-graph-trace node: the decode kernels then appear in cuda_kernels_nanoseconds by symbol and cuda_graphs_nanoseconds stays empty. See CUDA Profiling.

vLLM images may not include the CUPTI library. Install the matching Graphsignal CUPTI extra (e.g. graphsignal[cu12] for CUDA 12.x).

Example: a modified docker run that installs Graphsignal with CUPTI support and runs vLLM with it:

Terminal window
docker run --gpus all \
-p 8000:8000 \
--ipc=host \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--entrypoint bash \
vllm/vllm-openai:latest \
-lc 'pip install --no-cache-dir graphsignal[cu12] \
&& exec graphsignal-run vllm serve \
--model Qwen/Qwen2-VL-7B-Instruct \
--trust-remote-code'

To also read /signals from outside the container, publish the profiler’s port (-p 18259:18259) and bind the endpoint beyond loopback with graphsignal-run --listen-host 0.0.0.0 ... — the endpoint is unauthenticated, so expose it only on networks you trust. To upload signals to Graphsignal instead, add -e GRAPHSIGNAL_API_KEY=... — see the production feedback loop.