Skip to content

TensorRT-LLM Profiling

See the Quick Start guide on how to install Graphsignal.

graphsignal-run recognizes trtllm-serve, trtllm serve, and trtllm-llmapi-launch invocations and configures GPU profiling for you. It scrapes TensorRT-LLM’s Prometheus endpoint at /prometheus/metrics on the serving port (--port, default 8000). Everything measured is served at the local /signals endpoint while the server runs.

  • GPU profiling: cumulative time per CUDA kernel, per CUDA graph, and per memcpy/memset/synchronization kind, plus transfer byte counters. See CUDA Profiling for details on what’s captured GPU-side.
  • Engine metrics: request-level Prometheus metrics (trtllm_* counters, gauges, histograms) from /prometheus/metrics on the HTTP server. Graphsignal does not modify your trtllm-serve command; see Enabling engine metrics below if metrics are missing.
  • System metrics: CPU, host memory, and GPU telemetry via NVML — collected by the profiler sidecar regardless of engine.
  • Errors: warnings, errors, and tracebacks extracted from the server’s console output, including crashes.
Terminal window
graphsignal-run trtllm-serve Qwen/Qwen2.5-1.5B-Instruct --port 8000

For the PyTorch backend (common on newer GPUs):

Terminal window
graphsignal-run trtllm-serve Qwen/Qwen2.5-1.5B-Instruct \
--port 8000 \
--backend pytorch

While the server runs, read everything at:

Terminal window
curl -s http://127.0.0.1:18259/signals

The TensorRT-LLM PyTorch backend replays CUDA graphs during decode. Kernels launched eagerly appear in the cuda_kernels_nanoseconds profile by symbol; graph replays appear in the cuda_graphs_nanoseconds profile as one frame per unique graph structure, with cumulative replay time. To rank the individual kernels inside those replays instead, relaunch with --cuda-graph-trace node: the decode kernels then appear in cuda_kernels_nanoseconds by symbol and cuda_graphs_nanoseconds stays empty. See CUDA Profiling.

TensorRT-LLM exposes two different HTTP metrics endpoints:

EndpointFormatPurpose
/prometheus/metricsPrometheus text (trtllm_*)Request-level metrics scraped by Graphsignal
/metricsJSONPer-iteration stats (IterationStats); not used by Graphsignal

The /prometheus/metrics route is not enabled by default. It is registered only when return_perf_metrics is set to true in your server configuration. If engine metrics are missing, enable them in a YAML config and pass it to trtllm-serve:

trtllm-config.yaml
return_perf_metrics: true
Terminal window
graphsignal-run trtllm-serve Qwen/Qwen2.5-1.5B-Instruct \
--port 8000 \
--backend pytorch \
--config trtllm-config.yaml

After the server has finished loading and served at least one request, verify the endpoint locally:

Terminal window
curl http://localhost:8000/prometheus/metrics | head

You should see # HELP trtllm_... lines. For the full list of exported metrics and an end-to-end example, see NVIDIA’s Prometheus Metrics guide and the trtllm-serve metrics documentation.

Notes:

  • Match the scrape host to your --host flag (Graphsignal defaults to localhost, same as trtllm-serve).
  • --grpc mode does not expose an HTTP /prometheus/metrics endpoint; engine Prometheus metrics are not available in that mode. GPU profiling and system metrics work regardless.
  • On the PyTorch backend, iteration stats on /metrics are separate and may require enable_iter_perf_stats in config; that JSON endpoint is not scraped by Graphsignal.

Add Graphsignal to a TensorRT-LLM Docker image

Section titled “Add Graphsignal to a TensorRT-LLM Docker image”

TensorRT-LLM NGC release images may not include Graphsignal (or CUPTI). Install the matching Graphsignal CUPTI extra (e.g. graphsignal[cu12] for CUDA 12.x) at container startup and run the server through graphsignal-run.

Terminal window
docker run --gpus all \
-p 8000:8000 \
--ipc=host \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--entrypoint bash \
nvcr.io/nvidia/tensorrt-llm/release:latest \
-lc 'pip install --no-cache-dir graphsignal[cu12] \
&& exec graphsignal-run trtllm-serve \
Qwen/Qwen2.5-1.5B-Instruct \
--port 8000 \
--backend pytorch \
--config trtllm-config.yaml'

Ensure trtllm-config.yaml sets return_perf_metrics: true (see Enabling engine metrics). To upload signals to Graphsignal, add -e GRAPHSIGNAL_API_KEY=... — see the production feedback loop.