TensorRT-LLM Profiling
See the Quick Start guide on how to install Graphsignal.
graphsignal-run recognizes trtllm-serve, trtllm serve, and trtllm-llmapi-launch invocations and configures GPU profiling for you. It scrapes TensorRT-LLM’s Prometheus endpoint at /prometheus/metrics on the serving port (--port, default 8000). Everything measured is served at the local /signals endpoint while the server runs.
What’s captured
Section titled “What’s captured”- GPU profiling: cumulative time per CUDA kernel, per CUDA graph, and per memcpy/memset/synchronization kind, plus transfer byte counters. See CUDA Profiling for details on what’s captured GPU-side.
- Engine metrics: request-level Prometheus metrics (
trtllm_*counters, gauges, histograms) from/prometheus/metricson the HTTP server. Graphsignal does not modify yourtrtllm-servecommand; see Enabling engine metrics below if metrics are missing. - System metrics: CPU, host memory, and GPU telemetry via NVML — collected by the profiler sidecar regardless of engine.
- Errors: warnings, errors, and tracebacks extracted from the server’s console output, including crashes.
Run trtllm-serve with graphsignal-run
Section titled “Run trtllm-serve with graphsignal-run”graphsignal-run trtllm-serve Qwen/Qwen2.5-1.5B-Instruct --port 8000For the PyTorch backend (common on newer GPUs):
graphsignal-run trtllm-serve Qwen/Qwen2.5-1.5B-Instruct \ --port 8000 \ --backend pytorchWhile the server runs, read everything at:
curl -s http://127.0.0.1:18259/signalsCUDA graphs
Section titled “CUDA graphs”The TensorRT-LLM PyTorch backend replays CUDA graphs during decode. Kernels launched eagerly appear in the cuda_kernels_nanoseconds profile by symbol; graph replays appear in the cuda_graphs_nanoseconds profile as one frame per unique graph structure, with cumulative replay time. To rank the individual kernels inside those replays instead, relaunch with --cuda-graph-trace node: the decode kernels then appear in cuda_kernels_nanoseconds by symbol and cuda_graphs_nanoseconds stays empty. See CUDA Profiling.
Enabling engine metrics
Section titled “Enabling engine metrics”TensorRT-LLM exposes two different HTTP metrics endpoints:
| Endpoint | Format | Purpose |
|---|---|---|
/prometheus/metrics | Prometheus text (trtllm_*) | Request-level metrics scraped by Graphsignal |
/metrics | JSON | Per-iteration stats (IterationStats); not used by Graphsignal |
The /prometheus/metrics route is not enabled by default. It is registered only when return_perf_metrics is set to true in your server configuration. If engine metrics are missing, enable them in a YAML config and pass it to trtllm-serve:
return_perf_metrics: truegraphsignal-run trtllm-serve Qwen/Qwen2.5-1.5B-Instruct \ --port 8000 \ --backend pytorch \ --config trtllm-config.yamlAfter the server has finished loading and served at least one request, verify the endpoint locally:
curl http://localhost:8000/prometheus/metrics | headYou should see # HELP trtllm_... lines. For the full list of exported metrics and an end-to-end example, see NVIDIA’s Prometheus Metrics guide and the trtllm-serve metrics documentation.
Notes:
- Match the scrape host to your
--hostflag (Graphsignal defaults tolocalhost, same astrtllm-serve). --grpcmode does not expose an HTTP/prometheus/metricsendpoint; engine Prometheus metrics are not available in that mode. GPU profiling and system metrics work regardless.- On the PyTorch backend, iteration stats on
/metricsare separate and may requireenable_iter_perf_statsin config; that JSON endpoint is not scraped by Graphsignal.
Add Graphsignal to a TensorRT-LLM Docker image
Section titled “Add Graphsignal to a TensorRT-LLM Docker image”TensorRT-LLM NGC release images may not include Graphsignal (or CUPTI). Install the matching Graphsignal CUPTI extra (e.g. graphsignal[cu12] for CUDA 12.x) at container startup and run the server through graphsignal-run.
docker run --gpus all \ -p 8000:8000 \ --ipc=host \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --entrypoint bash \ nvcr.io/nvidia/tensorrt-llm/release:latest \ -lc 'pip install --no-cache-dir graphsignal[cu12] \ && exec graphsignal-run trtllm-serve \ Qwen/Qwen2.5-1.5B-Instruct \ --port 8000 \ --backend pytorch \ --config trtllm-config.yaml'Ensure trtllm-config.yaml sets return_perf_metrics: true (see Enabling engine metrics). To upload signals to Graphsignal, add -e GRAPHSIGNAL_API_KEY=... — see the production feedback loop.