Skip to content

CUDA Profiling

See the Quick Start guide on how to install Graphsignal.

Graphsignal profiles CUDA workloads via NVIDIA’s CUPTI activity API and observes NVIDIA GPUs via NVML — no code instrumentation required. Attach the profiler to any CUDA program (PyTorch, vLLM, SGLang, TensorRT-LLM, raw CUDA, custom Triton kernels) and the GPU side comes through automatically.

Everything lands in the local /signals endpoint as profile metrics — cumulative counters per named frame, sorted by value descending — plus counters and gauges:

  • cuda_kernels_nanoseconds — cumulative execution time per kernel, frames are the raw kernel symbols.
  • cuda_graphs_nanoseconds — cumulative replay time per unique CUDA-graph structure. Kernels inside a replayed graph are attributed to the graph, not listed individually; the graph identity is stable across replays, so a hot graph is one hot frame.
  • cuda_graph_trace_mode — which graph tracing granularity produced this payload: 0 = graph (default), 1 = node.
  • cuda_memcpy_nanoseconds — cumulative time per transfer kind (host↔device, device↔device, peer), with cuda_memcpy_bytes{kind} counters carrying the volumes.
  • cuda_memset_nanoseconds — cumulative time per memory kind, with cuda_memset_bytes{kind} counters.
  • cuda_sync_nanoseconds — cumulative host-blocking time per synchronization type (context, stream, event, stream-wait-event).
  • gpu_* — NVML telemetry per device: utilization, memory (used/free/total/reserved), temperature, power, clock speeds, throttling, PCIe and NVLink throughput, ECC errors, and XID error events (gpu_xid_critical_errors).

Engines that capture their decode step into a CUDA graph — vLLM, SGLang, TensorRT-LLM, llama.cpp — replay one graph per token, and what you get depends on --cuda-graph-trace:

  • graph (the default) — each replay is one timing in cuda_graphs_nanoseconds. cuda_kernels_nanoseconds then holds only eagerly launched kernels, typically the prefill path, and the decode kernels are not in it at all. This is the cheap mode and the one to leave on.
  • node — graphsignal-run --cuda-graph-trace node times the kernels inside the graph individually, so cuda_kernels_nanoseconds ranks every decode kernel by symbol and cuda_graphs_nanoseconds stays empty. Nothing else changes, and no elevated privileges are needed. CUPTI does more work per replay, so reach for it when you need the per-kernel breakdown, then go back to the default.

To instrument code and kernels — regions, ops, custom timings, whether added by hand or by an AI agent — add GPU probes; their values appear next to these metrics.

graphsignal-run injects a small activity library into the workload process via CUDA_INJECTION64_PATH; it collects CUPTI activity records with low overhead and publishes cumulative state once per second. Everything else — analysis, statistics, NVML polling, the HTTP endpoint — runs in the sidecar profiler process, never inside the workload.

Wrap your launch command with graphsignal-run:

Terminal window
graphsignal-run <my-app>

CUPTI activity collection and NVML metrics start automatically once a CUDA context exists in the workload. While it runs, read everything at http://127.0.0.1:18259/signals. See the Profiler CLI reference for options.