CUDA Profiling
See the Quick Start guide on how to install Graphsignal.
Graphsignal profiles CUDA workloads via NVIDIA’s CUPTI activity API and observes NVIDIA GPUs via NVML — no code instrumentation required. Attach the profiler to any CUDA program (PyTorch, vLLM, SGLang, TensorRT-LLM, raw CUDA, custom Triton kernels) and the GPU side comes through automatically.
What’s captured
Section titled “What’s captured”Everything lands in the local /signals endpoint as profile metrics — cumulative counters per named frame, sorted by value descending — plus counters and gauges:
cuda_kernels_nanoseconds— cumulative execution time per kernel, frames are the raw kernel symbols.cuda_graphs_nanoseconds— cumulative replay time per unique CUDA-graph structure. Kernels inside a replayed graph are attributed to the graph, not listed individually; the graph identity is stable across replays, so a hot graph is one hot frame.cuda_graph_trace_mode— which graph tracing granularity produced this payload:0= graph (default),1= node.cuda_memcpy_nanoseconds— cumulative time per transfer kind (host↔device, device↔device, peer), withcuda_memcpy_bytes{kind}counters carrying the volumes.cuda_memset_nanoseconds— cumulative time per memory kind, withcuda_memset_bytes{kind}counters.cuda_sync_nanoseconds— cumulative host-blocking time per synchronization type (context, stream, event, stream-wait-event).gpu_*— NVML telemetry per device: utilization, memory (used/free/total/reserved), temperature, power, clock speeds, throttling, PCIe and NVLink throughput, ECC errors, and XID error events (gpu_xid_critical_errors).
CUDA graphs: two granularities
Section titled “CUDA graphs: two granularities”Engines that capture their decode step into a CUDA graph — vLLM, SGLang, TensorRT-LLM, llama.cpp — replay one graph per token, and what you get depends on --cuda-graph-trace:
graph(the default) — each replay is one timing incuda_graphs_nanoseconds.cuda_kernels_nanosecondsthen holds only eagerly launched kernels, typically the prefill path, and the decode kernels are not in it at all. This is the cheap mode and the one to leave on.node—graphsignal-run --cuda-graph-trace nodetimes the kernels inside the graph individually, socuda_kernels_nanosecondsranks every decode kernel by symbol andcuda_graphs_nanosecondsstays empty. Nothing else changes, and no elevated privileges are needed. CUPTI does more work per replay, so reach for it when you need the per-kernel breakdown, then go back to the default.
To instrument code and kernels — regions, ops, custom timings, whether added by hand or by an AI agent — add GPU probes; their values appear next to these metrics.
How it works
Section titled “How it works”graphsignal-run injects a small activity library into the workload process via CUDA_INJECTION64_PATH; it collects CUPTI activity records with low overhead and publishes cumulative state once per second. Everything else — analysis, statistics, NVML polling, the HTTP endpoint — runs in the sidecar profiler process, never inside the workload.
Integration
Section titled “Integration”Wrap your launch command with graphsignal-run:
graphsignal-run <my-app>CUPTI activity collection and NVML metrics start automatically once a CUDA context exists in the workload. While it runs, read everything at http://127.0.0.1:18259/signals. See the Profiler CLI reference for options.