Skip to content

Profiler CLI

Graphsignal observes your workload from a sidecar process — the profiler. It runs out-of-process, never inside your workload. graphsignal-run launches a workload with the profiler attached.

Install the CLI as an isolated uv tool so it doesn’t pollute the workload environment:

Terminal window
UV_TOOL_BIN_DIR=/usr/local/bin uv tool install 'graphsignal[cu12]' # CUDA 12.x
# or
UV_TOOL_BIN_DIR=/usr/local/bin uv tool install 'graphsignal[cu13]' # CUDA 13.x

UV_TOOL_BIN_DIR=/usr/local/bin puts graphsignal-run in a directory that is already on PATH for every shell, including non-interactive scripts and containers. pip install 'graphsignal[cu12]' into the workload environment works too.

Wrap any launch command. graphsignal-run starts the profiler sidecar, enables GPU profiling, and launches your workload so process managers (init systems, container runtimes, etc.) see only the workload.

Terminal window
graphsignal-run <command> [args...]

Examples:

Terminal window
graphsignal-run vllm serve <model> --port 8001
graphsignal-run sglang serve --model-path <model>
graphsignal-run python -m sglang.launch_server --model-path <model>
graphsignal-run trtllm-serve <model> --port 8000
graphsignal-run --metrics-port 8000 trtllm-serve <model> --port 8000
graphsignal-run python myapp.py
graphsignal-run app.py

Options (must precede the command):

  • --version — Print the profiler version and exit.
  • --metrics-port PORT — Port to scrape the workload’s Prometheus /metrics endpoint on. Overrides the port derived from the engine’s --port flag or its default (e.g. 8000 for vLLM/TensorRT-LLM, 30000 for SGLang). Use this when metrics are exposed on a different port than the HTTP server. Not forwarded to the workload.
  • --listen-host HOST — Host to bind the /signals endpoint to (default 127.0.0.1, loopback only). Set e.g. 0.0.0.0 to expose the endpoint for remote access — the endpoint is unauthenticated, so anything that can reach that address can read it. Not forwarded to the workload.
  • --listen-port PORT — Port for the /signals endpoint (default 18259). Not forwarded to the workload.
  • --cuda-graph-trace {graph|node} — Granularity for CUDA graph launches (default graph). With graph, each graph replay is timed as a whole and reported in cuda_graphs_nanoseconds; the kernels inside it are not timed individually. With node, those kernels are timed individually and reported in cuda_kernels_nanoseconds by symbol, and cuda_graphs_nanoseconds stays empty. Use node to rank the kernels of an engine that replays CUDA graphs in decode (vLLM, SGLang, TensorRT-LLM, llama.cpp); it needs no extra privileges, but CUPTI does more work per replay, so prefer the default for long-running deployments. Ignored on ROCm, where rocprofiler-sdk reports every dispatch individually. Not forwarded to the workload.

Behavior:

  • Detects the engine from your command (vLLM, SGLang, TensorRT-LLM, or a generic fallback).
  • Scrapes Prometheus metrics from one known endpoint — http://127.0.0.1:<port>/metrics, or /prometheus/metrics for TensorRT-LLM — with the port resolved from --metrics-port, the engine’s --port, or the engine default. It never probes the workload’s other listening sockets; connecting to them blindly can corrupt internal IPC.
  • Collects GPU activity via CUPTI (or ROCm’s rocprofiler-sdk) as soon as the workload starts using the GPU.
  • Passes the workload command through byte-for-byte, with two metric-related adjustments: the SGLang launcher adds --enable-metrics (its Prometheus endpoint is off by default) and the vLLM launcher removes --disable-log-stats.
  • Traces CUDA graphs at graph granularity by default: a replay is one timing, and its kernels do not appear in the kernel profile. --cuda-graph-trace node switches to per-kernel granularity for the same graphs. The cuda_graph_trace_mode gauge in the payload always says which granularity produced it.
  • Serves everything measured at http://127.0.0.1:<listen-port>/signals while the workload runs.
  • Exits with the workload’s exact status.

The profiler reads its configuration from environment variables. Set these before invoking graphsignal-run.

VariablePurpose
GRAPHSIGNAL_LISTEN_HOSTHost to bind the /signals endpoint to (same as --listen-host; default 127.0.0.1).
GRAPHSIGNAL_LISTEN_PORTPort for the /signals endpoint (same as --listen-port).
GRAPHSIGNAL_TAG_<KEY>=<value>Arbitrary tag attached to all signals (e.g. GRAPHSIGNAL_TAG_DEPLOYMENT=us-prod).
GRAPHSIGNAL_CUDA_GRAPH_TRACEgraph (default) or node — CUDA graph tracing granularity, same as --cuda-graph-trace. The flag wins over this variable. An unrecognized value falls back to graph and is noted under GRAPHSIGNAL_DEBUG=1.
GRAPHSIGNAL_DEBUGSet to 1 to log the profiler’s own diagnostics to stderr.
GRAPHSIGNAL_API_KEY (optional)Enables the production feedback loop: signals are uploaded to Graphsignal. Without it, no profiling data is ever uploaded.
GRAPHSIGNAL_API_BASEOverride the upload endpoint (defaults to https://api.graphsignal.com).
GRAPHSIGNAL_DISABLE_VERSION_CHECKSet to 1 to turn off the once-per-run check for a newer release (see Security and Privacy).