Profiler CLI
Graphsignal observes your workload from a sidecar process — the profiler. It runs out-of-process, never inside your workload. graphsignal-run launches a workload with the profiler attached.
Install the CLI as an isolated uv tool so it doesn’t pollute the workload environment:
UV_TOOL_BIN_DIR=/usr/local/bin uv tool install 'graphsignal[cu12]' # CUDA 12.x# orUV_TOOL_BIN_DIR=/usr/local/bin uv tool install 'graphsignal[cu13]' # CUDA 13.xUV_TOOL_BIN_DIR=/usr/local/bin puts graphsignal-run in a directory that is already on PATH for every shell, including non-interactive scripts and containers. pip install 'graphsignal[cu12]' into the workload environment works too.
graphsignal-run
Section titled “graphsignal-run”Wrap any launch command. graphsignal-run starts the profiler sidecar, enables GPU profiling, and launches your workload so process managers (init systems, container runtimes, etc.) see only the workload.
graphsignal-run <command> [args...]Examples:
graphsignal-run vllm serve <model> --port 8001graphsignal-run sglang serve --model-path <model>graphsignal-run python -m sglang.launch_server --model-path <model>graphsignal-run trtllm-serve <model> --port 8000graphsignal-run --metrics-port 8000 trtllm-serve <model> --port 8000graphsignal-run python myapp.pygraphsignal-run app.pyOptions (must precede the command):
--version— Print the profiler version and exit.--metrics-port PORT— Port to scrape the workload’s Prometheus/metricsendpoint on. Overrides the port derived from the engine’s--portflag or its default (e.g. 8000 for vLLM/TensorRT-LLM, 30000 for SGLang). Use this when metrics are exposed on a different port than the HTTP server. Not forwarded to the workload.--listen-host HOST— Host to bind the/signalsendpoint to (default127.0.0.1, loopback only). Set e.g.0.0.0.0to expose the endpoint for remote access — the endpoint is unauthenticated, so anything that can reach that address can read it. Not forwarded to the workload.--listen-port PORT— Port for the/signalsendpoint (default18259). Not forwarded to the workload.--cuda-graph-trace {graph|node}— Granularity for CUDA graph launches (defaultgraph). Withgraph, each graph replay is timed as a whole and reported incuda_graphs_nanoseconds; the kernels inside it are not timed individually. Withnode, those kernels are timed individually and reported incuda_kernels_nanosecondsby symbol, andcuda_graphs_nanosecondsstays empty. Usenodeto rank the kernels of an engine that replays CUDA graphs in decode (vLLM, SGLang, TensorRT-LLM, llama.cpp); it needs no extra privileges, but CUPTI does more work per replay, so prefer the default for long-running deployments. Ignored on ROCm, whererocprofiler-sdkreports every dispatch individually. Not forwarded to the workload.
Behavior:
- Detects the engine from your command (vLLM, SGLang, TensorRT-LLM, or a generic fallback).
- Scrapes Prometheus metrics from one known endpoint —
http://127.0.0.1:<port>/metrics, or/prometheus/metricsfor TensorRT-LLM — with the port resolved from--metrics-port, the engine’s--port, or the engine default. It never probes the workload’s other listening sockets; connecting to them blindly can corrupt internal IPC. - Collects GPU activity via CUPTI (or ROCm’s
rocprofiler-sdk) as soon as the workload starts using the GPU. - Passes the workload command through byte-for-byte, with two metric-related adjustments: the SGLang launcher adds
--enable-metrics(its Prometheus endpoint is off by default) and the vLLM launcher removes--disable-log-stats. - Traces CUDA graphs at graph granularity by default: a replay is one timing, and its kernels do not appear in the kernel profile.
--cuda-graph-trace nodeswitches to per-kernel granularity for the same graphs. Thecuda_graph_trace_modegauge in the payload always says which granularity produced it. - Serves everything measured at
http://127.0.0.1:<listen-port>/signalswhile the workload runs. - Exits with the workload’s exact status.
Environment variables
Section titled “Environment variables”The profiler reads its configuration from environment variables. Set these before invoking graphsignal-run.
| Variable | Purpose |
|---|---|
GRAPHSIGNAL_LISTEN_HOST | Host to bind the /signals endpoint to (same as --listen-host; default 127.0.0.1). |
GRAPHSIGNAL_LISTEN_PORT | Port for the /signals endpoint (same as --listen-port). |
GRAPHSIGNAL_TAG_<KEY>=<value> | Arbitrary tag attached to all signals (e.g. GRAPHSIGNAL_TAG_DEPLOYMENT=us-prod). |
GRAPHSIGNAL_CUDA_GRAPH_TRACE | graph (default) or node — CUDA graph tracing granularity, same as --cuda-graph-trace. The flag wins over this variable. An unrecognized value falls back to graph and is noted under GRAPHSIGNAL_DEBUG=1. |
GRAPHSIGNAL_DEBUG | Set to 1 to log the profiler’s own diagnostics to stderr. |
GRAPHSIGNAL_API_KEY (optional) | Enables the production feedback loop: signals are uploaded to Graphsignal. Without it, no profiling data is ever uploaded. |
GRAPHSIGNAL_API_BASE | Override the upload endpoint (defaults to https://api.graphsignal.com). |
GRAPHSIGNAL_DISABLE_VERSION_CHECK | Set to 1 to turn off the once-per-run check for a newer release (see Security and Privacy). |