AI Optimization
The Graphsignal profiler is built for AI agents: everything it measures is served as one JSON document an agent can read, interpret, and act on. The loop is always the same — launch under the profiler, put load on the workload, read the signals, change something, measure again.
Give your agent the skill
Section titled “Give your agent the skill”The profiler repository ships a skill that teaches an agent the payload shape, the metric semantics, and how to interpret them.
Claude Code — install it into the personal skills directory:
mkdir -p ~/.claude/skills/graphsignal && curl -fsSL \ https://raw.githubusercontent.com/graphsignal/graphsignal/main/SKILL.md \ -o ~/.claude/skills/graphsignal/SKILL.mdOther agents can be pointed at SKILL.md in the repository directly — it is self-contained.
With the skill loaded, prompts like these work end to end:
- “Launch vLLM under graphsignal-run, run the benchmark, and tell me which kernels dominate.”
- “Profile this workload and find out why the GPU is underutilized.”
- “Compare the cuda_kernels_nanoseconds profile before and after this change.”
The local loop
Section titled “The local loop”The agent (or you) launches the workload under the profiler:
graphsignal-run vllm serve <model> --port 8001While the workload runs, the profiler serves everything it measures at a local endpoint:
curl -s http://127.0.0.1:18259/signalsAn effective read order:
errorsfirst — a crash or GPU XID error explains more than any metric.gpu_utilization_percentandgpu_memory_*per device — is the GPU busy, starved, or memory-bound?- The
cuda_kernels_nanosecondsandcuda_graphs_nanosecondsprofiles — rank the frames by cumulative time (they arrive sorted descending) and take the top few as the candidates to change. If most of the GPU time sits incuda_graphs_nanosecondsand the kernel profile looks thin, the engine replays a CUDA graph per decode step and you are ranking whole replays: relaunch withgraphsignal-run --cuda-graph-trace node <command>and the same profile ranks the kernels inside the graph by symbol. Thecuda_graph_trace_modegauge says which granularity a payload came from. - The
cuda_sync_nanosecondsprofile andcuda_memcpy_nanosecondsprofile/byte counters — heavy host synchronization or transfer volume signals CPU/IO bottlenecks. - Engine metrics (queue depth, running requests, token throughput) imported from the engine’s Prometheus endpoint.
Counters, histograms, and profiles are cumulative: the agent polls the endpoint periodically (every 10–60s) and diffs between polls to observe trends. Each read is the latest snapshot.
Once a kernel is named, the next question is which part of it — and that is where instrumentation takes over. Custom instrumentation joins the same loop: GPU probes let an agent (or you) add instruments to application code and kernels, and their values appear in /signals next to the built-in metrics.
The production loop
Section titled “The production loop”The local endpoint lives only as long as the workload, on the machine that runs it. For production fleets, set an API key and the profiler additionally uploads its signals to Graphsignal:
export GRAPHSIGNAL_API_KEY=<api-key>graphsignal-run vllm serve <model> --port 8001The agent then reads production behavior back from anywhere — typically from the development machine where it is changing flags or code:
# which instances have reportedcurl -s -H "X-API-Key: $GRAPHSIGNAL_API_KEY" \ "https://api.graphsignal.com/api/v1/instances"
# one instance's signals, same shape as the local /signals payloadcurl -s -H "X-API-Key: $GRAPHSIGNAL_API_KEY" \ "https://api.graphsignal.com/api/v1/signals?instance_id=<id>"This closes the larger loop: production workloads report what they actually did, and the agent uses it to tune the next deployment. See the Signals API reference for the endpoint details.
Example prompts against production data:
- “List the instances that reported in the last 24 hours and summarize their errors.”
- “Fetch the signals for instance
<id>and identify the main performance bottleneck.” - “Compare the kernel profiles of the two vLLM instances and explain the throughput difference.”