Docs Blog Pricing
Log in Sign up
Docs Blog Pricing GitHub

Inference Profiling

Graphsignal blog posts about Inference Profiling.

Oct 9, 2026

graphsignal-top: a live view of your inference GPUs

The open-source Graphsignal profiler serves everything it measures as JSON for agents. graphsignal-top is the same data as a live terminal screen: GPU telemetry, engine throughput and latency, and kernels ranked by time, locally or over the network.

Read full story

Sep 18, 2026

Graphsignal vs. torch.profiler and nsys

torch.profiler and nsys already give you per-kernel GPU time, so what does Graphsignal add? Comparable JSON snapshots, telemetry in the same document, GPU probes, overhead low enough to run in production, and a runner that wraps any engine.

Read full story

Jun 22, 2026

CUDA Profiler for Production Inference

Why dev-time CUDA profilers don't fit production inference, and what a profiler built for it looks like: low-overhead kernel attribution, host sync waits, and integrated telemetry.

Read full story

Mar 17, 2026

Traditional Observability Is Blind to Inference

Inference observability monitors inference systems at millisecond granularity, exposing internal runtime and GPU behavior hidden by second-level metrics.

Read full story

Mar 16, 2026

vLLM Production Observability: From Model to Hardware

Production-grade profiling and monitoring for vLLM: always-on vLLM, PyTorch and CUDA profiling with tracing, metrics and errors in one place.

Read full story

Footer

Product

  • Sign Up
  • Docs
  • Blog

Company

  • Contact Us
  • Terms of Service
  • Privacy Policy
  • Cookies Policy
LinkedIn X GitHub

© 2026 Graphsignal, Inc. All rights reserved.