AI agents

Debugging AI agents with NVIDIA NeMo Relay and OpenTelemetry

NVIDIA NeMo Relay adds observability to Hermes Agent, capturing detailed execution traces for debugging and performance evaluation.

NVIDIA has released a technical tutorial demonstrating how to trace and debug AI agent behavior using NeMo Relay. The guide focuses on the Hermes Agent framework, showing developers how to capture structured lifecycle events and visualize them in tools like Arize Phoenix. Published in late September 2026, this resource aims to solve the black-box problem inherent in complex autonomous systems.

What happened

The NVIDIA Technical Blog published a practical walkthrough for implementing observability in AI agents. The core subject is NeMo Relay, a layer that records ordered events and trajectories during agent execution. The tutorial uses Hermes Agent, which includes native support for this relay system, to generate three types of output: Agent Trajectory Observability Format (ATOF) logs, Agent Trajectory Interchange Format (ATIF) step-by-step records, and OpenTelemetry spans labeled with OpenInference standards.

The article details two specific experiments. The first involves a simple terminal task where the agent runs a script in an isolated Docker container to print "VALUE=42". This verifies the basic setup, ensuring the agent can invoke tools and produce trace files without network access. The second experiment is more complex, requiring the agent to read a travel record, search the web, verify conference details on an official site, and write a report identifying "COLT 2026". This multi-step process demonstrates how traces capture interactions across different tools and external services.

Beyond individual debugging, the post highlights a benchmark case study called Hermes ToolPerf. Researchers used NeMo Relay traces to evaluate changes to the agent harness across 108 runs. They compared a baseline version against a revised version with tool-layer fixes. The data revealed trade-offs that simple success metrics would miss, such as increased latency and token usage even when task completion rates improved.

How it works

NeMo Relay functions by intercepting the agent's workflow at key lifecycle points. It treats sessions, turns, model calls, and tool calls as hierarchical scopes. When work begins or ends, the relay records the event with precise timestamps and parent-child relationships. This structure allows developers to reconstruct the exact sequence of actions taken by the agent.

The system exports this data in formats suited for different analysis needs. ATOF provides a raw JSONL log ideal for auditing timing and error states. ATIF assembles these events into a readable narrative of the agent's decisions. For visual inspection, the relay uses an OpenInference exporter to send data via the OpenTelemetry Protocol (OTLP) to compatible backends. In the tutorial, this backend is Arize Phoenix, which displays the traces as interactive graphs showing model inputs, tool outputs, and duration metrics.

Key details

  • NeMo Relay captures three output formats: ATOF for event logs, ATIF for step-by-step trajectories, and OpenTelemetry spans for visualization.
  • The tutorial requires macOS or Linux, Docker, Git, and an NVIDIA Build API key for Nemotron 3.5 Lightning.
  • Hermes Agent version 0.21.1 and NeMo Relay version 0.8.3 are used in the provided examples.
  • The multi-tool research task successfully identified "COLT 2026" after reading files, searching the web, and verifying sources.
  • In the ToolPerf benchmark, Qwen Coder 30B improved task success from 70% to 81% with fixes but increased mean duration from 27s to 42s.
  • Traces can contain sensitive data like prompts and file paths, so the documentation advises reviewing them before sharing.

Why it matters

For engineers building autonomous agents, a correct final answer does not guarantee an efficient or safe process. An agent might succeed by luck, retrying failed searches multiple times or making redundant API calls. Without visibility into these intermediate steps, developers cannot optimize for cost, latency, or reliability. NeMo Relay provides the evidence layer needed to distinguish between robust performance and fragile hacks.

This observability is critical for production environments where agents interact with external tools and data. By capturing structured traces, teams can audit behavior for security compliance and debug failures that occur only under specific conditions. The ability to compare baseline and revised harnesses using consistent metrics allows for data-driven improvements rather than guesswork. As agents become more complex, understanding the "how" becomes as important as the "what."

What you can do

  • Clone the companion repository from GitHub to access the runnable examples and verifier scripts.
  • Set up an isolated runtime using the provided setup script to avoid conflicting with existing Python environments.
  • Run the simple terminal task first to verify that Docker, Hermes, and NeMo Relay are configured correctly.
  • Inspect the generated ATOF and ATIF files to understand the raw structure of agent events and trajectories.
  • Deploy Arize Phoenix locally and run the multi-tool research task to visualize OpenTelemetry spans in a graphical interface.
  • Use the trace data to evaluate any changes to your agent prompts or tool configurations before deploying them to production.

Tools from the Bytechap store

Keep reading

All stories