# How to Trace and Evaluate a Hermes Agent with NVIDIA NeMo Relay and Phoenix

## Executive summary

This tutorial demonstrates how to use NVIDIA NeMo Relay to trace and evaluate the complex execution path of a Hermes AI agent. By capturing the entire workflow—from initial requests and file-reading tool calls to the final response—developers can gain deep visibility into the agent's actions, including which tools were used, the data accessed, tokens consumed, and the overall cost of the run. This detailed tracing is crucial for debugging, evaluation, and improving AI agent reliability.

## Key takeaways

- Comprehensive Agent Tracing: NVIDIA NeMo Relay captures the full execution path of a Hermes Agent, providing visibility into tool usage, data access, and model activity beyond just the final answer.
- Reviewing the Trace in Phoenix: The captured trace can be opened in Phoenix, allowing developers to review the session, model activity, tool calls (e.g., reading files, searching the web, writing reports), and the specific data returned at each step.
- Security and Evaluation: The process includes a sanitizer that removes personally identifiable information (PII) before the trace leaves the machine, ensuring data privacy while maintaining full debug capability.

## Technical details

- Setup and Initialization: The process involves cloning the repository and setting up the environment, including configuring the NVIDIA build key and building a sandbox image that Hermes can use to run commands.
- Simple Agent Workflow Test: A simple task (e.g., running a script to print '42') is executed to demonstrate that NeMo Relay successfully records the agent's output and return value.
- Complex Agent Workflow Test: A complex task, such as generating a travel plan report (San Diego, June 29th to July 3rd), is run. This demonstrates the agent's ability to read files, search the web, verify information, and write a final report, all while being traced to Phoenix.

## Practical implications

- For build engineers, this capability is vital for debugging complex, multi-step AI workflows, ensuring that the system behaves predictably across different environments.
- The trace provides an auditable record of all inputs, tool calls, and intermediate data, which is essential for evaluating reliability and debugging failures in production AI agents.
- Understanding the full execution path helps optimize resource consumption by identifying which tools or steps consume the most time or tokens.

## Topics

AI Agent Workflow, Agent Tracing, NVIDIA NeMo Relay, Hermes Agent, AI Debugging, Tool Calling, GitHub: nemoclaw-community, Tech blog: Tracing Agent Harness Behavior, GitHub: NeMo-Relay

Source: https://www.youtube.com/watch?v=WWL99l93xsE
