# Evaluating Open Models with Artificial Analysis | Nemotron Labs

## Executive summary

This session provides a deep dive into the rigorous field of AI model evaluation, emphasizing that modern benchmarks must move beyond simple intelligence scores. The focus has shifted dramatically from basic chat capabilities to complex, multi-turn, agentic use cases. Experts detail how to evaluate models based on a composite of metrics—including cost, speed, and the number of turns required—and discuss the technical challenges of benchmarking, such as maintaining consistency across different hardware and model sizes.

## Key takeaways

- Shift to Agentic Benchmarking: The primary use case for LLMs has shifted from simple chat to complex agentic workflows. Evaluation must now test models' ability to 'do work' by executing multi-step tasks and using tools, rather than just generating single responses. This requires specialized infrastructure for running agent sandboxes.
- Multi-Metric Evaluation is Essential: Model selection requires balancing intelligence (capability), cost (tokens used), and speed (tokens per second). Benchmarks must break down metrics by tokens used per turn and the total number of turns/iterations required for a task.
- Hardware and Model Constraints: For local or constrained environments (e.g., 24 GB cards), running large models requires quantization (e.g., NVFP4) and careful consideration of total parameters versus active parameters. The optimal choice involves a trade-off between raw intelligence and operational efficiency (speed/turns).
- The Complexity of Benchmarking: To ensure scientific rigor, benchmarks must isolate the fewest variables possible. This involves techniques like replaying identical agentic trajectories across different hardware to ensure the same work is measured repeatedly, regardless of the underlying model or chip.

## Technical details

- Agentic Use Cases: The evaluation process now focuses on 'agentic use cases,' which involve models performing tasks autonomously over multiple turns. This contrasts with older benchmarks that only tested single tool calls or simple chat interactions.
- Quantization and Local Inference: The concept of 'near-lossless NVFP4 quantization' is discussed in the context of running large models (e.g., 27B parameters) on local, constrained hardware. The total model size must fit in memory, requiring high memory bandwidth for reading parameters during every token generation.
- Benchmarking Rigor: To compare models fairly, benchmarks aim to minimize variables. When testing hardware performance, the methodology involves forcing the model through the same recorded 'agentic trajectories' to ensure the same amount of work is done, allowing for consistent measurement of performance across different chips.

## Practical implications

- When designing agent pipelines, always evaluate model candidates using multi-turn, agentic benchmarks rather than relying solely on general intelligence scores.
- For resource-constrained deployments, prioritize models that offer the best balance between intelligence and operational efficiency (speed/low quantization requirements).
- When comparing models, understand that performance is a composite metric involving cost, speed, and task completion rate, not just a single score.
- Build engineers should utilize specialized tools (like AA Agent Perf Local) that simulate real-world hardware constraints to predict deployment performance.

## Topics

AI Benchmarking, Large Language Models (LLMs), Agentic AI, Model Quantization, Build Engineering, Inference Optimization, Artificial Analysis Website, NVIDIA Neotron GitHub, AA Agent Perf Local

Source: https://www.youtube.com/watch?v=afREwSZh1Yk
