# Post-Training at the Frontier: How You.com, NVIDIA, and CoreWeave Improved Nemotron’s Agentic Search

## Executive summary

You.com, NVIDIA, and CoreWeave collaborated to significantly improve Nemotron 3.5 Lightning's agentic web search and browsing capabilities through a rigorous post-training process. The effort focused on addressing the bottleneck of model tool usage proficiency. By leveraging You.com's proprietary web graph and index to generate specialized, decontaminated training data, the team applied Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) on CoreWeave's dedicated infrastructure. This resulted in an 8.5 absolute increase (23% relative increase) in accuracy on the BrowseComp benchmark and drastically improved efficiency, reducing average tool calls per task from 168 to 115.

## Key takeaways

- Post-Training Methodology: The post-training process involved generating specialized data using You.com's web graph (a super large data structure where pages are vertices and links are edges). This data was then filtered and decontaminated to ensure training tasks did not contain answers to the evaluation benchmark questions. The model was trained using SFT on curated traces and RL on tasks with variability in outcomes.
- Performance Gains: Post-training Nemotron 3.5 Lightning yielded an 8.5 absolute increase (23% relative increase) in accuracy on the BrowseComp benchmark. Furthermore, the model's efficiency improved dramatically, reducing average tool calls per task from 168 to 115, leading to significant token and latency savings.
- Infrastructure and Efficiency: CoreWeave provided dedicated RL and inference platforms, enabling rapid iteration. The use of CoreWeave's hotloading feature significantly reduced the time required for RL runs, allowing the team to complete the post-training process in a short timeframe.

## Technical details

- Model Architecture & Benchmarking: The base model was Nemotron 3.5 Lightning (a 30 billion parameter, 3 billion active parameter Mixture of Experts model). Evaluation was conducted using the BrowseComp benchmark, which requires sourcing facts and web browsing. The initial model achieved 37% on the benchmark.
- Failure Analysis: Initial model failures were attributed to weaker tool use, particularly with unique output schemas from web search tools. Patterns included 'second guessing' (over-searching after finding a correct answer) and 'death loops' (searching repeatedly without finding the correct information).
- Data Generation (Web Graph): You.com utilized its web graph to generate data by focusing on pages in the middle band of popularity (avoiding boilerplate or spammy URLs). They then used an LLM to generate verifiable question-answer pairs from these pages, creating a dataset of retrieval questions.
- Training Pipeline: The pipeline involved: 1) Filtering data by difficulty; 2) Decontaminating the dataset (ensuring training data did not contain answers to benchmark questions); 3) Running eight iterations of each question inside the Nemo harness with a binary reward judge; 4) Routing correct traces to SFT and remaining tasks to RL.

## Practical implications

- The process demonstrates that even a small, targeted dataset (5,000 SFT traces, 1,500 RL tasks) can yield outsized performance gains when applied to a powerful foundation model.
- Focusing on improving tool usage proficiency (e.g., web search API interaction) is a critical bottleneck for maximizing the value of large language models.
- The use of proprietary web graphs and specialized data generation techniques offers a blueprint for improving LLMs on real-world, time-sensitive search tasks.

## Topics

Large Language Models (LLMs), Reinforcement Learning (RL), Supervised Fine-Tuning (SFT), Web Search Infrastructure, Agentic AI, Nemotron 3.5 Lightning, BrowseComp, Nemo gym

Source: https://www.youtube.com/watch?v=ObzZcszkdSs
