Post-Training at the Frontier: How You.com, NVIDIA, and CoreWeave Improved Nemotron’s Agentic Search
Summary
You.com, NVIDIA, and CoreWeave collaborated to significantly improve Nemotron 3.5 Lightning's agentic web search and browsing capabilities through a rigorous post-training process. The effort focused on addressing the bottleneck of model tool usage proficiency. By leveraging You.com's proprietary web graph and index to generate specialized, decontaminated training data, the team applied Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) on CoreWeave's dedicated infrastructure. This resulted in an 8.5 absolute increase (23% relative increase) in accuracy on the BrowseComp benchmark and drastically improved efficiency, reducing average tool calls per task from 168 to 115.
Key takeaways
-
Post-Training Methodology
22:10
The post-training process involved generating specialized data using You.com's web graph (a super large data structure where pages are vertices and links are edges). This data was then filtered and decontaminated to ensure training tasks did not contain answers to the evaluation benchmark questions. The model was trained using SFT on curated traces and RL on tasks with variability in outcomes.
-
Performance Gains
18:50
Post-training Nemotron 3.5 Lightning yielded an 8.5 absolute increase (23% relative increase) in accuracy on the BrowseComp benchmark. Furthermore, the model's efficiency improved dramatically, reducing average tool calls per task from 168 to 115, leading to significant token and latency savings.
-
Infrastructure and Efficiency
16:00
CoreWeave provided dedicated RL and inference platforms, enabling rapid iteration. The use of CoreWeave's hotloading feature significantly reduced the time required for RL runs, allowing the team to complete the post-training process in a short timeframe.
Technical details
-
Model Architecture & Benchmarking
750s
The base model was Nemotron 3.5 Lightning (a 30 billion parameter, 3 billion active parameter Mixture of Experts model). Evaluation was conducted using the BrowseComp benchmark, which requires sourcing facts and web browsing. The initial model achieved 37% on the benchmark.
-
Failure Analysis
800s
Initial model failures were attributed to weaker tool use, particularly with unique output schemas from web search tools. Patterns included 'second guessing' (over-searching after finding a correct answer) and 'death loops' (searching repeatedly without finding the correct information).
-
Data Generation (Web Graph)
900s
You.com utilized its web graph to generate data by focusing on pages in the middle band of popularity (avoiding boilerplate or spammy URLs). They then used an LLM to generate verifiable question-answer pairs from these pages, creating a dataset of retrieval questions.
-
Training Pipeline
1000s
The pipeline involved: 1) Filtering data by difficulty; 2) Decontaminating the dataset (ensuring training data did not contain answers to benchmark questions); 3) Running eight iterations of each question inside the Nemo harness with a binary reward judge; 4) Routing correct traces to SFT and remaining tasks to RL.
Mentioned resources
- Nemotron 3.5 Lightning
- BrowseComp
- Nemo gym
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.