Post-Training at the Frontier: How You.com, NVIDIA, and CoreWeave Improved Nemotron’s Agentic Search
You.com, NVIDIA, and CoreWeave collaborated to significantly improve Nemotron 3.5 Lightning's agentic web search and browsing capabilities through a rigorous post-training process. The effort focused on addressing the bottleneck of model tool usage proficiency. By leveraging You.com's proprietary web graph and index to generate specialized, decontaminated training data, the team applied Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) on CoreWeave's dedicated infrastructure. This resulted in an 8.5 absolute increase (23% relative increase) in accuracy on the BrowseComp benchmark and drastically improved efficiency, reducing average tool calls per task from 168 to 115.
Key takeaways
-
Post-Training Methodology
22:10
The post-training process involved generating specialized data using You.com's web graph (a super large data structure where pages are vertices and links are edges). This data was then filtered and decontaminated to ensure training tasks did not contain answers to the evaluation benchmark questions. The model was trained using SFT on curated traces and RL on tasks with variability in outcomes.
-
Performance Gains
18:50
Post-training Nemotron 3.5 Lightning yielded an 8.5 absolute increase (23% relative increase) in accuracy on the BrowseComp benchmark. Furthermore, the model's efficiency improved dramatically, reducing average tool calls per task from 168 to 115, leading to significant token and latency savings.
-
Infrastructure and Efficiency
16:00
CoreWeave provided dedicated RL and inference platforms, enabling rapid iteration. The use of CoreWeave's hotloading feature significantly reduced the time required for RL runs, allowing the team to complete the post-training process in a short timeframe.