# CoreWeave Forge: Demo from Fully Connected 2026, with Corey Sanders

## Executive summary

This demo outlines the complete AI research and iteration loop using CoreWeave Forge, demonstrating how to systematically improve an AI agent (like an IT help desk agent) without leaving the platform. The process involves observing production failures (Agent Lens), curating failure patterns (Tags/LLM Judge), identifying systemic issues (Insights), training the model (Serverless RL in Weights & Biases Models), and validating the improvements (Evaluation and Lineage). The entire workflow is designed to move from a first agent to a best agent through continuous, measurable iteration.

## Key takeaways

- The AI Iteration Loop: The core workflow is a continuous cycle: Run $\rightarrow$ Observe $\rightarrow$ Curate $\rightarrow$ Improve $\rightarrow$ Evaluate $\rightarrow$ Run again. This allows for systematic, data-driven model improvement.
- Observability with Agent Lens: Agent Lens records every agent action in production, providing views of costs, tools, spans, and latency, allowing domain experts to understand agent behavior at a granular level.
- Automated Failure Detection: The Insights tab acts as an agent that detects subtle, systemic problems (e.g., 1.6% misprioritization) that would be impossible to spot by manual review.
- Scalable Model Training: Serverless Reinforcement Learning (RL) trains the agent by rewarding correct actions and penalizing incorrect ones. CoreWeave Sandboxes provide isolated, scalable environments for running these parallel RL jobs.
- Model Lineage and Validation: Weights & Biases Models tracks the entire run history. The lineage feature provides a beautiful tree view showing exactly how an improved model was derived from previous versions, ensuring traceability.

## Technical details

- CoreWeave Forge: The unified platform for managing the entire AI lifecycle, including observation, curation, improvement, and evaluation.
- CoreWeave Agent Lens: A tool that records every step an agent takes in production, grouping conversations that fail similarly to focus efforts on critical failure patterns.
- ARIA (AI Research & Iteration Agent): An AI research assistant that reads run traces and provides actionable plans, from suggesting code changes (diffs) to recommending post-training strategies.
- Serverless Reinforcement Learning (RL): A training method that allows an agent to learn through simulated runs by scoring rewards for correct actions and penalizing incorrect ones, without requiring a dedicated training cluster or hand-written reward function.
- CoreWeave Sandboxes: Provides isolated environments necessary for scaling RL jobs, allowing multiple training runs to execute in parallel.
- Model Lineage: A feature that tracks the complete history and derivation path of a model, showing where an improved version came from.

## Practical implications

- Engineers can automate the process of identifying and fixing production AI failures, moving beyond manual debugging.
- The platform integrates observability (Agent Lens) directly with model training (RL), closing the loop between production data and model improvement.
- The ability to compare new and old models side-by-side on real, historical traces minimizes deployment risk.
- The system provides full traceability (lineage) for all model changes, which is critical for regulated or complex build environments.

## Topics

AI/ML Ops, Reinforcement Learning, Observability, Build Engineering, LLMs, CI/CD, CoreWeave Forge, Weights & Biases Models

Source: https://www.youtube.com/watch?v=Q7UwZjEVFWA
