Weights & Biases

CoreWeave Forge: Demo from Fully Connected 2026, with Corey Sanders

Published 2026-10-06 · Duration 11:24

Summary

This demo outlines the complete AI research and iteration loop using CoreWeave Forge, demonstrating how to systematically improve an AI agent (like an IT help desk agent) without leaving the platform. The process involves observing production failures (Agent Lens), curating failure patterns (Tags/LLM Judge), identifying systemic issues (Insights), training the model (Serverless RL in Weights & Biases Models), and validating the improvements (Evaluation and Lineage). The entire workflow is designed to move from a first agent to a best agent through continuous, measurable iteration.

Download summary

Key takeaways

  1. The AI Iteration Loop

    The core workflow is a continuous cycle: Run $\rightarrow$ Observe $\rightarrow$ Curate $\rightarrow$ Improve $\rightarrow$ Evaluate $\rightarrow$ Run again. This allows for systematic, data-driven model improvement.

  2. Observability with Agent Lens 3:50

    Agent Lens records every agent action in production, providing views of costs, tools, spans, and latency, allowing domain experts to understand agent behavior at a granular level.

  3. Automated Failure Detection 7:26

    The Insights tab acts as an agent that detects subtle, systemic problems (e.g., 1.6% misprioritization) that would be impossible to spot by manual review.

  4. Scalable Model Training

    Serverless Reinforcement Learning (RL) trains the agent by rewarding correct actions and penalizing incorrect ones. CoreWeave Sandboxes provide isolated, scalable environments for running these parallel RL jobs.

  5. Model Lineage and Validation

    Weights & Biases Models tracks the entire run history. The lineage feature provides a beautiful tree view showing exactly how an improved model was derived from previous versions, ensuring traceability.

Technical details

  • CoreWeave Forge 0s

    The unified platform for managing the entire AI lifecycle, including observation, curation, improvement, and evaluation.

  • CoreWeave Agent Lens 230s

    A tool that records every step an agent takes in production, grouping conversations that fail similarly to focus efforts on critical failure patterns.

  • ARIA (AI Research & Iteration Agent) 338s

    An AI research assistant that reads run traces and provides actionable plans, from suggesting code changes (diffs) to recommending post-training strategies.

  • Serverless Reinforcement Learning (RL) 632s

    A training method that allows an agent to learn through simulated runs by scoring rewards for correct actions and penalizing incorrect ones, without requiring a dedicated training cluster or hand-written reward function.

  • CoreWeave Sandboxes

    Provides isolated environments necessary for scaling RL jobs, allowing multiple training runs to execute in parallel.

  • Model Lineage

    A feature that tracks the complete history and derivation path of a model, showing where an improved version came from.

Mentioned resources

  • CoreWeave Forge (Platform)
  • Weights & Biases Models (ML Tracking/Model Registry)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.