Topic

AI/ML Ops

All digests tagged AI/ML Ops

CoreWeave Forge: Demo from Fully Connected 2026, with Corey Sanders thumbnail

· 11:24

CoreWeave Forge: Demo from Fully Connected 2026, with Corey Sanders

This demo outlines the complete AI research and iteration loop using CoreWeave Forge, demonstrating how to systematically improve an AI agent (like an IT help desk agent) without leaving the platform. The process involves observing production failures (Agent Lens), curating failure patterns (Tags/LLM Judge), identifying systemic issues (Insights), training the model (Serverless RL in Weights & Biases Models), and validating the improvements (Evaluation and Lineage). The entire workflow is designed to move from a first agent to a best agent through continuous, measurable iteration.

Key takeaways

  1. The AI Iteration Loop

    The core workflow is a continuous cycle: Run $\rightarrow$ Observe $\rightarrow$ Curate $\rightarrow$ Improve $\rightarrow$ Evaluate $\rightarrow$ Run again. This allows for systematic, data-driven model improvement.

  2. Observability with Agent Lens 3:50

    Agent Lens records every agent action in production, providing views of costs, tools, spans, and latency, allowing domain experts to understand agent behavior at a granular level.

  3. Automated Failure Detection 7:26

    The Insights tab acts as an agent that detects subtle, systemic problems (e.g., 1.6% misprioritization) that would be impossible to spot by manual review.

  4. Scalable Model Training

    Serverless Reinforcement Learning (RL) trains the agent by rewarding correct actions and penalizing incorrect ones. CoreWeave Sandboxes provide isolated, scalable environments for running these parallel RL jobs.

  5. Model Lineage and Validation

    Weights & Biases Models tracks the entire run history. The lineage feature provides a beautiful tree view showing exactly how an improved model was derived from previous versions, ensuring traceability.

Watch on YouTube Full article