Weights & Biases

CoreWeave Arena: Serverless inference and serverless RL

Published 2026-09-29 · Duration 5:49

Summary

CoreWeave Arena provides a platform for validating AI workloads end-to-end—including pre-training, building agents, and post-training via Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL)—before committing to production. The presentation demonstrates how Serverless Inference and Serverless RL can be used to build and improve AI agents, specifically showing how RL can significantly boost agent performance and reliability, which is then deployed using CoreWeave Serverless Inference.

Download summary

Key takeaways

  1. CoreWeave Arena Purpose

    Arena is designed for validating AI workloads under production-like conditions, enabling users to run real workloads on CoreWeave and pair it with Coreweave Forge to close the AI development loop.

  2. Agent Development & Tracing 2:00

    Agents can be built using models like Quen 314b instruct via Serverless Inference. Performance is monitored by collecting agent traces and using custom signals (e.g., misalignment between user request and generated SQL statement) to identify behavioral issues.

  3. Serverless RL for Improvement 2:50

    Serverless RL is presented as a preferred method for improving AI application reliability by fine-tuning LLMs for agentic tasks. Coreweave addresses the challenges of RL by providing instant, elastic GPU capacity without provisioning.

  4. Evaluation and Deployment 4:20

    Agent improvement is validated using pre- and post-RL agent evaluations (e.g., spider charts). Once optimized, the model weights and necessary Python code for calling the model via Coreweave Serverless Inference are retrieved for production deployment.

Technical details

  • CoreWeave Arena 0s

    A platform for end-to-end validation of AI workloads, supporting pre-training, agent building, and post-training (SFT/RL).

  • Serverless Inference 80s

    Allows access to various open-source models for running agents (e.g., Quen 314b instruct) and generating outputs from natural language prompts.

  • Agent Tracing and Signals 120s

    Agent performance is tracked using collected traces. Custom signals can be configured to flag specific issues, such as 'misalignment between the user request and the SQL statement that generates the distribution list.'

  • Serverless RL Workflow 170s

    Users can run RL jobs, monitor progress via dashboards, and track best-performing model checkpoints. The process involves creating an environment, writing agent code, and defining training scenarios.

  • Model Deployment 280s

    After RL optimization, the specific model weights and a model URI are accessed via the 'artifacts' tab, providing sample Python code for integration using Coreweave Serverless Inference.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.