# CoreWeave Arena: Serverless inference and serverless RL

## Executive summary

CoreWeave Arena provides a platform for validating AI workloads end-to-end—including pre-training, building agents, and post-training via Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL)—before committing to production. The presentation demonstrates how Serverless Inference and Serverless RL can be used to build and improve AI agents, specifically showing how RL can significantly boost agent performance and reliability, which is then deployed using CoreWeave Serverless Inference.

## Key takeaways

- CoreWeave Arena Purpose: Arena is designed for validating AI workloads under production-like conditions, enabling users to run real workloads on CoreWeave and pair it with Coreweave Forge to close the AI development loop.
- Agent Development & Tracing: Agents can be built using models like Quen 314b instruct via Serverless Inference. Performance is monitored by collecting agent traces and using custom signals (e.g., misalignment between user request and generated SQL statement) to identify behavioral issues.
- Serverless RL for Improvement: Serverless RL is presented as a preferred method for improving AI application reliability by fine-tuning LLMs for agentic tasks. Coreweave addresses the challenges of RL by providing instant, elastic GPU capacity without provisioning.
- Evaluation and Deployment: Agent improvement is validated using pre- and post-RL agent evaluations (e.g., spider charts). Once optimized, the model weights and necessary Python code for calling the model via Coreweave Serverless Inference are retrieved for production deployment.

## Technical details

- CoreWeave Arena: A platform for end-to-end validation of AI workloads, supporting pre-training, agent building, and post-training (SFT/RL).
- Serverless Inference: Allows access to various open-source models for running agents (e.g., Quen 314b instruct) and generating outputs from natural language prompts.
- Agent Tracing and Signals: Agent performance is tracked using collected traces. Custom signals can be configured to flag specific issues, such as 'misalignment between the user request and the SQL statement that generates the distribution list.'
- Serverless RL Workflow: Users can run RL jobs, monitor progress via dashboards, and track best-performing model checkpoints. The process involves creating an environment, writing agent code, and defining training scenarios.
- Model Deployment: After RL optimization, the specific model weights and a model URI are accessed via the 'artifacts' tab, providing sample Python code for integration using Coreweave Serverless Inference.

## Practical implications

- Build engineers can integrate AI agent development and validation directly into the CI/CD pipeline using CoreWeave Arena.
- The use of Serverless RL allows for iterative, reliable improvement of LLM agents without requiring dedicated GPU infrastructure provisioning.
- The ability to compare 'before' and 'after' agent performance metrics (e.g., using spider charts) provides quantifiable proof of improvement for production readiness.
- The workflow streamlines the transition from experimental agent development to production deployment via model URI and sample code artifacts.

## Topics

AI Agents, Reinforcement Learning (RL), Large Language Models (LLMs), Serverless Computing, Prompt Engineering, Model Validation, CoreWeave ARENA, CoreWeave Serverless Inference, Serverless Inference docs

Source: https://www.youtube.com/watch?v=34w2NQsh974
