# Prompt engineering and serverless inference: closing the open model gap

## Executive summary

This video demonstrates the end-to-end workflow for building, validating, and improving AI agents using Coreweave's platform. The presenter showcases Coreweave Arena for running real-world, production-like workloads. The process involves using Serverless Inference for agent functionality (e.g., building a list builder using Qwen-3-14B-Instruct) and leveraging Serverless Reinforcement Learning (RL) to improve agent reliability and performance. The workflow allows teams to track performance using signals, compare model versions, and deploy improved models via their URI and sample code.

## Key takeaways

- AI Agent Validation with Coreweave Arena: Coreweave Arena allows users to validate AI workloads (pre-training, agent building, post-training) under production-like conditions before committing to deployment.
- Improving Agents via Serverless RL: Serverless RL is presented as a preferred method for fine-tuning Large Language Models (LLMs) for agentic tasks, improving overall AI application reliability by addressing mistakes identified during QA or production.
- End-to-End Deployment Workflow: After an RL job improves performance, the process involves pinning the best-performing model checkpoint, retrieving the model URI, and using sample Python code to call the model via Coreweave Serverless Inference.

## Technical details

- Agent Development & Inference: An agent was built using the Qwen-3-14B-Instruct model served through Coreweave's serverless inference. The agent was designed to generate distribution lists based on natural language requests.
- Performance Monitoring: Agent performance is monitored using 'signals' (e.g., low-quality response, misalignment between user request and generated SQL statement) logged in the 'Signals' tab.
- Serverless RL Implementation: Coreweave Serverless RL provides instant access to GPU capacity with elastic scaling, eliminating the need for complex provisioning and skilled developers to manage RL scripts.
- Model Comparison: Agent evaluation allows users to compare model versions (e.g., 'blue' vs. 'pink' results) across specific metrics using spider charts and column charts to quantify performance improvements.

## Practical implications

- Build engineers can establish a complete MLOps loop for AI agents: build (Serverless Inference) $\rightarrow$ test/monitor (Arena Signals) $\rightarrow$ improve (Serverless RL) $\rightarrow$ deploy (Model URI).
- The platform minimizes infrastructure overhead by providing elastic GPU scaling for resource-intensive tasks like Reinforcement Learning.
- The ability to run 'before and after' evaluations provides clear, quantifiable metrics for justifying model upgrades and performance improvements.

## Topics

AI Agents, LLMs, Reinforcement Learning, Serverless Computing, MLOps, Coreweave ARENA, Coreweave Serverless Inference, Serverless Inference docs

Source: https://www.youtube.com/watch?v=SDvIdNumJkc
