Prompt engineering and serverless inference: closing the open model gap
Summary
This video demonstrates the end-to-end workflow for building, validating, and improving AI agents using Coreweave's platform. The presenter showcases Coreweave Arena for running real-world, production-like workloads. The process involves using Serverless Inference for agent functionality (e.g., building a list builder using Qwen-3-14B-Instruct) and leveraging Serverless Reinforcement Learning (RL) to improve agent reliability and performance. The workflow allows teams to track performance using signals, compare model versions, and deploy improved models via their URI and sample code.
Key takeaways
-
AI Agent Validation with Coreweave Arena
Coreweave Arena allows users to validate AI workloads (pre-training, agent building, post-training) under production-like conditions before committing to deployment.
-
Improving Agents via Serverless RL
2:02
Serverless RL is presented as a preferred method for fine-tuning Large Language Models (LLMs) for agentic tasks, improving overall AI application reliability by addressing mistakes identified during QA or production.
-
End-to-End Deployment Workflow
4:20
After an RL job improves performance, the process involves pinning the best-performing model checkpoint, retrieving the model URI, and using sample Python code to call the model via Coreweave Serverless Inference.
Technical details
-
Agent Development & Inference
70s
An agent was built using the Qwen-3-14B-Instruct model served through Coreweave's serverless inference. The agent was designed to generate distribution lists based on natural language requests.
-
Performance Monitoring
100s
Agent performance is monitored using 'signals' (e.g., low-quality response, misalignment between user request and generated SQL statement) logged in the 'Signals' tab.
-
Serverless RL Implementation
130s
Coreweave Serverless RL provides instant access to GPU capacity with elastic scaling, eliminating the need for complex provisioning and skilled developers to manage RL scripts.
-
Model Comparison
220s
Agent evaluation allows users to compare model versions (e.g., 'blue' vs. 'pink' results) across specific metrics using spider charts and column charts to quantify performance improvements.
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.