Topic

Serverless Inference docs

All digests tagged Serverless Inference docs

Prompt engineering and serverless inference: closing the open model gap thumbnail

· 5:49

Prompt engineering and serverless inference: closing the open model gap

This video demonstrates the end-to-end workflow for building, validating, and improving AI agents using Coreweave's platform. The presenter showcases Coreweave Arena for running real-world, production-like workloads. The process involves using Serverless Inference for agent functionality (e.g., building a list builder using Qwen-3-14B-Instruct) and leveraging Serverless Reinforcement Learning (RL) to improve agent reliability and performance. The workflow allows teams to track performance using signals, compare model versions, and deploy improved models via their URI and sample code.

Key takeaways

  1. AI Agent Validation with Coreweave Arena

    Coreweave Arena allows users to validate AI workloads (pre-training, agent building, post-training) under production-like conditions before committing to deployment.

  2. Improving Agents via Serverless RL 2:02

    Serverless RL is presented as a preferred method for fine-tuning Large Language Models (LLMs) for agentic tasks, improving overall AI application reliability by addressing mistakes identified during QA or production.

  3. End-to-End Deployment Workflow 4:20

    After an RL job improves performance, the process involves pinning the best-performing model checkpoint, retrieving the model URI, and using sample Python code to call the model via Coreweave Serverless Inference.

Watch on YouTube Full article