Topic

Reserved Instances

All digests tagged Reserved Instances

Optimize capacity, control cost, and scale economically thumbnail

· 40:53

Optimize capacity, control cost, and scale economically

This session details strategies for optimizing AI compute capacity, controlling costs, and scaling infrastructure economically. CoreWeave provides a suite of tools—including Flex Reservations, Spot Instances, and the Mission Control platform—to help teams manage diverse AI workloads (training, fine-tuning, inference). Key focus areas include achieving granular visibility into usage via the Billing Insights UI and the Focus API, and leveraging AI-driven insights to identify underutilized resources and forecast future spending.

Key takeaways

  1. Optimizing Capacity with Tiered Plans 4:30

    Instead of relying on single capacity models, customers should bundle plans: Reserved Instances for 24/7 guaranteed capacity; Flex Reservations for peak needs with pay-as-you-go savings during idle periods; Spot Instances for stateless, fault-tolerant, one-off jobs; and On-Demand for unpredictable needs. This maximizes ROI across the AI lifecycle.

  2. Achieving Usage Visibility with Focus API 19:10

    The Focus API provides programmatic, standardized access to billing data using the PHOPS Foundation schema (supporting versions 1.2 and 2). This allows external tools to ingest usage details, grouped by resource type or capacity plan, and even request output in CSV format for accounting teams.

  3. Actionable Insights via Mission Control 24:10

    Mission Control combines Focus API billing data with real-time fleet telemetry. Users can ask natural language questions (e.g., 'Am I using my 96 H100s all month?') to identify overages, pinpoint underutilized nodes (e.g., 72% GPU utilization), recommend capacity plan adjustments (e.g., shifting to Flex Reservations), and forecast future costs.

Watch on YouTube Full article