# Optimize capacity, control cost, and scale economically

## Executive summary

This session details strategies for optimizing AI compute capacity, controlling costs, and scaling infrastructure economically. CoreWeave provides a suite of tools—including Flex Reservations, Spot Instances, and the Mission Control platform—to help teams manage diverse AI workloads (training, fine-tuning, inference). Key focus areas include achieving granular visibility into usage via the Billing Insights UI and the Focus API, and leveraging AI-driven insights to identify underutilized resources and forecast future spending.

## Key takeaways

- Optimizing Capacity with Tiered Plans: Instead of relying on single capacity models, customers should bundle plans: Reserved Instances for 24/7 guaranteed capacity; Flex Reservations for peak needs with pay-as-you-go savings during idle periods; Spot Instances for stateless, fault-tolerant, one-off jobs; and On-Demand for unpredictable needs. This maximizes ROI across the AI lifecycle.
- Achieving Usage Visibility with Focus API: The Focus API provides programmatic, standardized access to billing data using the PHOPS Foundation schema (supporting versions 1.2 and 2). This allows external tools to ingest usage details, grouped by resource type or capacity plan, and even request output in CSV format for accounting teams.
- Actionable Insights via Mission Control: Mission Control combines Focus API billing data with real-time fleet telemetry. Users can ask natural language questions (e.g., 'Am I using my 96 H100s all month?') to identify overages, pinpoint underutilized nodes (e.g., 72% GPU utilization), recommend capacity plan adjustments (e.g., shifting to Flex Reservations), and forecast future costs.

## Technical details

- Flex Reservations Mechanics: Flex Reservations provide guaranteed capacity for peak needs while allowing cost savings during idle time. Billing involves a continuous 'holding fee' (like reserved instances) plus a variable 'usage rate' (pay-as-you-go) that only incurs charges when nodes are actively used.
- Spot Instances and Preemption: Spot Instances are pay-as-you-go, non-committal capacity that can be reclaimed by the provider. They are optimized for AI workloads by offering a 7-minute preemption window, giving users time to save state and minimize lost work. They are best suited for stateless, fault-tolerant jobs.
- Focus API and PHOPS Standard: The Focus API provides a standardized, consumable format for billing information, adhering to the PHOPS Foundation schema. It supports detailed grouping by resource type and capacity plan, ensuring data consistency across multiple platforms.

## Practical implications

- For build engineers managing CI/CD or ML pipelines, leveraging Spot Instances for stateless, fault-tolerant jobs can drastically reduce compute costs.
- Using Mission Control to analyze utilization gaps (e.g., identifying nodes with only 72% GPU utilization) allows teams to optimize resource allocation and prevent paying for idle capacity.
- The Focus API enables integration of billing data into custom financial or resource management dashboards, providing a single source of truth for cost tracking across multiple cloud services.

## Topics

AI Compute, Cost Optimization, Cloud Infrastructure, Capacity Planning, MLOps, CoreWeave, Reserved Instances, On-Demand Instances, Flex Reservations, Spot Instances, Capacity Finder, Billing Insights UI, Focus API, Mission Control

Source: https://www.youtube.com/watch?v=fF3mQiy1yjA
