Weights & Biases

Optimize capacity, control cost, and scale economically

Published 2026-10-09 · Duration 40:53

Summary

This session details strategies for optimizing AI compute capacity, controlling costs, and scaling infrastructure economically. CoreWeave provides a suite of tools—including Flex Reservations, Spot Instances, and the Mission Control platform—to help teams manage diverse AI workloads (training, fine-tuning, inference). Key focus areas include achieving granular visibility into usage via the Billing Insights UI and the Focus API, and leveraging AI-driven insights to identify underutilized resources and forecast future spending.

Download summary

Key takeaways

  1. Optimizing Capacity with Tiered Plans 4:30

    Instead of relying on single capacity models, customers should bundle plans: Reserved Instances for 24/7 guaranteed capacity; Flex Reservations for peak needs with pay-as-you-go savings during idle periods; Spot Instances for stateless, fault-tolerant, one-off jobs; and On-Demand for unpredictable needs. This maximizes ROI across the AI lifecycle.

  2. Achieving Usage Visibility with Focus API 19:10

    The Focus API provides programmatic, standardized access to billing data using the PHOPS Foundation schema (supporting versions 1.2 and 2). This allows external tools to ingest usage details, grouped by resource type or capacity plan, and even request output in CSV format for accounting teams.

  3. Actionable Insights via Mission Control 24:10

    Mission Control combines Focus API billing data with real-time fleet telemetry. Users can ask natural language questions (e.g., 'Am I using my 96 H100s all month?') to identify overages, pinpoint underutilized nodes (e.g., 72% GPU utilization), recommend capacity plan adjustments (e.g., shifting to Flex Reservations), and forecast future costs.

Technical details

  • Flex Reservations Mechanics 320s

    Flex Reservations provide guaranteed capacity for peak needs while allowing cost savings during idle time. Billing involves a continuous 'holding fee' (like reserved instances) plus a variable 'usage rate' (pay-as-you-go) that only incurs charges when nodes are actively used.

  • Spot Instances and Preemption 420s

    Spot Instances are pay-as-you-go, non-committal capacity that can be reclaimed by the provider. They are optimized for AI workloads by offering a 7-minute preemption window, giving users time to save state and minimize lost work. They are best suited for stateless, fault-tolerant jobs.

  • Focus API and PHOPS Standard 1150s

    The Focus API provides a standardized, consumable format for billing information, adhering to the PHOPS Foundation schema. It supports detailed grouping by resource type and capacity plan, ensuring data consistency across multiple platforms.

Mentioned resources

  • CoreWeave (Cloud Compute Platform)
  • Reserved Instances (Capacity Plan)
  • On-Demand Instances (Capacity Plan)
  • Flex Reservations (Capacity Plan)
  • Spot Instances (Capacity Plan)
  • Capacity Finder (Tooling)
  • Billing Insights UI (Monitoring Tool)
  • Focus API (API/Data Standard)
  • Mission Control (AI/Analytics Tool)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.