# Stop Overpaying for Intelligence | DevDay 2026

## Executive summary

This session outlines advanced strategies for optimizing AI operational costs, shifting the focus from merely minimizing input/output tokens to maximizing 'cost per task.' Key recommendations include adopting 'cost per task' as the primary metric, systematically right-sizing models using Pareto frontier analysis, and leveraging advanced API features like Prompt Caching, Programmatic Tool Calling, and Batch APIs to drastically reduce overhead while maintaining quality.

## Key takeaways

- Shift Metric from Tokens to Tasks: The most critical metric for cost optimization is 'cost per task,' not 'cost per token.' A cheaper model may fail to complete the task, requiring human intervention, which drastically increases the true cost (time + tokens).
- Model Selection via Pareto Frontier: To right-size a model, define the task and the required accuracy threshold. Plotting the Pareto curve (Accuracy vs. Cost) helps identify the optimal model configuration and harness that meets the required accuracy at the lowest possible cost.
- Advanced Cost Levers Beyond the Model: Beyond model choice, developers can optimize costs using four levers: Prompt Caching (reusing pre-processing of shared instructions), Programmatic Tool Calling (offloading data processing to code rather than passing all results to the LLM), Reasoning Effort (starting with lower levels for routing tasks), and Batch/Flex API processing (for non-urgent, high-volume tasks).
- Optimizing Prompt Cache Hit Rate: To maximize prompt caching efficiency, keep the beginning of the prompt consistent. For greater control, utilize explicit breakpoints or 'explicit only mode' when using newer model families (e.g., GPT-4o/GPT-5.6 families) to define precise cache boundaries.

## Technical details

- Cost Metric Definition: The true cost of an AI application must account for human intervention time and the cost of tokens, making 'cost per task' the superior metric over simple token counts.
- Model Right-Sizing: The process involves defining the task and expected accuracy, then using the Pareto curve concept to find the model configuration that achieves the required accuracy threshold at the minimum cost.
- Programmatic Tool Calling: This technique allows the model to run a small program that coordinates tool calls and processes results, eliminating the need to send all intermediary data back to the LLM. This significantly reduces the number of tokens processed by the model.
- Prompt Caching Optimization: Consistency in the initial portion of the prompt is key for cache reuse. Newer model families offer explicit breakpoints and 'explicit only mode' for granular control over which parts of the prompt are cached.

## Practical implications

- Always measure and optimize for 'cost per task' rather than just token count.
- Use the Pareto frontier approach to systematically test and select the most cost-effective model for a given accuracy requirement.
- Implement programmatic tool calling to offload data processing and context management from the LLM, reducing context window overhead.
- For non-urgent, high-volume workflows, utilize the Batch API for significant cost savings (up to 50% lower token pricing).

## Topics

AI Cost Optimization, LLM Architecture, Prompt Engineering, Programmatic Tool Calling, API Design, Model Evaluation, GPT-4o / GPT-5.6 Families, Luna, Astra, Batch API, Diagnostics API

Source: https://www.youtube.com/watch?v=_nmlHbSB8kM
