# The Caveman Prompting Challenge

## Executive summary

The shift toward AI-driven capabilities is fundamentally changing how organizations interact with data, moving away from traditional UIs and dashboards toward API/CLI-based interactions. Implementing AI requires rigorous governance, treating agents like code (Policy as Code), and establishing clear unit metrics (North Star) to measure value beyond simple cost savings. The core challenge involves balancing model cost, speed, and accuracy while managing the complexity of exploratory (R&D) versus production workloads.

## Key takeaways

- The Shift from UI to API/CLI Interaction: The modern UI is becoming obsolete; the future involves interacting with systems via APIs, CLIs, or MCP apps, allowing an agent to pull information from multiple systems (e.g., InfoSec, FinOps, Cloud Config) simultaneously, rather than requiring manual dashboard navigation.
- Defining AI Value with Unit Metrics: To measure AI value, organizations must define a 'unit metric' or 'North Star' that aligns to a core business KPI, rather than relying on broad goals like 'driving business outcomes faster.' This allows for a quantifiable conversation: 'To achieve one unit of work, it costs us $X.'
- Governance for AI Agents: Governance must be applied to agents using proactive blocks and reactive checks, defining rules via 'Policy as Code.' This ensures that if a human cannot perform an action (e.g., creating a public bucket), an agent cannot either.

## Technical details

- Model Selection and FinOps: The FinOps Foundation AI working group addresses the trade-off between model cost, speed, and accuracy for specific workloads. The goal is to determine the optimal model for a given task, moving beyond simple model choice to holistic resource management.
- Optimization and Agent Harnessing: A common optimization strategy is to start with the smartest, most expensive model and then step down to a smaller, cheaper model (e.g., 70B parameter model) while augmenting its capability with a 'harness agent' to maintain output quality.
- Cost Attribution (R&D vs. Production): It is critical to distinguish between 'R&D tokens' (exploratory work done locally or in testing) and actual production workload costs. CFOs typically do not classify exploratory work as a production expense.

## Practical implications

- Implement Policy as Code (PaC) to govern AI agents, ensuring that all actions are auditable and constrained by defined organizational policies.
- Shift focus from measuring total cost to measuring the efficiency of the 'unit of value' (the North Star metric) that AI enables.
- Develop internal frameworks to separate and budget for exploratory (R&D) AI work from mission-critical production workloads.
- Prioritize building API/CLI access points over relying on complex, outdated web dashboards to maximize agent utility and minimize friction.

## Topics

Generative AI, FinOps, AI Governance, Software Architecture, Prompt Engineering, System Optimization, Demetrios Brinkmann, James Barney

Source: https://www.youtube.com/watch?v=2QHE75oWT2w
