# How a Logistics Giant Keeps AI Data Locked Down

## Executive summary

This discussion provides a deep dive into FinOps for AI, detailing how a logistics giant (C.H. Robinson) manages and optimizes AI usage across complex workflows. Key strategies include tagging all AI resources for ownership tracking, utilizing platforms like Azure OpenAI and Vertex AI to maintain data locality, and implementing advanced observability metrics. The speaker emphasizes moving beyond simple total spend to measure efficiency using metrics like 'cost per thought,' 'reasoning ratio,' and 'cache hit rate' to accurately assess ROI.

## Key takeaways

- Resource Tagging is Foundational: The first step in managing AI spend is tagging every AI resource (at the model or workload level) to establish clear ownership and accountability.
- Value Over Spend: When reporting AI value, focusing solely on 'spend' is insufficient. It is more valuable to calculate ROI by measuring 'cost per hour saved' or 'cost per order,' demonstrating efficiency gains.
- Advanced Observability Metrics: To accurately track AI value, advanced metrics are necessary. These include 'prompt bloat' (excessive token usage), 'context starvation' (insufficient context leading to retries), 'reasoning ratio,' and 'cache hit rate.'

## Technical details

- AI Use Cases & Platforms: C.H. Robinson uses AI for end-to-end logistics processes, including order entry, pricing, quoting, booking, and tracking. They utilize Azure OpenAI and conduct testing with GCP's Vertex AI, prioritizing platforms that keep data locked down within their environment.
- Observability and Cost Metrics: The speaker uses an LLM observability platform to find metrics like 'verbosity' (how chatty an API call is) and tracks token consumption across full traces. This allows for identifying issues like prompt bloat (excessive tokens) or context starvation (insufficient context leading to retries).
- Efficiency Metrics: For complex, non-discrete tasks, the speaker proposes 'cost per thought' (total cost divided by the trace, considering RAGs and conversation) as an efficiency metric. Other advanced metrics include 'reasoning ratio' (reasoning tokens vs. non-reasoning tokens) and 'cache hit rate' (percentage of prompts utilizing a cache).

## Practical implications

- Implement granular tagging of AI resources (by model or workload) to track ownership and cost.
- Shift reporting focus from total AI spend to value-based metrics (e.g., cost per order, cost per automated task).
- Monitor for efficiency issues like prompt bloat and context starvation, which often indicate underlying code bugs or workflow inefficiencies.
- For complex agentic workflows, separate baselines are needed for agentic vs. conversational workloads to accurately measure true efficiency.

## Topics

FinOps for AI, AI Observability, LLM Cost Management, Logistics Technology, Agentic Workflows, Azure OpenAI, Vertex AI, GitHub Copilot

Source: https://www.youtube.com/watch?v=oYlb1Wv6Vkk
