How a Logistics Giant Keeps AI Data Locked Down
Summary
This discussion provides a deep dive into FinOps for AI, detailing how a logistics giant (C.H. Robinson) manages and optimizes AI usage across complex workflows. Key strategies include tagging all AI resources for ownership tracking, utilizing platforms like Azure OpenAI and Vertex AI to maintain data locality, and implementing advanced observability metrics. The speaker emphasizes moving beyond simple total spend to measure efficiency using metrics like 'cost per thought,' 'reasoning ratio,' and 'cache hit rate' to accurately assess ROI.
Key takeaways
-
Resource Tagging is Foundational
10:33
The first step in managing AI spend is tagging every AI resource (at the model or workload level) to establish clear ownership and accountability.
-
Value Over Spend
12:24
When reporting AI value, focusing solely on 'spend' is insufficient. It is more valuable to calculate ROI by measuring 'cost per hour saved' or 'cost per order,' demonstrating efficiency gains.
-
Advanced Observability Metrics
7:01
To accurately track AI value, advanced metrics are necessary. These include 'prompt bloat' (excessive token usage), 'context starvation' (insufficient context leading to retries), 'reasoning ratio,' and 'cache hit rate.'
Technical details
-
AI Use Cases & Platforms
118s
C.H. Robinson uses AI for end-to-end logistics processes, including order entry, pricing, quoting, booking, and tracking. They utilize Azure OpenAI and conduct testing with GCP's Vertex AI, prioritizing platforms that keep data locked down within their environment.
-
Observability and Cost Metrics
340s
The speaker uses an LLM observability platform to find metrics like 'verbosity' (how chatty an API call is) and tracks token consumption across full traces. This allows for identifying issues like prompt bloat (excessive tokens) or context starvation (insufficient context leading to retries).
-
Efficiency Metrics
1035s
For complex, non-discrete tasks, the speaker proposes 'cost per thought' (total cost divided by the trace, considering RAGs and conversation) as an efficiency metric. Other advanced metrics include 'reasoning ratio' (reasoning tokens vs. non-reasoning tokens) and 'cache hit rate' (percentage of prompts utilizing a cache).
Mentioned resources
- Azure OpenAI
- Vertex AI
- GitHub Copilot
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.