MLOps Community

How a Logistics Giant Keeps AI Data Locked Down

Published 2026-10-05 · Duration 20:35

Summary

This discussion provides a deep dive into FinOps for AI, detailing how a logistics giant (C.H. Robinson) manages and optimizes AI usage across complex workflows. Key strategies include tagging all AI resources for ownership tracking, utilizing platforms like Azure OpenAI and Vertex AI to maintain data locality, and implementing advanced observability metrics. The speaker emphasizes moving beyond simple total spend to measure efficiency using metrics like 'cost per thought,' 'reasoning ratio,' and 'cache hit rate' to accurately assess ROI.

Download summary

Key takeaways

  1. Resource Tagging is Foundational 10:33

    The first step in managing AI spend is tagging every AI resource (at the model or workload level) to establish clear ownership and accountability.

  2. Value Over Spend 12:24

    When reporting AI value, focusing solely on 'spend' is insufficient. It is more valuable to calculate ROI by measuring 'cost per hour saved' or 'cost per order,' demonstrating efficiency gains.

  3. Advanced Observability Metrics 7:01

    To accurately track AI value, advanced metrics are necessary. These include 'prompt bloat' (excessive token usage), 'context starvation' (insufficient context leading to retries), 'reasoning ratio,' and 'cache hit rate.'

Technical details

  • AI Use Cases & Platforms 118s

    C.H. Robinson uses AI for end-to-end logistics processes, including order entry, pricing, quoting, booking, and tracking. They utilize Azure OpenAI and conduct testing with GCP's Vertex AI, prioritizing platforms that keep data locked down within their environment.

  • Observability and Cost Metrics 340s

    The speaker uses an LLM observability platform to find metrics like 'verbosity' (how chatty an API call is) and tracks token consumption across full traces. This allows for identifying issues like prompt bloat (excessive tokens) or context starvation (insufficient context leading to retries).

  • Efficiency Metrics 1035s

    For complex, non-discrete tasks, the speaker proposes 'cost per thought' (total cost divided by the trace, considering RAGs and conversation) as an efficiency metric. Other advanced metrics include 'reasoning ratio' (reasoning tokens vs. non-reasoning tokens) and 'cache hit rate' (percentage of prompts utilizing a cache).

Mentioned resources

  • Azure OpenAI (AI Platform)
  • Vertex AI (AI Platform)
  • GitHub Copilot (Developer Tool)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.