# The 6 Pillars of an Agentic Harness for Production — Varun Krovvidi, Resolve AI

## Executive summary

While AI excels at generating modular, single-domain code, the majority of engineering time (70%) is spent on the complex, multi-domain tasks of running and fixing production systems. The speaker outlines that scaling AI for production requires moving beyond simple model calls and implementing a robust 'agentic harness.' This harness must incorporate six critical pillars—Model Orchestration, Context Engineering, Causal Reasoning, Governed Actions, Learning Systems, and Evals—to handle real-world complexity, prevent hallucination, and ensure reliable root cause analysis.

## Key takeaways

- The Shift from Code Generation to Production Operations: Code is self-documenting, modular, and single-domain, making it easy for AI. However, production systems require managing multiple domains (code, infrastructure, telemetry, teams), which is fundamentally harder for AI to solve.
- The Failure Modes of Production Agents: As AI scales, common failure modes include anchoring bias, the model treadmill (needing continuous model updates), context window issues (over- or under-exploration), lack of causal reasoning, missing guardrails, and failure to learn across investigations.
- The Six Pillars of an Agentic Harness: A robust system requires: 1) Model Orchestration (matching the best model for the task); 2) Context Engineering (defining the precise amount of context needed); 3) Causal Reasoning (establishing a causal chain of evidence); 4) Governed Actions (defining least-privilege guardrails); 5) Learning Systems; and 6) Evals (systematic evaluation across multiple levels).

## Technical details

- Agent Architecture: Resolve AI utilizes three agent types: On-call agents (for daily fixes), Incident agents (for driving complex root cause analysis), and Ambient background agents (for monitoring/analysis).
- Causal Reasoning: AI systems must be designed to provide a causal chain of evidence (like a detective) rather than just a coherent answer, especially when diagnosing production incidents.
- Model Orchestration: This involves two layers: keeping up with the 'model treadmill' (new models) and matching the best model for the specific task (e.g., Gemini for image reasoning, OpenAI for deterministic steps).
- Context Engineering: It is not just about large context windows. It requires combining techniques (e.g., Graph RAG) and defining precise tool calls (for logs, metrics, dashboards) to prevent token waste and ensure focus.

## Practical implications

- Engineering teams must shift focus from AI-assisted code generation to building robust, multi-layered agentic systems capable of handling complex, multi-domain production incidents.
- When building AI for production, prioritize establishing a causal chain of evidence and implementing strict guardrails (least privilege access) over relying solely on model coherence.
- The concept of 'Evals' must be integrated into the architecture as a continuous process to test and calibrate the system against new models and use cases.

## Topics

AI Engineering, Site Reliability Engineering (SRE), Agentic Systems, MLOps, Root Cause Analysis, Resolve AI

Source: https://www.youtube.com/watch?v=eXA2tjRZIbY
