# Building Engine v2: How LangChain approaches agent improvement

## Executive summary

LangChain's LangSmith Engine is an agent designed for agent engineering, automating the complex process of building, evaluating, and improving AI agents. The presentation details the advancements in Engine v2, focusing on enhanced capabilities like detecting cross-trace metric trends, identifying inefficient agent paths, and introducing 'validated fixes' and 'red teaming.' These features aim to move agent development toward a continual learning cycle, allowing for proactive bug detection and model optimization before deployment.

## Key takeaways

- Engine's Core Functionality: Engine identifies problematic traces, clusters them by root cause, writes readable issue descriptions, and proposes fixes and ready-to-use evaluation (evals) datasets to prevent regressions. Since its launch, it has scanned over 70 million traces and created over 21,000 issues.
- Meta Engine for Self-Monitoring: A significant improvement is running Engine on its own output (the 'meta Engine'). This allows the team to monitor for suboptimal responses or failures within Engine's own generated traces, serving as a primary source for identifying growth opportunities.
- Validated Fixes and Red Teaming (v2): Engine v2 introduces 'validated fixes,' where the proposed fix is tested against the original breaking inputs and a broader dataset on a preview deployment. Additionally, 'red teaming' proactively generates hypotheses for potential inputs that might break the agent, allowing for pre-deployment testing.
- Cross-Trace and Efficiency Analysis: Engine v2 can now detect complex issues that are not visible in a single trace, such as tracking metric trends (e.g., massive increases in tool calls) or identifying inefficient agent paths (side quests) that consume resources unnecessarily.
- Future Vision (Engine v3): Future plans include expanding validated fixes and red teaming to different deployment stacks, creating a baseline set of evals for general regression testing, and providing verified model recommendations (e.g., confirming a cheaper, smaller model performs better against specific evals).

## Technical details

- Agent Benchmarking: To ensure performance, Engine requires unique benchmarks for each task (e.g., identifying offending traces vs. clustering by root cause). The process involves generating benchmarks with seeded problems and dogfooding the tool on internal agents (go-to-market and coding agents) to ensure benchmarks are faithful to production reality.
- Model Selection Strategy: The team found that for specific, 'dumb' tasks, using smaller, non-frontier models can be more effective and cheaper than relying on state-of-the-art, high-cost models.
- Regression Testing Flow: The validated fix process involves: 1) Confirming the bug using production inputs; 2) Proposing a fix in a preview branch; 3) Running the breaking inputs against the fix; and 4) Running broader evals against the LangSmith dataset to ensure no regression occurs.
- Issue Identification Scope: Engine has evolved from finding visible single-trace errors (e.g., hallucination, broken tool calls) to identifying complex, cross-trace issues, such as subtle changes in resource consumption or tool call frequency over time.

## Practical implications

- Engine automates the traditionally manual and challenging agent development lifecycle, from root cause analysis to fix deployment.
- Engine provides proactive safety nets (Red Teaming) that allow teams to test agents against potential failure modes before any user interaction.
- The ability to validate fixes and run evals on a preview deployment significantly reduces deployment risk and accelerates the path to production.
- Engine helps optimize operational costs by recommending model swaps based on performance and cost metrics against defined evals.

## Topics

Agent Engineering, LLM Evaluation, MLOps, Automated Testing, LangChain, System Design, LangSmith Engine, LangSmith Engine docs, LangSmith Deployment

Source: https://www.youtube.com/watch?v=S3l7U7gOsyY
