Topic

LangSmith Engine docs

All digests tagged LangSmith Engine docs

Building Engine v2: How LangChain approaches agent improvement thumbnail

· 20:19

Building Engine v2: How LangChain approaches agent improvement

LangChain's LangSmith Engine is an agent designed for agent engineering, automating the complex process of building, evaluating, and improving AI agents. The presentation details the advancements in Engine v2, focusing on enhanced capabilities like detecting cross-trace metric trends, identifying inefficient agent paths, and introducing 'validated fixes' and 'red teaming.' These features aim to move agent development toward a continual learning cycle, allowing for proactive bug detection and model optimization before deployment.

Key takeaways

  1. Engine's Core Functionality

    Engine identifies problematic traces, clusters them by root cause, writes readable issue descriptions, and proposes fixes and ready-to-use evaluation (evals) datasets to prevent regressions. Since its launch, it has scanned over 70 million traces and created over 21,000 issues.

  2. Meta Engine for Self-Monitoring 9:17

    A significant improvement is running Engine on its own output (the 'meta Engine'). This allows the team to monitor for suboptimal responses or failures within Engine's own generated traces, serving as a primary source for identifying growth opportunities.

  3. Validated Fixes and Red Teaming (v2)

    Engine v2 introduces 'validated fixes,' where the proposed fix is tested against the original breaking inputs and a broader dataset on a preview deployment. Additionally, 'red teaming' proactively generates hypotheses for potential inputs that might break the agent, allowing for pre-deployment testing.

  4. Cross-Trace and Efficiency Analysis

    Engine v2 can now detect complex issues that are not visible in a single trace, such as tracking metric trends (e.g., massive increases in tool calls) or identifying inefficient agent paths (side quests) that consume resources unnecessarily.

  5. Future Vision (Engine v3)

    Future plans include expanding validated fixes and red teaming to different deployment stacks, creating a baseline set of evals for general regression testing, and providing verified model recommendations (e.g., confirming a cheaper, smaller model performs better against specific evals).

Watch on YouTube Full article