# The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse

## Executive summary

Marc Klingen discusses the emerging reference stack for building self-improving AI agents, emphasizing the shift from manual agent refinement to automated loops. The core concept is integrating online tracing/monitoring (how users interact with the agent) with offline components (datasets, experiments, and evaluations). This allows agents to automatically propose fixes and maintain evaluation criteria, significantly reducing manual labor while maintaining human oversight for setting high-level goals and direction. The talk concludes with a demo showing a coding agent autonomously improving a changelog writer by identifying and fixing internal jargon leaks.

## Key takeaways

- The Online + Offline Agent Loop: Building agents requires integrating online tracing and monitoring (user interaction) with offline components (datasets, experiments, evals). This loop is necessary because benchmarking on inaccurate datasets or monitoring production data without offline evaluation leads to incomplete understanding.
- Layers of Improvement Loops: Agent capability is represented by multiple nested loops, ranging from the lowest level (next-token prediction) up to the highest level (human goal-setting). Automation is possible at higher levels, but humans must still review dataset amendments and proposed fix changes to prevent overfitting or misalignment.
- Automating Improvement Cycles: AI can be used to propose fixes and maintain the evaluation criteria and datasets. This includes aligning datasets with actual user behavior (e.g., in a customer support application) and proposing new evaluators based on observed error patterns.
- Data Ownership and Scalability: As agents process massive amounts of data, the data system shifts from being write-intensive (ingesting traces) to read-intensive. It is critical for teams to own their data layer to ensure long-term retention and prevent sampling limitations.

## Technical details

- Agent Architecture: The agent stack connects online tracing/monitoring with offline components (datasets, experiments, evals). This structure is necessary to close the loop between production usage and structured evaluation.
- Agent Improvement Workflow: The process involves using production traces and established eval criteria/datasets to 'hill climb' toward success. AI is primarily used to propose improvements to the agent's implementation, which are then back-tested against updated datasets.
- Self-Correction Demo: A coding agent was used to identify that the changelog writer leaked internal jargon. The agent then suggested changes to the datasets and evaluators to improve clarity and user-facing language, followed by back-testing V1 vs V2.
- Signal Integration: Agents can utilize implicit user signals (e.g., edits, approvals, or user feedback) as input to the improvement loop, allowing the system to act on continuous streams of data without manual intervention.

## Practical implications

- Teams can automate the tedious process of agent improvement by letting AI manage the loop between tracing, dataset maintenance, and evaluation.
- Focusing on owning the data layer is crucial for long-term agent development, especially given the shift to read-intensive data processing.
- The best practice involves automating the low-level, tedious tasks while retaining human involvement for setting high-level goals and defining the direction of change.

## Topics

AI Agents, LLM Development, MLOps, Observability, Data Pipelines, Langfuse, Langfuse on GitHub, Loopcraft article

Source: https://www.youtube.com/watch?v=TeErpYBUIeM
