The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse
Summary
Marc Klingen discusses the emerging reference stack for building self-improving AI agents, emphasizing the shift from manual agent refinement to automated loops. The core concept is integrating online tracing/monitoring (how users interact with the agent) with offline components (datasets, experiments, and evaluations). This allows agents to automatically propose fixes and maintain evaluation criteria, significantly reducing manual labor while maintaining human oversight for setting high-level goals and direction. The talk concludes with a demo showing a coding agent autonomously improving a changelog writer by identifying and fixing internal jargon leaks.
Key takeaways
-
The Online + Offline Agent Loop
2:02
Building agents requires integrating online tracing and monitoring (user interaction) with offline components (datasets, experiments, evals). This loop is necessary because benchmarking on inaccurate datasets or monitoring production data without offline evaluation leads to incomplete understanding.
-
Layers of Improvement Loops
4:02
Agent capability is represented by multiple nested loops, ranging from the lowest level (next-token prediction) up to the highest level (human goal-setting). Automation is possible at higher levels, but humans must still review dataset amendments and proposed fix changes to prevent overfitting or misalignment.
-
Automating Improvement Cycles
7:30
AI can be used to propose fixes and maintain the evaluation criteria and datasets. This includes aligning datasets with actual user behavior (e.g., in a customer support application) and proposing new evaluators based on observed error patterns.
-
Data Ownership and Scalability
As agents process massive amounts of data, the data system shifts from being write-intensive (ingesting traces) to read-intensive. It is critical for teams to own their data layer to ensure long-term retention and prevent sampling limitations.
Technical details
-
Agent Architecture
122s
The agent stack connects online tracing/monitoring with offline components (datasets, experiments, evals). This structure is necessary to close the loop between production usage and structured evaluation.
-
Agent Improvement Workflow
380s
The process involves using production traces and established eval criteria/datasets to 'hill climb' toward success. AI is primarily used to propose improvements to the agent's implementation, which are then back-tested against updated datasets.
-
Self-Correction Demo
A coding agent was used to identify that the changelog writer leaked internal jargon. The agent then suggested changes to the datasets and evaluators to improve clarity and user-facing language, followed by back-testing V1 vs V2.
-
Signal Integration
1004s
Agents can utilize implicit user signals (e.g., edits, approvals, or user feedback) as input to the improvement loop, allowing the system to act on continuous streams of data without manual intervention.
Mentioned resources
- Langfuse
- Langfuse on GitHub
- Loopcraft article
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.