Topic

Langfuse on GitHub

All digests tagged Langfuse on GitHub

The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse thumbnail

· 17:07

The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse

Marc Klingen discusses the emerging reference stack for building self-improving AI agents, emphasizing the shift from manual agent refinement to automated loops. The core concept is integrating online tracing/monitoring (how users interact with the agent) with offline components (datasets, experiments, and evaluations). This allows agents to automatically propose fixes and maintain evaluation criteria, significantly reducing manual labor while maintaining human oversight for setting high-level goals and direction. The talk concludes with a demo showing a coding agent autonomously improving a changelog writer by identifying and fixing internal jargon leaks.

Key takeaways

  1. The Online + Offline Agent Loop 2:02

    Building agents requires integrating online tracing and monitoring (user interaction) with offline components (datasets, experiments, evals). This loop is necessary because benchmarking on inaccurate datasets or monitoring production data without offline evaluation leads to incomplete understanding.

  2. Layers of Improvement Loops 4:02

    Agent capability is represented by multiple nested loops, ranging from the lowest level (next-token prediction) up to the highest level (human goal-setting). Automation is possible at higher levels, but humans must still review dataset amendments and proposed fix changes to prevent overfitting or misalignment.

  3. Automating Improvement Cycles 7:30

    AI can be used to propose fixes and maintain the evaluation criteria and datasets. This includes aligning datasets with actual user behavior (e.g., in a customer support application) and proposing new evaluators based on observed error patterns.

  4. Data Ownership and Scalability

    As agents process massive amounts of data, the data system shifts from being write-intensive (ingesting traces) to read-intensive. It is critical for teams to own their data layer to ensure long-term retention and prevent sampling limitations.

Watch on YouTube Full article