# Jacob Lauritzen: Token Maxing Is the New Lines of Code

## Executive summary

Jacob Lauritzen, CTO of Legora, discusses how AI coding agents are transforming legal tech and software development. He details Legora's engineering practices, including running multiple concurrent agents and building custom orchestration tools to manage shared resources (e.g., Postgres, Redis). The discussion emphasizes that while AI can automate menial tasks and handle verifiable processes (like math or coding tests), human judgment remains critical for tasks lacking objective truth (e.g., litigation strategy). Furthermore, the conversation critiques 'token maxing' as a poor metric for AI adoption, advocating instead for measuring high output and the efficiency of the overall system.

## Key takeaways

- Scaling Agentic Workflows: Legora runs full agentic loops that don't involve human developers, such as a bug report in Slack triggering a cloud agent to reproduce the bug, fix it, write a regression test, and open a PR. This process is then reviewed by a human engineer.
- The Pitfalls of Token Maxing: Measuring AI adoption by tokens burned is ineffective. The focus should shift to setting high expectations for output and measuring the speed and quality of results, which is considered the 'new lines of code.'
- Decision Models vs. LLMs: Decision models (like Jev) are highly valuable because they provide an answer with a confidence score, potentially replacing large amounts of decision-making and workflow that do not require full autoregressive LLM capabilities.
- The Role of Human Judgment: Human judgment is irreplaceable when a problem lacks an objective truth (e.g., determining the best litigation strategy among multiple options).

## Technical details

- Agent Orchestration & Parallelization: The team built custom tooling to orchestrate many agents, including a Kanban board for agents and a custom tool to share single instances of dependencies (Postgres, Redis, observability stack) across multiple workspaces, allowing for efficient local development.
- System Evaluation and Testing: To ensure reliability, Legora runs weekly model benchmarks and harness evaluations (regression testing). They must test not only models but also the 'harness'—the task decomposition, compaction, and search logic—to prevent regressions.
- Model Decomposition and Efficiency: To achieve 'efficient AI,' tasks are decomposed into different sub-agents and subcomponents, allowing the use of different models for different tasks, optimizing for Pareto efficiency (speed and cost).
- Data Verification and Auditability: For critical applications, the system must be auditable. Legora uses multi-hop citations, allowing users to trace the source of any output down to the original document or source document.

## Practical implications

- Focus on building robust, auditable systems that track the chain of reasoning (multi-hop citations) rather than just the final output.
- Treat AI enablement as a core organizational investment, aiming to make all engineers 5-10% more effective.
- Design interfaces that move beyond simple chat threads, utilizing structured output (e.g., spreadsheet-like views) for complex analysis.
- When building agentic systems, prioritize the ability to decompose tasks and select the optimal model/strategy for each sub-task to maximize efficiency (Pareto frontier).

## Topics

AI Coding Agents, Agentic Workflows, LLM Architecture, Software Development Lifecycle, Legal Tech, Legora, Tessl

Source: https://www.youtube.com/watch?v=mbdDv1qwLdk
