Jacob Lauritzen: Token Maxing Is the New Lines of Code
Summary
Jacob Lauritzen, CTO of Legora, discusses how AI coding agents are transforming legal tech and software development. He details Legora's engineering practices, including running multiple concurrent agents and building custom orchestration tools to manage shared resources (e.g., Postgres, Redis). The discussion emphasizes that while AI can automate menial tasks and handle verifiable processes (like math or coding tests), human judgment remains critical for tasks lacking objective truth (e.g., litigation strategy). Furthermore, the conversation critiques 'token maxing' as a poor metric for AI adoption, advocating instead for measuring high output and the efficiency of the overall system.
Key takeaways
-
Scaling Agentic Workflows
5:34
Legora runs full agentic loops that don't involve human developers, such as a bug report in Slack triggering a cloud agent to reproduce the bug, fix it, write a regression test, and open a PR. This process is then reviewed by a human engineer.
-
The Pitfalls of Token Maxing
16:50
Measuring AI adoption by tokens burned is ineffective. The focus should shift to setting high expectations for output and measuring the speed and quality of results, which is considered the 'new lines of code.'
-
Decision Models vs. LLMs
23:33
Decision models (like Jev) are highly valuable because they provide an answer with a confidence score, potentially replacing large amounts of decision-making and workflow that do not require full autoregressive LLM capabilities.
-
The Role of Human Judgment
33:00
Human judgment is irreplaceable when a problem lacks an objective truth (e.g., determining the best litigation strategy among multiple options).
Technical details
-
Agent Orchestration & Parallelization
380s
The team built custom tooling to orchestrate many agents, including a Kanban board for agents and a custom tool to share single instances of dependencies (Postgres, Redis, observability stack) across multiple workspaces, allowing for efficient local development.
-
System Evaluation and Testing
2422s
To ensure reliability, Legora runs weekly model benchmarks and harness evaluations (regression testing). They must test not only models but also the 'harness'—the task decomposition, compaction, and search logic—to prevent regressions.
-
Model Decomposition and Efficiency
To achieve 'efficient AI,' tasks are decomposed into different sub-agents and subcomponents, allowing the use of different models for different tasks, optimizing for Pareto efficiency (speed and cost).
-
Data Verification and Auditability
For critical applications, the system must be auditable. Legora uses multi-hop citations, allowing users to trace the source of any output down to the original document or source document.
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.