Justin Cormack - When Tests Lie: Using Observability to Keep AI Honest - AI Native DevCon June 2026
Summary
The talk explores the challenges of using AI to build large-scale, complex distributed systems, exemplified by building an AWS S3 compatible object storage system in Rust. While testing is crucial, relying solely on achieving 100% test coverage is insufficient for complex systems. The speaker emphasizes that observability, robust test articles (like external services), and a 'human-in-the-loop' approach are necessary to enforce correctness, discover edge cases, and manage issues like race conditions and flaky tests in AI-assisted development.
Key takeaways
-
Observability is Critical for Large Systems
17:47
For complex distributed systems, the public API often doesn't cover all background behaviors. Observability techniques (like tracing) are necessary to infer or observe invisible behaviors that standard APIs cannot expose.
-
Test Articles Provide Grounding
21:00
Using an existing system, such as AWS S3, as a 'test article' provides a crucial behavioral baseline. This is more valuable than relying on documentation, which may be inaccurate.
-
Beyond 100% Test Coverage
13:44
Achieving 100% test coverage can lead to writing trivial or unhelpful tests. The focus should instead be on expanding the scope of testing and thinking like a QA professional to find edge cases.
-
Flaky Tests Must Be Fixed
22:00
The speaker asserts that flaky tests must be fixed immediately, as AI models may incorrectly suggest ignoring them based on training data. Running repeated test suites helps identify these issues.
Technical details
-
AI-Assisted Development Scale
246s
The speaker built a distributed system (S3 compatible object storage) with an initial codebase of 350,000 lines of Rust. The process required continuous human intervention and architectural oversight to prevent the project from descending into chaos.
-
Testing Methodologies
1420s
Advanced testing techniques include property-based testing, fuzz testing, and running repeated test suites (overnight runs) to find rare errors and race conditions. The speaker also noted that type systems can enforce security checks (e.g., an 'authorized request' type), reducing the need for explicit tests.
-
Debugging with Tracing
1590s
Implementing a hand-built tracing framework, even without connecting it to production, is highly useful. It allows AI to reproduce rare bugs or errors by providing concrete traces rather than just descriptions.
-
Security Review
1740s
Using tools like CodeX security and conducting manual code reviews (state-of-the-code review) in conjunction with AI findings proved valuable for identifying major, missed issues.
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.