AI Native Dev

Justin Cormack - When Tests Lie: Using Observability to Keep AI Honest - AI Native DevCon June 2026

Published 2026-07-11 · Duration 32:04

Summary

The talk explores the challenges of using AI to build large-scale, complex distributed systems, exemplified by building an AWS S3 compatible object storage system in Rust. While testing is crucial, relying solely on achieving 100% test coverage is insufficient for complex systems. The speaker emphasizes that observability, robust test articles (like external services), and a 'human-in-the-loop' approach are necessary to enforce correctness, discover edge cases, and manage issues like race conditions and flaky tests in AI-assisted development.

Download summary

Key takeaways

  1. Observability is Critical for Large Systems 17:47

    For complex distributed systems, the public API often doesn't cover all background behaviors. Observability techniques (like tracing) are necessary to infer or observe invisible behaviors that standard APIs cannot expose.

  2. Test Articles Provide Grounding 21:00

    Using an existing system, such as AWS S3, as a 'test article' provides a crucial behavioral baseline. This is more valuable than relying on documentation, which may be inaccurate.

  3. Beyond 100% Test Coverage 13:44

    Achieving 100% test coverage can lead to writing trivial or unhelpful tests. The focus should instead be on expanding the scope of testing and thinking like a QA professional to find edge cases.

  4. Flaky Tests Must Be Fixed 22:00

    The speaker asserts that flaky tests must be fixed immediately, as AI models may incorrectly suggest ignoring them based on training data. Running repeated test suites helps identify these issues.

Technical details

  • AI-Assisted Development Scale 246s

    The speaker built a distributed system (S3 compatible object storage) with an initial codebase of 350,000 lines of Rust. The process required continuous human intervention and architectural oversight to prevent the project from descending into chaos.

  • Testing Methodologies 1420s

    Advanced testing techniques include property-based testing, fuzz testing, and running repeated test suites (overnight runs) to find rare errors and race conditions. The speaker also noted that type systems can enforce security checks (e.g., an 'authorized request' type), reducing the need for explicit tests.

  • Debugging with Tracing 1590s

    Implementing a hand-built tracing framework, even without connecting it to production, is highly useful. It allows AI to reproduce rare bugs or errors by providing concrete traces rather than just descriptions.

  • Security Review 1740s

    Using tools like CodeX security and conducting manual code reviews (state-of-the-code review) in conjunction with AI findings proved valuable for identifying major, missed issues.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.