AI Engineer

State of Data — Sean Cai, Independent / State of Data

Published 2026-07-26 · Duration 18:22

Summary

The data market is undergoing a structural shift from relying on sheer quantity of annotated images (the 'least interesting part') to capturing high-quality, process-based expertise. Data's value lies in the trajectory and reasoning trace—not just the final output. The speaker argues that while model improvement requires balancing compute, data, and talent, data remains the most underfunded leg. Successful companies must pivot from being mere 'data businesses' to becoming infrastructure providers (neo-labs) that build robust pipelines into real-world work.

Download summary

Key takeaways

  1. Data Shift: From State to Process 2:08

    The most valuable data is process-based data—the reasoning trace or sequence of decisions, rather than state-based data (e.g., rows in an ERP). Type one data (pure capture of real workflows like GitHub commits) offers superior realism compared to type two data (contrived examples manufactured by experts).

  2. The Importance of Verifiability 5:50

    A task's ease of training is proportional to its verifiability, which depends on three axes: asymmetry of verification (decomposability into checkable steps), veracity of verification (consensus on what 'correct' means), and proliferation of verification (availability of fresh examples). Coding scored highly because it solved all three.

  3. The Builder's Moat is the Pipeline 13:10

    For data companies, the durable value accrues to the services and application layer of actual work. The true moat for builders is not the raw data itself, but the pipeline into real-world work and the infrastructure required to keep retraining as models improve.

Technical details

  • Data Types 128s

    Distinguishes between state-based data (final output, like ERP rows) and process-based data (the trajectory/reasoning trace). Further differentiates Type One data (pure capture of real workflows, e.g., GitHub commits) from Type Two data (contrived examples manufactured in an arbitrary setting).

  • Model Improvement Inputs 205s

    Model performance is a function of three inputs: Compute, Data, and Talent. The speaker notes that data is often the underfunded leg, creating an imbalance penalty parameter opportunity.

  • Verification Metrics 350s

    Verifiability is assessed using three axes: Asymmetry of verification (decomposability), Veracity of verification (consensus on correctness), and Proliferation of verification (frequency of real-world examples).

  • Infrastructure Needs 750s

    Future data companies must build abstraction layers for: 1) Serving and routing small models; 2) Managing RL datasets across base model migrations (automatic post-training rerun); and 3) Implementing 'Antikythera mechanisms' to translate messy business context into Evals.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.