AI Engineer

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Published 2026-07-31 · Duration 19:05

Summary

In an era of increasing compute scarcity—evidenced by rising H100 prices and skyrocketing token usage—data quality has emerged as the critical 'compute multiplier' for model training. The presentation outlines a systematic approach to data enhancement through four stages: Clean, Curate, Create, and Compose. By maximizing the signal per token (marginal information gain), organizations can achieve performance levels comparable to models trained with vastly more compute budgets. Practical applications include improving Vision Language Models (VLMs) and enhancing multilingual capabilities using proprietary or public datasets.

Download summary

Key takeaways

  1. Compute Scarcity Drives Data Focus

    The availability of compute is becoming increasingly constrained, leading to market actions like Google capping Meta's Gemini usage and OpenAI selling token futures. This necessitates a shift in focus from raw compute power to data quality.

  2. Data Quality as Compute Multiplier 3:39

    Improving data quality allows for dramatically better performance (blue curve) compared to training with the same limited compute budget (gray curve), effectively simulating much larger compute investments.

  3. The Four C's of Data Enhancement 5:48

    Data improvement is achieved through a pipeline: Clean (heuristic filters, decontamination), Curate (quality classifiers, redundancy reduction), Create (synthetic data generation/rephrasing), and Compose (sequencing across multiple training stages).

  4. Cross-Lingual Benefits from Curation 15:24

    Curating English data can positively benefit non-English performance, demonstrating cross-lingual transfer. Similarly, curating non-English data benefits English performance.

Technical details

  • Data Optimization Goal 280s

    The fundamental goal is to maximize the marginal information gain per data point shown to the model, ensuring relevance to specific target use cases.

  • VLM Performance Gains 690s

    Curating a dataset (e.g., using the Mammoth dataset) can yield an absolute 14 percentage point improvement in VLM performance while requiring significantly less training compute compared to public benchmarks like Quen 3.5b.

  • Multilingual Scaling

    Curating data is critical for non-Western use cases, allowing models to achieve strong multilingual MMLU performance using only a small fraction of the total tokens (e.g., 8% of data used as multilingual tokens).

  • Synthetic Data Generation

    The technique of 'rephrasing' converts proprietary documents into multiple formats (e.g., true/false questions), increasing diversity and allowing training on high-quality, structured data points without model collapse.

Mentioned resources

  • DatologyAI (Company)
  • H100 (Hardware/Compute Unit)
  • Gemini (Model/API)
  • Meta's Gemini usage cap (Industry Constraint)
  • Quen 3.5b (Model Benchmark)
  • Thomson Reuters (Customer/Domain)
  • RCI Trendy Large (Model/Client Project)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.