AI Engineer

Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI

Published 2026-10-02 · Duration 20:32

Summary

This talk details the advanced engineering lessons required to scale synthetic data generation for massive AI pre-training runs, specifically achieving 12 trillion synthetic tokens across web, math, and code. The speaker outlines the transition from siloed research setups (Slurm) to a unified, scalable cloud infrastructure using Ray, KubeRay, and vLLM on EKS/HyperPod. Key optimizations addressed include batching S3 metadata fetches, implementing checkpointing for GPU failure recovery, achieving atomic cross-cluster scheduling of CPU/GPU resources, and optimizing VLM inference parameters for significant throughput gains.

Download summary

Key takeaways

  1. Synthetic Data Recipe (BeyondWeb) 6:51

    The BeyondWeb recipe generates synthetic data by identifying high-quality data points in existing datasets and rephrasing or restructuring them, which can allow models to achieve performance using significantly fewer tokens (e.g., matching performance at 3B parameters using data equivalent to 8B parameters).

  2. S3 Metadata Bottleneck Solution 15:40

    Fetching metadata for trillions of tokens from S3 was initially a multi-day process (9–11 days) due to rate limits and slow individual requests. The solution involves batching metadata requests using S3 APIs, reducing the processing time to approximately two hours.

  3. GPU Failure Resilience 20:29

    To prevent lost compute time from GPU instability, the pipeline must implement checkpointing and/or right-size partitions. This allows for recoverable retries, minimizing downtime from multi-day job failures.

  4. Atomic Cross-Cluster Orchestration

    Scaling requires orchestrating jobs across multiple, disparate clusters (e.g., dedicated Spark clusters and H100 GPU clusters). The solution involves dedicated resource pools and scheduling both CPU and GPU requirements atomically to prevent scheduling bottlenecks.

  5. Inference Throughput Optimization

    Significant throughput gains (up to 40%) can be achieved by rigorously benchmarking and optimizing VLM parameters (e.g., batch size, speculative decoding) during batch inference, rather than relying on standard online inference.

Technical details

  • Data Scaling and Need 12s

    Due to scaling laws, pre-training requires exponentially more data. Since the web has a limited amount of usable text (estimated at 30 trillion tokens), high-quality synthetic data is necessary to supplement training sets.

  • System Architecture Evolution 551s

    The system evolved from a slow, error-prone two-track system (Slurm for research, Kubernetes for product) to a unified stack running on EKS/HyperPod. The core components are Ray, KubeRay, and vLLM, orchestrated within a single Kubernetes environment.

  • Pipeline Workflow 840s

    The end-to-end workflow involves distinct stages: curate (Spark jobs), synthesize (Ray jobs), train, and evaluate. These stages are orchestrated together in a single, cohesive pipeline.

  • Resource Management

    The architecture utilizes dedicated resource pools within Kubernetes to ensure atomic scheduling of both CPU and GPU requirements, which is critical for efficient cross-cluster job execution.

Mentioned resources

  • DatologyAI (Company)
  • BeyondWeb (Synthetic Data Recipe)
  • Ray, KubeRay, vLLM (ML/Orchestration Frameworks)
  • EKS on HyperPod (Cloud Infrastructure)
  • H100 (GPU Hardware)
  • S3 (Storage Layer)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.