Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI
Summary
This talk details the advanced engineering lessons required to scale synthetic data generation for massive AI pre-training runs, specifically achieving 12 trillion synthetic tokens across web, math, and code. The speaker outlines the transition from siloed research setups (Slurm) to a unified, scalable cloud infrastructure using Ray, KubeRay, and vLLM on EKS/HyperPod. Key optimizations addressed include batching S3 metadata fetches, implementing checkpointing for GPU failure recovery, achieving atomic cross-cluster scheduling of CPU/GPU resources, and optimizing VLM inference parameters for significant throughput gains.
Key takeaways
-
Synthetic Data Recipe (BeyondWeb)
6:51
The BeyondWeb recipe generates synthetic data by identifying high-quality data points in existing datasets and rephrasing or restructuring them, which can allow models to achieve performance using significantly fewer tokens (e.g., matching performance at 3B parameters using data equivalent to 8B parameters).
-
S3 Metadata Bottleneck Solution
15:40
Fetching metadata for trillions of tokens from S3 was initially a multi-day process (9–11 days) due to rate limits and slow individual requests. The solution involves batching metadata requests using S3 APIs, reducing the processing time to approximately two hours.
-
GPU Failure Resilience
20:29
To prevent lost compute time from GPU instability, the pipeline must implement checkpointing and/or right-size partitions. This allows for recoverable retries, minimizing downtime from multi-day job failures.
-
Atomic Cross-Cluster Orchestration
Scaling requires orchestrating jobs across multiple, disparate clusters (e.g., dedicated Spark clusters and H100 GPU clusters). The solution involves dedicated resource pools and scheduling both CPU and GPU requirements atomically to prevent scheduling bottlenecks.
-
Inference Throughput Optimization
Significant throughput gains (up to 40%) can be achieved by rigorously benchmarking and optimizing VLM parameters (e.g., batch size, speculative decoding) during batch inference, rather than relying on standard online inference.
Technical details
-
Data Scaling and Need
12s
Due to scaling laws, pre-training requires exponentially more data. Since the web has a limited amount of usable text (estimated at 30 trillion tokens), high-quality synthetic data is necessary to supplement training sets.
-
System Architecture Evolution
551s
The system evolved from a slow, error-prone two-track system (Slurm for research, Kubernetes for product) to a unified stack running on EKS/HyperPod. The core components are Ray, KubeRay, and vLLM, orchestrated within a single Kubernetes environment.
-
Pipeline Workflow
840s
The end-to-end workflow involves distinct stages: curate (Spark jobs), synthesize (Ray jobs), train, and evaluate. These stages are orchestrated together in a single, cohesive pipeline.
-
Resource Management
The architecture utilizes dedicated resource pools within Kubernetes to ensure atomic scheduling of both CPU and GPU requirements, which is critical for efficient cross-cluster job execution.
Mentioned resources
- DatologyAI
- BeyondWeb
- Ray, KubeRay, vLLM
- EKS on HyperPod
- H100
- S3
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.