# Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI

## Executive summary

This talk details the advanced engineering lessons required to scale synthetic data generation for massive AI pre-training runs, specifically achieving 12 trillion synthetic tokens across web, math, and code. The speaker outlines the transition from siloed research setups (Slurm) to a unified, scalable cloud infrastructure using Ray, KubeRay, and vLLM on EKS/HyperPod. Key optimizations addressed include batching S3 metadata fetches, implementing checkpointing for GPU failure recovery, achieving atomic cross-cluster scheduling of CPU/GPU resources, and optimizing VLM inference parameters for significant throughput gains.

## Key takeaways

- Synthetic Data Recipe (BeyondWeb): The BeyondWeb recipe generates synthetic data by identifying high-quality data points in existing datasets and rephrasing or restructuring them, which can allow models to achieve performance using significantly fewer tokens (e.g., matching performance at 3B parameters using data equivalent to 8B parameters).
- S3 Metadata Bottleneck Solution: Fetching metadata for trillions of tokens from S3 was initially a multi-day process (9–11 days) due to rate limits and slow individual requests. The solution involves batching metadata requests using S3 APIs, reducing the processing time to approximately two hours.
- GPU Failure Resilience: To prevent lost compute time from GPU instability, the pipeline must implement checkpointing and/or right-size partitions. This allows for recoverable retries, minimizing downtime from multi-day job failures.
- Atomic Cross-Cluster Orchestration: Scaling requires orchestrating jobs across multiple, disparate clusters (e.g., dedicated Spark clusters and H100 GPU clusters). The solution involves dedicated resource pools and scheduling both CPU and GPU requirements atomically to prevent scheduling bottlenecks.
- Inference Throughput Optimization: Significant throughput gains (up to 40%) can be achieved by rigorously benchmarking and optimizing VLM parameters (e.g., batch size, speculative decoding) during batch inference, rather than relying on standard online inference.

## Technical details

- Data Scaling and Need: Due to scaling laws, pre-training requires exponentially more data. Since the web has a limited amount of usable text (estimated at 30 trillion tokens), high-quality synthetic data is necessary to supplement training sets.
- System Architecture Evolution: The system evolved from a slow, error-prone two-track system (Slurm for research, Kubernetes for product) to a unified stack running on EKS/HyperPod. The core components are Ray, KubeRay, and vLLM, orchestrated within a single Kubernetes environment.
- Pipeline Workflow: The end-to-end workflow involves distinct stages: curate (Spark jobs), synthesize (Ray jobs), train, and evaluate. These stages are orchestrated together in a single, cohesive pipeline.
- Resource Management: The architecture utilizes dedicated resource pools within Kubernetes to ensure atomic scheduling of both CPU and GPU requirements, which is critical for efficient cross-cluster job execution.

## Practical implications

- For build engineers managing large-scale ML pipelines, adopting a unified orchestration layer (like Ray/KubeRay) across diverse hardware (H100, general GPUs) is crucial for maximizing resource utilization.
- Implementing robust data governance and metadata management (e.g., batching S3 fetches) is necessary to prevent I/O bottlenecks from stalling expensive compute jobs.
- Pipeline design must incorporate checkpointing and failure recovery mechanisms to handle inevitable hardware failures in long-running, multi-day training jobs.
- Resource scheduling must move beyond single-resource allocation, treating CPU and GPU requirements as atomic units for optimal cross-cluster efficiency.

## Topics

AI Engineering, Data Scaling, Distributed Computing, Cloud Infrastructure, MLOps, DatologyAI, BeyondWeb, Ray, KubeRay, vLLM, EKS on HyperPod, H100, S3

Source: https://www.youtube.com/watch?v=FQwTqUmcbRg
