Topic

BeyondWeb

All digests tagged BeyondWeb

Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI thumbnail

· 20:32

Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI

This talk details the advanced engineering lessons required to scale synthetic data generation for massive AI pre-training runs, specifically achieving 12 trillion synthetic tokens across web, math, and code. The speaker outlines the transition from siloed research setups (Slurm) to a unified, scalable cloud infrastructure using Ray, KubeRay, and vLLM on EKS/HyperPod. Key optimizations addressed include batching S3 metadata fetches, implementing checkpointing for GPU failure recovery, achieving atomic cross-cluster scheduling of CPU/GPU resources, and optimizing VLM inference parameters for significant throughput gains.

Key takeaways

  1. Synthetic Data Recipe (BeyondWeb) 6:51

    The BeyondWeb recipe generates synthetic data by identifying high-quality data points in existing datasets and rephrasing or restructuring them, which can allow models to achieve performance using significantly fewer tokens (e.g., matching performance at 3B parameters using data equivalent to 8B parameters).

  2. S3 Metadata Bottleneck Solution 15:40

    Fetching metadata for trillions of tokens from S3 was initially a multi-day process (9–11 days) due to rate limits and slow individual requests. The solution involves batching metadata requests using S3 APIs, reducing the processing time to approximately two hours.

  3. GPU Failure Resilience 20:29

    To prevent lost compute time from GPU instability, the pipeline must implement checkpointing and/or right-size partitions. This allows for recoverable retries, minimizing downtime from multi-day job failures.

  4. Atomic Cross-Cluster Orchestration

    Scaling requires orchestrating jobs across multiple, disparate clusters (e.g., dedicated Spark clusters and H100 GPU clusters). The solution involves dedicated resource pools and scheduling both CPU and GPU requirements atomically to prevent scheduling bottlenecks.

  5. Inference Throughput Optimization

    Significant throughput gains (up to 40%) can be achieved by rigorously benchmarking and optimizing VLM parameters (e.g., batch size, speculative decoding) during batch inference, rather than relying on standard online inference.

Watch on YouTube Full article