Weights & Biases

40 Trillion Tokens a Day (Yes, More Than OpenAI) | Lin Qiao, CEO of Fireworks

Published 2026-08-04 · Duration 1:19:05

Summary

The discussion centers on the industry shift from general-purpose AI models (like those from OpenAI/Anthropic) toward specialized intelligence. Lin Qiao of Fireworks argues that true innovation lies in leveraging proprietary, locked-in enterprise data—the 'alpha'—to build customized models. She asserts that this specialization is necessary because generalized models cannot capture a company's unique knowledge or judgment. Technically, the conversation details advanced training methods (SFT, DPO, KTO, RL) and emphasizes platform control, noting that Fireworks achieves bitwise equivalence between training and inference results to ensure maximum quality while optimizing for cost and speed.

Download summary

Key takeaways

  1. The Rise of Specialized Intelligence 1:08:55

    Lin Qiao argues that the future belongs to specialized intelligence—customized models built on private company data—rather than general-purpose AGI. She believes every company is unique, making it difficult for a single general model to capture proprietary knowledge (41:35).

  2. Fireworks' Scale and Focus 22:16

    Fireworks claims to process over 40 trillion tokens daily, stating that 95% of this traffic comes from customized model inference deployment, not off-the-shelf APIs. This volume surpasses both OpenAI API and Gemini API usage (13:36).

  3. Open vs. Closed Models for Security 1:18:20

    Lin Qiao suggests that open models are better suited to strike a balance in the security debate, encouraging broader community participation to increase defensive complexity against potential cyber threats (47:00).

Technical details

  • Model Customization and Training Paradigms 2336s

    The process of specialization must be continuous. Techniques discussed include Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), KTO, and Reinforcement Learning (RL). RL is highlighted as a method where the model learns by interacting with a product or simulation environment and receiving a 'reward' signal, rather than just being shown ground truth (23:36).

  • Platform Engineering & Optimization

    Fireworks emphasizes its proprietary platform design to maximize quality while optimizing for speed and cost. A key technical achievement is reaching bitwise equivalence between training and inference results, which is crucial for maintaining precision during deployment (50:51).

  • Self-Serve Training Infrastructure 2853s

    Fireworks offers a self-serve training SDK that allows developers to plug in custom loss functions and tweak algorithms. This enables customers with varying levels of AI expertise to conduct complex experiments without needing full expert intervention (28:53).

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.