40 Trillion Tokens a Day (Yes, More Than OpenAI) | Lin Qiao, CEO of Fireworks
Summary
The discussion centers on the industry shift from general-purpose AI models (like those from OpenAI/Anthropic) toward specialized intelligence. Lin Qiao of Fireworks argues that true innovation lies in leveraging proprietary, locked-in enterprise data—the 'alpha'—to build customized models. She asserts that this specialization is necessary because generalized models cannot capture a company's unique knowledge or judgment. Technically, the conversation details advanced training methods (SFT, DPO, KTO, RL) and emphasizes platform control, noting that Fireworks achieves bitwise equivalence between training and inference results to ensure maximum quality while optimizing for cost and speed.
Key takeaways
-
The Rise of Specialized Intelligence
1:08:55
Lin Qiao argues that the future belongs to specialized intelligence—customized models built on private company data—rather than general-purpose AGI. She believes every company is unique, making it difficult for a single general model to capture proprietary knowledge (41:35).
-
Fireworks' Scale and Focus
22:16
Fireworks claims to process over 40 trillion tokens daily, stating that 95% of this traffic comes from customized model inference deployment, not off-the-shelf APIs. This volume surpasses both OpenAI API and Gemini API usage (13:36).
-
Open vs. Closed Models for Security
1:18:20
Lin Qiao suggests that open models are better suited to strike a balance in the security debate, encouraging broader community participation to increase defensive complexity against potential cyber threats (47:00).
Technical details
-
Model Customization and Training Paradigms
2336s
The process of specialization must be continuous. Techniques discussed include Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), KTO, and Reinforcement Learning (RL). RL is highlighted as a method where the model learns by interacting with a product or simulation environment and receiving a 'reward' signal, rather than just being shown ground truth (23:36).
-
Platform Engineering & Optimization
Fireworks emphasizes its proprietary platform design to maximize quality while optimizing for speed and cost. A key technical achievement is reaching bitwise equivalence between training and inference results, which is crucial for maintaining precision during deployment (50:51).
-
Self-Serve Training Infrastructure
2853s
Fireworks offers a self-serve training SDK that allows developers to plug in custom loss functions and tweak algorithms. This enables customers with varying levels of AI expertise to conduct complex experiments without needing full expert intervention (28:53).
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.