Fueling Innovation at Scale: Inside Pinterest’s Machine Learning Platform
Summary
Pinterest detailed its journey toward mature Model Lifecycle Management (MLOps), focusing on migrating its core ML platform from MLflow to Weights & Biases (W&B). The discussion covers the massive scale of Pinterest's ML operations (supporting 30,000+ workloads/week and 500 model deployments/week) and the critical need for a reliable platform. The migration strategy emphasized achieving system parity, zero downtime, and adding innovation through a parallel deployment approach, which significantly mitigated risk.
Key takeaways
-
MLOps Maturity Goals
11:12
The migration was guided by three main goals: 1) System parity (no change in model quality/performance), 2) No downtime (due to the critical nature of the infrastructure), and 3) Adding innovation and improvements beyond simple maintenance.
-
Parallel Migration Strategy
12:10
Pinterest implemented a parallel migration, supporting both MLflow and Weights & Biases simultaneously. This mitigated risk by allowing a fallback to the old system at any time. Techniques included adding automatic logging to MLflow functions to also log to W&B, and providing self-serve tools for data transfer.
-
Serving Stack Validation and Parity
Before deploying a W&B model, a rigorous validation process was implemented. This involved verifying file accessibility, ensuring metadata parity (version, stages, aliases, tags) with the MLflow equivalent, and confirming that all required serving files were present. This parity check acted as a 'green light' for migration.
-
Governance and Dev/Prod Separation
To improve reliability and governance, Pinterest implemented features like entity-based access (separating Dev and Prod environments) and promotion guardrails. These guardrails prevent model promotion if specific production rules are not met (e.g., missing files, failing load tests).
Technical details
-
Platform Scale and Scope
240s
The platform supports over 580 million monthly active users, 390 billion saved pins, and runs ML workloads on thousands of GPUs. The ML platform supports over 400 ML engineers running 30,000+ ML workloads weekly and deploying 500 models weekly.
-
ML Training Infrastructure
320s
The infrastructure is split into a compute layer (on AWS/K8s for GPUs, job management, auto-scaling) and an ML runtime layer (internal libraries on Ray and PyTorch). Experiment tracking and the model registry are critical components.
-
Migration Challenges (MLflow)
440s
The previous MLflow fork was unreliable, stuck on an outdated version, and suffered from numerous outages. A major bottleneck was the logging operation, which physically copied every file to S3, causing slow timeouts, especially for large models like LLMs.
-
Automation and Efficiency Gains
850s
To speed up the migration, Pinterest employed advanced techniques: 1) Building an agent to identify and apply code changes to custom logging scripts, and 2) Training an on-call agent (using Pinclaw) to automate support and onboarding questions.
-
Serving Stack Integration
The serving stack uses a ranking service (Scorpion) that calls a model server (Scorpion Model Server). W&B integrates at the build stage (packaging models), the validation stage (checking parity with MLflow), and the monitoring stage. The build logic was updated to create both MLflow and W&B builds, sharing the same models.
Mentioned resources
- Weights & Biases (W&B)
- MLflow
- AWS/K8s
- Ray and PyTorch
- Scorpion
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.