Fueling Innovation at Scale: Inside Pinterest’s Machine Learning Platform
Pinterest detailed its journey toward mature Model Lifecycle Management (MLOps), focusing on migrating its core ML platform from MLflow to Weights & Biases (W&B). The discussion covers the massive scale of Pinterest's ML operations (supporting 30,000+ workloads/week and 500 model deployments/week) and the critical need for a reliable platform. The migration strategy emphasized achieving system parity, zero downtime, and adding innovation through a parallel deployment approach, which significantly mitigated risk.
Key takeaways
-
MLOps Maturity Goals
11:12
The migration was guided by three main goals: 1) System parity (no change in model quality/performance), 2) No downtime (due to the critical nature of the infrastructure), and 3) Adding innovation and improvements beyond simple maintenance.
-
Parallel Migration Strategy
12:10
Pinterest implemented a parallel migration, supporting both MLflow and Weights & Biases simultaneously. This mitigated risk by allowing a fallback to the old system at any time. Techniques included adding automatic logging to MLflow functions to also log to W&B, and providing self-serve tools for data transfer.
-
Serving Stack Validation and Parity
Before deploying a W&B model, a rigorous validation process was implemented. This involved verifying file accessibility, ensuring metadata parity (version, stages, aliases, tags) with the MLflow equivalent, and confirming that all required serving files were present. This parity check acted as a 'green light' for migration.
-
Governance and Dev/Prod Separation
To improve reliability and governance, Pinterest implemented features like entity-based access (separating Dev and Prod environments) and promotion guardrails. These guardrails prevent model promotion if specific production rules are not met (e.g., missing files, failing load tests).