# Fueling Innovation at Scale: Inside Pinterest’s Machine Learning Platform

## Executive summary

Pinterest detailed its journey toward mature Model Lifecycle Management (MLOps), focusing on migrating its core ML platform from MLflow to Weights & Biases (W&B). The discussion covers the massive scale of Pinterest's ML operations (supporting 30,000+ workloads/week and 500 model deployments/week) and the critical need for a reliable platform. The migration strategy emphasized achieving system parity, zero downtime, and adding innovation through a parallel deployment approach, which significantly mitigated risk.

## Key takeaways

- MLOps Maturity Goals: The migration was guided by three main goals: 1) System parity (no change in model quality/performance), 2) No downtime (due to the critical nature of the infrastructure), and 3) Adding innovation and improvements beyond simple maintenance.
- Parallel Migration Strategy: Pinterest implemented a parallel migration, supporting both MLflow and Weights & Biases simultaneously. This mitigated risk by allowing a fallback to the old system at any time. Techniques included adding automatic logging to MLflow functions to also log to W&B, and providing self-serve tools for data transfer.
- Serving Stack Validation and Parity: Before deploying a W&B model, a rigorous validation process was implemented. This involved verifying file accessibility, ensuring metadata parity (version, stages, aliases, tags) with the MLflow equivalent, and confirming that all required serving files were present. This parity check acted as a 'green light' for migration.
- Governance and Dev/Prod Separation: To improve reliability and governance, Pinterest implemented features like entity-based access (separating Dev and Prod environments) and promotion guardrails. These guardrails prevent model promotion if specific production rules are not met (e.g., missing files, failing load tests).

## Technical details

- Platform Scale and Scope: The platform supports over 580 million monthly active users, 390 billion saved pins, and runs ML workloads on thousands of GPUs. The ML platform supports over 400 ML engineers running 30,000+ ML workloads weekly and deploying 500 models weekly.
- ML Training Infrastructure: The infrastructure is split into a compute layer (on AWS/K8s for GPUs, job management, auto-scaling) and an ML runtime layer (internal libraries on Ray and PyTorch). Experiment tracking and the model registry are critical components.
- Migration Challenges (MLflow): The previous MLflow fork was unreliable, stuck on an outdated version, and suffered from numerous outages. A major bottleneck was the logging operation, which physically copied every file to S3, causing slow timeouts, especially for large models like LLMs.
- Automation and Efficiency Gains: To speed up the migration, Pinterest employed advanced techniques: 1) Building an agent to identify and apply code changes to custom logging scripts, and 2) Training an on-call agent (using Pinclaw) to automate support and onboarding questions.
- Serving Stack Integration: The serving stack uses a ranking service (Scorpion) that calls a model server (Scorpion Model Server). W&B integrates at the build stage (packaging models), the validation stage (checking parity with MLflow), and the monitoring stage. The build logic was updated to create both MLflow and W&B builds, sharing the same models.

## Practical implications

- For build engineers managing critical, high-scale services, adopting a parallel migration strategy (supporting both old and new systems) is a highly effective risk mitigation technique.
- Implementing robust governance layers, such as promotion guardrails and entity-based access, is crucial for maintaining reliability and compliance when modernizing ML infrastructure.
- The concept of 'parity checking'—ensuring the new system's output matches the old system's output before cutover—is a best practice for minimizing risk during major platform upgrades.

## Topics

MLOps, Model Lifecycle Management, CI/CD, Build Engineering, Distributed Systems, Machine Learning, Weights & Biases (W&B), MLflow, AWS/K8s, Ray and PyTorch, Scorpion

Source: https://www.youtube.com/watch?v=Mva-kQayVck
