Weights & Biases

Fueling Innovation at Scale: Inside Pinterest’s Machine Learning Platform

Published 2026-10-07 · Duration 17:28

Summary

Pinterest detailed its journey toward mature Model Lifecycle Management (MLOps), focusing on migrating its core ML platform from MLflow to Weights & Biases (W&B). The discussion covers the massive scale of Pinterest's ML operations (supporting 30,000+ workloads/week and 500 model deployments/week) and the critical need for a reliable platform. The migration strategy emphasized achieving system parity, zero downtime, and adding innovation through a parallel deployment approach, which significantly mitigated risk.

Download summary

Key takeaways

  1. MLOps Maturity Goals 11:12

    The migration was guided by three main goals: 1) System parity (no change in model quality/performance), 2) No downtime (due to the critical nature of the infrastructure), and 3) Adding innovation and improvements beyond simple maintenance.

  2. Parallel Migration Strategy 12:10

    Pinterest implemented a parallel migration, supporting both MLflow and Weights & Biases simultaneously. This mitigated risk by allowing a fallback to the old system at any time. Techniques included adding automatic logging to MLflow functions to also log to W&B, and providing self-serve tools for data transfer.

  3. Serving Stack Validation and Parity

    Before deploying a W&B model, a rigorous validation process was implemented. This involved verifying file accessibility, ensuring metadata parity (version, stages, aliases, tags) with the MLflow equivalent, and confirming that all required serving files were present. This parity check acted as a 'green light' for migration.

  4. Governance and Dev/Prod Separation

    To improve reliability and governance, Pinterest implemented features like entity-based access (separating Dev and Prod environments) and promotion guardrails. These guardrails prevent model promotion if specific production rules are not met (e.g., missing files, failing load tests).

Technical details

  • Platform Scale and Scope 240s

    The platform supports over 580 million monthly active users, 390 billion saved pins, and runs ML workloads on thousands of GPUs. The ML platform supports over 400 ML engineers running 30,000+ ML workloads weekly and deploying 500 models weekly.

  • ML Training Infrastructure 320s

    The infrastructure is split into a compute layer (on AWS/K8s for GPUs, job management, auto-scaling) and an ML runtime layer (internal libraries on Ray and PyTorch). Experiment tracking and the model registry are critical components.

  • Migration Challenges (MLflow) 440s

    The previous MLflow fork was unreliable, stuck on an outdated version, and suffered from numerous outages. A major bottleneck was the logging operation, which physically copied every file to S3, causing slow timeouts, especially for large models like LLMs.

  • Automation and Efficiency Gains 850s

    To speed up the migration, Pinterest employed advanced techniques: 1) Building an agent to identify and apply code changes to custom logging scripts, and 2) Training an on-call agent (using Pinclaw) to automate support and onboarding questions.

  • Serving Stack Integration

    The serving stack uses a ranking service (Scorpion) that calls a model server (Scorpion Model Server). W&B integrates at the build stage (packaging models), the validation stage (checking parity with MLflow), and the monitoring stage. The build logic was updated to create both MLflow and W&B builds, sharing the same models.

Mentioned resources

  • Weights & Biases (W&B) (MLOps Platform)
  • MLflow (MLOps Platform (Legacy))
  • AWS/K8s (Cloud Infrastructure)
  • Ray and PyTorch (ML Frameworks)
  • Scorpion (Ranking Service)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.