Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub
Summary
The discussion critiques the notion that AI has solved protein folding, arguing that current models like AlphaFold primarily replicate structures found in the PDB rather than understanding the full dynamics or true ground state of proteins. The panel emphasizes that progress requires moving beyond simply scaling compute and data, advocating instead for a focus on finding the correct biological scaling laws, incorporating deep scientific intuition (inductive biases), and developing holistic models of living systems, such as the 'virtual cell.'
Key takeaways
-
Protein Folding is Not Solved
Current models, including AlphaFold, are highly effective at predicting structures based on existing data (PDB), but this does not equate to understanding all protein dynamics or the true ground state of proteins. The models are useful, but the problem remains fundamentally open.
-
Rethinking the Scaling Law (The Bitter Lesson for Data)
2:20
The traditional 'Bitter Lesson' (that scaling methods wins) must be refined for biology. The challenge is not merely adding data or compute, but identifying the specific scaling law that governs the problem. The ability to find this law is the critical bottleneck.
-
Systemic Modeling is the Next Frontier
7:30
To achieve major breakthroughs, modeling must shift from focusing on individual proteins to understanding complex biological systems (e.g., building a 'virtual cell'). This requires fundamentally different datasets and a multi-disciplinary approach.
-
Prioritizing Understanding and Trustworthiness
21:40
While predictive power is valuable, the focus must also be on model interpretability and understanding the model's limitations (e.g., uncertainty calibration). Trustworthy AI requires knowing *what* the model can do and, crucially, *what* it cannot do.
Technical details
-
AlphaFold Limitations
1100s
AlphaFold 2 was a 'work of art' utilizing carefully handcrafted features and scientific intuition from biophysics and biochemistry. Its success is limited to replicating known structures, and its pLDDT score must be interpreted with caution, as calibration of uncertainty is paramount.
-
Data Modalities and Bias
240s
Training protein language models (PLMs) on non-pristine data, such as metagenomic sequences, can significantly improve performance for designing functional proteins. This demonstrates that 'good data' (contextually relevant, even if low-quality) is often more valuable than simply having 'more data.'
-
Model Architecture and Scaling
1200s
The evolution of modeling involves moving beyond simple scaling. While the Transformer architecture is foundational, future progress requires developing bespoke architectures that are highly fit-for-purpose and can handle internet-scale data.
-
Biological Constraints
320s
The process of advancing science requires identifying the core challenge (the problem) first, rather than being limited by the available data or the current modeling technology. This requires a multi-disciplinary approach that integrates scientific expertise.
Mentioned resources
- PDB (Protein Data Bank)
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.