Early Career Seminar Series #1: Prof. Mohammed AlQuraishi

Early Career Seminar Series #1: Prof. Mohammed AlQuraishi

Introduction to the Series

Overview of the Series

  • This series invites early-career faculty members and PIs to share their insights and visions in protein design and molecular machine learning.
  • The aim is to highlight innovative ideas and high-quality work that inspire advancements in these fields.

Motivation for Speakers

  • Many speakers are emerging PhD students and postdocs preparing for the job market, providing them a platform for practice and exposure.
  • The series has been planned extensively, aiming to be both useful and exciting for attendees.

Introduction of Professor Muhammad Al Kureshi

Speaker Background

  • Dr. Muhammad Al Kureshi is an assistant professor at Columbia University, specializing in systems biology, machine learning, and biophysics.
  • He holds an MS in statistics and a PhD in genetics from Stanford University, with prior experience at Harvard Medical School developing models for protein structure learning.

Recent Contributions

  • His recent work includes OpenFold, an open-source extension of AlphaFold, and Genie, a generative model for protein design.

Key Themes of the Talk

Focus Areas

  • Dr. Al Kureshi will discuss how AlphaFold-based systems make predictions about protein structures during inference processes.
  • He will also explore how these systems learn from data during training phases, acquiring knowledge about proteins they model.

Importance of Understanding Models

  • The talk aims to address the level of physical understanding that machine learning models have regarding proteins and other biological modalities.

Revolution in Protein Structure Prediction

Historical Context

  • Protein structure prediction has seen limited success historically but underwent significant progress with AlphaFold's introduction around 2019–2020.
  • The accuracy of predictions has improved dramatically since then, with results often indistinguishable from experimental structures.

Current Challenges

  • While advancements have been made primarily in protein modeling, extending these successes to other biomolecules remains challenging due to discrepancies observed in predictive performance across different types of molecules.

AlphaFold's Learning Process

Mechanisms of Learning

  • Dr. Al Kureshi will focus on what AlphaFold learns about physical properties while predicting structures as well as how it learns throughout its training process using various observations from his lab's research as well as others'.

Observations on Physical Understanding

  1. Multiple Sequence Alignment (MSA): MSA shapes the energy landscape guiding model predictions by providing structural hypotheses based on evolutionary data.
  1. Model Performance: Removing MSA leads to reliance solely on scoring functions; this impacts confidence scores assigned by AlphaFold when making predictions without strong priors.
  1. Layered Predictions: Analysis shows that deeper layers do not always yield better predictions if initial conditions (like MSA depth) are favorable; stability can be achieved early on.

Training Dynamics

  1. Inductive Prior: Recycling outputs within AlphaFold induces a refinement process where varying degrees of structural completion are presented.
  1. Structural Probes: Using probes allows visualization of how model understanding evolves through network layers during training iterations.

Conclusions Drawn from Experiments

Multimer Complex Predictions

  • In experiments focusing on multimer complexes at CASP15, some groups achieved superior results without pairing MSAs compared to traditional methods relying heavily on co-evolutionary signals between sequences.

This indicates potential flexibility or alternative strategies within predictive modeling frameworks that could enhance outcomes beyond established methodologies used previously.


These notes encapsulate key discussions from Professor Muhammad Al Kureshi’s talk while linking back directly to specific timestamps for further exploration or review within the context provided by the transcript content shared above.

Understanding the Role of MSAs in Protein Structure Prediction

The Impact of Not Pairing MSAs

  • The absence of paired Multiple Sequence Alignments (MSAs) leads to a loss of inter-protein co-variation information, although intra-protein co-variation remains intact.
  • Despite this limitation, models can still perform well by relying on their implicit physical understanding to assemble structures without the co-variation signal.
  • Observations suggest that models may have developed some level of physical understanding, as they remain robust even when breaking the co-variation signal.

Generalization to Unseen Structural Spaces

  • The discussion transitions towards how models generalize to unseen regions of structural space, indicating potential for broader applicability beyond memorized data.
  • Evidence will be presented showing that while models can generalize effectively, there are limits and specific adversarial cases where performance declines significantly.

OpenFold and Its Development

  • OpenFold was initiated as an effort to reproduce AlphaFold 2's capabilities and provide a platform for new developments in protein structure prediction.
  • Current efforts are underway to reproduce AlphaFold 3 with OpenFold 3, aiming for improved functionalities based on industry-academic collaboration.

Evaluating Model Performance Against AlphaFold

  • A comparison between predictions from AlphaFold 2 and OpenFold shows high correlation in performance metrics despite differences in underlying architecture.
  • The CASP 14 competition highlighted challenges faced by AlphaFold due to structurally dissimilar targets, raising questions about its generalization capabilities.

Training Data Subsampling Experiments

  • Initial experiments involved subsampling the Protein Data Bank (PDB), revealing that even with reduced data availability, model performance remained relatively strong compared to earlier versions like AlphaFold 1.

Structural Stratification Analysis

  • Further experiments stratified PDB data based on structural classifications (e.g., CATH), testing model performance against topologically distinct datasets.

Generalization Across Different Topologies

  • Results indicate that models trained on limited topology subsets can still perform competitively when tested against disjoint topologies, suggesting effective generalization at various structural levels.

Challenges in Predicting Mutant Forms

Limitations with Structurally Disruptive Mutations

  • Models struggle significantly with predicting outcomes for mutant forms where single amino acid changes lead to substantial structural disruptions.

Training Biases Affecting Predictions

  • The training process emphasizes maintaining structural integrity within families represented by MSAs; thus, disruptive mutations challenge this learned behavior leading to poor predictions.

Insights into Learning Processes

Observations from Partial Training Models

  • An exploration into how undertrained models progress through different stages during training reveals insights into their learning dynamics and optimization landscapes.

Implications of Physical Priors in Modeling

  • Introducing physical priors may complicate optimization landscapes; nonphysical intermediate states observed during training suggest caution when integrating such constraints into model design.

Understanding Protein Structure Prediction Models

Hypotheses on Model Learning Dimensions

  • The model aims to learn information across 1D, 2D, and 3D aspects of protein structures, extracting the most accessible data in each dimension.
  • A comparison is made between full predictions and lower-dimensional projections (2D and 1D), indicating a loss of information when reducing dimensions.
  • If all spatial dimensions are learned simultaneously, one would expect a consistent drop in prediction accuracy across dimensions before flattening out due to inherent dimensional limitations.
  • An alternative hypothesis suggests that the model learns 1D information first, followed by 2D and then 3D, leading to staggered learning phases observable in empirical results.
  • Empirical observations lean towards the staggered approach rather than simultaneous extraction of all dimensional information.

Implications of Learning Behavior

  • The learning behavior observed may be specific to the AlphaFold architecture; previous models exhibited different learning patterns without intermediate stages.
  • Early predictions from other models often resulted in unrealistic helical structures before converging towards more globular forms.
  • The model's training process shows a rapid acquisition phase where approximately 90% of final accuracy is achieved within the first few thousand steps.

Training Dynamics and Model Performance

  • After initial rapid improvement, subsequent training yields diminishing returns regarding overall accuracy but may enhance implicit physical understanding.
  • Experiments were conducted to assess how well partially trained models rank hypotheses based on their implicit physical understanding as training progresses.
  • By the tenth thousand iteration mark, correlation with experimental rankings improves significantly but still lags behind final performance metrics.

Insights into MSA vs. Implicit Physics

  • While additional training may not improve overall accuracy significantly, it plays a crucial role in enhancing implicit scoring functions during model evaluation.

Application of AlphaFold for Designed Proteins

  • Predictions using single sequences generally yield poor results unless there’s an underlying optimization process like Rosetta energy function optimization for designed proteins.
  • Designed proteins often exhibit characteristics that allow single-sequence predictions to perform adequately due to their rigid structure properties derived from design processes.

Challenges with Chemical Space and Generalization

  • The larger chemical space presents challenges for generalization compared to protein prediction due to lack of multiple sequence alignment (MSA).

Future Directions in Research

  • Innovations in high-throughput crystallography could provide new opportunities for better data collection and improved modeling capabilities over time.
Video description

Dr. Mohammed AlQuraishi is an Assistant Professor in the Department of Systems Biology and a member of Columbia’s Program for Mathematical Genomics, where he works at the intersection of machine learning, biophysics, and systems biology. He earned an MS in statistics and a PhD in genetics from Stanford University. He subsequently joined the Systems Biology Department at Harvard Medical School as a Departmental Fellow and a Fellow in Systems Pharmacology, where he developed the first end-to-end differentiable model for learning protein structure from data. Prior to starting his academic career, Dr. AlQuraishi spent three years founding two startups in the mobile computing space. He joined the Columbia faculty in 2020. Prof. AlQuraishi’s recent work includes OpenFold, a widely used open-source reproduction and extension of AlphaFold, and Genie, a diffusion-based generative model for protein design. A pioneer in the field of computational protein modeling and design, it is an honor to open our series with him!