{
  "id": 460403,
  "title": "[3rd Place Solution] AlphaFold Style Twin Tower Architecture + Squeezeformer",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460403",
  "author_name": "dan4o",
  "post_date": "2023-12-09T04:14:09.013000",
  "votes": 24,
  "comment_count": 19,
  "views": 0,
  "content": "<p>A big thanks to the organizers for making this competition happen and putting their efforts in solving hard problems, such as RNA structure prediction. This sentence stuck with me throughout the whole competition: “… without being able to understand how RNA molecules fold, we are missing a deeper <strong>understanding of how nature works, how life began</strong>, and how we can design…“</p>\n<p>Many will glance over this statement but it is a huge mystery that sometimes kept me awake at night. I mean it is SO strange… the evolution/creation of 4 nucleotides with specific physical/chemical properties that can code and take on different structures to fulfill specific actions in the body, and we still don’t understand the origin/purpose of these molecules… Hopefully we are in the right direction in finding out the truth and reason behind it or even “what” created them. </p>\n<p>Solution is not based on dozens of ensemble models, but rather 2 strong independent models. There are 2 models since I wasn’t sure if I will be able to produce an Alpha fold style viable solution in the allotted time, so I focused on a smaller “safe” version model first and a bigger “riskier” model later in the competition. But the ensemble of the 2 provided crucial in gaining the 3rd place.</p>\n<h2>TLDR</h2>\n<p>Blend of 2 Independent models. The smaller “safer” version model is based on augmented Squeezeformer architecture, which consists of RelativeMultiheadSelfAttention, Convolution and FeedFoward modules. Learnable BPP’s through 2d convolution are added to the attention scores, and augmented low Signal-to-Noise data is additionally used in training apart from clean training data. Bigger twin-tower model is based on augmented Alpha Fold style architecture, that consists of MSA stack representation and Pair stack representation, that communicate in a criss-cross fashion through Outer Product Mean and Pair Representation Bias.</p>\n<h2>Data Preprocessing/Cross Validation</h2>\n<p>Training data was split in 4 folds similar to this <a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb/notebook\" target=\"_blank\">notebook</a>. There was a stable correlation between local cv vs lb. BPP’s were processed and cached as .npz files. Most of the time, models were evaluated only on fold0. <br>\nTwin-tower model was trained only on clean training data and it has the power to predict its confidence/error_estimate similar to pLDDT in AlphaFold. After getting a good generalization from the twin-tower model on the clean data, new training data was created by augmenting the noisy low SNR data with the model’s confidence/error_estimate. For a particular nucleotide position in the low SNR dataset the idea is to combine how confident the model's prediction is with the position's reactivity error from the experiment and thus “fix” the noisy data. This gave significant improvements in the smaller “safer” model which was trained on clean dataset + improved low snr dataset. Even with this useful new data being available for the bigger twin-tower model, it was never trained on this extra data because of time constraints and only the clean dataset was used for final submission. Twin-tower experiments are currently underway on the full dataset.</p>\n<h2>Squeezeformer Model</h2>\n<p>The smaller “safer” model is based on Squeezeformer architecture which was used in previous Google American Sign Language competition. The inputs to the model were tokenized RNA sequences and BPP matrices. The model consists of 14xSqueezeformer blocks and an output projection layer that predicts chemical reactivities at each position. One Squeezeformer block consists of three modules: Relative MultiHeadSelfAttention module, Convolution module, and FeedForward module.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F55cab60a14c35cb028ecc1de472c6566%2FSqueezeFormerJPG-01.jpg?generation=1702092589333558&amp;alt=media\" alt=\"\"></p>\n<p>Instead of absolute positional encodings, relative encodings are used for generalization on longer sequences. Attention scores apart from relative pos scores are further affected by BasePairProbability matrices, even though I was reluctant to use them at first since they are created with software that is not capable of detecting long range pseudo-knots and this bias is unfortunately induced in the model. The attention scores in the transformer are calculated in this manner: (content score + relative position score)/sqrt(head_dim) + bpp_bias_score. Bpp_bias_score is obtained by passing the BPP matrices through a 2D convolution block.</p>\n<h2>AlphaFold style Twin-tower Model</h2>\n<p>This model was inspired from Google's AlphaFold and its derivatives OpenComplex/RhoFold. The original AlphaFold architecture relies on two different input representations for its predictions. It jointly uses MSA(Multiple Sequence Alignment) and Pair Representation features. The MSA representation uses row-wise attention to find intra-sequence features, while the column-wise attention is used to obtain inter-sequence evolutionary signals from the MSA stack. Since MSA for this competition wasn’t helpful in extracting evolutionary information, because of its synthetic nature of the RNA sequences, just the tokenized input sequence was used. Because MSA was not used, the axial-self attention was replaced with relative multi head attention and convolution.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F99a3d1c0bbb83caf86c5a82c2d2a7627%2FChemFoldJPG-01.jpg?generation=1702092825428051&amp;alt=media\" alt=\"\"><br>\n(At the time of submission, only the Single representation branch was used for predictions, the pair features were completely ignored due to time restrictions)</p>\n<h3>The workflow of the data from the input all the way to the prediction is as follows:</h3>\n<p>1) Input sequence gets tokenized, sequence and pair masks are generated</p>\n<p>2) Tokenized sequence are embedded through embedding networks: MSANet and PairNet</p>\n<ul>\n<li>PairRepresentation features are embedded using Relative2D positional encodings to provide information about position of residues. Maximum position is clipped at 32, and each position afterwards is considered as “far” away. This inductive bias trains the model not to rely heavily on nucleotide positions and generalize better to any length, as stated in the alpha fold supplemental materials.</li>\n<li>MSA representation is passed through a simple Embedding layer (since positional encodings are added later in the transformer layer).</li>\n</ul>\n<p>3) Embedded MSA and Pair features are passed through the main trunk of the network, which consists of 8 Chemformer blocks illustrated above. Some key points are:</p>\n<ul>\n<li>MSA representation is updated within the Squeezeformer attention <br>\ninstead of axial self-attention proposed in the original paper.</li>\n<li>Pair representation is updated only with triangular multiplicative updates. Triangular self-attention wasn’t used in this model because of memory constraints but using both should give an increase in score.</li>\n<li>The MSA representation updates the pair representation through an element-wise outer product that is summed over the MSA sequence dimension.</li>\n<li>The Pair representation updates the MSA representation through a projection of additional logits from the pair stack to bias the MSA attention scores.</li>\n<li>Both representations are passed through a 2-layer MLP that acts as transition before the communication.</li>\n<li>This communication is repeated within each block, so 8 times in total.</li>\n<li>Residual connections and Row-wise/Column-wise dropouts are used.</li>\n</ul>\n<p>4) The processed representations MSA Features and Pair Features are then passed through various heads to extract meaningful information such as confidence/error_estimate, chemical reactivity, base_pair probabilities etc.</p>\n<h2>Training</h2>\n<p>Training procedure is as follows:</p>\n<p>First the Twin-tower model is trained on the clean dataset for 60 epochs with lr of 1e-3. Batch size of 8 per gpu was used with gradient accumulation of 4, leading to an effective batch size of 64 on 2 GPUs. The optimizer used was AdamW with cosine scheduling, weight decay parameter is 0.05 and warmup is 0.5. Model was trained on 2x4090 for 30 hours total. No extensive parameter tuning was done for this model.<br>\nThis leads to a 0.13746 public/0.14398 private single model score. After this a synthetic dataset was created by blending the model’s predictions with a subset of noisy data that was between 0.35&lt;data&lt;1 SNR. Selecting a lower SNR threshold increased the training examples but lowered the quality. Current synthetic dataset is created with simple weighing, because of time constraints even though the model was trained and outputs valid plddt/error_estimates, they were not fully utilized. I didn’t have time to experiment with the formula for combining plddt score with experimental reactivity_error on a per nucleotide basis which is guaranteed to produce better synthetic data. <br>\nThe smaller “safer” model was trained for 200 epochs with lr of 7e-4 a batch size of 64. Optimizer and lr scheduling was the same as before, however no grad accumulation was used. This model was trained both on clean dataset and the synthetically created dataset which doubled the training examples. This single model achieves 0.13865 public/0.14256 private score.</p>\n<p>A day before the competition ended I decided to just run longer epochs on the Twin-tower model, so I continued the same training procedure with starting weights from epoch 30 of the previous iteration. This gave improvements just by running the model longer which led to the final 0.13706 public/0.14366 private score of this single model. The final score is a blend of the twin-tower model which was trained only on the input sequence and tried to learn all of the interactions on its own and the squeeze model which contributed with the BPP’s and the synthetic data to the final prediction.</p>\n<h2>Not utilized/Can be improved for Twin-tower model</h2>\n<ul>\n<li>Model was trained only with sequences as input. </li>\n<li>BPP’s and other supplemental data can be used. </li>\n<li>Synthetic dataset which improved the smaller model should be used.</li>\n<li>With my tests, by increasing the depth (8-&gt;12) blocks, there is an increase in score, submission is only 8 blocks long. </li>\n<li>Deeper model 16,24 blocks will be tested soon by utilizing checkpoiting/rematerialization to counter the memory problem and interleaving the convolution blocks after several MHSA blocks to aid training for this deeper model. </li>\n<li>Recycling from the original paper was not utilized.</li>\n<li>Only triangular multiplicative updates are used in the Pair Representation, triangular self-attention was not utilized.</li>\n<li>Since RNA can take on multiple conformations, dropout can be utilized at inference time and results averaged to get a better estimate of the RNA’s structure</li>\n</ul>\n<p>All of these adjustments are very likely to provide benefits to the model. Some of the tests are currently underway.</p>\n<h2>Comments</h2>\n<p>The twin-tower model was a large task, partially because I was competing solo. The model at submission time was trained only on clean dataset, and only the MSA Feature pathway was used in the predictions, the other Pair Feature pathway was completely ignored, but it can be used to try and recreate the BPP’s which should increase the score (tests are underway as im writing). Best submission of the model without BPP’s, loop types or any pre/post processing is 0.13709 public and 0.14366 private. However, the twin-tower model has a bigger gap in generalization compared to the smaller “safer” squeezeformer model which scored 0.13865 public but 0.14265 private, so currently investigating the reason behind it.</p>\n<p>I am new to machine learning. I started to learn the field in May of this year, so I am sure there will be a lot of mistakes in the code and in my approach and sorry if my explanation is all over the place, all this is new to me and I am still learning.</p>\n<p>Open Sourced Code:<br>\n<a href=\"https://github.com/GosUxD/OpenChemFold\" target=\"_blank\">https://github.com/GosUxD/OpenChemFold</a></p>\n<h2>References:</h2>\n<p>[1] Squeezeformer: An Efficient Transformer for Automatic Speech Recognition<br>\nSehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W. Mahoney, Kurt Keutzer arXiv:2206.00888 [eess.AS] <a href=\"https://doi.org/10.48550/arXiv.2206.00888\" target=\"_blank\">https://doi.org/10.48550/arXiv.2206.00888</a></p>\n<p>[2] Winner of Google American Sign Language Fingerspelling Competition<br>\n<a href=\"https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution</a></p>\n<p>[3] Jumper, J., Evans, R., Pritzel, A. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). <a href=\"https://doi.org/10.1038/s41586-021-03819-2\" target=\"_blank\">https://doi.org/10.1038/s41586-021-03819-2</a></p>\n<p>[4] OpenComplex github code repository:<br>\n<a href=\"https://github.com/baaihealth/OpenComplex\" target=\"_blank\">https://github.com/baaihealth/OpenComplex</a></p>\n<p>[5] E2Efold-3D: End-to-End Deep Learning Method for accurate de novo RNA 3D Structure Prediction<br>\nTao Shen, Zhihang Hu, Zhangzhi Peng, Jiayang Chen, Peng Xiong, Liang Hong, Liangzhen Zheng, Yixuan Wang, Irwin King, Sheng Wang, Siqi Sun, Yu Li. arXiv:2207.01586 [q-bio.QM] <a href=\"https://doi.org/10.48550/arXiv.2207.01586\" target=\"_blank\">https://doi.org/10.48550/arXiv.2207.01586</a></p>",
  "messages": [
    {
      "id": 2554334,
      "postDate": "2023-12-09T04:14:09.013Z",
      "content": "<p>A big thanks to the organizers for making this competition happen and putting their efforts in solving hard problems, such as RNA structure prediction. This sentence stuck with me throughout the whole competition: “… without being able to understand how RNA molecules fold, we are missing a deeper <strong>understanding of how nature works, how life began</strong>, and how we can design…“</p>\n<p>Many will glance over this statement but it is a huge mystery that sometimes kept me awake at night. I mean it is SO strange… the evolution/creation of 4 nucleotides with specific physical/chemical properties that can code and take on different structures to fulfill specific actions in the body, and we still don’t understand the origin/purpose of these molecules… Hopefully we are in the right direction in finding out the truth and reason behind it or even “what” created them. </p>\n<p>Solution is not based on dozens of ensemble models, but rather 2 strong independent models. There are 2 models since I wasn’t sure if I will be able to produce an Alpha fold style viable solution in the allotted time, so I focused on a smaller “safe” version model first and a bigger “riskier” model later in the competition. But the ensemble of the 2 provided crucial in gaining the 3rd place.</p>\n<h2>TLDR</h2>\n<p>Blend of 2 Independent models. The smaller “safer” version model is based on augmented Squeezeformer architecture, which consists of RelativeMultiheadSelfAttention, Convolution and FeedFoward modules. Learnable BPP’s through 2d convolution are added to the attention scores, and augmented low Signal-to-Noise data is additionally used in training apart from clean training data. Bigger twin-tower model is based on augmented Alpha Fold style architecture, that consists of MSA stack representation and Pair stack representation, that communicate in a criss-cross fashion through Outer Product Mean and Pair Representation Bias.</p>\n<h2>Data Preprocessing/Cross Validation</h2>\n<p>Training data was split in 4 folds similar to this <a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb/notebook\" target=\"_blank\">notebook</a>. There was a stable correlation between local cv vs lb. BPP’s were processed and cached as .npz files. Most of the time, models were evaluated only on fold0. <br>\nTwin-tower model was trained only on clean training data and it has the power to predict its confidence/error_estimate similar to pLDDT in AlphaFold. After getting a good generalization from the twin-tower model on the clean data, new training data was created by augmenting the noisy low SNR data with the model’s confidence/error_estimate. For a particular nucleotide position in the low SNR dataset the idea is to combine how confident the model's prediction is with the position's reactivity error from the experiment and thus “fix” the noisy data. This gave significant improvements in the smaller “safer” model which was trained on clean dataset + improved low snr dataset. Even with this useful new data being available for the bigger twin-tower model, it was never trained on this extra data because of time constraints and only the clean dataset was used for final submission. Twin-tower experiments are currently underway on the full dataset.</p>\n<h2>Squeezeformer Model</h2>\n<p>The smaller “safer” model is based on Squeezeformer architecture which was used in previous Google American Sign Language competition. The inputs to the model were tokenized RNA sequences and BPP matrices. The model consists of 14xSqueezeformer blocks and an output projection layer that predicts chemical reactivities at each position. One Squeezeformer block consists of three modules: Relative MultiHeadSelfAttention module, Convolution module, and FeedForward module.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F55cab60a14c35cb028ecc1de472c6566%2FSqueezeFormerJPG-01.jpg?generation=1702092589333558&amp;alt=media\" alt=\"\"></p>\n<p>Instead of absolute positional encodings, relative encodings are used for generalization on longer sequences. Attention scores apart from relative pos scores are further affected by BasePairProbability matrices, even though I was reluctant to use them at first since they are created with software that is not capable of detecting long range pseudo-knots and this bias is unfortunately induced in the model. The attention scores in the transformer are calculated in this manner: (content score + relative position score)/sqrt(head_dim) + bpp_bias_score. Bpp_bias_score is obtained by passing the BPP matrices through a 2D convolution block.</p>\n<h2>AlphaFold style Twin-tower Model</h2>\n<p>This model was inspired from Google's AlphaFold and its derivatives OpenComplex/RhoFold. The original AlphaFold architecture relies on two different input representations for its predictions. It jointly uses MSA(Multiple Sequence Alignment) and Pair Representation features. The MSA representation uses row-wise attention to find intra-sequence features, while the column-wise attention is used to obtain inter-sequence evolutionary signals from the MSA stack. Since MSA for this competition wasn’t helpful in extracting evolutionary information, because of its synthetic nature of the RNA sequences, just the tokenized input sequence was used. Because MSA was not used, the axial-self attention was replaced with relative multi head attention and convolution.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F99a3d1c0bbb83caf86c5a82c2d2a7627%2FChemFoldJPG-01.jpg?generation=1702092825428051&amp;alt=media\" alt=\"\"><br>\n(At the time of submission, only the Single representation branch was used for predictions, the pair features were completely ignored due to time restrictions)</p>\n<h3>The workflow of the data from the input all the way to the prediction is as follows:</h3>\n<p>1) Input sequence gets tokenized, sequence and pair masks are generated</p>\n<p>2) Tokenized sequence are embedded through embedding networks: MSANet and PairNet</p>\n<ul>\n<li>PairRepresentation features are embedded using Relative2D positional encodings to provide information about position of residues. Maximum position is clipped at 32, and each position afterwards is considered as “far” away. This inductive bias trains the model not to rely heavily on nucleotide positions and generalize better to any length, as stated in the alpha fold supplemental materials.</li>\n<li>MSA representation is passed through a simple Embedding layer (since positional encodings are added later in the transformer layer).</li>\n</ul>\n<p>3) Embedded MSA and Pair features are passed through the main trunk of the network, which consists of 8 Chemformer blocks illustrated above. Some key points are:</p>\n<ul>\n<li>MSA representation is updated within the Squeezeformer attention <br>\ninstead of axial self-attention proposed in the original paper.</li>\n<li>Pair representation is updated only with triangular multiplicative updates. Triangular self-attention wasn’t used in this model because of memory constraints but using both should give an increase in score.</li>\n<li>The MSA representation updates the pair representation through an element-wise outer product that is summed over the MSA sequence dimension.</li>\n<li>The Pair representation updates the MSA representation through a projection of additional logits from the pair stack to bias the MSA attention scores.</li>\n<li>Both representations are passed through a 2-layer MLP that acts as transition before the communication.</li>\n<li>This communication is repeated within each block, so 8 times in total.</li>\n<li>Residual connections and Row-wise/Column-wise dropouts are used.</li>\n</ul>\n<p>4) The processed representations MSA Features and Pair Features are then passed through various heads to extract meaningful information such as confidence/error_estimate, chemical reactivity, base_pair probabilities etc.</p>\n<h2>Training</h2>\n<p>Training procedure is as follows:</p>\n<p>First the Twin-tower model is trained on the clean dataset for 60 epochs with lr of 1e-3. Batch size of 8 per gpu was used with gradient accumulation of 4, leading to an effective batch size of 64 on 2 GPUs. The optimizer used was AdamW with cosine scheduling, weight decay parameter is 0.05 and warmup is 0.5. Model was trained on 2x4090 for 30 hours total. No extensive parameter tuning was done for this model.<br>\nThis leads to a 0.13746 public/0.14398 private single model score. After this a synthetic dataset was created by blending the model’s predictions with a subset of noisy data that was between 0.35&lt;data&lt;1 SNR. Selecting a lower SNR threshold increased the training examples but lowered the quality. Current synthetic dataset is created with simple weighing, because of time constraints even though the model was trained and outputs valid plddt/error_estimates, they were not fully utilized. I didn’t have time to experiment with the formula for combining plddt score with experimental reactivity_error on a per nucleotide basis which is guaranteed to produce better synthetic data. <br>\nThe smaller “safer” model was trained for 200 epochs with lr of 7e-4 a batch size of 64. Optimizer and lr scheduling was the same as before, however no grad accumulation was used. This model was trained both on clean dataset and the synthetically created dataset which doubled the training examples. This single model achieves 0.13865 public/0.14256 private score.</p>\n<p>A day before the competition ended I decided to just run longer epochs on the Twin-tower model, so I continued the same training procedure with starting weights from epoch 30 of the previous iteration. This gave improvements just by running the model longer which led to the final 0.13706 public/0.14366 private score of this single model. The final score is a blend of the twin-tower model which was trained only on the input sequence and tried to learn all of the interactions on its own and the squeeze model which contributed with the BPP’s and the synthetic data to the final prediction.</p>\n<h2>Not utilized/Can be improved for Twin-tower model</h2>\n<ul>\n<li>Model was trained only with sequences as input. </li>\n<li>BPP’s and other supplemental data can be used. </li>\n<li>Synthetic dataset which improved the smaller model should be used.</li>\n<li>With my tests, by increasing the depth (8-&gt;12) blocks, there is an increase in score, submission is only 8 blocks long. </li>\n<li>Deeper model 16,24 blocks will be tested soon by utilizing checkpoiting/rematerialization to counter the memory problem and interleaving the convolution blocks after several MHSA blocks to aid training for this deeper model. </li>\n<li>Recycling from the original paper was not utilized.</li>\n<li>Only triangular multiplicative updates are used in the Pair Representation, triangular self-attention was not utilized.</li>\n<li>Since RNA can take on multiple conformations, dropout can be utilized at inference time and results averaged to get a better estimate of the RNA’s structure</li>\n</ul>\n<p>All of these adjustments are very likely to provide benefits to the model. Some of the tests are currently underway.</p>\n<h2>Comments</h2>\n<p>The twin-tower model was a large task, partially because I was competing solo. The model at submission time was trained only on clean dataset, and only the MSA Feature pathway was used in the predictions, the other Pair Feature pathway was completely ignored, but it can be used to try and recreate the BPP’s which should increase the score (tests are underway as im writing). Best submission of the model without BPP’s, loop types or any pre/post processing is 0.13709 public and 0.14366 private. However, the twin-tower model has a bigger gap in generalization compared to the smaller “safer” squeezeformer model which scored 0.13865 public but 0.14265 private, so currently investigating the reason behind it.</p>\n<p>I am new to machine learning. I started to learn the field in May of this year, so I am sure there will be a lot of mistakes in the code and in my approach and sorry if my explanation is all over the place, all this is new to me and I am still learning.</p>\n<p>Open Sourced Code:<br>\n<a href=\"https://github.com/GosUxD/OpenChemFold\" target=\"_blank\">https://github.com/GosUxD/OpenChemFold</a></p>\n<h2>References:</h2>\n<p>[1] Squeezeformer: An Efficient Transformer for Automatic Speech Recognition<br>\nSehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W. Mahoney, Kurt Keutzer arXiv:2206.00888 [eess.AS] <a href=\"https://doi.org/10.48550/arXiv.2206.00888\" target=\"_blank\">https://doi.org/10.48550/arXiv.2206.00888</a></p>\n<p>[2] Winner of Google American Sign Language Fingerspelling Competition<br>\n<a href=\"https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution</a></p>\n<p>[3] Jumper, J., Evans, R., Pritzel, A. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). <a href=\"https://doi.org/10.1038/s41586-021-03819-2\" target=\"_blank\">https://doi.org/10.1038/s41586-021-03819-2</a></p>\n<p>[4] OpenComplex github code repository:<br>\n<a href=\"https://github.com/baaihealth/OpenComplex\" target=\"_blank\">https://github.com/baaihealth/OpenComplex</a></p>\n<p>[5] E2Efold-3D: End-to-End Deep Learning Method for accurate de novo RNA 3D Structure Prediction<br>\nTao Shen, Zhihang Hu, Zhangzhi Peng, Jiayang Chen, Peng Xiong, Liang Hong, Liangzhen Zheng, Yixuan Wang, Irwin King, Sheng Wang, Siqi Sun, Yu Li. arXiv:2207.01586 [q-bio.QM] <a href=\"https://doi.org/10.48550/arXiv.2207.01586\" target=\"_blank\">https://doi.org/10.48550/arXiv.2207.01586</a></p>",
      "rawMarkdown": "A big thanks to the organizers for making this competition happen and putting their efforts in solving hard problems, such as RNA structure prediction. This sentence stuck with me throughout the whole competition: “... without being able to understand how RNA molecules fold, we are missing a deeper **understanding of how nature works, how life began**, and how we can design…“\n\nMany will glance over this statement but it is a huge mystery that sometimes kept me awake at night. I mean it is SO strange… the evolution/creation of 4 nucleotides with specific physical/chemical properties that can code and take on different structures to fulfill specific actions in the body, and we still don’t understand the origin/purpose of these molecules… Hopefully we are in the right direction in finding out the truth and reason behind it or even “what” created them. \n\nSolution is not based on dozens of ensemble models, but rather 2 strong independent models. There are 2 models since I wasn’t sure if I will be able to produce an Alpha fold style viable solution in the allotted time, so I focused on a smaller “safe” version model first and a bigger “riskier” model later in the competition. But the ensemble of the 2 provided crucial in gaining the 3rd place.\n\n## TLDR\nBlend of 2 Independent models. The smaller “safer” version model is based on augmented Squeezeformer architecture, which consists of RelativeMultiheadSelfAttention, Convolution and FeedFoward modules. Learnable BPP’s through 2d convolution are added to the attention scores, and augmented low Signal-to-Noise data is additionally used in training apart from clean training data. Bigger twin-tower model is based on augmented Alpha Fold style architecture, that consists of MSA stack representation and Pair stack representation, that communicate in a criss-cross fashion through Outer Product Mean and Pair Representation Bias.\n\n##Data Preprocessing/Cross Validation\n\nTraining data was split in 4 folds similar to this [notebook](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb/notebook). There was a stable correlation between local cv vs lb. BPP’s were processed and cached as .npz files. Most of the time, models were evaluated only on fold0. \nTwin-tower model was trained only on clean training data and it has the power to predict its confidence/error_estimate similar to pLDDT in AlphaFold. After getting a good generalization from the twin-tower model on the clean data, new training data was created by augmenting the noisy low SNR data with the model’s confidence/error_estimate. For a particular nucleotide position in the low SNR dataset the idea is to combine how confident the model's prediction is with the position's reactivity error from the experiment and thus “fix” the noisy data. This gave significant improvements in the smaller “safer” model which was trained on clean dataset + improved low snr dataset. Even with this useful new data being available for the bigger twin-tower model, it was never trained on this extra data because of time constraints and only the clean dataset was used for final submission. Twin-tower experiments are currently underway on the full dataset.\n\n##Squeezeformer Model\nThe smaller “safer” model is based on Squeezeformer architecture which was used in previous Google American Sign Language competition. The inputs to the model were tokenized RNA sequences and BPP matrices. The model consists of 14xSqueezeformer blocks and an output projection layer that predicts chemical reactivities at each position. One Squeezeformer block consists of three modules: Relative MultiHeadSelfAttention module, Convolution module, and FeedForward module.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F55cab60a14c35cb028ecc1de472c6566%2FSqueezeFormerJPG-01.jpg?generation=1702092589333558&alt=media)\n\nInstead of absolute positional encodings, relative encodings are used for generalization on longer sequences. Attention scores apart from relative pos scores are further affected by BasePairProbability matrices, even though I was reluctant to use them at first since they are created with software that is not capable of detecting long range pseudo-knots and this bias is unfortunately induced in the model. The attention scores in the transformer are calculated in this manner: (content score + relative position score)/sqrt(head_dim) + bpp_bias_score. Bpp_bias_score is obtained by passing the BPP matrices through a 2D convolution block.\n\n## AlphaFold style Twin-tower Model\nThis model was inspired from Google's AlphaFold and its derivatives OpenComplex/RhoFold. The original AlphaFold architecture relies on two different input representations for its predictions. It jointly uses MSA(Multiple Sequence Alignment) and Pair Representation features. The MSA representation uses row-wise attention to find intra-sequence features, while the column-wise attention is used to obtain inter-sequence evolutionary signals from the MSA stack. Since MSA for this competition wasn’t helpful in extracting evolutionary information, because of its synthetic nature of the RNA sequences, just the tokenized input sequence was used. Because MSA was not used, the axial-self attention was replaced with relative multi head attention and convolution.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F99a3d1c0bbb83caf86c5a82c2d2a7627%2FChemFoldJPG-01.jpg?generation=1702092825428051&alt=media)\n(At the time of submission, only the Single representation branch was used for predictions, the pair features were completely ignored due to time restrictions)\n\n###The workflow of the data from the input all the way to the prediction is as follows:\n\n1) Input sequence gets tokenized, sequence and pair masks are generated\n\n2) Tokenized sequence are embedded through embedding networks: MSANet and PairNet\n- PairRepresentation features are embedded using Relative2D positional encodings to provide information about position of residues. Maximum position is clipped at 32, and each position afterwards is considered as “far” away. This inductive bias trains the model not to rely heavily on nucleotide positions and generalize better to any length, as stated in the alpha fold supplemental materials.\n- MSA representation is passed through a simple Embedding layer (since positional encodings are added later in the transformer layer).\n\n3) Embedded MSA and Pair features are passed through the main trunk of the network, which consists of 8 Chemformer blocks illustrated above. Some key points are:\n- MSA representation is updated within the Squeezeformer attention \ninstead of axial self-attention proposed in the original paper.\n- Pair representation is updated only with triangular multiplicative updates. Triangular self-attention wasn’t used in this model because of memory constraints but using both should give an increase in score.\n- The MSA representation updates the pair representation through an element-wise outer product that is summed over the MSA sequence dimension.\n- The Pair representation updates the MSA representation through a projection of additional logits from the pair stack to bias the MSA attention scores.\n- Both representations are passed through a 2-layer MLP that acts as transition before the communication.\n- This communication is repeated within each block, so 8 times in total.\n- Residual connections and Row-wise/Column-wise dropouts are used.\n\n4) The processed representations MSA Features and Pair Features are then passed through various heads to extract meaningful information such as confidence/error_estimate, chemical reactivity, base_pair probabilities etc.\n\n## Training\n\nTraining procedure is as follows:\n\nFirst the Twin-tower model is trained on the clean dataset for 60 epochs with lr of 1e-3. Batch size of 8 per gpu was used with gradient accumulation of 4, leading to an effective batch size of 64 on 2 GPUs. The optimizer used was AdamW with cosine scheduling, weight decay parameter is 0.05 and warmup is 0.5. Model was trained on 2x4090 for 30 hours total. No extensive parameter tuning was done for this model.\nThis leads to a 0.13746 public/0.14398 private single model score. After this a synthetic dataset was created by blending the model’s predictions with a subset of noisy data that was between 0.35<data<1 SNR. Selecting a lower SNR threshold increased the training examples but lowered the quality. Current synthetic dataset is created with simple weighing, because of time constraints even though the model was trained and outputs valid plddt/error_estimates, they were not fully utilized. I didn’t have time to experiment with the formula for combining plddt score with experimental reactivity_error on a per nucleotide basis which is guaranteed to produce better synthetic data. \nThe smaller “safer” model was trained for 200 epochs with lr of 7e-4 a batch size of 64. Optimizer and lr scheduling was the same as before, however no grad accumulation was used. This model was trained both on clean dataset and the synthetically created dataset which doubled the training examples. This single model achieves 0.13865 public/0.14256 private score.\n\nA day before the competition ended I decided to just run longer epochs on the Twin-tower model, so I continued the same training procedure with starting weights from epoch 30 of the previous iteration. This gave improvements just by running the model longer which led to the final 0.13706 public/0.14366 private score of this single model. The final score is a blend of the twin-tower model which was trained only on the input sequence and tried to learn all of the interactions on its own and the squeeze model which contributed with the BPP’s and the synthetic data to the final prediction.\n\n\n##Not utilized/Can be improved for Twin-tower model\n\n- Model was trained only with sequences as input. \n- BPP’s and other supplemental data can be used. \n- Synthetic dataset which improved the smaller model should be used.\n- With my tests, by increasing the depth (8->12) blocks, there is an increase in score, submission is only 8 blocks long. \n- Deeper model 16,24 blocks will be tested soon by utilizing checkpoiting/rematerialization to counter the memory problem and interleaving the convolution blocks after several MHSA blocks to aid training for this deeper model. \n- Recycling from the original paper was not utilized.\n- Only triangular multiplicative updates are used in the Pair Representation, triangular self-attention was not utilized.\n- Since RNA can take on multiple conformations, dropout can be utilized at inference time and results averaged to get a better estimate of the RNA’s structure\n\nAll of these adjustments are very likely to provide benefits to the model. Some of the tests are currently underway.\n\n\n##Comments \n\nThe twin-tower model was a large task, partially because I was competing solo. The model at submission time was trained only on clean dataset, and only the MSA Feature pathway was used in the predictions, the other Pair Feature pathway was completely ignored, but it can be used to try and recreate the BPP’s which should increase the score (tests are underway as im writing). Best submission of the model without BPP’s, loop types or any pre/post processing is 0.13709 public and 0.14366 private. However, the twin-tower model has a bigger gap in generalization compared to the smaller “safer” squeezeformer model which scored 0.13865 public but 0.14265 private, so currently investigating the reason behind it.\n\nI am new to machine learning. I started to learn the field in May of this year, so I am sure there will be a lot of mistakes in the code and in my approach and sorry if my explanation is all over the place, all this is new to me and I am still learning.\n\nOpen Sourced Code:\nhttps://github.com/GosUxD/OpenChemFold\n\n## References:\n[1] Squeezeformer: An Efficient Transformer for Automatic Speech Recognition\nSehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W. Mahoney, Kurt Keutzer arXiv:2206.00888 [eess.AS] https://doi.org/10.48550/arXiv.2206.00888\n\n[2] Winner of Google American Sign Language Fingerspelling Competition\nhttps://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution\n\n[3] Jumper, J., Evans, R., Pritzel, A. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). https://doi.org/10.1038/s41586-021-03819-2\n\n[4] OpenComplex github code repository:\nhttps://github.com/baaihealth/OpenComplex\n\n[5] E2Efold-3D: End-to-End Deep Learning Method for accurate de novo RNA 3D Structure Prediction\nTao Shen, Zhihang Hu, Zhangzhi Peng, Jiayang Chen, Peng Xiong, Liang Hong, Liangzhen Zheng, Yixuan Wang, Irwin King, Sheng Wang, Siqi Sun, Yu Li. arXiv:2207.01586 [q-bio.QM] https://doi.org/10.48550/arXiv.2207.01586\n",
      "votes": 24
    },
    {
      "id": 3146160,
      "postDate": "2025-03-10T15:17:53.157Z",
      "content": "<p>Congrats on the great result.</p>\n<p>I'm new to rna folding and I'm trying out the new rna 3d competition.</p>\n<p>I've been trying to understad the pairs representation matrix, but it's hard to find any description of how it's made.</p>\n<p>Could you explain or help with some references?</p>",
      "rawMarkdown": "Congrats on the great result.\n\nI'm new to rna folding and I'm trying out the new rna 3d competition.\n\nI've been trying to understad the pairs representation matrix, but it's hard to find any description of how it's made.\n\nCould you explain or help with some references?"
    },
    {
      "id": 2852291,
      "postDate": "2024-06-03T07:43:12.393Z",
      "content": "<p>Congratulations! A non-contest related question. Which tools did you use to draw the schematic diagram of the model structure?</p>",
      "rawMarkdown": "Congratulations! A non-contest related question. Which tools did you use to draw the schematic diagram of the model structure?",
      "replies": [
        {
          "id": 2853127,
          "postDate": "2024-06-03T15:58:50.197Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2560429,
      "postDate": "2023-12-13T16:16:44.467Z",
      "content": "<p><a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a> I'm wondering if you would be willing to give a presentation on your model and your approach to the competition to the Eterna players? It usually is only a small group and quite informal. Please send me a DM or reply here and I'll give you my email address for details.</p>",
      "rawMarkdown": "@dankrstev I'm wondering if you would be willing to give a presentation on your model and your approach to the competition to the Eterna players? It usually is only a small group and quite informal. Please send me a DM or reply here and I'll give you my email address for details.",
      "replies": [
        {
          "id": 2560698,
          "postDate": "2023-12-13T23:24:19.370Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/digitalembrace\" target=\"_blank\">@digitalembrace</a>, sure lets talk via email</p>",
          "rawMarkdown": "Hey @digitalembrace, sure lets talk via email",
          "replies": [
            {
              "id": 2560760,
              "postDate": "2023-12-14T01:31:35.120Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2554710,
      "postDate": "2023-12-09T10:42:43.513Z",
      "content": "<p>Thanks for the write-up! </p>\n<p>Our team tried recycling [one cycle], however it decreased train loss, but increased val loss compared to the same setup with no recycling.</p>",
      "rawMarkdown": "Thanks for the write-up! \n\nOur team tried recycling [one cycle], however it decreased train loss, but increased val loss compared to the same setup with no recycling.",
      "replies": [
        {
          "id": 2554743,
          "postDate": "2023-12-09T11:16:39.517Z",
          "content": "<p>Hmm that is interesting to know. How did you guys setup the recycling in your model?</p>",
          "rawMarkdown": "Hmm that is interesting to know. How did you guys setup the recycling in your model?",
          "replies": [
            {
              "id": 2554760,
              "postDate": "2023-12-09T11:26:20.807Z",
              "content": "<p>The models we tried recycling on were homogeneous starting from the second transformer block, so for the 24-block model we did it like that<br>\n1) do forward pass from 1st to 24th block [whole model]<br>\n2) take last hidden state, apply forward from 2nd to 24th block [only homogeneous blocks]<br>\n3) apply regression head</p>",
              "rawMarkdown": "The models we tried recycling on were homogeneous starting from the second transformer block, so for the 24-block model we did it like that\n1) do forward pass from 1st to 24th block [whole model]\n2) take last hidden state, apply forward from 2nd to 24th block [only homogeneous blocks]\n3) apply regression head",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2554583,
      "postDate": "2023-12-09T08:18:15.463Z",
      "content": "<p>Have you used Squeezeformer with “squeezing” idea or without, like in the  Google American Sign Language competition? </p>",
      "rawMarkdown": "Have you used Squeezeformer with “squeezing” idea or without, like in the  Google American Sign Language competition? ",
      "replies": [
        {
          "id": 2554703,
          "postDate": "2023-12-09T10:32:27.723Z",
          "content": "<p>The actual \"squeeze\" idea was not used, but the general architecture and blocks</p>",
          "rawMarkdown": "The actual \"squeeze\" idea was not used, but the general architecture and blocks",
          "votes": 1
        }
      ]
    },
    {
      "id": 2554580,
      "postDate": "2023-12-09T08:13:44.773Z",
      "content": "<p>Very interesting solution with very strong results. If I am not mistaken, the loss for the Chem head is a simple mae with the ground trurh, right? However, can you please explain what is the loss in pcdt head, i.e. how exactly does it predict confidence? Thank you. I'm waiting to see your code.</p>",
      "rawMarkdown": "Very interesting solution with very strong results. If I am not mistaken, the loss for the Chem head is a simple mae with the ground trurh, right? However, can you please explain what is the loss in pcdt head, i.e. how exactly does it predict confidence? Thank you. I'm waiting to see your code.",
      "replies": [
        {
          "id": 2554739,
          "postDate": "2023-12-09T11:09:47.130Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/greySnow\" target=\"_blank\">@greySnow</a>, this is the general idea. The final single representation obtained from the MSA Features is ran through the pCDT Head (predicted Chemical Difference Test) which is analogous to the pLDDT (predicted Local Distance Difference Test) in Alpha Fold. </p>\n<p>The pCDT head, projects the final single representation into 50 bins, so (batch, seq_len, 2, 50) where each bin is covering a 0.02 chemical distance difference range between 0-1 [*Testing to lower this range between (0-0.5) because most of the errors the model makes are less than 0.5, so half of the bins are not used most of the time]. </p>\n<p>Next the final chemical reactivity predictions are scored against the ground truth. The L1 difference between the prediction and ground truth is discretized into 50 bins (with values greater than 1 being placed in the final bin) and used as the target value for the cross-entropy loss. Simply stated the model is trying to guess in which of the 50 bins the current error should be placed. </p>\n<p>For evaluation, a pcdt score/confidence percentage is obtained by calculating the expected value of the per-residue pCDT distribution and converted into percentages.</p>",
          "rawMarkdown": "Hey @greySnow, this is the general idea. The final single representation obtained from the MSA Features is ran through the pCDT Head (predicted Chemical Difference Test) which is analogous to the pLDDT (predicted Local Distance Difference Test) in Alpha Fold. \n\nThe pCDT head, projects the final single representation into 50 bins, so (batch, seq_len, 2, 50) where each bin is covering a 0.02 chemical distance difference range between 0-1 [*Testing to lower this range between (0-0.5) because most of the errors the model makes are less than 0.5, so half of the bins are not used most of the time]. \n\nNext the final chemical reactivity predictions are scored against the ground truth. The L1 difference between the prediction and ground truth is discretized into 50 bins (with values greater than 1 being placed in the final bin) and used as the target value for the cross-entropy loss. Simply stated the model is trying to guess in which of the 50 bins the current error should be placed. \n\nFor evaluation, a pcdt score/confidence percentage is obtained by calculating the expected value of the per-residue pCDT distribution and converted into percentages.",
          "votes": 5
        }
      ]
    },
    {
      "id": 2554358,
      "postDate": "2023-12-09T04:51:34.717Z",
      "content": "<p>If it's not a secret, could you tell your affiliation? Very interesting</p>",
      "rawMarkdown": "If it's not a secret, could you tell your affiliation? Very interesting",
      "replies": [
        {
          "id": 2554360,
          "postDate": "2023-12-09T04:58:08.317Z",
          "content": "<p>Affiliation?</p>",
          "rawMarkdown": "Affiliation?",
          "replies": [
            {
              "id": 2554361,
              "postDate": "2023-12-09T05:00:58.913Z",
              "content": "<p>University/Company) </p>",
              "rawMarkdown": "University/Company) ",
              "votes": 1
            },
            {
              "id": 2554699,
              "postDate": "2023-12-09T10:28:45.880Z",
              "content": "<p>i am not working at a company or at any university, just a very curious guy i suppose</p>",
              "rawMarkdown": "i am not working at a company or at any university, just a very curious guy i suppose",
              "votes": 7
            }
          ]
        }
      ]
    },
    {
      "id": 2554356,
      "postDate": "2023-12-09T04:50:36.343Z",
      "content": "<p>Nice explanation.</p>\n<p>About higher generalization gap for larger model - I recommend to try  submitting your model predictions with zeroed 13% sequences occuring in training data. Maybe the issue is with model memorizing this ones (absolute value of public score is a complete mess because ot that and I completely sure that this 13% explains many shake-downs)</p>\n<p>I really like the use of relative positional encoding from the AlphaFold2 paper -- we have failed to take it into consideration and definitely should try it</p>\n<p>Am I right that for larger model you used BPP as one of the targets to be predicted? This is different from \"organizers-training-data-only model\" as information from bpps (actually -- the eternafold parameters used to calculate it) will definitely be used by the model at training time. <br>\nWe also want to try it for our model. <br>\nHave you tried to use that loss in Squeezeformer Model (without providing it with bpps and providing instead with distance or diagonal matrix) and see if it would increase the model performance?</p>\n<p>Also, what is the intuition behind using Squeezeformer? I see many participants used it.. </p>",
      "rawMarkdown": "Nice explanation.\n\nAbout higher generalization gap for larger model - I recommend to try  submitting your model predictions with zeroed 13% sequences occuring in training data. Maybe the issue is with model memorizing this ones (absolute value of public score is a complete mess because ot that and I completely sure that this 13% explains many shake-downs)\n\nI really like the use of relative positional encoding from the AlphaFold2 paper -- we have failed to take it into consideration and definitely should try it\n\nAm I right that for larger model you used BPP as one of the targets to be predicted? This is different from \"organizers-training-data-only model\" as information from bpps (actually -- the eternafold parameters used to calculate it) will definitely be used by the model at training time. \nWe also want to try it for our model. \nHave you tried to use that loss in Squeezeformer Model (without providing it with bpps and providing instead with distance or diagonal matrix) and see if it would increase the model performance?\n\nAlso, what is the intuition behind using Squeezeformer? I see many participants used it.. ",
      "replies": [
        {
          "id": 2554363,
          "postDate": "2023-12-09T05:05:49.280Z",
          "content": "<blockquote>\n  <p>Am I right that for larger model you used BPP as one of the targets to be predicted?</p>\n</blockquote>\n<p>The submission was with nothing other than the sequences. No BPPs were used as inputs, or targets. Pair Features branch was completely cut off and BPP pipeline wasn't even implemented in the final submission of the model. <br>\nBut currently testing the new version with BPP's as the models target. Will publish results once they are ready.</p>\n<blockquote>\n  <p>Have you tried to use that loss in Squeezeformer Model?</p>\n</blockquote>\n<p>Didn't work as much on the smaller model, it was just a back up plan. But it is an interesting experiment how the results will vary!</p>",
          "rawMarkdown": ">Am I right that for larger model you used BPP as one of the targets to be predicted?\n\nThe submission was with nothing other than the sequences. No BPPs were used as inputs, or targets. Pair Features branch was completely cut off and BPP pipeline wasn't even implemented in the final submission of the model. \nBut currently testing the new version with BPP's as the models target. Will publish results once they are ready.\n\n>Have you tried to use that loss in Squeezeformer Model?\n\nDidn't work as much on the smaller model, it was just a back up plan. But it is an interesting experiment how the results will vary!",
          "votes": 3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3146160,
      "author_name": "CJ ROCKBALL",
      "author_url": "",
      "post_date": "2025-03-10T15:17:53.157000",
      "content": "<p>Congrats on the great result.</p>\n<p>I'm new to rna folding and I'm trying out the new rna 3d competition.</p>\n<p>I've been trying to understad the pairs representation matrix, but it's hard to find any description of how it's made.</p>\n<p>Could you explain or help with some references?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2852291,
      "author_name": "DECEM",
      "author_url": "",
      "post_date": "2024-06-03T07:43:12.393000",
      "content": "<p>Congratulations! A non-contest related question. Which tools did you use to draw the schematic diagram of the model structure?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2853127,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-06-03T15:58:50.197000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2560429,
      "author_name": "DigitalEmbrace",
      "author_url": "",
      "post_date": "2023-12-13T16:16:44.467000",
      "content": "<p><a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a> I'm wondering if you would be willing to give a presentation on your model and your approach to the competition to the Eterna players? It usually is only a small group and quite informal. Please send me a DM or reply here and I'll give you my email address for details.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2560698,
          "author_name": "dan4o",
          "author_url": "",
          "post_date": "2023-12-13T23:24:19.370000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/digitalembrace\" target=\"_blank\">@digitalembrace</a>, sure lets talk via email</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2560760,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-12-14T01:31:35.120000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2554710,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2023-12-09T10:42:43.513000",
      "content": "<p>Thanks for the write-up! </p>\n<p>Our team tried recycling [one cycle], however it decreased train loss, but increased val loss compared to the same setup with no recycling.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2554743,
          "author_name": "dan4o",
          "author_url": "",
          "post_date": "2023-12-09T11:16:39.517000",
          "content": "<p>Hmm that is interesting to know. How did you guys setup the recycling in your model?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2554760,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2023-12-09T11:26:20.807000",
              "content": "<p>The models we tried recycling on were homogeneous starting from the second transformer block, so for the 24-block model we did it like that<br>\n1) do forward pass from 1st to 24th block [whole model]<br>\n2) take last hidden state, apply forward from 2nd to 24th block [only homogeneous blocks]<br>\n3) apply regression head</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2554583,
      "author_name": "Penzar Dmitry",
      "author_url": "",
      "post_date": "2023-12-09T08:18:15.463000",
      "content": "<p>Have you used Squeezeformer with “squeezing” idea or without, like in the  Google American Sign Language competition? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2554703,
          "author_name": "dan4o",
          "author_url": "",
          "post_date": "2023-12-09T10:32:27.723000",
          "content": "<p>The actual \"squeeze\" idea was not used, but the general architecture and blocks</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2554580,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2023-12-09T08:13:44.773000",
      "content": "<p>Very interesting solution with very strong results. If I am not mistaken, the loss for the Chem head is a simple mae with the ground trurh, right? However, can you please explain what is the loss in pcdt head, i.e. how exactly does it predict confidence? Thank you. I'm waiting to see your code.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2554739,
          "author_name": "dan4o",
          "author_url": "",
          "post_date": "2023-12-09T11:09:47.130000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/greySnow\" target=\"_blank\">@greySnow</a>, this is the general idea. The final single representation obtained from the MSA Features is ran through the pCDT Head (predicted Chemical Difference Test) which is analogous to the pLDDT (predicted Local Distance Difference Test) in Alpha Fold. </p>\n<p>The pCDT head, projects the final single representation into 50 bins, so (batch, seq_len, 2, 50) where each bin is covering a 0.02 chemical distance difference range between 0-1 [*Testing to lower this range between (0-0.5) because most of the errors the model makes are less than 0.5, so half of the bins are not used most of the time]. </p>\n<p>Next the final chemical reactivity predictions are scored against the ground truth. The L1 difference between the prediction and ground truth is discretized into 50 bins (with values greater than 1 being placed in the final bin) and used as the target value for the cross-entropy loss. Simply stated the model is trying to guess in which of the 50 bins the current error should be placed. </p>\n<p>For evaluation, a pcdt score/confidence percentage is obtained by calculating the expected value of the per-residue pCDT distribution and converted into percentages.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 2554358,
      "author_name": "Penzar Dmitry",
      "author_url": "",
      "post_date": "2023-12-09T04:51:34.717000",
      "content": "<p>If it's not a secret, could you tell your affiliation? Very interesting</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2554360,
          "author_name": "dan4o",
          "author_url": "",
          "post_date": "2023-12-09T04:58:08.317000",
          "content": "<p>Affiliation?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2554361,
              "author_name": "Penzar Dmitry",
              "author_url": "",
              "post_date": "2023-12-09T05:00:58.913000",
              "content": "<p>University/Company) </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2554699,
              "author_name": "dan4o",
              "author_url": "",
              "post_date": "2023-12-09T10:28:45.880000",
              "content": "<p>i am not working at a company or at any university, just a very curious guy i suppose</p>",
              "votes": 7,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2554356,
      "author_name": "Penzar Dmitry",
      "author_url": "",
      "post_date": "2023-12-09T04:50:36.343000",
      "content": "<p>Nice explanation.</p>\n<p>About higher generalization gap for larger model - I recommend to try  submitting your model predictions with zeroed 13% sequences occuring in training data. Maybe the issue is with model memorizing this ones (absolute value of public score is a complete mess because ot that and I completely sure that this 13% explains many shake-downs)</p>\n<p>I really like the use of relative positional encoding from the AlphaFold2 paper -- we have failed to take it into consideration and definitely should try it</p>\n<p>Am I right that for larger model you used BPP as one of the targets to be predicted? This is different from \"organizers-training-data-only model\" as information from bpps (actually -- the eternafold parameters used to calculate it) will definitely be used by the model at training time. <br>\nWe also want to try it for our model. <br>\nHave you tried to use that loss in Squeezeformer Model (without providing it with bpps and providing instead with distance or diagonal matrix) and see if it would increase the model performance?</p>\n<p>Also, what is the intuition behind using Squeezeformer? I see many participants used it.. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2554363,
          "author_name": "dan4o",
          "author_url": "",
          "post_date": "2023-12-09T05:05:49.280000",
          "content": "<blockquote>\n  <p>Am I right that for larger model you used BPP as one of the targets to be predicted?</p>\n</blockquote>\n<p>The submission was with nothing other than the sequences. No BPPs were used as inputs, or targets. Pair Features branch was completely cut off and BPP pipeline wasn't even implemented in the final submission of the model. <br>\nBut currently testing the new version with BPP's as the models target. Will publish results once they are ready.</p>\n<blockquote>\n  <p>Have you tried to use that loss in Squeezeformer Model?</p>\n</blockquote>\n<p>Didn't work as much on the smaller model, it was just a back up plan. But it is an interesting experiment how the results will vary!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2554334": "A big thanks to the organizers for making this competition happen and putting their efforts in solving hard problems, such as RNA structure prediction. This sentence stuck with me throughout the whole competition: “... without being able to understand how RNA molecules fold, we are missing a deeper **understanding of how nature works, how life began**, and how we can design…“\n\nMany will glance over this statement but it is a huge mystery that sometimes kept me awake at night. I mean it is SO strange… the evolution/creation of 4 nucleotides with specific physical/chemical properties that can code and take on different structures to fulfill specific actions in the body, and we still don’t understand the origin/purpose of these molecules… Hopefully we are in the right direction in finding out the truth and reason behind it or even “what” created them. \n\nSolution is not based on dozens of ensemble models, but rather 2 strong independent models. There are 2 models since I wasn’t sure if I will be able to produce an Alpha fold style viable solution in the allotted time, so I focused on a smaller “safe” version model first and a bigger “riskier” model later in the competition. But the ensemble of the 2 provided crucial in gaining the 3rd place.\n\n## TLDR\nBlend of 2 Independent models. The smaller “safer” version model is based on augmented Squeezeformer architecture, which consists of RelativeMultiheadSelfAttention, Convolution and FeedFoward modules. Learnable BPP’s through 2d convolution are added to the attention scores, and augmented low Signal-to-Noise data is additionally used in training apart from clean training data. Bigger twin-tower model is based on augmented Alpha Fold style architecture, that consists of MSA stack representation and Pair stack representation, that communicate in a criss-cross fashion through Outer Product Mean and Pair Representation Bias.\n\n##Data Preprocessing/Cross Validation\n\nTraining data was split in 4 folds similar to this [notebook](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb/notebook). There was a stable correlation between local cv vs lb. BPP’s were processed and cached as .npz files. Most of the time, models were evaluated only on fold0. \nTwin-tower model was trained only on clean training data and it has the power to predict its confidence/error_estimate similar to pLDDT in AlphaFold. After getting a good generalization from the twin-tower model on the clean data, new training data was created by augmenting the noisy low SNR data with the model’s confidence/error_estimate. For a particular nucleotide position in the low SNR dataset the idea is to combine how confident the model's prediction is with the position's reactivity error from the experiment and thus “fix” the noisy data. This gave significant improvements in the smaller “safer” model which was trained on clean dataset + improved low snr dataset. Even with this useful new data being available for the bigger twin-tower model, it was never trained on this extra data because of time constraints and only the clean dataset was used for final submission. Twin-tower experiments are currently underway on the full dataset.\n\n##Squeezeformer Model\nThe smaller “safer” model is based on Squeezeformer architecture which was used in previous Google American Sign Language competition. The inputs to the model were tokenized RNA sequences and BPP matrices. The model consists of 14xSqueezeformer blocks and an output projection layer that predicts chemical reactivities at each position. One Squeezeformer block consists of three modules: Relative MultiHeadSelfAttention module, Convolution module, and FeedForward module.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F55cab60a14c35cb028ecc1de472c6566%2FSqueezeFormerJPG-01.jpg?generation=1702092589333558&alt=media)\n\nInstead of absolute positional encodings, relative encodings are used for generalization on longer sequences. Attention scores apart from relative pos scores are further affected by BasePairProbability matrices, even though I was reluctant to use them at first since they are created with software that is not capable of detecting long range pseudo-knots and this bias is unfortunately induced in the model. The attention scores in the transformer are calculated in this manner: (content score + relative position score)/sqrt(head_dim) + bpp_bias_score. Bpp_bias_score is obtained by passing the BPP matrices through a 2D convolution block.\n\n## AlphaFold style Twin-tower Model\nThis model was inspired from Google's AlphaFold and its derivatives OpenComplex/RhoFold. The original AlphaFold architecture relies on two different input representations for its predictions. It jointly uses MSA(Multiple Sequence Alignment) and Pair Representation features. The MSA representation uses row-wise attention to find intra-sequence features, while the column-wise attention is used to obtain inter-sequence evolutionary signals from the MSA stack. Since MSA for this competition wasn’t helpful in extracting evolutionary information, because of its synthetic nature of the RNA sequences, just the tokenized input sequence was used. Because MSA was not used, the axial-self attention was replaced with relative multi head attention and convolution.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F99a3d1c0bbb83caf86c5a82c2d2a7627%2FChemFoldJPG-01.jpg?generation=1702092825428051&alt=media)\n(At the time of submission, only the Single representation branch was used for predictions, the pair features were completely ignored due to time restrictions)\n\n###The workflow of the data from the input all the way to the prediction is as follows:\n\n1) Input sequence gets tokenized, sequence and pair masks are generated\n\n2) Tokenized sequence are embedded through embedding networks: MSANet and PairNet\n- PairRepresentation features are embedded using Relative2D positional encodings to provide information about position of residues. Maximum position is clipped at 32, and each position afterwards is considered as “far” away. This inductive bias trains the model not to rely heavily on nucleotide positions and generalize better to any length, as stated in the alpha fold supplemental materials.\n- MSA representation is passed through a simple Embedding layer (since positional encodings are added later in the transformer layer).\n\n3) Embedded MSA and Pair features are passed through the main trunk of the network, which consists of 8 Chemformer blocks illustrated above. Some key points are:\n- MSA representation is updated within the Squeezeformer attention \ninstead of axial self-attention proposed in the original paper.\n- Pair representation is updated only with triangular multiplicative updates. Triangular self-attention wasn’t used in this model because of memory constraints but using both should give an increase in score.\n- The MSA representation updates the pair representation through an element-wise outer product that is summed over the MSA sequence dimension.\n- The Pair representation updates the MSA representation through a projection of additional logits from the pair stack to bias the MSA attention scores.\n- Both representations are passed through a 2-layer MLP that acts as transition before the communication.\n- This communication is repeated within each block, so 8 times in total.\n- Residual connections and Row-wise/Column-wise dropouts are used.\n\n4) The processed representations MSA Features and Pair Features are then passed through various heads to extract meaningful information such as confidence/error_estimate, chemical reactivity, base_pair probabilities etc.\n\n## Training\n\nTraining procedure is as follows:\n\nFirst the Twin-tower model is trained on the clean dataset for 60 epochs with lr of 1e-3. Batch size of 8 per gpu was used with gradient accumulation of 4, leading to an effective batch size of 64 on 2 GPUs. The optimizer used was AdamW with cosine scheduling, weight decay parameter is 0.05 and warmup is 0.5. Model was trained on 2x4090 for 30 hours total. No extensive parameter tuning was done for this model.\nThis leads to a 0.13746 public/0.14398 private single model score. After this a synthetic dataset was created by blending the model’s predictions with a subset of noisy data that was between 0.35<data<1 SNR. Selecting a lower SNR threshold increased the training examples but lowered the quality. Current synthetic dataset is created with simple weighing, because of time constraints even though the model was trained and outputs valid plddt/error_estimates, they were not fully utilized. I didn’t have time to experiment with the formula for combining plddt score with experimental reactivity_error on a per nucleotide basis which is guaranteed to produce better synthetic data. \nThe smaller “safer” model was trained for 200 epochs with lr of 7e-4 a batch size of 64. Optimizer and lr scheduling was the same as before, however no grad accumulation was used. This model was trained both on clean dataset and the synthetically created dataset which doubled the training examples. This single model achieves 0.13865 public/0.14256 private score.\n\nA day before the competition ended I decided to just run longer epochs on the Twin-tower model, so I continued the same training procedure with starting weights from epoch 30 of the previous iteration. This gave improvements just by running the model longer which led to the final 0.13706 public/0.14366 private score of this single model. The final score is a blend of the twin-tower model which was trained only on the input sequence and tried to learn all of the interactions on its own and the squeeze model which contributed with the BPP’s and the synthetic data to the final prediction.\n\n\n##Not utilized/Can be improved for Twin-tower model\n\n- Model was trained only with sequences as input. \n- BPP’s and other supplemental data can be used. \n- Synthetic dataset which improved the smaller model should be used.\n- With my tests, by increasing the depth (8->12) blocks, there is an increase in score, submission is only 8 blocks long. \n- Deeper model 16,24 blocks will be tested soon by utilizing checkpoiting/rematerialization to counter the memory problem and interleaving the convolution blocks after several MHSA blocks to aid training for this deeper model. \n- Recycling from the original paper was not utilized.\n- Only triangular multiplicative updates are used in the Pair Representation, triangular self-attention was not utilized.\n- Since RNA can take on multiple conformations, dropout can be utilized at inference time and results averaged to get a better estimate of the RNA’s structure\n\nAll of these adjustments are very likely to provide benefits to the model. Some of the tests are currently underway.\n\n\n##Comments \n\nThe twin-tower model was a large task, partially because I was competing solo. The model at submission time was trained only on clean dataset, and only the MSA Feature pathway was used in the predictions, the other Pair Feature pathway was completely ignored, but it can be used to try and recreate the BPP’s which should increase the score (tests are underway as im writing). Best submission of the model without BPP’s, loop types or any pre/post processing is 0.13709 public and 0.14366 private. However, the twin-tower model has a bigger gap in generalization compared to the smaller “safer” squeezeformer model which scored 0.13865 public but 0.14265 private, so currently investigating the reason behind it.\n\nI am new to machine learning. I started to learn the field in May of this year, so I am sure there will be a lot of mistakes in the code and in my approach and sorry if my explanation is all over the place, all this is new to me and I am still learning.\n\nOpen Sourced Code:\nhttps://github.com/GosUxD/OpenChemFold\n\n## References:\n[1] Squeezeformer: An Efficient Transformer for Automatic Speech Recognition\nSehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W. Mahoney, Kurt Keutzer arXiv:2206.00888 [eess.AS] https://doi.org/10.48550/arXiv.2206.00888\n\n[2] Winner of Google American Sign Language Fingerspelling Competition\nhttps://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution\n\n[3] Jumper, J., Evans, R., Pritzel, A. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). https://doi.org/10.1038/s41586-021-03819-2\n\n[4] OpenComplex github code repository:\nhttps://github.com/baaihealth/OpenComplex\n\n[5] E2Efold-3D: End-to-End Deep Learning Method for accurate de novo RNA 3D Structure Prediction\nTao Shen, Zhihang Hu, Zhangzhi Peng, Jiayang Chen, Peng Xiong, Liang Hong, Liangzhen Zheng, Yixuan Wang, Irwin King, Sheng Wang, Siqi Sun, Yu Li. arXiv:2207.01586 [q-bio.QM] https://doi.org/10.48550/arXiv.2207.01586\n",
    "3146160": "Congrats on the great result.\n\nI'm new to rna folding and I'm trying out the new rna 3d competition.\n\nI've been trying to understad the pairs representation matrix, but it's hard to find any description of how it's made.\n\nCould you explain or help with some references?",
    "2852291": "Congratulations! A non-contest related question. Which tools did you use to draw the schematic diagram of the model structure?",
    "2560429": "@dankrstev I'm wondering if you would be willing to give a presentation on your model and your approach to the competition to the Eterna players? It usually is only a small group and quite informal. Please send me a DM or reply here and I'll give you my email address for details.",
    "2554710": "Thanks for the write-up! \n\nOur team tried recycling [one cycle], however it decreased train loss, but increased val loss compared to the same setup with no recycling.",
    "2554583": "Have you used Squeezeformer with “squeezing” idea or without, like in the  Google American Sign Language competition? ",
    "2554580": "Very interesting solution with very strong results. If I am not mistaken, the loss for the Chem head is a simple mae with the ground trurh, right? However, can you please explain what is the loss in pcdt head, i.e. how exactly does it predict confidence? Thank you. I'm waiting to see your code.",
    "2554358": "If it's not a secret, could you tell your affiliation? Very interesting",
    "2554356": "Nice explanation.\n\nAbout higher generalization gap for larger model - I recommend to try  submitting your model predictions with zeroed 13% sequences occuring in training data. Maybe the issue is with model memorizing this ones (absolute value of public score is a complete mess because ot that and I completely sure that this 13% explains many shake-downs)\n\nI really like the use of relative positional encoding from the AlphaFold2 paper -- we have failed to take it into consideration and definitely should try it\n\nAm I right that for larger model you used BPP as one of the targets to be predicted? This is different from \"organizers-training-data-only model\" as information from bpps (actually -- the eternafold parameters used to calculate it) will definitely be used by the model at training time. \nWe also want to try it for our model. \nHave you tried to use that loss in Squeezeformer Model (without providing it with bpps and providing instead with distance or diagonal matrix) and see if it would increase the model performance?\n\nAlso, what is the intuition behind using Squeezeformer? I see many participants used it.. "
  }
}