{
  "id": 460315,
  "title": "44th Place Solution for the Stanford Ribonanza RNA Folding Competition",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460315",
  "author_name": "Leandro Bugnon",
  "post_date": "2023-12-08T17:05:51.684000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>Context section</h2>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/overview\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data</a></li>\n</ul>\n<h2>Overview of the Approach</h2>\n<p>This is my first write-up (also the first time in the medal section 😁). Thanks to the organizers and participants! My approach was based on work we are doing in our lab for RNA secondary structure prediction [1] (this article includes source code). This is a  ResNet-based model that takes the input sequences as one-hot representations and outputs the interaction matrix prediction (similar to the bpp provided). The model has a first 1D feature extraction stage, then it is converted to 2D with a simple matmul, and undergoes a second stage of 2D feature extraction to reach final prediction. For this competition, the output was \"flattened\" using a sum by columns, thus arriving to a representation of the activation of each nucleotide in the sequence. The inner 2D representation allow us to add bpp information as an additional channel to learn features from.</p>\n<p>In terms of data it was rather a simple approach, the best model used the training data with SNR&gt;0.5. Train/test splits were performed using clustered sequences using cdhit-est (80% threshold) [2]. This is important to not overfit patterns on similar sequences. Provided BPP were also used. </p>\n<p><img src=\"https://github.com/sinc-lab/sincFold/blob/main/abstract.png?raw=true\" alt=\"\"></p>\n<h2>Details of the submission</h2>\n<p>I wanted to try different approaches but focused on improving the one detailed above. As training was time consuming, using only the medioids of each sequence cluster proved to reach competitive results (compared to my best submission) thus it was the approach during development. It would be nice to filter leaked test sequences but it seems that my best solution in public LB is the same in private, so good news for the model. Final results were averaging 5 models, 3 using only medioids because I ran out of time. Also didn't use test data for anything except looking at the public LB </p>\n<p>An interesting take is the use of BPP information. As it is based on classical RNA structure prediction methods and do not model pseudoknots, it could affect model prediction in those cases. I trained models without BPP reaching to far worse avg results (public LB ~.16). I did not tried in private yet, but analyzing a sample case as described in [3] it can be seen that the final ensemble model (middle image) miss the predictions pointed by the arrow in the reference prediction (upper image), while the model without BPP (bottom image) have some resemblance. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F208067%2F77df07e5804b090f0301093fe0dceb73%2Fplot_reffinalsinbpp.png?generation=1702052560298402&amp;alt=media\" alt=\"\"></p>\n<p>Things that didn't work: </p>\n<ul>\n<li>Tried white noise and flip augmentation</li>\n<li>Using sequences with less than .5 snr</li>\n<li>Tried to use errors during training but didn't reach to a converging method. It could work though </li>\n</ul>\n<p>Hope we see more works on bio sequences!</p>\n<h2>Sources</h2>\n<p>[1] <a href=\"https://www.biorxiv.org/content/10.1101/2023.10.10.561771v1\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2023.10.10.561771v1</a><br>\n[2] <a href=\"https://sites.google.com/view/cd-hit\" target=\"_blank\">https://sites.google.com/view/cd-hit</a><br>\n[3] <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653</a></p>",
  "messages": [
    {
      "id": 2553978,
      "postDate": "2023-12-08T17:05:51.683Z",
      "content": "<h2>Context section</h2>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/overview\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data</a></li>\n</ul>\n<h2>Overview of the Approach</h2>\n<p>This is my first write-up (also the first time in the medal section 😁). Thanks to the organizers and participants! My approach was based on work we are doing in our lab for RNA secondary structure prediction [1] (this article includes source code). This is a  ResNet-based model that takes the input sequences as one-hot representations and outputs the interaction matrix prediction (similar to the bpp provided). The model has a first 1D feature extraction stage, then it is converted to 2D with a simple matmul, and undergoes a second stage of 2D feature extraction to reach final prediction. For this competition, the output was \"flattened\" using a sum by columns, thus arriving to a representation of the activation of each nucleotide in the sequence. The inner 2D representation allow us to add bpp information as an additional channel to learn features from.</p>\n<p>In terms of data it was rather a simple approach, the best model used the training data with SNR&gt;0.5. Train/test splits were performed using clustered sequences using cdhit-est (80% threshold) [2]. This is important to not overfit patterns on similar sequences. Provided BPP were also used. </p>\n<p><img src=\"https://github.com/sinc-lab/sincFold/blob/main/abstract.png?raw=true\" alt=\"\"></p>\n<h2>Details of the submission</h2>\n<p>I wanted to try different approaches but focused on improving the one detailed above. As training was time consuming, using only the medioids of each sequence cluster proved to reach competitive results (compared to my best submission) thus it was the approach during development. It would be nice to filter leaked test sequences but it seems that my best solution in public LB is the same in private, so good news for the model. Final results were averaging 5 models, 3 using only medioids because I ran out of time. Also didn't use test data for anything except looking at the public LB </p>\n<p>An interesting take is the use of BPP information. As it is based on classical RNA structure prediction methods and do not model pseudoknots, it could affect model prediction in those cases. I trained models without BPP reaching to far worse avg results (public LB ~.16). I did not tried in private yet, but analyzing a sample case as described in [3] it can be seen that the final ensemble model (middle image) miss the predictions pointed by the arrow in the reference prediction (upper image), while the model without BPP (bottom image) have some resemblance. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F208067%2F77df07e5804b090f0301093fe0dceb73%2Fplot_reffinalsinbpp.png?generation=1702052560298402&amp;alt=media\" alt=\"\"></p>\n<p>Things that didn't work: </p>\n<ul>\n<li>Tried white noise and flip augmentation</li>\n<li>Using sequences with less than .5 snr</li>\n<li>Tried to use errors during training but didn't reach to a converging method. It could work though </li>\n</ul>\n<p>Hope we see more works on bio sequences!</p>\n<h2>Sources</h2>\n<p>[1] <a href=\"https://www.biorxiv.org/content/10.1101/2023.10.10.561771v1\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2023.10.10.561771v1</a><br>\n[2] <a href=\"https://sites.google.com/view/cd-hit\" target=\"_blank\">https://sites.google.com/view/cd-hit</a><br>\n[3] <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653</a></p>",
      "rawMarkdown": "## Context section\n- Business context: https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/overview\n- Data context: https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\n\n## Overview of the Approach\n\nThis is my first write-up (also the first time in the medal section 😁). Thanks to the organizers and participants! My approach was based on work we are doing in our lab for RNA secondary structure prediction [1] (this article includes source code). This is a  ResNet-based model that takes the input sequences as one-hot representations and outputs the interaction matrix prediction (similar to the bpp provided). The model has a first 1D feature extraction stage, then it is converted to 2D with a simple matmul, and undergoes a second stage of 2D feature extraction to reach final prediction. For this competition, the output was \"flattened\" using a sum by columns, thus arriving to a representation of the activation of each nucleotide in the sequence. The inner 2D representation allow us to add bpp information as an additional channel to learn features from.\n\nIn terms of data it was rather a simple approach, the best model used the training data with SNR>0.5. Train/test splits were performed using clustered sequences using cdhit-est (80% threshold) [2]. This is important to not overfit patterns on similar sequences. Provided BPP were also used. \n\n\n![](https://github.com/sinc-lab/sincFold/blob/main/abstract.png?raw=true)\n\n## Details of the submission\n\nI wanted to try different approaches but focused on improving the one detailed above. As training was time consuming, using only the medioids of each sequence cluster proved to reach competitive results (compared to my best submission) thus it was the approach during development. It would be nice to filter leaked test sequences but it seems that my best solution in public LB is the same in private, so good news for the model. Final results were averaging 5 models, 3 using only medioids because I ran out of time. Also didn't use test data for anything except looking at the public LB \n\nAn interesting take is the use of BPP information. As it is based on classical RNA structure prediction methods and do not model pseudoknots, it could affect model prediction in those cases. I trained models without BPP reaching to far worse avg results (public LB ~.16). I did not tried in private yet, but analyzing a sample case as described in [3] it can be seen that the final ensemble model (middle image) miss the predictions pointed by the arrow in the reference prediction (upper image), while the model without BPP (bottom image) have some resemblance. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F208067%2F77df07e5804b090f0301093fe0dceb73%2Fplot_reffinalsinbpp.png?generation=1702052560298402&alt=media)\n\nThings that didn't work: \n- Tried white noise and flip augmentation\n- Using sequences with less than .5 snr\n- Tried to use errors during training but didn't reach to a converging method. It could work though \n\nHope we see more works on bio sequences!\n\n## Sources\n[1] https://www.biorxiv.org/content/10.1101/2023.10.10.561771v1\n[2] https://sites.google.com/view/cd-hit\n[3] https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653\n",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2553978": "## Context section\n- Business context: https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/overview\n- Data context: https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\n\n## Overview of the Approach\n\nThis is my first write-up (also the first time in the medal section 😁). Thanks to the organizers and participants! My approach was based on work we are doing in our lab for RNA secondary structure prediction [1] (this article includes source code). This is a  ResNet-based model that takes the input sequences as one-hot representations and outputs the interaction matrix prediction (similar to the bpp provided). The model has a first 1D feature extraction stage, then it is converted to 2D with a simple matmul, and undergoes a second stage of 2D feature extraction to reach final prediction. For this competition, the output was \"flattened\" using a sum by columns, thus arriving to a representation of the activation of each nucleotide in the sequence. The inner 2D representation allow us to add bpp information as an additional channel to learn features from.\n\nIn terms of data it was rather a simple approach, the best model used the training data with SNR>0.5. Train/test splits were performed using clustered sequences using cdhit-est (80% threshold) [2]. This is important to not overfit patterns on similar sequences. Provided BPP were also used. \n\n\n![](https://github.com/sinc-lab/sincFold/blob/main/abstract.png?raw=true)\n\n## Details of the submission\n\nI wanted to try different approaches but focused on improving the one detailed above. As training was time consuming, using only the medioids of each sequence cluster proved to reach competitive results (compared to my best submission) thus it was the approach during development. It would be nice to filter leaked test sequences but it seems that my best solution in public LB is the same in private, so good news for the model. Final results were averaging 5 models, 3 using only medioids because I ran out of time. Also didn't use test data for anything except looking at the public LB \n\nAn interesting take is the use of BPP information. As it is based on classical RNA structure prediction methods and do not model pseudoknots, it could affect model prediction in those cases. I trained models without BPP reaching to far worse avg results (public LB ~.16). I did not tried in private yet, but analyzing a sample case as described in [3] it can be seen that the final ensemble model (middle image) miss the predictions pointed by the arrow in the reference prediction (upper image), while the model without BPP (bottom image) have some resemblance. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F208067%2F77df07e5804b090f0301093fe0dceb73%2Fplot_reffinalsinbpp.png?generation=1702052560298402&alt=media)\n\nThings that didn't work: \n- Tried white noise and flip augmentation\n- Using sequences with less than .5 snr\n- Tried to use errors during training but didn't reach to a converging method. It could work though \n\nHope we see more works on bio sequences!\n\n## Sources\n[1] https://www.biorxiv.org/content/10.1101/2023.10.10.561771v1\n[2] https://sites.google.com/view/cd-hit\n[3] https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653\n"
  }
}