{
  "id": 460301,
  "title": "Host solution (RNAdegformer public/private 0.1437/0.1458) and some insights",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460301",
  "author_name": "Shujun",
  "post_date": "2023-12-08T16:12:41.237000",
  "votes": 23,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi Kagglers, it has been fun to be behind the scenes this time as a host and I’m glad to see many top solutions posted already. Here I’d like to share my solution and some other insights. </p>\n<h1>Best MAE model</h1>\n<p>My best MAE model is a fold single fold RNAdegformer (published at <a href=\"https://academic.oup.com/bib/article/24/1/bbac581/6986359\" target=\"_blank\">https://academic.oup.com/bib/article/24/1/bbac581/6986359</a>) that scores<br>\nPublic: 0.1437<br>\nPrivate: 0.1458<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F535628c4e75489ab3c6529213f7aba30%2Frnadegformer.png?generation=1702051686843214&amp;alt=media\" alt=\"\"></p>\n<p>It takes bpp and sequence as features and is nearly exactly the same as what I describe in the paper with some small changes (fewer heads, only inverse distance matrix as pos encoding, and no MFE/structure/bpRNA features). If only using train data where both 2A3/DMS SN&gt;1, this model scores Private: 0.1471 Public: 0.1450 and later on I trained another one that uses all profiles with SN&gt;1 (even if one of 2A3/DMS has SN&lt;=1) while masking SN &lt;1 profiles (one of 2A3/DMS), it scores Public: 0.1437\nPrivate: 0.1458. Only using SN&gt;1 takes about 3 hours while using  all profiles with SN&gt;1 takes 5 hours on 2x3090. </p>\n<p>Note that the difference between public/private is quite small in my case. From a modeling standpoint, I think this is due to adding bpp into self-attention as a relatively large bias (I set gamma=32 in my model) results in a sparse attention pattern where row-wise entropy is low and therefore generalizes well to longer sequences. But from another perspective, it’s probably also due to the fact that I’m not a competitor and can see the private data. </p>\n<p>From this point on, I will refer to this model as the bpp model. <br>\nI also have another model where I simply took out bpp and only use sequence as input (referred to as sequence only model) that scores at best Public: 0.1496 Private: 0.1507. This model was more or less trained to figure out if we have enough data that allows the model to not rely on bpp and outperform the bpp model. It turns out we’re not there yet. </p>\n<h1>Flip augmentation</h1>\n<p>Flip augmentation is very useful for me, for the bpp model it has around 0.003 boost, while  for the sequence only model, the boost is much bigger, about 0.01. I’ve seen some competitors say flip aug does not help but I’m not sure if they are doing it in the exact same way as me, which is</p>\n<ol>\n<li>during training: 50% of the time, flip sequence/labels/bpp</li>\n<li>during inference: do TTA, where pred=(model(input)+flip(model(flip(inputs)))/2</li>\n</ol>\n<h1>Positional encoding</h1>\n<p>Positional encoding is important for generalization. Generally, you want to use a relative positional encoding that doesn’t distinguish positions at long ranges. I use a 2D convolved inverse distance matrix as an an attention bias, stacked with bpp (for bpp model), but there are other ways to do it as mentioned in other top solutions.  </p>\n<h1>About pseudo-knots</h1>\n<p>While the bpp model scores well, there are some issues. Notably, early on when <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> inspected the M2 (mutate and map) predictions of some the pseudo-knotted design sequences, the bpp model is completely missing the pseudo-knots (denoted by “[]”), which are hallmarks of RNA 3D structure and not predicted by eternafold. This is not good, because if we want to use models in this competition to predict 3D structures, missing the pseudo-knots can mean predicting entirely wrong 3D folding patterns.  </p>\n<p>However, interestingly, the sequence only model, does recognize the pseudo-knots better. In the plot below you can see that although the bpp model scores better on MAE, it misses the pseudo-knot (circled in red). <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F13ff17d2393757e8bf904087bfc220c5%2Fslide1.png?generation=1702051754381827&amp;alt=media\" alt=\"\"></p>\n<p>But sometimes the bpp model does predict pseudo knots, it’s actually just more conservative overall on pseudo-knots.  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2Fd4083dc420823458030436ad1396d27e%2Fslide2.png?generation=1702051803175730&amp;alt=media\" alt=\"\"></p>\n<h1>long sequence generalization</h1>\n<p>Early on, we saw some competitors’ predictions on the sequence in in this post <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653</a> look unreasonable, which I thought was due to using absolute positional encoding that generalizes poorly beyond length longer than is available in train. A lot of discussion happened there and I think it served as a good sanity check for competitors. The picture I used is actually a 50/50 ensemble of the bpp model and the sequence model, which combines the bpp model’s strength in predicting nested secondary structure and the sequence model’s ability to predict pseudo-knots. </p>\n<p>Interestingly, for the sequence I posted later (R1138/7PTL), although my sequence only model was missing the short stems circled in red, it does catch the pseudo-knot in cyan, if I do column wise standardization (i.e. Z-score). </p>\n<p>&lt;img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F2b61a9d5e22f6a107019c3c293388ca6%2Fdegformer_kissing_loop.png?generation=1702051865166909&amp;alt=media\" alt=\"!<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F5196e7259178126e75ef13ec513791b8%2FR1138_2ND.png?generation=1702051860388554&amp;alt=media\" target=\"_blank\">\" /&gt;</a></p>\n<h1>Data scaling</h1>\n<p>I did some experiments to figure out how model performance scales with training data size by using subsets of train (0.01/0.1/1 fraction-1.4k/14k/140k) with a few different conditions. Also, I wanted to figure out if it looks like the sequence only model can overtake the bpp model. Interestingly, it does seem the gap is closing, as delta MAE of sequence model and bpp model becomes smaller as I used more data. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F5ae4cb6fc6c3d18ba09c3f2cc63140bf%2Fscaling.png?generation=1702051895415046&amp;alt=media\" alt=\"\"></p>\n<p>Very nicely, I also saw that the model starts to learn how to do M2 even at 0.1 fraction of the training data, but obviously becomes better full train data. This shows that there may be some emerging properties to our problem. Notably, the model is completely unable to correctly respond to single mutations (aside for the mutated position it self), when using 0.01 data (1.4k), which is similar to the scale of OpenVaccine. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F65a05d1c6fc83b23ab9a2cd544ebae25%2Fm2_scaling.png?generation=1702051926393030&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2553926,
      "postDate": "2023-12-08T16:12:41.237Z",
      "content": "<p>Hi Kagglers, it has been fun to be behind the scenes this time as a host and I’m glad to see many top solutions posted already. Here I’d like to share my solution and some other insights. </p>\n<h1>Best MAE model</h1>\n<p>My best MAE model is a fold single fold RNAdegformer (published at <a href=\"https://academic.oup.com/bib/article/24/1/bbac581/6986359\" target=\"_blank\">https://academic.oup.com/bib/article/24/1/bbac581/6986359</a>) that scores<br>\nPublic: 0.1437<br>\nPrivate: 0.1458<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F535628c4e75489ab3c6529213f7aba30%2Frnadegformer.png?generation=1702051686843214&amp;alt=media\" alt=\"\"></p>\n<p>It takes bpp and sequence as features and is nearly exactly the same as what I describe in the paper with some small changes (fewer heads, only inverse distance matrix as pos encoding, and no MFE/structure/bpRNA features). If only using train data where both 2A3/DMS SN&gt;1, this model scores Private: 0.1471 Public: 0.1450 and later on I trained another one that uses all profiles with SN&gt;1 (even if one of 2A3/DMS has SN&lt;=1) while masking SN &lt;1 profiles (one of 2A3/DMS), it scores Public: 0.1437\nPrivate: 0.1458. Only using SN&gt;1 takes about 3 hours while using  all profiles with SN&gt;1 takes 5 hours on 2x3090. </p>\n<p>Note that the difference between public/private is quite small in my case. From a modeling standpoint, I think this is due to adding bpp into self-attention as a relatively large bias (I set gamma=32 in my model) results in a sparse attention pattern where row-wise entropy is low and therefore generalizes well to longer sequences. But from another perspective, it’s probably also due to the fact that I’m not a competitor and can see the private data. </p>\n<p>From this point on, I will refer to this model as the bpp model. <br>\nI also have another model where I simply took out bpp and only use sequence as input (referred to as sequence only model) that scores at best Public: 0.1496 Private: 0.1507. This model was more or less trained to figure out if we have enough data that allows the model to not rely on bpp and outperform the bpp model. It turns out we’re not there yet. </p>\n<h1>Flip augmentation</h1>\n<p>Flip augmentation is very useful for me, for the bpp model it has around 0.003 boost, while  for the sequence only model, the boost is much bigger, about 0.01. I’ve seen some competitors say flip aug does not help but I’m not sure if they are doing it in the exact same way as me, which is</p>\n<ol>\n<li>during training: 50% of the time, flip sequence/labels/bpp</li>\n<li>during inference: do TTA, where pred=(model(input)+flip(model(flip(inputs)))/2</li>\n</ol>\n<h1>Positional encoding</h1>\n<p>Positional encoding is important for generalization. Generally, you want to use a relative positional encoding that doesn’t distinguish positions at long ranges. I use a 2D convolved inverse distance matrix as an an attention bias, stacked with bpp (for bpp model), but there are other ways to do it as mentioned in other top solutions.  </p>\n<h1>About pseudo-knots</h1>\n<p>While the bpp model scores well, there are some issues. Notably, early on when <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> inspected the M2 (mutate and map) predictions of some the pseudo-knotted design sequences, the bpp model is completely missing the pseudo-knots (denoted by “[]”), which are hallmarks of RNA 3D structure and not predicted by eternafold. This is not good, because if we want to use models in this competition to predict 3D structures, missing the pseudo-knots can mean predicting entirely wrong 3D folding patterns.  </p>\n<p>However, interestingly, the sequence only model, does recognize the pseudo-knots better. In the plot below you can see that although the bpp model scores better on MAE, it misses the pseudo-knot (circled in red). <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F13ff17d2393757e8bf904087bfc220c5%2Fslide1.png?generation=1702051754381827&amp;alt=media\" alt=\"\"></p>\n<p>But sometimes the bpp model does predict pseudo knots, it’s actually just more conservative overall on pseudo-knots.  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2Fd4083dc420823458030436ad1396d27e%2Fslide2.png?generation=1702051803175730&amp;alt=media\" alt=\"\"></p>\n<h1>long sequence generalization</h1>\n<p>Early on, we saw some competitors’ predictions on the sequence in in this post <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653</a> look unreasonable, which I thought was due to using absolute positional encoding that generalizes poorly beyond length longer than is available in train. A lot of discussion happened there and I think it served as a good sanity check for competitors. The picture I used is actually a 50/50 ensemble of the bpp model and the sequence model, which combines the bpp model’s strength in predicting nested secondary structure and the sequence model’s ability to predict pseudo-knots. </p>\n<p>Interestingly, for the sequence I posted later (R1138/7PTL), although my sequence only model was missing the short stems circled in red, it does catch the pseudo-knot in cyan, if I do column wise standardization (i.e. Z-score). </p>\n<p>&lt;img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F2b61a9d5e22f6a107019c3c293388ca6%2Fdegformer_kissing_loop.png?generation=1702051865166909&amp;alt=media\" alt=\"!<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F5196e7259178126e75ef13ec513791b8%2FR1138_2ND.png?generation=1702051860388554&amp;alt=media\" target=\"_blank\">\" /&gt;</a></p>\n<h1>Data scaling</h1>\n<p>I did some experiments to figure out how model performance scales with training data size by using subsets of train (0.01/0.1/1 fraction-1.4k/14k/140k) with a few different conditions. Also, I wanted to figure out if it looks like the sequence only model can overtake the bpp model. Interestingly, it does seem the gap is closing, as delta MAE of sequence model and bpp model becomes smaller as I used more data. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F5ae4cb6fc6c3d18ba09c3f2cc63140bf%2Fscaling.png?generation=1702051895415046&amp;alt=media\" alt=\"\"></p>\n<p>Very nicely, I also saw that the model starts to learn how to do M2 even at 0.1 fraction of the training data, but obviously becomes better full train data. This shows that there may be some emerging properties to our problem. Notably, the model is completely unable to correctly respond to single mutations (aside for the mutated position it self), when using 0.01 data (1.4k), which is similar to the scale of OpenVaccine. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F65a05d1c6fc83b23ab9a2cd544ebae25%2Fm2_scaling.png?generation=1702051926393030&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi Kagglers, it has been fun to be behind the scenes this time as a host and I’m glad to see many top solutions posted already. Here I’d like to share my solution and some other insights. \n\n# Best MAE model\n\nMy best MAE model is a fold single fold RNAdegformer (published at https://academic.oup.com/bib/article/24/1/bbac581/6986359) that scores\nPublic: 0.1437\nPrivate: 0.1458\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F535628c4e75489ab3c6529213f7aba30%2Frnadegformer.png?generation=1702051686843214&alt=media)\n\nIt takes bpp and sequence as features and is nearly exactly the same as what I describe in the paper with some small changes (fewer heads, only inverse distance matrix as pos encoding, and no MFE/structure/bpRNA features). If only using train data where both 2A3/DMS SN>1, this model scores Private: 0.1471 Public: 0.1450 and later on I trained another one that uses all profiles with SN>1 (even if one of 2A3/DMS has SN<=1) while masking SN <1 profiles (one of 2A3/DMS), it scores Public: 0.1437\nPrivate: 0.1458. Only using SN>1 takes about 3 hours while using  all profiles with SN>1 takes 5 hours on 2x3090. \n\nNote that the difference between public/private is quite small in my case. From a modeling standpoint, I think this is due to adding bpp into self-attention as a relatively large bias (I set gamma=32 in my model) results in a sparse attention pattern where row-wise entropy is low and therefore generalizes well to longer sequences. But from another perspective, it’s probably also due to the fact that I’m not a competitor and can see the private data. \n\nFrom this point on, I will refer to this model as the bpp model. \nI also have another model where I simply took out bpp and only use sequence as input (referred to as sequence only model) that scores at best Public: 0.1496 Private: 0.1507. This model was more or less trained to figure out if we have enough data that allows the model to not rely on bpp and outperform the bpp model. It turns out we’re not there yet. \n\n# Flip augmentation\n\nFlip augmentation is very useful for me, for the bpp model it has around 0.003 boost, while  for the sequence only model, the boost is much bigger, about 0.01. I’ve seen some competitors say flip aug does not help but I’m not sure if they are doing it in the exact same way as me, which is\n\n1. during training: 50% of the time, flip sequence/labels/bpp\n2. during inference: do TTA, where pred=(model(input)+flip(model(flip(inputs)))/2\n\n# Positional encoding\n\nPositional encoding is important for generalization. Generally, you want to use a relative positional encoding that doesn’t distinguish positions at long ranges. I use a 2D convolved inverse distance matrix as an an attention bias, stacked with bpp (for bpp model), but there are other ways to do it as mentioned in other top solutions.  \n\n\n# About pseudo-knots\n\nWhile the bpp model scores well, there are some issues. Notably, early on when @rhijudas inspected the M2 (mutate and map) predictions of some the pseudo-knotted design sequences, the bpp model is completely missing the pseudo-knots (denoted by “[]”), which are hallmarks of RNA 3D structure and not predicted by eternafold. This is not good, because if we want to use models in this competition to predict 3D structures, missing the pseudo-knots can mean predicting entirely wrong 3D folding patterns.  \n\nHowever, interestingly, the sequence only model, does recognize the pseudo-knots better. In the plot below you can see that although the bpp model scores better on MAE, it misses the pseudo-knot (circled in red). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F13ff17d2393757e8bf904087bfc220c5%2Fslide1.png?generation=1702051754381827&alt=media)\n\nBut sometimes the bpp model does predict pseudo knots, it’s actually just more conservative overall on pseudo-knots.  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2Fd4083dc420823458030436ad1396d27e%2Fslide2.png?generation=1702051803175730&alt=media)\n\n# long sequence generalization \n\nEarly on, we saw some competitors’ predictions on the sequence in in this post https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653 look unreasonable, which I thought was due to using absolute positional encoding that generalizes poorly beyond length longer than is available in train. A lot of discussion happened there and I think it served as a good sanity check for competitors. The picture I used is actually a 50/50 ensemble of the bpp model and the sequence model, which combines the bpp model’s strength in predicting nested secondary structure and the sequence model’s ability to predict pseudo-knots. \n\n\nInterestingly, for the sequence I posted later (R1138/7PTL), although my sequence only model was missing the short stems circled in red, it does catch the pseudo-knot in cyan, if I do column wise standardization (i.e. Z-score). \n\n![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F2b61a9d5e22f6a107019c3c293388ca6%2Fdegformer_kissing_loop.png?generation=1702051865166909&alt=media)](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F5196e7259178126e75ef13ec513791b8%2FR1138_2ND.png?generation=1702051860388554&alt=media)\n\n# Data scaling\n\nI did some experiments to figure out how model performance scales with training data size by using subsets of train (0.01/0.1/1 fraction-1.4k/14k/140k) with a few different conditions. Also, I wanted to figure out if it looks like the sequence only model can overtake the bpp model. Interestingly, it does seem the gap is closing, as delta MAE of sequence model and bpp model becomes smaller as I used more data. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F5ae4cb6fc6c3d18ba09c3f2cc63140bf%2Fscaling.png?generation=1702051895415046&alt=media)\n\nVery nicely, I also saw that the model starts to learn how to do M2 even at 0.1 fraction of the training data, but obviously becomes better full train data. This shows that there may be some emerging properties to our problem. Notably, the model is completely unable to correctly respond to single mutations (aside for the mutated position it self), when using 0.01 data (1.4k), which is similar to the scale of OpenVaccine. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F65a05d1c6fc83b23ab9a2cd544ebae25%2Fm2_scaling.png?generation=1702051926393030&alt=media)",
      "votes": 23
    },
    {
      "id": 2555796,
      "postDate": "2023-12-10T07:58:16.277Z",
      "content": "<p>the host should give himself for a kaggle swag prize for the best host solution ranking and writeup in history!!!</p>",
      "rawMarkdown": "the host should give himself for a kaggle swag prize for the best host solution ranking and writeup in history!!!",
      "votes": 1
    },
    {
      "id": 2554704,
      "postDate": "2023-12-09T10:34:28.167Z",
      "content": "<p>Great to see host competition!</p>",
      "rawMarkdown": "Great to see host competition!",
      "votes": 1
    },
    {
      "id": 2554053,
      "postDate": "2023-12-08T18:31:18.500Z",
      "content": "<p>Thanks for the post! <br>\nYes, unfortunately we realized that using bpp could hide pseudoknots only at the last week of the competition.<br>\nIt's really interesting because many pseudoknot predicting algorithms use BPP as an input</p>",
      "rawMarkdown": "Thanks for the post! \nYes, unfortunately we realized that using bpp could hide pseudoknots only at the last week of the competition.\nIt's really interesting because many pseudoknot predicting algorithms use BPP as an input",
      "votes": 1,
      "replies": [
        {
          "id": 2554990,
          "postDate": "2023-12-09T15:39:51.723Z",
          "content": "<p>Yes but these pseudoknot predicting algorithms use BPP are not that accurate. Notably in pretesting, I tested them on a smaller subset of the full training, when it was still computationally feasible to use them, boost was very small, and the best one was -0.001 MAE. about half of them even made MAE worse. but in my pretesting, bpp was still very useful (w/ eternafold being the best), so we decided to generate bpps for the competitors</p>",
          "rawMarkdown": "Yes but these pseudoknot predicting algorithms use BPP are not that accurate. Notably in pretesting, I tested them on a smaller subset of the full training, when it was still computationally feasible to use them, boost was very small, and the best one was -0.001 MAE. about half of them even made MAE worse. but in my pretesting, bpp was still very useful (w/ eternafold being the best), so we decided to generate bpps for the competitors",
          "votes": 2
        }
      ]
    },
    {
      "id": 2554666,
      "postDate": "2023-12-09T09:40:41.070Z",
      "content": "<p>Thanks for the fun competition! <br>\nOur solution was also greatly influenced by your RNAdegformer.</p>\n<p>We tried flip augmentation without TTA and it didn't work. Will your CV improve without TTA?</p>",
      "rawMarkdown": "Thanks for the fun competition! \nOur solution was also greatly influenced by your RNAdegformer.\n\nWe tried flip augmentation without TTA and it didn't work. Will your CV improve without TTA?",
      "replies": [
        {
          "id": 2554731,
          "postDate": "2023-12-09T11:03:49.727Z",
          "content": "<p>Thanks and glad my previous work is useful. I'm not too sure if CV improves without TTA, cuz I always did val with TTA. But I think when I first added flip aug it did not have val TTA and val score was more or less the same </p>",
          "rawMarkdown": "Thanks and glad my previous work is useful. I'm not too sure if CV improves without TTA, cuz I always did val with TTA. But I think when I first added flip aug it did not have val TTA and val score was more or less the same ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2554021,
      "postDate": "2023-12-08T17:53:36.103Z",
      "content": "<p>We are not there yet, but considering the fact that <a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a> could get <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460172\" target=\"_blank\">very good score</a> with only train data, single model and steep filtering, I feel that we are almost there. Actually, with your 1M test seqs, it's very possible that you will really outperform BPP models with only train data…really exciting times :)</p>",
      "rawMarkdown": "We are not there yet, but considering the fact that @dankrstev could get [very good score](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460172) with only train data, single model and steep filtering, I feel that we are almost there. Actually, with your 1M test seqs, it's very possible that you will really outperform BPP models with only train data...really exciting times :)",
      "replies": [
        {
          "id": 2554735,
          "postDate": "2023-12-09T11:04:52.823Z",
          "content": "<p>Yes it looks like we're close than I thought!</p>",
          "rawMarkdown": "Yes it looks like we're close than I thought!"
        }
      ]
    },
    {
      "id": 2554002,
      "postDate": "2023-12-08T17:33:50.250Z",
      "content": "<p>Thanks for sharing. It's great to see the host posting their own solution! It looks like it will be an prize solution easily with some basic ensembles… 👍</p>",
      "rawMarkdown": "Thanks for sharing. It's great to see the host posting their own solution! It looks like it will be an prize solution easily with some basic ensembles... 👍",
      "replies": [
        {
          "id": 2554733,
          "postDate": "2023-12-09T11:04:30.950Z",
          "content": "<p>Thanks and congrats on your top finish!</p>",
          "rawMarkdown": "Thanks and congrats on your top finish!"
        }
      ]
    },
    {
      "id": 2554773,
      "postDate": "2023-12-09T11:45:15.193Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2555796,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-12-10T07:58:16.277000",
      "content": "<p>the host should give himself for a kaggle swag prize for the best host solution ranking and writeup in history!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2554704,
      "author_name": "Michal Bogacz",
      "author_url": "",
      "post_date": "2023-12-09T10:34:28.167000",
      "content": "<p>Great to see host competition!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2554053,
      "author_name": "Penzar Dmitry",
      "author_url": "",
      "post_date": "2023-12-08T18:31:18.500000",
      "content": "<p>Thanks for the post! <br>\nYes, unfortunately we realized that using bpp could hide pseudoknots only at the last week of the competition.<br>\nIt's really interesting because many pseudoknot predicting algorithms use BPP as an input</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2554990,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2023-12-09T15:39:51.723000",
          "content": "<p>Yes but these pseudoknot predicting algorithms use BPP are not that accurate. Notably in pretesting, I tested them on a smaller subset of the full training, when it was still computationally feasible to use them, boost was very small, and the best one was -0.001 MAE. about half of them even made MAE worse. but in my pretesting, bpp was still very useful (w/ eternafold being the best), so we decided to generate bpps for the competitors</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2554666,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "2023-12-09T09:40:41.070000",
      "content": "<p>Thanks for the fun competition! <br>\nOur solution was also greatly influenced by your RNAdegformer.</p>\n<p>We tried flip augmentation without TTA and it didn't work. Will your CV improve without TTA?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2554731,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2023-12-09T11:03:49.727000",
          "content": "<p>Thanks and glad my previous work is useful. I'm not too sure if CV improves without TTA, cuz I always did val with TTA. But I think when I first added flip aug it did not have val TTA and val score was more or less the same </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2554021,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2023-12-08T17:53:36.103000",
      "content": "<p>We are not there yet, but considering the fact that <a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a> could get <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460172\" target=\"_blank\">very good score</a> with only train data, single model and steep filtering, I feel that we are almost there. Actually, with your 1M test seqs, it's very possible that you will really outperform BPP models with only train data…really exciting times :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2554735,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2023-12-09T11:04:52.823000",
          "content": "<p>Yes it looks like we're close than I thought!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2554002,
      "author_name": "hoyso48",
      "author_url": "",
      "post_date": "2023-12-08T17:33:50.250000",
      "content": "<p>Thanks for sharing. It's great to see the host posting their own solution! It looks like it will be an prize solution easily with some basic ensembles… 👍</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2554733,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2023-12-09T11:04:30.950000",
          "content": "<p>Thanks and congrats on your top finish!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2554773,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-09T11:45:15.193000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2553926": "Hi Kagglers, it has been fun to be behind the scenes this time as a host and I’m glad to see many top solutions posted already. Here I’d like to share my solution and some other insights. \n\n# Best MAE model\n\nMy best MAE model is a fold single fold RNAdegformer (published at https://academic.oup.com/bib/article/24/1/bbac581/6986359) that scores\nPublic: 0.1437\nPrivate: 0.1458\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F535628c4e75489ab3c6529213f7aba30%2Frnadegformer.png?generation=1702051686843214&alt=media)\n\nIt takes bpp and sequence as features and is nearly exactly the same as what I describe in the paper with some small changes (fewer heads, only inverse distance matrix as pos encoding, and no MFE/structure/bpRNA features). If only using train data where both 2A3/DMS SN>1, this model scores Private: 0.1471 Public: 0.1450 and later on I trained another one that uses all profiles with SN>1 (even if one of 2A3/DMS has SN<=1) while masking SN <1 profiles (one of 2A3/DMS), it scores Public: 0.1437\nPrivate: 0.1458. Only using SN>1 takes about 3 hours while using  all profiles with SN>1 takes 5 hours on 2x3090. \n\nNote that the difference between public/private is quite small in my case. From a modeling standpoint, I think this is due to adding bpp into self-attention as a relatively large bias (I set gamma=32 in my model) results in a sparse attention pattern where row-wise entropy is low and therefore generalizes well to longer sequences. But from another perspective, it’s probably also due to the fact that I’m not a competitor and can see the private data. \n\nFrom this point on, I will refer to this model as the bpp model. \nI also have another model where I simply took out bpp and only use sequence as input (referred to as sequence only model) that scores at best Public: 0.1496 Private: 0.1507. This model was more or less trained to figure out if we have enough data that allows the model to not rely on bpp and outperform the bpp model. It turns out we’re not there yet. \n\n# Flip augmentation\n\nFlip augmentation is very useful for me, for the bpp model it has around 0.003 boost, while  for the sequence only model, the boost is much bigger, about 0.01. I’ve seen some competitors say flip aug does not help but I’m not sure if they are doing it in the exact same way as me, which is\n\n1. during training: 50% of the time, flip sequence/labels/bpp\n2. during inference: do TTA, where pred=(model(input)+flip(model(flip(inputs)))/2\n\n# Positional encoding\n\nPositional encoding is important for generalization. Generally, you want to use a relative positional encoding that doesn’t distinguish positions at long ranges. I use a 2D convolved inverse distance matrix as an an attention bias, stacked with bpp (for bpp model), but there are other ways to do it as mentioned in other top solutions.  \n\n\n# About pseudo-knots\n\nWhile the bpp model scores well, there are some issues. Notably, early on when @rhijudas inspected the M2 (mutate and map) predictions of some the pseudo-knotted design sequences, the bpp model is completely missing the pseudo-knots (denoted by “[]”), which are hallmarks of RNA 3D structure and not predicted by eternafold. This is not good, because if we want to use models in this competition to predict 3D structures, missing the pseudo-knots can mean predicting entirely wrong 3D folding patterns.  \n\nHowever, interestingly, the sequence only model, does recognize the pseudo-knots better. In the plot below you can see that although the bpp model scores better on MAE, it misses the pseudo-knot (circled in red). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F13ff17d2393757e8bf904087bfc220c5%2Fslide1.png?generation=1702051754381827&alt=media)\n\nBut sometimes the bpp model does predict pseudo knots, it’s actually just more conservative overall on pseudo-knots.  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2Fd4083dc420823458030436ad1396d27e%2Fslide2.png?generation=1702051803175730&alt=media)\n\n# long sequence generalization \n\nEarly on, we saw some competitors’ predictions on the sequence in in this post https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653 look unreasonable, which I thought was due to using absolute positional encoding that generalizes poorly beyond length longer than is available in train. A lot of discussion happened there and I think it served as a good sanity check for competitors. The picture I used is actually a 50/50 ensemble of the bpp model and the sequence model, which combines the bpp model’s strength in predicting nested secondary structure and the sequence model’s ability to predict pseudo-knots. \n\n\nInterestingly, for the sequence I posted later (R1138/7PTL), although my sequence only model was missing the short stems circled in red, it does catch the pseudo-knot in cyan, if I do column wise standardization (i.e. Z-score). \n\n![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F2b61a9d5e22f6a107019c3c293388ca6%2Fdegformer_kissing_loop.png?generation=1702051865166909&alt=media)](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F5196e7259178126e75ef13ec513791b8%2FR1138_2ND.png?generation=1702051860388554&alt=media)\n\n# Data scaling\n\nI did some experiments to figure out how model performance scales with training data size by using subsets of train (0.01/0.1/1 fraction-1.4k/14k/140k) with a few different conditions. Also, I wanted to figure out if it looks like the sequence only model can overtake the bpp model. Interestingly, it does seem the gap is closing, as delta MAE of sequence model and bpp model becomes smaller as I used more data. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F5ae4cb6fc6c3d18ba09c3f2cc63140bf%2Fscaling.png?generation=1702051895415046&alt=media)\n\nVery nicely, I also saw that the model starts to learn how to do M2 even at 0.1 fraction of the training data, but obviously becomes better full train data. This shows that there may be some emerging properties to our problem. Notably, the model is completely unable to correctly respond to single mutations (aside for the mutated position it self), when using 0.01 data (1.4k), which is similar to the scale of OpenVaccine. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F65a05d1c6fc83b23ab9a2cd544ebae25%2Fm2_scaling.png?generation=1702051926393030&alt=media)",
    "2555796": "the host should give himself for a kaggle swag prize for the best host solution ranking and writeup in history!!!",
    "2554704": "Great to see host competition!",
    "2554053": "Thanks for the post! \nYes, unfortunately we realized that using bpp could hide pseudoknots only at the last week of the competition.\nIt's really interesting because many pseudoknot predicting algorithms use BPP as an input",
    "2554666": "Thanks for the fun competition! \nOur solution was also greatly influenced by your RNAdegformer.\n\nWe tried flip augmentation without TTA and it didn't work. Will your CV improve without TTA?",
    "2554021": "We are not there yet, but considering the fact that @dankrstev could get [very good score](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460172) with only train data, single model and steep filtering, I feel that we are almost there. Actually, with your 1M test seqs, it's very possible that you will really outperform BPP models with only train data...really exciting times :)",
    "2554002": "Thanks for sharing. It's great to see the host posting their own solution! It looks like it will be an prize solution easily with some basic ensembles... 👍",
    "2554773": ""
  }
}