{
  "id": 460316,
  "title": "2nd place solution - Squeezeformer + BPP Conv2D Attention",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460316",
  "author_name": "hoyso48",
  "post_date": "2023-12-08T17:08:00.069000",
  "votes": 42,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Thanks to Kaggle and the hosts for organizing this competition.  It was truly inspiring and challenging, and I learned a lot from this one.👍</p>\n<h3>Code: <a href=\"https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution\" target=\"_blank\">https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution</a></h3>\n<h1>TLDR</h1>\n<p><strong>Keypoints:</strong></p>\n<ul>\n<li>Squeezeformer[1] + GRU head.</li>\n<li>Simple Conv2DNet for bpp, adding it as a bias to the attention matrix.</li>\n<li>ALiBi positional encoding[2] for robust generalization on longer sequences.</li>\n<li>Weighted loss with signal_to_noise, with longer epochs.</li>\n<li>Additional features for minor score improvements.</li>\n</ul>\n<p>I adopted Squeezeformer, which I became familiar with after the ASL fingerspelling competition.Thanks to <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/goldenlock\" target=\"_blank\">@goldenlock</a> for their solutions in the last ASL competition. The most crucial part of my solution is how to utilize the bpp matrix. I applied a simple shallow Conv2DNet to bpp and directly added it to the attention matrix.</p>\n<p><strong>Features:</strong></p>\n<p>I used some features found useful in the OpenVaccine Challenge, to help fast initial convergence. These included:</p>\n<ul>\n<li>CapR looptype.</li>\n<li>eternafold mfe.</li>\n<li>predicted Looptype with eternafold mfe.</li>\n<li>bpp features (sum, nzero, max).</li>\n</ul>\n<p>However, unlike in the OpenVaccine challenge, these features only marginally helped (about -0.0005). Therefore, I believe these features should be removed in the future for the simplicity.</p>\n<h1>Model</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Fb3c99f0b3ace13f5c4381c094b899d46%2F-3.drawio-2.png?generation=1702054547887335&amp;alt=media\" alt=\"\"></p>\n<p><strong>Squeezeformer Encoder:</strong></p>\n<p>I chose Squeezeformer with minor modifications (BN after conv1d, SwiGLU in FFN, etc.), which mixes Conv1D blocks with Transformer. While I tried other recent Conv-Transformer Hybrid architectures, Squeezeformer was the most efficient. Compared to a Vanilla Transformer, Squeezeformer showed strong performance early in training and consistently showed faster convergence.</p>\n<p>The models used the following parameters: dim=192, num_heads=4, kernel_size=17, num_layers=12.</p>\n<p><strong>GRU head:</strong></p>\n<p>Adding a single GRU layer after the encoder yielded minor improvements.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F3191b22f70fe4b38f04b6664871d2b96%2F.drawio-10.png?generation=1702054568362614&amp;alt=media\" alt=\"\"></p>\n<p><strong>ALiBi positional encoding:</strong></p>\n<p>I adopted AliBi positional encoding as it claimed to generalize better over long sequences than other methods.</p>\n<p><strong>BPP as Attention Bias:</strong></p>\n<p>The bpp matrix (using only the provided one) was added as a bias in the attention matrix after multiplied by per-head predefined scales, significantly improved performance (around -0.0025).</p>\n<p><strong>BPP 2DConvNet:</strong></p>\n<p>Using Bpp directly as an attention bias was a good start, but I felt it needed more flexibility(I felt it was too sparse). Among various options, adding a 2D CNN on top of the BPP matrix proved very helpful (-0.002). However, multiple 2D CNNs applied to the BPP matrix (usually 206 x 206) were inefficient in terms of training/inference time. Thus, I just used a simple shallow 2-layer 2DCNN, with the output matrix shared across all Transformer block layers.</p>\n<h1>Training</h1>\n<ul>\n<li>Epochs: 200.</li>\n<li>Batch size: 256.</li>\n<li>Learning rate: 2e-3, with Cosine Decay and warmup.</li>\n<li>Optimizer: AdamW, weight decay = 0.01.</li>\n<li>Loss: Weighted MAE (weight = log1p(signal_to_noise).clip(0,10)).</li>\n</ul>\n<p>The single model CV (K-fold, k=5) scored 0.119, with a public LB of 0.140 and a private LB of 0.142. After ensemble on different seed I got public LB of 0.135 and private LB of 0.140.<br>\nAlthough there were some questionable correlations between CV/LB in certain submissions, generally, they aligned well with the CV.<br>\nWith the above setup, training a single model took around 30 hours on a single RTX 4090. </p>\n<h2>Discarded ideas &amp; Thoughts:</h2>\n<ul>\n<li><strong>Self-Supervised Learning (SSL)</strong>: At first I motivated to participate in this competition as it would be really nice if any SSL method could be successfully applied without using any features other than the sequence.  Initial trials with Data2Vec and BERT-like SSL methods showed inconsistent improvements. Due to the additional training time required, I did not consider SSL further. However, I believe there is still huge potential in this idea.</li>\n<li><strong>Large Models</strong>: Attempts to train larger models (dim &gt; 512) with proper regularizations were unsuccessful. I think this and SSL failure suggests that the primary challenge lies in the inherent noise within the training dataset.</li>\n<li><strong>Augmentations</strong>: Most augmentation methods I tried had no effect.</li>\n<li><strong>Pseudo Labels</strong>: While pseudo labeling might help in LB, it didn't improve CV in my case, so I didn't use it for safety&amp;training time. However, after seeing the correlation between public and private LB, I think it might have been slightly beneficial in both public and private LB.</li>\n</ul>\n<p>For me, this competition was a series of choices regarding whether to experiment with or adopt some promising ideas, especially when there were only 2-3 weeks left. Some of the ideas I thought might be helpful were abandoned without further consideration because they required more time for implementation and training. I think that this strategy may have made my solution somewhat suboptimal or redundant, but overall I see it worked quite well as my solution appeared to capture most of the crucial aspects of other teams' solutions.</p>\n<h2>References</h2>\n<p>[1]Sehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W. Mahoney, and Kurt Keutzer. 2022. Squeezeformer: An Efficient Transformer for Automatic Speech Recognition. arXiv:2206.00888 [eess.AS]. <a href=\"url\" target=\"_blank\">https://doi.org/10.48550/arXiv.2206.00888</a></p>\n<p>[2]Ofir Press, Noah A. Smith, Mike Lewis. 2022. Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. arXiv:2108.12409 [cs.CL]. <a href=\"url\" target=\"_blank\">https://doi.org/10.48550/arXiv.2108.12409</a></p>",
  "messages": [
    {
      "id": 2553982,
      "postDate": "2023-12-08T17:08:00.070Z",
      "content": "<p>Thanks to Kaggle and the hosts for organizing this competition.  It was truly inspiring and challenging, and I learned a lot from this one.👍</p>\n<h3>Code: <a href=\"https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution\" target=\"_blank\">https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution</a></h3>\n<h1>TLDR</h1>\n<p><strong>Keypoints:</strong></p>\n<ul>\n<li>Squeezeformer[1] + GRU head.</li>\n<li>Simple Conv2DNet for bpp, adding it as a bias to the attention matrix.</li>\n<li>ALiBi positional encoding[2] for robust generalization on longer sequences.</li>\n<li>Weighted loss with signal_to_noise, with longer epochs.</li>\n<li>Additional features for minor score improvements.</li>\n</ul>\n<p>I adopted Squeezeformer, which I became familiar with after the ASL fingerspelling competition.Thanks to <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/goldenlock\" target=\"_blank\">@goldenlock</a> for their solutions in the last ASL competition. The most crucial part of my solution is how to utilize the bpp matrix. I applied a simple shallow Conv2DNet to bpp and directly added it to the attention matrix.</p>\n<p><strong>Features:</strong></p>\n<p>I used some features found useful in the OpenVaccine Challenge, to help fast initial convergence. These included:</p>\n<ul>\n<li>CapR looptype.</li>\n<li>eternafold mfe.</li>\n<li>predicted Looptype with eternafold mfe.</li>\n<li>bpp features (sum, nzero, max).</li>\n</ul>\n<p>However, unlike in the OpenVaccine challenge, these features only marginally helped (about -0.0005). Therefore, I believe these features should be removed in the future for the simplicity.</p>\n<h1>Model</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Fb3c99f0b3ace13f5c4381c094b899d46%2F-3.drawio-2.png?generation=1702054547887335&amp;alt=media\" alt=\"\"></p>\n<p><strong>Squeezeformer Encoder:</strong></p>\n<p>I chose Squeezeformer with minor modifications (BN after conv1d, SwiGLU in FFN, etc.), which mixes Conv1D blocks with Transformer. While I tried other recent Conv-Transformer Hybrid architectures, Squeezeformer was the most efficient. Compared to a Vanilla Transformer, Squeezeformer showed strong performance early in training and consistently showed faster convergence.</p>\n<p>The models used the following parameters: dim=192, num_heads=4, kernel_size=17, num_layers=12.</p>\n<p><strong>GRU head:</strong></p>\n<p>Adding a single GRU layer after the encoder yielded minor improvements.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F3191b22f70fe4b38f04b6664871d2b96%2F.drawio-10.png?generation=1702054568362614&amp;alt=media\" alt=\"\"></p>\n<p><strong>ALiBi positional encoding:</strong></p>\n<p>I adopted AliBi positional encoding as it claimed to generalize better over long sequences than other methods.</p>\n<p><strong>BPP as Attention Bias:</strong></p>\n<p>The bpp matrix (using only the provided one) was added as a bias in the attention matrix after multiplied by per-head predefined scales, significantly improved performance (around -0.0025).</p>\n<p><strong>BPP 2DConvNet:</strong></p>\n<p>Using Bpp directly as an attention bias was a good start, but I felt it needed more flexibility(I felt it was too sparse). Among various options, adding a 2D CNN on top of the BPP matrix proved very helpful (-0.002). However, multiple 2D CNNs applied to the BPP matrix (usually 206 x 206) were inefficient in terms of training/inference time. Thus, I just used a simple shallow 2-layer 2DCNN, with the output matrix shared across all Transformer block layers.</p>\n<h1>Training</h1>\n<ul>\n<li>Epochs: 200.</li>\n<li>Batch size: 256.</li>\n<li>Learning rate: 2e-3, with Cosine Decay and warmup.</li>\n<li>Optimizer: AdamW, weight decay = 0.01.</li>\n<li>Loss: Weighted MAE (weight = log1p(signal_to_noise).clip(0,10)).</li>\n</ul>\n<p>The single model CV (K-fold, k=5) scored 0.119, with a public LB of 0.140 and a private LB of 0.142. After ensemble on different seed I got public LB of 0.135 and private LB of 0.140.<br>\nAlthough there were some questionable correlations between CV/LB in certain submissions, generally, they aligned well with the CV.<br>\nWith the above setup, training a single model took around 30 hours on a single RTX 4090. </p>\n<h2>Discarded ideas &amp; Thoughts:</h2>\n<ul>\n<li><strong>Self-Supervised Learning (SSL)</strong>: At first I motivated to participate in this competition as it would be really nice if any SSL method could be successfully applied without using any features other than the sequence.  Initial trials with Data2Vec and BERT-like SSL methods showed inconsistent improvements. Due to the additional training time required, I did not consider SSL further. However, I believe there is still huge potential in this idea.</li>\n<li><strong>Large Models</strong>: Attempts to train larger models (dim &gt; 512) with proper regularizations were unsuccessful. I think this and SSL failure suggests that the primary challenge lies in the inherent noise within the training dataset.</li>\n<li><strong>Augmentations</strong>: Most augmentation methods I tried had no effect.</li>\n<li><strong>Pseudo Labels</strong>: While pseudo labeling might help in LB, it didn't improve CV in my case, so I didn't use it for safety&amp;training time. However, after seeing the correlation between public and private LB, I think it might have been slightly beneficial in both public and private LB.</li>\n</ul>\n<p>For me, this competition was a series of choices regarding whether to experiment with or adopt some promising ideas, especially when there were only 2-3 weeks left. Some of the ideas I thought might be helpful were abandoned without further consideration because they required more time for implementation and training. I think that this strategy may have made my solution somewhat suboptimal or redundant, but overall I see it worked quite well as my solution appeared to capture most of the crucial aspects of other teams' solutions.</p>\n<h2>References</h2>\n<p>[1]Sehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W. Mahoney, and Kurt Keutzer. 2022. Squeezeformer: An Efficient Transformer for Automatic Speech Recognition. arXiv:2206.00888 [eess.AS]. <a href=\"url\" target=\"_blank\">https://doi.org/10.48550/arXiv.2206.00888</a></p>\n<p>[2]Ofir Press, Noah A. Smith, Mike Lewis. 2022. Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. arXiv:2108.12409 [cs.CL]. <a href=\"url\" target=\"_blank\">https://doi.org/10.48550/arXiv.2108.12409</a></p>",
      "rawMarkdown": "Thanks to Kaggle and the hosts for organizing this competition.  It was truly inspiring and challenging, and I learned a lot from this one.👍\n\n###Code: https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution\n\n# TLDR\n\n**Keypoints:**\n\n- Squeezeformer[1] + GRU head.\n- Simple Conv2DNet for bpp, adding it as a bias to the attention matrix.\n- ALiBi positional encoding[2] for robust generalization on longer sequences.\n- Weighted loss with signal_to_noise, with longer epochs.\n- Additional features for minor score improvements.\n\nI adopted Squeezeformer, which I became familiar with after the ASL fingerspelling competition.Thanks to @christofhenkel and @goldenlock for their solutions in the last ASL competition. The most crucial part of my solution is how to utilize the bpp matrix. I applied a simple shallow Conv2DNet to bpp and directly added it to the attention matrix.\n\n**Features:**\n\nI used some features found useful in the OpenVaccine Challenge, to help fast initial convergence. These included:\n\n- CapR looptype.\n- eternafold mfe.\n- predicted Looptype with eternafold mfe.\n- bpp features (sum, nzero, max).\n\nHowever, unlike in the OpenVaccine challenge, these features only marginally helped (about -0.0005). Therefore, I believe these features should be removed in the future for the simplicity.\n\n#Model\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Fb3c99f0b3ace13f5c4381c094b899d46%2F-3.drawio-2.png?generation=1702054547887335&alt=media)\n\n**Squeezeformer Encoder:**\n\nI chose Squeezeformer with minor modifications (BN after conv1d, SwiGLU in FFN, etc.), which mixes Conv1D blocks with Transformer. While I tried other recent Conv-Transformer Hybrid architectures, Squeezeformer was the most efficient. Compared to a Vanilla Transformer, Squeezeformer showed strong performance early in training and consistently showed faster convergence.\n\nThe models used the following parameters: dim=192, num_heads=4, kernel_size=17, num_layers=12.\n\n**GRU head:**\n\nAdding a single GRU layer after the encoder yielded minor improvements.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F3191b22f70fe4b38f04b6664871d2b96%2F.drawio-10.png?generation=1702054568362614&alt=media)\n\n**ALiBi positional encoding:**\n\nI adopted AliBi positional encoding as it claimed to generalize better over long sequences than other methods.\n\n**BPP as Attention Bias:**\n\nThe bpp matrix (using only the provided one) was added as a bias in the attention matrix after multiplied by per-head predefined scales, significantly improved performance (around -0.0025).\n\n**BPP 2DConvNet:**\n\nUsing Bpp directly as an attention bias was a good start, but I felt it needed more flexibility(I felt it was too sparse). Among various options, adding a 2D CNN on top of the BPP matrix proved very helpful (-0.002). However, multiple 2D CNNs applied to the BPP matrix (usually 206 x 206) were inefficient in terms of training/inference time. Thus, I just used a simple shallow 2-layer 2DCNN, with the output matrix shared across all Transformer block layers.\n\n#Training\n\n- Epochs: 200.\n- Batch size: 256.\n- Learning rate: 2e-3, with Cosine Decay and warmup.\n- Optimizer: AdamW, weight decay = 0.01.\n- Loss: Weighted MAE (weight = log1p(signal_to_noise).clip(0,10)).\n\n\nThe single model CV (K-fold, k=5) scored 0.119, with a public LB of 0.140 and a private LB of 0.142. After ensemble on different seed I got public LB of 0.135 and private LB of 0.140.\nAlthough there were some questionable correlations between CV/LB in certain submissions, generally, they aligned well with the CV.\nWith the above setup, training a single model took around 30 hours on a single RTX 4090. \n\n##Discarded ideas & Thoughts:\n\n- **Self-Supervised Learning (SSL)**: At first I motivated to participate in this competition as it would be really nice if any SSL method could be successfully applied without using any features other than the sequence.  Initial trials with Data2Vec and BERT-like SSL methods showed inconsistent improvements. Due to the additional training time required, I did not consider SSL further. However, I believe there is still huge potential in this idea.\n- **Large Models**: Attempts to train larger models (dim > 512) with proper regularizations were unsuccessful. I think this and SSL failure suggests that the primary challenge lies in the inherent noise within the training dataset.\n- **Augmentations**: Most augmentation methods I tried had no effect.\n- **Pseudo Labels**: While pseudo labeling might help in LB, it didn't improve CV in my case, so I didn't use it for safety&training time. However, after seeing the correlation between public and private LB, I think it might have been slightly beneficial in both public and private LB.\n\n\nFor me, this competition was a series of choices regarding whether to experiment with or adopt some promising ideas, especially when there were only 2-3 weeks left. Some of the ideas I thought might be helpful were abandoned without further consideration because they required more time for implementation and training. I think that this strategy may have made my solution somewhat suboptimal or redundant, but overall I see it worked quite well as my solution appeared to capture most of the crucial aspects of other teams' solutions.\n\n##References\n[1]Sehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W. Mahoney, and Kurt Keutzer. 2022. Squeezeformer: An Efficient Transformer for Automatic Speech Recognition. arXiv:2206.00888 [eess.AS]. [https://doi.org/10.48550/arXiv.2206.00888](url)\n\n[2]Ofir Press, Noah A. Smith, Mike Lewis. 2022. Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. arXiv:2108.12409 [cs.CL]. [https://doi.org/10.48550/arXiv.2108.12409](url)\n\n",
      "votes": 42
    },
    {
      "id": 2554728,
      "postDate": "2023-12-09T10:58:05.453Z",
      "content": "<p>Thanks for the details and congrats on 2nd position, </p>\n<p>BPP is not available during test time, so how did the test inference differ?</p>",
      "rawMarkdown": "Thanks for the details and congrats on 2nd position, \n\nBPP is not available during test time, so how did the test inference differ?",
      "votes": 1,
      "replies": [
        {
          "id": 2557039,
          "postDate": "2023-12-11T06:43:02.430Z",
          "content": "<p>Hi, there are BPPs for the test sequences in a folder Ribonanza_bpp_files.</p>",
          "rawMarkdown": "Hi, there are BPPs for the test sequences in a folder Ribonanza_bpp_files.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2554056,
      "postDate": "2023-12-08T18:33:10.947Z",
      "content": "<p>Oh, one more question, if I may: have you tried AWP? I could not make it work, but I thought that since you are the master of AWP, maybe you could make it work or at least explain why it is less useful here…</p>",
      "rawMarkdown": "Oh, one more question, if I may: have you tried AWP? I could not make it work, but I thought that since you are the master of AWP, maybe you could make it work or at least explain why it is less useful here...",
      "votes": 1,
      "replies": [
        {
          "id": 2554062,
          "postDate": "2023-12-08T18:45:09.080Z",
          "content": "<p>I tried it once, but it didn't work. I can't fully explain it, but I think it's because there's already a certain amount of noise in the dataset. (Adding any kind of noise to the dataset in the form of augmentation also didn't work in my case)</p>",
          "rawMarkdown": "I tried it once, but it didn't work. I can't fully explain it, but I think it's because there's already a certain amount of noise in the dataset. (Adding any kind of noise to the dataset in the form of augmentation also didn't work in my case)",
          "votes": 2,
          "replies": [
            {
              "id": 2558197,
              "postDate": "2023-12-12T01:21:42.340Z",
              "content": "<p>just a little comment, I use AWP in various tasks with some success and it didn't work for me here</p>",
              "rawMarkdown": "just a little comment, I use AWP in various tasks with some success and it didn't work for me here",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2553998,
      "postDate": "2023-12-08T17:31:07.760Z",
      "content": "<p>Congrats on achieving such a high rank again! </p>\n<p>Looking forward to learning from your code ♥️</p>",
      "rawMarkdown": "Congrats on achieving such a high rank again! \n\nLooking forward to learning from your code ♥️",
      "votes": 1,
      "replies": [
        {
          "id": 2554006,
          "postDate": "2023-12-08T17:37:50.730Z",
          "content": "<p>Thank you! 🫡</p>",
          "rawMarkdown": "Thank you! 🫡",
          "votes": 1
        }
      ]
    },
    {
      "id": 2554009,
      "postDate": "2023-12-08T17:39:23.417Z",
      "content": "<p>When I saw that you joined the competition, I was sure you would end up at the top due to the similarities to ASL and your expertise.<br>\nI used your architecture from ASL with small modifications, and it worked pretty well. Maybe next time, I will also try finally Squeezeformer (if you publish tensorflow code, it will be the best, hehe).<br>\nI want to ask one question… Is Alibi PE really necessary? You, of all people, should be familiar with the fact that 1d conv enables the models to encode positions 'natively.' I chose the no PE path, which I learned from you, and it was really great. If Alibi helps even further, it's good to know.</p>",
      "rawMarkdown": "When I saw that you joined the competition, I was sure you would end up at the top due to the similarities to ASL and your expertise.\nI used your architecture from ASL with small modifications, and it worked pretty well. Maybe next time, I will also try finally Squeezeformer (if you publish tensorflow code, it will be the best, hehe).\nI want to ask one question… Is Alibi PE really necessary? You, of all people, should be familiar with the fact that 1d conv enables the models to encode positions 'natively.' I chose the no PE path, which I learned from you, and it was really great. If Alibi helps even further, it's good to know.",
      "votes": 2,
      "replies": [
        {
          "id": 2554033,
          "postDate": "2023-12-08T18:10:41.007Z",
          "content": "<p>That's a great question! In fact, it is indeed not necessary in terms of CV(+0.0001 in my case without AliBi). However, the reason I used it in the model is that I thought alibi bias would indeed generalize well for long sequences. As you can see, alibi bias simply adds negative values to the attention matrix for distant frames. It's a very simple method, but the reason it provides robustness for long sequences, as I understand it, is simply because it just <strong>blocks attention between frames that are too far apart</strong>. </p>\n<p>If you look at the table, you will see that absolute positional encoding (sinusoidal), unlike relative positional encoding, has significantly lower robustness for long sequences. As I understand it, this is because when the model receives a long sequence it has not seen during training, for each query, there is a relatively higher likelihood of attending to all positions(even when very distant) of the key compared to the relative method, which ultimately means that due to the normalization by softmax, it is very likely to get a less sparse attention matrix compared to when a short sequence is input. </p>\n<p>Even if you use conv1d instead of positional encoding, the same issue may arise in the Transformer's attention as there is no constraint on the score between long frames. I haven't read the alibi paper in detail or conducted any experiments, but I decided to just add the alibi bias based on this hypothesis. It is just naive thoughts on it but I hope this answer is helpful…!<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F3950e09c5ad2e1d18d6125571f2b60fa%2Fimage.png?generation=1702058837308276&amp;alt=media\" alt=\"\"> </p>",
          "rawMarkdown": "That's a great question! In fact, it is indeed not necessary in terms of CV(+0.0001 in my case without AliBi). However, the reason I used it in the model is that I thought alibi bias would indeed generalize well for long sequences. As you can see, alibi bias simply adds negative values to the attention matrix for distant frames. It's a very simple method, but the reason it provides robustness for long sequences, as I understand it, is simply because it just **blocks attention between frames that are too far apart**. \n\nIf you look at the table, you will see that absolute positional encoding (sinusoidal), unlike relative positional encoding, has significantly lower robustness for long sequences. As I understand it, this is because when the model receives a long sequence it has not seen during training, for each query, there is a relatively higher likelihood of attending to all positions(even when very distant) of the key compared to the relative method, which ultimately means that due to the normalization by softmax, it is very likely to get a less sparse attention matrix compared to when a short sequence is input. \n\nEven if you use conv1d instead of positional encoding, the same issue may arise in the Transformer's attention as there is no constraint on the score between long frames. I haven't read the alibi paper in detail or conducted any experiments, but I decided to just add the alibi bias based on this hypothesis. It is just naive thoughts on it but I hope this answer is helpful...!![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F3950e09c5ad2e1d18d6125571f2b60fa%2Fimage.png?generation=1702058837308276&alt=media) ",
          "votes": 4,
          "replies": [
            {
              "id": 2554043,
              "postDate": "2023-12-08T18:19:46.960Z",
              "content": "<p>Thank you for the answer. I see your reasoning. On the other hand, unlike in sentences, distant rna nucleotides can have a strong interaction (remember the pseudo-knot provided by the host), so by your explanation, Alibi may actually be detrimental. Furthermore, 1DConv should natively 'learn' to assign less importance for larger distance if less importance is true. I guess that further experiments are necessary to know for sure.</p>",
              "rawMarkdown": "Thank you for the answer. I see your reasoning. On the other hand, unlike in sentences, distant rna nucleotides can have a strong interaction (remember the pseudo-knot provided by the host), so by your explanation, Alibi may actually be detrimental. Furthermore, 1DConv should natively 'learn' to assign less importance for larger distance if less importance is true. I guess that further experiments are necessary to know for sure.",
              "votes": 3
            },
            {
              "id": 2554061,
              "postDate": "2023-12-08T18:39:15.930Z",
              "content": "<p>yep, this is just the hypothesis and definitely need some experiments to check it's also applicable in RNA sequences. But as far as I think Transformer can always capture distant relationships very well(whenever it can) other than any other architectures with multiple layers. and I put robustness on top of everything so I felt adding it is more safe.</p>",
              "rawMarkdown": "yep, this is just the hypothesis and definitely need some experiments to check it's also applicable in RNA sequences. But as far as I think Transformer can always capture distant relationships very well(whenever it can) other than any other architectures with multiple layers. and I put robustness on top of everything so I felt adding it is more safe.",
              "votes": 1
            },
            {
              "id": 2557299,
              "postDate": "2023-12-11T11:45:25.557Z",
              "content": "<p>Thank you for the discussion. If I am following right:</p>\n<ul>\n<li>Conceptually we could use either Alibi or conv-bpp bias?</li>\n<li>The reason you use both is Alibi has good robustness to long seq length and conv-bpp-bias is a learned bias encoding properties of BPP?</li>\n<li>Both of these were added to the attention values similar to using Alibi in its paper?</li>\n</ul>\n<p>Thanks!</p>",
              "rawMarkdown": "Thank you for the discussion. If I am following right:\n- Conceptually we could use either Alibi or conv-bpp bias?\n- The reason you use both is Alibi has good robustness to long seq length and conv-bpp-bias is a learned bias encoding properties of BPP?\n- Both of these were added to the attention values similar to using Alibi in its paper?\n\nThanks!"
            },
            {
              "id": 2557328,
              "postDate": "2023-12-11T12:16:38.270Z",
              "content": "<p>I did not use bpp bias (at all). Just stuck 1dconv blocks between the transformers. This allows the model to be PE-free. Look for <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> solution for American isolated sign language competition to get the general idea.</p>",
              "rawMarkdown": "I did not use bpp bias (at all). Just stuck 1dconv blocks between the transformers. This allows the model to be PE-free. Look for @hoyso48 solution for American isolated sign language competition to get the general idea."
            },
            {
              "id": 2558231,
              "postDate": "2023-12-12T02:27:14.980Z",
              "content": "<p>thank you <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> <br>\nWhat is the intuition on adding conv1D? If we do not want a PE the maybe just stack more Encoder Layers could perform better?</p>",
              "rawMarkdown": "thank you @shlomoron \nWhat is the intuition on adding conv1D? If we do not want a PE the maybe just stack more Encoder Layers could perform better?"
            },
            {
              "id": 2558420,
              "postDate": "2023-12-12T06:43:01.667Z",
              "content": "<p>You have to have either PE or conv. Study the architecture of transformers and you will see.</p>",
              "rawMarkdown": "You have to have either PE or conv. Study the architecture of transformers and you will see.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2554789,
      "postDate": "2023-12-09T12:06:27.853Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2557048,
          "postDate": "2023-12-11T06:51:27.150Z",
          "content": "<p>Ablation study results may vary depending on the final model, but if I remember correctly, GRU gave approximately -0.0003. And the features you pointed out were the most important. I believe both CapR and eternafold mfe are -0.0002.</p>\n<p>Congratulations on third place too! I'm very excited for your next move.👍</p>",
          "rawMarkdown": "Ablation study results may vary depending on the final model, but if I remember correctly, GRU gave approximately -0.0003. And the features you pointed out were the most important. I believe both CapR and eternafold mfe are -0.0002.\n\nCongratulations on third place too! I'm very excited for your next move.👍"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2554728,
      "author_name": "FdotRK",
      "author_url": "",
      "post_date": "2023-12-09T10:58:05.453000",
      "content": "<p>Thanks for the details and congrats on 2nd position, </p>\n<p>BPP is not available during test time, so how did the test inference differ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2557039,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-12-11T06:43:02.430000",
          "content": "<p>Hi, there are BPPs for the test sequences in a folder Ribonanza_bpp_files.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2554056,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2023-12-08T18:33:10.947000",
      "content": "<p>Oh, one more question, if I may: have you tried AWP? I could not make it work, but I thought that since you are the master of AWP, maybe you could make it work or at least explain why it is less useful here…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2554062,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-12-08T18:45:09.080000",
          "content": "<p>I tried it once, but it didn't work. I can't fully explain it, but I think it's because there's already a certain amount of noise in the dataset. (Adding any kind of noise to the dataset in the form of augmentation also didn't work in my case)</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2558197,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2023-12-12T01:21:42.340000",
              "content": "<p>just a little comment, I use AWP in various tasks with some success and it didn't work for me here</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2553998,
      "author_name": "Yu Wu",
      "author_url": "",
      "post_date": "2023-12-08T17:31:07.760000",
      "content": "<p>Congrats on achieving such a high rank again! </p>\n<p>Looking forward to learning from your code ♥️</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2554006,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-12-08T17:37:50.730000",
          "content": "<p>Thank you! 🫡</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2554009,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2023-12-08T17:39:23.417000",
      "content": "<p>When I saw that you joined the competition, I was sure you would end up at the top due to the similarities to ASL and your expertise.<br>\nI used your architecture from ASL with small modifications, and it worked pretty well. Maybe next time, I will also try finally Squeezeformer (if you publish tensorflow code, it will be the best, hehe).<br>\nI want to ask one question… Is Alibi PE really necessary? You, of all people, should be familiar with the fact that 1d conv enables the models to encode positions 'natively.' I chose the no PE path, which I learned from you, and it was really great. If Alibi helps even further, it's good to know.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2554033,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-12-08T18:10:41.007000",
          "content": "<p>That's a great question! In fact, it is indeed not necessary in terms of CV(+0.0001 in my case without AliBi). However, the reason I used it in the model is that I thought alibi bias would indeed generalize well for long sequences. As you can see, alibi bias simply adds negative values to the attention matrix for distant frames. It's a very simple method, but the reason it provides robustness for long sequences, as I understand it, is simply because it just <strong>blocks attention between frames that are too far apart</strong>. </p>\n<p>If you look at the table, you will see that absolute positional encoding (sinusoidal), unlike relative positional encoding, has significantly lower robustness for long sequences. As I understand it, this is because when the model receives a long sequence it has not seen during training, for each query, there is a relatively higher likelihood of attending to all positions(even when very distant) of the key compared to the relative method, which ultimately means that due to the normalization by softmax, it is very likely to get a less sparse attention matrix compared to when a short sequence is input. </p>\n<p>Even if you use conv1d instead of positional encoding, the same issue may arise in the Transformer's attention as there is no constraint on the score between long frames. I haven't read the alibi paper in detail or conducted any experiments, but I decided to just add the alibi bias based on this hypothesis. It is just naive thoughts on it but I hope this answer is helpful…!<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F3950e09c5ad2e1d18d6125571f2b60fa%2Fimage.png?generation=1702058837308276&amp;alt=media\" alt=\"\"> </p>",
          "votes": 4,
          "replies": [
            {
              "id": 2554043,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2023-12-08T18:19:46.960000",
              "content": "<p>Thank you for the answer. I see your reasoning. On the other hand, unlike in sentences, distant rna nucleotides can have a strong interaction (remember the pseudo-knot provided by the host), so by your explanation, Alibi may actually be detrimental. Furthermore, 1DConv should natively 'learn' to assign less importance for larger distance if less importance is true. I guess that further experiments are necessary to know for sure.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2554061,
              "author_name": "hoyso48",
              "author_url": "",
              "post_date": "2023-12-08T18:39:15.930000",
              "content": "<p>yep, this is just the hypothesis and definitely need some experiments to check it's also applicable in RNA sequences. But as far as I think Transformer can always capture distant relationships very well(whenever it can) other than any other architectures with multiple layers. and I put robustness on top of everything so I felt adding it is more safe.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2557299,
              "author_name": "FdotRK",
              "author_url": "",
              "post_date": "2023-12-11T11:45:25.557000",
              "content": "<p>Thank you for the discussion. If I am following right:</p>\n<ul>\n<li>Conceptually we could use either Alibi or conv-bpp bias?</li>\n<li>The reason you use both is Alibi has good robustness to long seq length and conv-bpp-bias is a learned bias encoding properties of BPP?</li>\n<li>Both of these were added to the attention values similar to using Alibi in its paper?</li>\n</ul>\n<p>Thanks!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2557328,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2023-12-11T12:16:38.270000",
              "content": "<p>I did not use bpp bias (at all). Just stuck 1dconv blocks between the transformers. This allows the model to be PE-free. Look for <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> solution for American isolated sign language competition to get the general idea.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2558231,
              "author_name": "FdotRK",
              "author_url": "",
              "post_date": "2023-12-12T02:27:14.980000",
              "content": "<p>thank you <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> <br>\nWhat is the intuition on adding conv1D? If we do not want a PE the maybe just stack more Encoder Layers could perform better?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2558420,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2023-12-12T06:43:01.667000",
              "content": "<p>You have to have either PE or conv. Study the architecture of transformers and you will see.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2554789,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-09T12:06:27.853000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2557048,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-12-11T06:51:27.150000",
          "content": "<p>Ablation study results may vary depending on the final model, but if I remember correctly, GRU gave approximately -0.0003. And the features you pointed out were the most important. I believe both CapR and eternafold mfe are -0.0002.</p>\n<p>Congratulations on third place too! I'm very excited for your next move.👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2553982": "Thanks to Kaggle and the hosts for organizing this competition.  It was truly inspiring and challenging, and I learned a lot from this one.👍\n\n###Code: https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution\n\n# TLDR\n\n**Keypoints:**\n\n- Squeezeformer[1] + GRU head.\n- Simple Conv2DNet for bpp, adding it as a bias to the attention matrix.\n- ALiBi positional encoding[2] for robust generalization on longer sequences.\n- Weighted loss with signal_to_noise, with longer epochs.\n- Additional features for minor score improvements.\n\nI adopted Squeezeformer, which I became familiar with after the ASL fingerspelling competition.Thanks to @christofhenkel and @goldenlock for their solutions in the last ASL competition. The most crucial part of my solution is how to utilize the bpp matrix. I applied a simple shallow Conv2DNet to bpp and directly added it to the attention matrix.\n\n**Features:**\n\nI used some features found useful in the OpenVaccine Challenge, to help fast initial convergence. These included:\n\n- CapR looptype.\n- eternafold mfe.\n- predicted Looptype with eternafold mfe.\n- bpp features (sum, nzero, max).\n\nHowever, unlike in the OpenVaccine challenge, these features only marginally helped (about -0.0005). Therefore, I believe these features should be removed in the future for the simplicity.\n\n#Model\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Fb3c99f0b3ace13f5c4381c094b899d46%2F-3.drawio-2.png?generation=1702054547887335&alt=media)\n\n**Squeezeformer Encoder:**\n\nI chose Squeezeformer with minor modifications (BN after conv1d, SwiGLU in FFN, etc.), which mixes Conv1D blocks with Transformer. While I tried other recent Conv-Transformer Hybrid architectures, Squeezeformer was the most efficient. Compared to a Vanilla Transformer, Squeezeformer showed strong performance early in training and consistently showed faster convergence.\n\nThe models used the following parameters: dim=192, num_heads=4, kernel_size=17, num_layers=12.\n\n**GRU head:**\n\nAdding a single GRU layer after the encoder yielded minor improvements.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F3191b22f70fe4b38f04b6664871d2b96%2F.drawio-10.png?generation=1702054568362614&alt=media)\n\n**ALiBi positional encoding:**\n\nI adopted AliBi positional encoding as it claimed to generalize better over long sequences than other methods.\n\n**BPP as Attention Bias:**\n\nThe bpp matrix (using only the provided one) was added as a bias in the attention matrix after multiplied by per-head predefined scales, significantly improved performance (around -0.0025).\n\n**BPP 2DConvNet:**\n\nUsing Bpp directly as an attention bias was a good start, but I felt it needed more flexibility(I felt it was too sparse). Among various options, adding a 2D CNN on top of the BPP matrix proved very helpful (-0.002). However, multiple 2D CNNs applied to the BPP matrix (usually 206 x 206) were inefficient in terms of training/inference time. Thus, I just used a simple shallow 2-layer 2DCNN, with the output matrix shared across all Transformer block layers.\n\n#Training\n\n- Epochs: 200.\n- Batch size: 256.\n- Learning rate: 2e-3, with Cosine Decay and warmup.\n- Optimizer: AdamW, weight decay = 0.01.\n- Loss: Weighted MAE (weight = log1p(signal_to_noise).clip(0,10)).\n\n\nThe single model CV (K-fold, k=5) scored 0.119, with a public LB of 0.140 and a private LB of 0.142. After ensemble on different seed I got public LB of 0.135 and private LB of 0.140.\nAlthough there were some questionable correlations between CV/LB in certain submissions, generally, they aligned well with the CV.\nWith the above setup, training a single model took around 30 hours on a single RTX 4090. \n\n##Discarded ideas & Thoughts:\n\n- **Self-Supervised Learning (SSL)**: At first I motivated to participate in this competition as it would be really nice if any SSL method could be successfully applied without using any features other than the sequence.  Initial trials with Data2Vec and BERT-like SSL methods showed inconsistent improvements. Due to the additional training time required, I did not consider SSL further. However, I believe there is still huge potential in this idea.\n- **Large Models**: Attempts to train larger models (dim > 512) with proper regularizations were unsuccessful. I think this and SSL failure suggests that the primary challenge lies in the inherent noise within the training dataset.\n- **Augmentations**: Most augmentation methods I tried had no effect.\n- **Pseudo Labels**: While pseudo labeling might help in LB, it didn't improve CV in my case, so I didn't use it for safety&training time. However, after seeing the correlation between public and private LB, I think it might have been slightly beneficial in both public and private LB.\n\n\nFor me, this competition was a series of choices regarding whether to experiment with or adopt some promising ideas, especially when there were only 2-3 weeks left. Some of the ideas I thought might be helpful were abandoned without further consideration because they required more time for implementation and training. I think that this strategy may have made my solution somewhat suboptimal or redundant, but overall I see it worked quite well as my solution appeared to capture most of the crucial aspects of other teams' solutions.\n\n##References\n[1]Sehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W. Mahoney, and Kurt Keutzer. 2022. Squeezeformer: An Efficient Transformer for Automatic Speech Recognition. arXiv:2206.00888 [eess.AS]. [https://doi.org/10.48550/arXiv.2206.00888](url)\n\n[2]Ofir Press, Noah A. Smith, Mike Lewis. 2022. Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. arXiv:2108.12409 [cs.CL]. [https://doi.org/10.48550/arXiv.2108.12409](url)\n\n",
    "2554728": "Thanks for the details and congrats on 2nd position, \n\nBPP is not available during test time, so how did the test inference differ?",
    "2554056": "Oh, one more question, if I may: have you tried AWP? I could not make it work, but I thought that since you are the master of AWP, maybe you could make it work or at least explain why it is less useful here...",
    "2553998": "Congrats on achieving such a high rank again! \n\nLooking forward to learning from your code ♥️",
    "2554009": "When I saw that you joined the competition, I was sure you would end up at the top due to the similarities to ASL and your expertise.\nI used your architecture from ASL with small modifications, and it worked pretty well. Maybe next time, I will also try finally Squeezeformer (if you publish tensorflow code, it will be the best, hehe).\nI want to ask one question… Is Alibi PE really necessary? You, of all people, should be familiar with the fact that 1d conv enables the models to encode positions 'natively.' I chose the no PE path, which I learned from you, and it was really great. If Alibi helps even further, it's good to know.",
    "2554789": ""
  }
}