{
  "id": 460250,
  "title": "5th Place Solution",
  "url": "/competitions/stanford-ribonanza-rna-folding/writeups/r-l-j-5th-place-solution",
  "author_name": "",
  "post_date": "2024-09-04T16:50:30.707Z",
  "votes": 17,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I would like to congratulate my teammates <a href=\"https://www.kaggle.com/ayaanjang\" target=\"_blank\">@ayaanjang</a> and <a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> on this hard earned (master triplet) win! We each had our own takes on the competition and they were combined for the final ensemble. Below we present the most important performance contributors; each of our transformer based models in the ensemble used a different combination of the following methods.</p>\n<h2>Code</h2>\n<ul>\n<li>Roger: <a href=\"https://github.com/s-rog/StanfordRibonanza2023\" target=\"_blank\">https://github.com/s-rog/StanfordRibonanza2023</a></li>\n<li>Ayaan: <a href=\"https://github.com/ehdgnsdl/2023-Stanford-Ribonanza-RNA-Folding\" target=\"_blank\">https://github.com/ehdgnsdl/2023-Stanford-Ribonanza-RNA-Folding</a></li>\n<li>Junseong: <a href=\"https://github.com/JunSeongLee1102/2023-Kaggle-Stanford-Ribonanza-RNA-Folding\" target=\"_blank\">https://github.com/JunSeongLee1102/2023-Kaggle-Stanford-Ribonanza-RNA-Folding</a></li>\n<li>Docker <a href=\"https://hub.docker.com/layers/junseonglee1102/ngc-custom/xformer-tmux/images/sha256-c98a8b37c2134b268056f5a77797b8f7e847938530d1fa603a4e1422153e1ada?context=repo\" target=\"_blank\">Image</a></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2F6ececb933c804b2a3bcd0f5c1bba7845%2FRibonanza.png?generation=1702124640894983&amp;alt=media\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2F69cd9cdfc5a3f955c5d9c903a170f112%2Fimage.png?generation=1702223532569356&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2Fea9f3b82fec44453abb87485457db0fd%2Fjunseong_model_figure.png?generation=1702267969051338&amp;alt=media\"></p>\n<h2>TLDR</h2>\n<p>We started with <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>'s baseline model and made the following changes:</p>\n<ul>\n<li>Utilize lower SN samples</li>\n<li>Utilize Eterna BPP</li>\n<li>Replace layer norm with RMS norm</li>\n<li>Replace absolute positional embedding with relative positional bias</li>\n<li>Replace encoder layer pre-norm layout with ResiDual norm layout</li>\n<li>Remove bias from both QKV and FFN layers</li>\n<li>Increase head size from 32 to 48</li>\n<li>Add a single layer bidirectional GRU at the end of each encoder layer</li>\n<li>Multi-stage pseudo label training</li>\n<li>Misc: inverse square root LR schedule, 0.1 weight decay</li>\n</ul>\n<h2>Data Sampling</h2>\n<p>I set a fixed number of sequences per epoch and used <code>WeightedRandomSampler</code> to draw samples without replacement weighted by SN. A bias term is added to the SN as a hyperparameter and is set on a schedule, increasing the amount of data used and regularization throughout training. The schedule hyperparameter can look something like:<br>\n<code>{(i * 10): v / 10 for i, v in enumerate(range(-12, -7))}</code></p>\n<h2>Data Augmentation</h2>\n<p>We were able to utilize flip to an extent. While training with flip usually resulted in longer training and marginally worse CV, there is a boost to be had through flip TTA.<br>\nAugmenting BPP with gaussian noise also improved CV and long range knot visibility.</p>\n<h2>Norm</h2>\n<p>I experimented with pre-norm, post-norm with <a href=\"https://arxiv.org/abs/2004.08249\" target=\"_blank\">Admin</a> and <a href=\"https://arxiv.org/abs/2304.14802\" target=\"_blank\">ResiDual</a>-norm. ResiDual gave the best CV as well as the best training stability. Changing from layer to RMS norm gave another slight boost.</p>\n<h2>Attention Bias: Positional</h2>\n<p>Relative positional bias is used in all 12 layers of the encoder during attention (all 12 layers use the same positional bias). Alibi and dynamic positional biases were also tested, but relative performed the best with our architecture. Rotary embedding was also worse than relative. Credits to <a href=\"https://www.kaggle.com/lucidrains\" target=\"_blank\">@lucidrains</a> for the <a href=\"https://github.com/lucidrains/x-transformers/blob/main/x_transformers/x_transformers.py#L252\" target=\"_blank\">implementation</a>.</p>\n<h2>GRU</h2>\n<p><a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> added a GRU at the end of each encoder layer which improved perf when we could not scale the attention layers further (the residual output from ResiDual does not interact with the GRU). I included this in my model and experimented with many layouts but it seems like simple is best. I also tried replacing the GRU with a 1D FusedMBConv block, which improved CV but worsened LB.</p>\n<h2>Multi-head RNN and LSTM</h2>\n<p>Similar to the multi-head attention structure, multiple RNN layers can replace single RNN layers by halving their hidden dimension (increasing # of layers and decreasing the hidden dimension offset each other). <a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> tested 2, 3, 4, 8 heads instead of single RNN layer. Among them, 2 head RNN structures combining a GRU layer and a LSTM layer gave the best performance.</p>\n<h2>BPP BMM Convolutional Block</h2>\n<p><a href=\"https://www.kaggle.com/ayaanjang\" target=\"_blank\">@ayaanjang</a> used <a href=\"https://www.kaggle.com/code/nyanpn/6th-place-cnn-gcn\" target=\"_blank\">\"1D conv + ResidualBPPAttention (using bmm)\"</a> to integrate BPP, see the following code:</p>\n<pre><code> :\n    def :\n        super.\n        self.conv1 = \n        self.conv2 = \n        self.relu = nn.\n\n    def forward(self, src, attn):\n        h = self.conv2(self.conv1(torch.bmm(src, attn)))\n        return self.relu(src + h)\n\n :\n    def :\n        super.\n        self.conv = nn.\n        self.bn = nn.\n        self.relu = nn.\n        self.dropout = nn.\n\n    def forward(self, src):\n        return self.dropout(self.relu(self.bn(self.conv(src))))\n</code></pre>\n<p>This BPP attention block was added between the FFN and the GRU, also not affecting the residual output from ResiDual. The inclusion of BPP greatly reduced CV.</p>\n<h2>Attention Bias: BPP</h2>\n<p>I used a different method to add BPP. In the first 6 layers of the encoder, BPP is added to the existing positional/mask bias after a convolution (every layer uses the same BPP derived bias). With 6 heads, the single channel BPP is passed to a 1x1 convolution layer that outputs 6 channels. Reducing the number of layers BPP is used in improved CV as well as long range knot visibility (12 -&gt; 6 layers). Reducing the number of BPP heads, adding Contra BPP, using 3x3 conv all yielded no improvements.</p>\n<h2>Psuedo Labels: 2 Stage Pipeline</h2>\n<p>To generate psuedo labels, a 5 fold ensemble was trained with flip (and inferenced with TTA) then filtered by stdev quantile 0.75. Filtered valued have their reactivities set to nan, credits to <a href=\"https://www.kaggle.com/ayaanjang\" target=\"_blank\">@ayaanjang</a> for this pipeline. This parquet is then used as train to train a single pretrained model, which is then used to train 5 folds again on the original competition train set. This method from <a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> is quite effective and efficient time wise. The final stage took under 2 hours for each fold on a single 4090.</p>\n<h2>Psuedo Labels: 3 Stage Pipeline</h2>\n<p><a href=\"https://www.kaggle.com/ayaanjang\" target=\"_blank\">@ayaanjang</a> used a different pipeline as follows: train dataset only -&gt; train + pseudo dataset -&gt; train dataset only. The LRs for the stages are: 2e-3 -&gt; 2e-4 -&gt; 2e-5. The Pseudo labels were also filtered with the same method above, without filtering <a href=\"https://www.kaggle.com/ayaanjang\" target=\"_blank\">@ayaanjang</a> found there was no increase in performance. The 3rd stage finetuning serves to reduce overfitting.</p>\n<h2>GRU Mystery</h2>\n<p>Now a mystery… Here is my GRU code:</p>\n<pre><code></code></pre>\n<p>At first I added the transpose and flattern operations thinking that they would delegate each direction in the GRU to half the (3) heads or the other way around, mix them between heads. Upon further inspection I think it does neither… but it noticeably improves performance. If you have a satisfactory explanation please comment!</p>\n<h2>Thanks</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/ayaanjang\" target=\"_blank\">@ayaanjang</a> and <a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> again for their hard work</li>\n<li><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> for his excellent baseline, taught me a lot as this is my first transformer model</li>\n<li><a href=\"https://www.kaggle.com/lucidrains\" target=\"_blank\">@lucidrains</a> for his various transformer implementations that I referred to excessively</li>\n<li>The devs behind <a href=\"https://github.com/facebookresearch/xformers\" target=\"_blank\">xformers</a> and <a href=\"https://github.com/NVIDIA/apex\" target=\"_blank\">apex</a> for greatly speeding up training</li>\n<li>Authors of the numerous papers that helped guide our methods</li>\n<li>All kagglers who participated in discussions with us</li>\n<li>Competition hosts for everything!</li>\n</ul>",
  "messages": [
    {
      "id": "2553636",
      "postDate": "12/08/2023 12:12:09",
      "content": "<p>I would like to congratulate my teammates <a href=\"https://www.kaggle.com/ayaanjang\" target=\"_blank\">@ayaanjang</a> and <a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> on this hard earned (master triplet) win! We each had our own takes on the competition and they were combined for the final ensemble. Below we present the most important performance contributors; each of our transformer based models in the ensemble used a different combination of the following methods.</p>\n<h2>Code</h2>\n<ul>\n<li>Roger: <a href=\"https://github.com/s-rog/StanfordRibonanza2023\" target=\"_blank\">https://github.com/s-rog/StanfordRibonanza2023</a></li>\n<li>Ayaan: <a href=\"https://github.com/ehdgnsdl/2023-Stanford-Ribonanza-RNA-Folding\" target=\"_blank\">https://github.com/ehdgnsdl/2023-Stanford-Ribonanza-RNA-Folding</a></li>\n<li>Junseong: <a href=\"https://github.com/JunSeongLee1102/2023-Kaggle-Stanford-Ribonanza-RNA-Folding\" target=\"_blank\">https://github.com/JunSeongLee1102/2023-Kaggle-Stanford-Ribonanza-RNA-Folding</a></li>\n<li>Docker <a href=\"https://hub.docker.com/layers/junseonglee1102/ngc-custom/xformer-tmux/images/sha256-c98a8b37c2134b268056f5a77797b8f7e847938530d1fa603a4e1422153e1ada?context=repo\" target=\"_blank\">Image</a></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2F6ececb933c804b2a3bcd0f5c1bba7845%2FRibonanza.png?generation=1702124640894983&amp;alt=media\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2F69cd9cdfc5a3f955c5d9c903a170f112%2Fimage.png?generation=1702223532569356&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2Fea9f3b82fec44453abb87485457db0fd%2Fjunseong_model_figure.png?generation=1702267969051338&amp;alt=media\"></p>\n<h2>TLDR</h2>\n<p>We started with <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>'s baseline model and made the following changes:</p>\n<ul>\n<li>Utilize lower SN samples</li>\n<li>Utilize Eterna BPP</li>\n<li>Replace layer norm with RMS norm</li>\n<li>Replace absolute positional embedding with relative positional bias</li>\n<li>Replace encoder layer pre-norm layout with ResiDual norm layout</li>\n<li>Remove bias from both QKV and FFN layers</li>\n<li>Increase head size from 32 to 48</li>\n<li>Add a single layer bidirectional GRU at the end of each encoder layer</li>\n<li>Multi-stage pseudo label training</li>\n<li>Misc: inverse square root LR schedule, 0.1 weight decay</li>\n</ul>\n<h2>Data Sampling</h2>\n<p>I set a fixed number of sequences per epoch and used <code>WeightedRandomSampler</code> to draw samples without replacement weighted by SN. A bias term is added to the SN as a hyperparameter and is set on a schedule, increasing the amount of data used and regularization throughout training. The schedule hyperparameter can look something like:<br>\n<code>{(i * 10): v / 10 for i, v in enumerate(range(-12, -7))}</code></p>\n<h2>Data Augmentation</h2>\n<p>We were able to utilize flip to an extent. While training with flip usually resulted in longer training and marginally worse CV, there is a boost to be had through flip TTA.<br>\nAugmenting BPP with gaussian noise also improved CV and long range knot visibility.</p>\n<h2>Norm</h2>\n<p>I experimented with pre-norm, post-norm with <a href=\"https://arxiv.org/abs/2004.08249\" target=\"_blank\">Admin</a> and <a href=\"https://arxiv.org/abs/2304.14802\" target=\"_blank\">ResiDual</a>-norm. ResiDual gave the best CV as well as the best training stability. Changing from layer to RMS norm gave another slight boost.</p>\n<h2>Attention Bias: Positional</h2>\n<p>Relative positional bias is used in all 12 layers of the encoder during attention (all 12 layers use the same positional bias). Alibi and dynamic positional biases were also tested, but relative performed the best with our architecture. Rotary embedding was also worse than relative. Credits to <a href=\"https://www.kaggle.com/lucidrains\" target=\"_blank\">@lucidrains</a> for the <a href=\"https://github.com/lucidrains/x-transformers/blob/main/x_transformers/x_transformers.py#L252\" target=\"_blank\">implementation</a>.</p>\n<h2>GRU</h2>\n<p><a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> added a GRU at the end of each encoder layer which improved perf when we could not scale the attention layers further (the residual output from ResiDual does not interact with the GRU). I included this in my model and experimented with many layouts but it seems like simple is best. I also tried replacing the GRU with a 1D FusedMBConv block, which improved CV but worsened LB.</p>\n<h2>Multi-head RNN and LSTM</h2>\n<p>Similar to the multi-head attention structure, multiple RNN layers can replace single RNN layers by halving their hidden dimension (increasing # of layers and decreasing the hidden dimension offset each other). <a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> tested 2, 3, 4, 8 heads instead of single RNN layer. Among them, 2 head RNN structures combining a GRU layer and a LSTM layer gave the best performance.</p>\n<h2>BPP BMM Convolutional Block</h2>\n<p><a href=\"https://www.kaggle.com/ayaanjang\" target=\"_blank\">@ayaanjang</a> used <a href=\"https://www.kaggle.com/code/nyanpn/6th-place-cnn-gcn\" target=\"_blank\">\"1D conv + ResidualBPPAttention (using bmm)\"</a> to integrate BPP, see the following code:</p>\n<pre><code> :\n    def :\n        super.\n        self.conv1 = \n        self.conv2 = \n        self.relu = nn.\n\n    def forward(self, src, attn):\n        h = self.conv2(self.conv1(torch.bmm(src, attn)))\n        return self.relu(src + h)\n\n :\n    def :\n        super.\n        self.conv = nn.\n        self.bn = nn.\n        self.relu = nn.\n        self.dropout = nn.\n\n    def forward(self, src):\n        return self.dropout(self.relu(self.bn(self.conv(src))))\n</code></pre>\n<p>This BPP attention block was added between the FFN and the GRU, also not affecting the residual output from ResiDual. The inclusion of BPP greatly reduced CV.</p>\n<h2>Attention Bias: BPP</h2>\n<p>I used a different method to add BPP. In the first 6 layers of the encoder, BPP is added to the existing positional/mask bias after a convolution (every layer uses the same BPP derived bias). With 6 heads, the single channel BPP is passed to a 1x1 convolution layer that outputs 6 channels. Reducing the number of layers BPP is used in improved CV as well as long range knot visibility (12 -&gt; 6 layers). Reducing the number of BPP heads, adding Contra BPP, using 3x3 conv all yielded no improvements.</p>\n<h2>Psuedo Labels: 2 Stage Pipeline</h2>\n<p>To generate psuedo labels, a 5 fold ensemble was trained with flip (and inferenced with TTA) then filtered by stdev quantile 0.75. Filtered valued have their reactivities set to nan, credits to <a href=\"https://www.kaggle.com/ayaanjang\" target=\"_blank\">@ayaanjang</a> for this pipeline. This parquet is then used as train to train a single pretrained model, which is then used to train 5 folds again on the original competition train set. This method from <a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> is quite effective and efficient time wise. The final stage took under 2 hours for each fold on a single 4090.</p>\n<h2>Psuedo Labels: 3 Stage Pipeline</h2>\n<p><a href=\"https://www.kaggle.com/ayaanjang\" target=\"_blank\">@ayaanjang</a> used a different pipeline as follows: train dataset only -&gt; train + pseudo dataset -&gt; train dataset only. The LRs for the stages are: 2e-3 -&gt; 2e-4 -&gt; 2e-5. The Pseudo labels were also filtered with the same method above, without filtering <a href=\"https://www.kaggle.com/ayaanjang\" target=\"_blank\">@ayaanjang</a> found there was no increase in performance. The 3rd stage finetuning serves to reduce overfitting.</p>\n<h2>GRU Mystery</h2>\n<p>Now a mystery… Here is my GRU code:</p>\n<pre><code></code></pre>\n<p>At first I added the transpose and flattern operations thinking that they would delegate each direction in the GRU to half the (3) heads or the other way around, mix them between heads. Upon further inspection I think it does neither… but it noticeably improves performance. If you have a satisfactory explanation please comment!</p>\n<h2>Thanks</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/ayaanjang\" target=\"_blank\">@ayaanjang</a> and <a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> again for their hard work</li>\n<li><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> for his excellent baseline, taught me a lot as this is my first transformer model</li>\n<li><a href=\"https://www.kaggle.com/lucidrains\" target=\"_blank\">@lucidrains</a> for his various transformer implementations that I referred to excessively</li>\n<li>The devs behind <a href=\"https://github.com/facebookresearch/xformers\" target=\"_blank\">xformers</a> and <a href=\"https://github.com/NVIDIA/apex\" target=\"_blank\">apex</a> for greatly speeding up training</li>\n<li>Authors of the numerous papers that helped guide our methods</li>\n<li>All kagglers who participated in discussions with us</li>\n<li>Competition hosts for everything!</li>\n</ul>",
      "rawMarkdown": "I would like to congratulate my teammates @ayaanjang and @junseonglee11 on this hard earned (master triplet) win! We each had our own takes on the competition and they were combined for the final ensemble. Below we present the most important performance contributors; each of our transformer based models in the ensemble used a different combination of the following methods.\n\n## Code\n- Roger: https://github.com/s-rog/StanfordRibonanza2023\n- Ayaan: https://github.com/ehdgnsdl/2023-Stanford-Ribonanza-RNA-Folding\n- Junseong: https://github.com/JunSeongLee1102/2023-Kaggle-Stanford-Ribonanza-RNA-Folding\n- Docker [Image](https://hub.docker.com/layers/junseonglee1102/ngc-custom/xformer-tmux/images/sha256-c98a8b37c2134b268056f5a77797b8f7e847938530d1fa603a4e1422153e1ada?context=repo)\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2F6ececb933c804b2a3bcd0f5c1bba7845%2FRibonanza.png?generation=1702124640894983&alt=media\" height=\"520\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2F69cd9cdfc5a3f955c5d9c903a170f112%2Fimage.png?generation=1702223532569356&alt=media\" height=\"520\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2Fea9f3b82fec44453abb87485457db0fd%2Fjunseong_model_figure.png?generation=1702267969051338&alt=media\" height=\"520\">\n\n## TLDR\nWe started with @iafoss's baseline model and made the following changes:\n- Utilize lower SN samples\n- Utilize Eterna BPP\n- Replace layer norm with RMS norm\n- Replace absolute positional embedding with relative positional bias\n- Replace encoder layer pre-norm layout with ResiDual norm layout\n- Remove bias from both QKV and FFN layers\n- Increase head size from 32 to 48\n- Add a single layer bidirectional GRU at the end of each encoder layer\n- Multi-stage pseudo label training\n- Misc: inverse square root LR schedule, 0.1 weight decay\n\n## Data Sampling\nI set a fixed number of sequences per epoch and used `WeightedRandomSampler` to draw samples without replacement weighted by SN. A bias term is added to the SN as a hyperparameter and is set on a schedule, increasing the amount of data used and regularization throughout training. The schedule hyperparameter can look something like:\n`{(i * 10): v / 10 for i, v in enumerate(range(-12, -7))}`\n\n## Data Augmentation\nWe were able to utilize flip to an extent. While training with flip usually resulted in longer training and marginally worse CV, there is a boost to be had through flip TTA.\nAugmenting BPP with gaussian noise also improved CV and long range knot visibility.\n\n## Norm\nI experimented with pre-norm, post-norm with [Admin](https://arxiv.org/abs/2004.08249) and [ResiDual](https://arxiv.org/abs/2304.14802)-norm. ResiDual gave the best CV as well as the best training stability. Changing from layer to RMS norm gave another slight boost.\n\n## Attention Bias: Positional\nRelative positional bias is used in all 12 layers of the encoder during attention (all 12 layers use the same positional bias). Alibi and dynamic positional biases were also tested, but relative performed the best with our architecture. Rotary embedding was also worse than relative. Credits to @lucidrains for the [implementation](https://github.com/lucidrains/x-transformers/blob/main/x_transformers/x_transformers.py#L252).\n\n## GRU\n@junseonglee11 added a GRU at the end of each encoder layer which improved perf when we could not scale the attention layers further (the residual output from ResiDual does not interact with the GRU). I included this in my model and experimented with many layouts but it seems like simple is best. I also tried replacing the GRU with a 1D FusedMBConv block, which improved CV but worsened LB.\n\n## Multi-head RNN and LSTM\nSimilar to the multi-head attention structure, multiple RNN layers can replace single RNN layers by halving their hidden dimension (increasing # of layers and decreasing the hidden dimension offset each other). @junseonglee11 tested 2, 3, 4, 8 heads instead of single RNN layer. Among them, 2 head RNN structures combining a GRU layer and a LSTM layer gave the best performance.\n\n## BPP BMM Convolutional Block\n@ayaanjang used [\"1D conv + ResidualBPPAttention (using bmm)\"](https://www.kaggle.com/code/nyanpn/6th-place-cnn-gcn) to integrate BPP, see the following code:\n```\nclass ResidualBPPAttention(nn.Module):\n    def __init__(self, d_model:int, kernel_size:int, dropout:float):\n        super().__init__()\n        self.conv1 = Conv(d_model, d_model, kernel_size=kernel_size, dropout=dropout)\n        self.conv2 = Conv(d_model, d_model, kernel_size=kernel_size, dropout=dropout)\n        self.relu = nn.ReLU()\n\n    def forward(self, src, attn):\n        h = self.conv2(self.conv1(torch.bmm(src, attn)))\n        return self.relu(src + h)\n\nclass Conv(nn.Module):\n    def __init__(self, d_in:int, d_out:int, kernel_size:int, dropout=0.1):\n        super().__init__()\n        self.conv = nn.Conv1d(d_in, d_out, kernel_size=kernel_size, padding=kernel_size // 2)\n        self.bn = nn.BatchNorm1d(d_out)\n        self.relu = nn.ReLU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, src):\n        return self.dropout(self.relu(self.bn(self.conv(src))))\n```\nThis BPP attention block was added between the FFN and the GRU, also not affecting the residual output from ResiDual. The inclusion of BPP greatly reduced CV.\n\n## Attention Bias: BPP\nI used a different method to add BPP. In the first 6 layers of the encoder, BPP is added to the existing positional/mask bias after a convolution (every layer uses the same BPP derived bias). With 6 heads, the single channel BPP is passed to a 1x1 convolution layer that outputs 6 channels. Reducing the number of layers BPP is used in improved CV as well as long range knot visibility (12 -> 6 layers). Reducing the number of BPP heads, adding Contra BPP, using 3x3 conv all yielded no improvements.\n\n## Psuedo Labels: 2 Stage Pipeline\nTo generate psuedo labels, a 5 fold ensemble was trained with flip (and inferenced with TTA) then filtered by stdev quantile 0.75. Filtered valued have their reactivities set to nan, credits to @ayaanjang for this pipeline. This parquet is then used as train to train a single pretrained model, which is then used to train 5 folds again on the original competition train set. This method from @junseonglee11 is quite effective and efficient time wise. The final stage took under 2 hours for each fold on a single 4090.\n\n## Psuedo Labels: 3 Stage Pipeline\n@ayaanjang used a different pipeline as follows: train dataset only -> train + pseudo dataset -> train dataset only. The LRs for the stages are: 2e-3 -> 2e-4 -> 2e-5. The Pseudo labels were also filtered with the same method above, without filtering @ayaanjang found there was no increase in performance. The 3rd stage finetuning serves to reduce overfitting.\n\n\n## GRU Mystery\nNow a mystery... Here is my GRU code:\n```\nclass GRU(nn.Module):\n    def __init__(self, d_model: int, p_dropout: float):\n        super().__init__()\n        self.gru = nn.GRU(d_model, d_model // 2, batch_first=True, bidirectional=True)\n        self.dropout = nn.Dropout(p_dropout)\n\n    def forward(self, x: Tensor) -> Tensor:\n        B, L, D = *x.shape[:2], x.size(-1) // 2\n        x = x.view(B, L, D, 2).transpose(2, 3).flatten(2)\n        x = self.gru(x)[0]\n        x = x.view(B, L, 2, D).transpose(2, 3).flatten(2)\n        return self.dropout(x)\n```\nAt first I added the transpose and flattern operations thinking that they would delegate each direction in the GRU to half the (3) heads or the other way around, mix them between heads. Upon further inspection I think it does neither... but it noticeably improves performance. If you have a satisfactory explanation please comment!\n\n## Thanks\n- @ayaanjang and @junseonglee11 again for their hard work\n- @iafoss for his excellent baseline, taught me a lot as this is my first transformer model\n- @lucidrains for his various transformer implementations that I referred to excessively\n- The devs behind [xformers](https://github.com/facebookresearch/xformers) and [apex](https://github.com/NVIDIA/apex) for greatly speeding up training\n- Authors of the numerous papers that helped guide our methods\n- All kagglers who participated in discussions with us\n- Competition hosts for everything!",
      "votes": null
    },
    {
      "id": "2553778",
      "postDate": "12/08/2023 14:04:45",
      "content": "<p>Two hours per fold looks amazing, how did you implement attention bias in your MHSA layers so it behaved efficient? </p>",
      "rawMarkdown": "Two hours per fold looks amazing, how did you implement attention bias in your MHSA layers so it behaved efficient?",
      "votes": null
    },
    {
      "id": "2553992",
      "postDate": "12/08/2023 17:22:29",
      "content": "<p>I used cutlass in xformer's memory efficient attention. Not much training was needed in the final stage, around 20 epochs (each epoch being 20k seqs) as the model overfits quickly starting from the pretrained checkpoint.</p>",
      "rawMarkdown": "I used cutlass in xformer's memory efficient attention. Not much training was needed in the final stage, around 20 epochs (each epoch being 20k seqs) as the model overfits quickly starting from the pretrained checkpoint.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2553778,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "12/08/2023 14:04:45",
      "content": "<p>Two hours per fold looks amazing, how did you implement attention bias in your MHSA layers so it behaved efficient? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2553992,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "12/08/2023 17:22:29",
          "content": "<p>I used cutlass in xformer's memory efficient attention. Not much training was needed in the final stage, around 20 epochs (each epoch being 20k seqs) as the model overfits quickly starting from the pretrained checkpoint.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2553636": "I would like to congratulate my teammates @ayaanjang and @junseonglee11 on this hard earned (master triplet) win! We each had our own takes on the competition and they were combined for the final ensemble. Below we present the most important performance contributors; each of our transformer based models in the ensemble used a different combination of the following methods.\n\n## Code\n- Roger: https://github.com/s-rog/StanfordRibonanza2023\n- Ayaan: https://github.com/ehdgnsdl/2023-Stanford-Ribonanza-RNA-Folding\n- Junseong: https://github.com/JunSeongLee1102/2023-Kaggle-Stanford-Ribonanza-RNA-Folding\n- Docker [Image](https://hub.docker.com/layers/junseonglee1102/ngc-custom/xformer-tmux/images/sha256-c98a8b37c2134b268056f5a77797b8f7e847938530d1fa603a4e1422153e1ada?context=repo)\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2F6ececb933c804b2a3bcd0f5c1bba7845%2FRibonanza.png?generation=1702124640894983&alt=media\" height=\"520\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2F69cd9cdfc5a3f955c5d9c903a170f112%2Fimage.png?generation=1702223532569356&alt=media\" height=\"520\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2Fea9f3b82fec44453abb87485457db0fd%2Fjunseong_model_figure.png?generation=1702267969051338&alt=media\" height=\"520\">\n\n## TLDR\nWe started with @iafoss's baseline model and made the following changes:\n- Utilize lower SN samples\n- Utilize Eterna BPP\n- Replace layer norm with RMS norm\n- Replace absolute positional embedding with relative positional bias\n- Replace encoder layer pre-norm layout with ResiDual norm layout\n- Remove bias from both QKV and FFN layers\n- Increase head size from 32 to 48\n- Add a single layer bidirectional GRU at the end of each encoder layer\n- Multi-stage pseudo label training\n- Misc: inverse square root LR schedule, 0.1 weight decay\n\n## Data Sampling\nI set a fixed number of sequences per epoch and used `WeightedRandomSampler` to draw samples without replacement weighted by SN. A bias term is added to the SN as a hyperparameter and is set on a schedule, increasing the amount of data used and regularization throughout training. The schedule hyperparameter can look something like:\n`{(i * 10): v / 10 for i, v in enumerate(range(-12, -7))}`\n\n## Data Augmentation\nWe were able to utilize flip to an extent. While training with flip usually resulted in longer training and marginally worse CV, there is a boost to be had through flip TTA.\nAugmenting BPP with gaussian noise also improved CV and long range knot visibility.\n\n## Norm\nI experimented with pre-norm, post-norm with [Admin](https://arxiv.org/abs/2004.08249) and [ResiDual](https://arxiv.org/abs/2304.14802)-norm. ResiDual gave the best CV as well as the best training stability. Changing from layer to RMS norm gave another slight boost.\n\n## Attention Bias: Positional\nRelative positional bias is used in all 12 layers of the encoder during attention (all 12 layers use the same positional bias). Alibi and dynamic positional biases were also tested, but relative performed the best with our architecture. Rotary embedding was also worse than relative. Credits to @lucidrains for the [implementation](https://github.com/lucidrains/x-transformers/blob/main/x_transformers/x_transformers.py#L252).\n\n## GRU\n@junseonglee11 added a GRU at the end of each encoder layer which improved perf when we could not scale the attention layers further (the residual output from ResiDual does not interact with the GRU). I included this in my model and experimented with many layouts but it seems like simple is best. I also tried replacing the GRU with a 1D FusedMBConv block, which improved CV but worsened LB.\n\n## Multi-head RNN and LSTM\nSimilar to the multi-head attention structure, multiple RNN layers can replace single RNN layers by halving their hidden dimension (increasing # of layers and decreasing the hidden dimension offset each other). @junseonglee11 tested 2, 3, 4, 8 heads instead of single RNN layer. Among them, 2 head RNN structures combining a GRU layer and a LSTM layer gave the best performance.\n\n## BPP BMM Convolutional Block\n@ayaanjang used [\"1D conv + ResidualBPPAttention (using bmm)\"](https://www.kaggle.com/code/nyanpn/6th-place-cnn-gcn) to integrate BPP, see the following code:\n```\nclass ResidualBPPAttention(nn.Module):\n    def __init__(self, d_model:int, kernel_size:int, dropout:float):\n        super().__init__()\n        self.conv1 = Conv(d_model, d_model, kernel_size=kernel_size, dropout=dropout)\n        self.conv2 = Conv(d_model, d_model, kernel_size=kernel_size, dropout=dropout)\n        self.relu = nn.ReLU()\n\n    def forward(self, src, attn):\n        h = self.conv2(self.conv1(torch.bmm(src, attn)))\n        return self.relu(src + h)\n\nclass Conv(nn.Module):\n    def __init__(self, d_in:int, d_out:int, kernel_size:int, dropout=0.1):\n        super().__init__()\n        self.conv = nn.Conv1d(d_in, d_out, kernel_size=kernel_size, padding=kernel_size // 2)\n        self.bn = nn.BatchNorm1d(d_out)\n        self.relu = nn.ReLU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, src):\n        return self.dropout(self.relu(self.bn(self.conv(src))))\n```\nThis BPP attention block was added between the FFN and the GRU, also not affecting the residual output from ResiDual. The inclusion of BPP greatly reduced CV.\n\n## Attention Bias: BPP\nI used a different method to add BPP. In the first 6 layers of the encoder, BPP is added to the existing positional/mask bias after a convolution (every layer uses the same BPP derived bias). With 6 heads, the single channel BPP is passed to a 1x1 convolution layer that outputs 6 channels. Reducing the number of layers BPP is used in improved CV as well as long range knot visibility (12 -> 6 layers). Reducing the number of BPP heads, adding Contra BPP, using 3x3 conv all yielded no improvements.\n\n## Psuedo Labels: 2 Stage Pipeline\nTo generate psuedo labels, a 5 fold ensemble was trained with flip (and inferenced with TTA) then filtered by stdev quantile 0.75. Filtered valued have their reactivities set to nan, credits to @ayaanjang for this pipeline. This parquet is then used as train to train a single pretrained model, which is then used to train 5 folds again on the original competition train set. This method from @junseonglee11 is quite effective and efficient time wise. The final stage took under 2 hours for each fold on a single 4090.\n\n## Psuedo Labels: 3 Stage Pipeline\n@ayaanjang used a different pipeline as follows: train dataset only -> train + pseudo dataset -> train dataset only. The LRs for the stages are: 2e-3 -> 2e-4 -> 2e-5. The Pseudo labels were also filtered with the same method above, without filtering @ayaanjang found there was no increase in performance. The 3rd stage finetuning serves to reduce overfitting.\n\n\n## GRU Mystery\nNow a mystery... Here is my GRU code:\n```\nclass GRU(nn.Module):\n    def __init__(self, d_model: int, p_dropout: float):\n        super().__init__()\n        self.gru = nn.GRU(d_model, d_model // 2, batch_first=True, bidirectional=True)\n        self.dropout = nn.Dropout(p_dropout)\n\n    def forward(self, x: Tensor) -> Tensor:\n        B, L, D = *x.shape[:2], x.size(-1) // 2\n        x = x.view(B, L, D, 2).transpose(2, 3).flatten(2)\n        x = self.gru(x)[0]\n        x = x.view(B, L, 2, D).transpose(2, 3).flatten(2)\n        return self.dropout(x)\n```\nAt first I added the transpose and flattern operations thinking that they would delegate each direction in the GRU to half the (3) heads or the other way around, mix them between heads. Upon further inspection I think it does neither... but it noticeably improves performance. If you have a satisfactory explanation please comment!\n\n## Thanks\n- @ayaanjang and @junseonglee11 again for their hard work\n- @iafoss for his excellent baseline, taught me a lot as this is my first transformer model\n- @lucidrains for his various transformer implementations that I referred to excessively\n- The devs behind [xformers](https://github.com/facebookresearch/xformers) and [apex](https://github.com/NVIDIA/apex) for greatly speeding up training\n- Authors of the numerous papers that helped guide our methods\n- All kagglers who participated in discussions with us\n- Competition hosts for everything!",
    "2553778": "Two hours per fold looks amazing, how did you implement attention bias in your MHSA layers so it behaved efficient?",
    "2553992": "I used cutlass in xformer's memory efficient attention. Not much training was needed in the final stage, around 20 epochs (each epoch being 20k seqs) as the model overfits quickly starting from the pretrained checkpoint."
  },
  "source": "meta"
}