{
  "id": 460203,
  "title": "4th place solution",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460203",
  "author_name": "tattaka",
  "post_date": "2023-12-08T08:16:55.909000",
  "votes": 29,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First of all, a big thank you to Kaggle staff and host for providing a fun competition.  <br>\n<a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">tattaka</a> and <a href=\"https://www.kaggle.com/fuumin621\" target=\"_blank\">monnu</a> were able to complete their journey to Competitions Grandmaster with this result.</p>\n<h2>Summary</h2>\n<ul>\n<li>Various models ensemble<ul>\n<li>Modified RNAdegformer proposed by Shujun</li>\n<li>1D Conv &amp; Residual BPP Attention</li>\n<li>Transformer with BPP Attention Bias</li></ul></li>\n<li>bp_matrix generated by eternafold (partially contrafold)<ul>\n<li>I didn't realize until halfway through that it was being provided.</li></ul></li>\n<li>finetuning with higher s/n threshold</li>\n<li>Pseudo Labeling using reactivity error prediction</li>\n</ul>\n<h2>Modified RNAdegformer</h2>\n<p>code: <a href=\"https://github.com/tattaka/stanford-ribonanza-rna-folding-public\" target=\"_blank\">https://github.com/tattaka/stanford-ribonanza-rna-folding-public</a> </p>\n<h3>Input</h3>\n<ul>\n<li>Sequence</li>\n<li>BPP Matrix (by EternaFold and Contrafold)</li>\n</ul>\n<h3>Architecture</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fd89808656d85bd143d94055a543a5a55%2F2023-12-09%2018.25.55.png?generation=1702114143461138&amp;alt=media\" alt=\"\"><br>\nMade some changes to the <a href=\"https://academic.oup.com/bib/article/24/1/bbac581/6986359\" target=\"_blank\">RNAdegformer</a>(<a href=\"https://github.com/Shujun-He/RNAdegformer\" target=\"_blank\">https://github.com/Shujun-He/RNAdegformer</a>) proposed by <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189564\" target=\"_blank\">Shujun</a>, including the order of layers.</p>\n<ul>\n<li>kernel_size = 7 except for the last transformer, which is 1</li>\n<li>postnorm</li>\n<li>GLU family activation</li>\n<li><a href=\"https://arxiv.org/abs/2108.12409\" target=\"_blank\">ALiBi</a> positional encoding is applied separately head from bp_matrix</li>\n<li>Other minor changes ensemble</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>model name</th>\n<th>error prediction for pseudo labeling</th>\n<th>use pseudo label</th>\n<th>act_fn for feedforward</th>\n<th>norm layer</th>\n<th>add norm and act for conv1d</th>\n<th>use contrafold(second bpps)</th>\n<th>connect attn_weight to bpps bias</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>exp064</td>\n<td>yes</td>\n<td>no</td>\n<td>SwiGLU</td>\n<td>Layernorm</td>\n<td>no</td>\n<td>no</td>\n<td>no</td>\n<td>0.12087</td>\n</tr>\n<tr>\n<td>exp070</td>\n<td>no</td>\n<td>exp064</td>\n<td>SwiGLU</td>\n<td>RMSnorm</td>\n<td>yes</td>\n<td>no</td>\n<td>no</td>\n<td>0.1199 / tiny: 0.12143</td>\n</tr>\n<tr>\n<td>exp071</td>\n<td>yes</td>\n<td>no</td>\n<td>GeGLU</td>\n<td>RMSnorm</td>\n<td>yes</td>\n<td>yes</td>\n<td>no</td>\n<td>0.12146</td>\n</tr>\n<tr>\n<td>exp072</td>\n<td>no</td>\n<td>exp064 + exp071</td>\n<td>GeGLU</td>\n<td>RMSnorm</td>\n<td>yes</td>\n<td>no</td>\n<td>yes</td>\n<td>0.11976</td>\n</tr>\n</tbody>\n</table>\n<h3>Training</h3>\n<ul>\n<li>simple kfold (k = 5)</li>\n<li>1st stage: First train with sn &gt; 0.5 (300epoch)</li>\n<li>2nd stage: Then finetune with sn &gt; 1.0 for a short number of epochs (15epoch)</li>\n<li>When training with train dataset only, output errors for each nucleotide in addition to reactivity</li>\n<li>For pseudo labels, use sn_pred&gt;0.75 for the 1st stage, sn_pred&gt;1.0 for the 2nd stage, and future=1 only<ul>\n<li>Train from scratch with the pseudo labels added</li></ul></li>\n<li>lr=1e-3, bs=256, AdamW(eps=1e-6), with warmup for 1st stage</li>\n</ul>\n<h3>Score for the single model</h3>\n<p>best model: exp072  </p>\n<ul>\n<li>CV (k-fold): 0.11976</li>\n<li>Public Score: 0.13681</li>\n<li>Private Score: 0.14124</li>\n</ul>\n<h2>1D Conv &amp; Residual BPP Attention</h2>\n<p>code: <a href=\"https://github.com/fuumin621/stanford-ribonanza-rna-folding-4th\" target=\"_blank\">https://github.com/fuumin621/stanford-ribonanza-rna-folding-4th</a></p>\n<p>Based on the 1D Conv + BPP Attention architecture proposed by <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189241\" target=\"_blank\">nyanp</a>, we made improvements.</p>\n<h3>Architecture</h3>\n<ol>\n<li>Sequence Embedding</li>\n<li>Conv 1d</li>\n<li>SE Residual &amp; Residual BPP Attention x 12 layers<ul>\n<li>Using BPP as attention weight</li>\n<li>Reduced the kernel size of 1D convolution as the layer depth increased</li></ul></li>\n<li>Bi-LSTM x 2 layers </li>\n<li>Linear</li>\n</ol>\n<h3>Input</h3>\n<ul>\n<li>Sequence</li>\n<li>BPP Matrix (by EternaFold)</li>\n</ul>\n<h3>Parameter</h3>\n<ul>\n<li>Drop out rate: 0.1</li>\n<li>n_dim: 256</li>\n<li>Kernel size: 9, 7, 5, 3 (decreases as the layer depth increases)</li>\n<li>Learning rate: 4e-3 with cosine scheduler</li>\n<li>Batch size: 64</li>\n</ul>\n<h3>Training</h3>\n<p>The training strategy essentially adopted the same approach as the above Modified RNAdegformer</p>\n<h3>Score for the single model</h3>\n<ul>\n<li>CV (k-fold): 0.12161</li>\n<li>Public Score: 0.13889</li>\n<li>Private Score: 0.1425</li>\n</ul>\n<h3>What Didn't Work</h3>\n<ul>\n<li>BPP packages other than EternaFold (contrafold_2)</li>\n<li>Distance Matrix</li>\n<li>Structure</li>\n<li>BPP Feature Engineering (max, sum, nb_count)</li>\n<li>Sample Weight by SN</li>\n<li>etc…</li>\n</ul>\n<h2>Transformer with BPP Attention Bias (@ren4yu's Part)</h2>\n<p>code: <a href=\"https://github.com/yu4u/kaggle-stanford-ribonanza-rna-folding-4th-place-solution\" target=\"_blank\">https://github.com/yu4u/kaggle-stanford-ribonanza-rna-folding-4th-place-solution</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F7eea07453fa8011a1435c012533a084d%2Ftransformer.png?generation=1702099940846948&amp;alt=media\"></p>\n<p>This model borrows the bpp attention bias idea from the RNAdegformer, but is closer to the original Transformer architecture.<br>\nInstead of using positional embedding, Conv1D was used in FFN to give relative positional information.</p>\n<h3>Training Procedure</h3>\n<ul>\n<li>Train model with AdamW LR=2e-3 to 2e-4, BS=128, use s/n filter threshold=0.5</li>\n<li>Finetune model with AdamW LR=2e-4 to 0, BS=256, use s/n filter threshold=1.0</li>\n</ul>\n<h3>Score for the single model</h3>\n<ul>\n<li>CV (k-fold): 0.12188</li>\n<li>Public Score: 0.13948</li>\n<li>Private Score: 0.14267</li>\n</ul>\n<h3>Does not work for me</h3>\n<ul>\n<li>Increasing dimension or number of layers</li>\n<li>UNet-like hierarchical architecture</li>\n</ul>",
  "messages": [
    {
      "id": 2553394,
      "postDate": "2023-12-08T08:16:55.910Z",
      "content": "<p>First of all, a big thank you to Kaggle staff and host for providing a fun competition.  <br>\n<a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">tattaka</a> and <a href=\"https://www.kaggle.com/fuumin621\" target=\"_blank\">monnu</a> were able to complete their journey to Competitions Grandmaster with this result.</p>\n<h2>Summary</h2>\n<ul>\n<li>Various models ensemble<ul>\n<li>Modified RNAdegformer proposed by Shujun</li>\n<li>1D Conv &amp; Residual BPP Attention</li>\n<li>Transformer with BPP Attention Bias</li></ul></li>\n<li>bp_matrix generated by eternafold (partially contrafold)<ul>\n<li>I didn't realize until halfway through that it was being provided.</li></ul></li>\n<li>finetuning with higher s/n threshold</li>\n<li>Pseudo Labeling using reactivity error prediction</li>\n</ul>\n<h2>Modified RNAdegformer</h2>\n<p>code: <a href=\"https://github.com/tattaka/stanford-ribonanza-rna-folding-public\" target=\"_blank\">https://github.com/tattaka/stanford-ribonanza-rna-folding-public</a> </p>\n<h3>Input</h3>\n<ul>\n<li>Sequence</li>\n<li>BPP Matrix (by EternaFold and Contrafold)</li>\n</ul>\n<h3>Architecture</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fd89808656d85bd143d94055a543a5a55%2F2023-12-09%2018.25.55.png?generation=1702114143461138&amp;alt=media\" alt=\"\"><br>\nMade some changes to the <a href=\"https://academic.oup.com/bib/article/24/1/bbac581/6986359\" target=\"_blank\">RNAdegformer</a>(<a href=\"https://github.com/Shujun-He/RNAdegformer\" target=\"_blank\">https://github.com/Shujun-He/RNAdegformer</a>) proposed by <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189564\" target=\"_blank\">Shujun</a>, including the order of layers.</p>\n<ul>\n<li>kernel_size = 7 except for the last transformer, which is 1</li>\n<li>postnorm</li>\n<li>GLU family activation</li>\n<li><a href=\"https://arxiv.org/abs/2108.12409\" target=\"_blank\">ALiBi</a> positional encoding is applied separately head from bp_matrix</li>\n<li>Other minor changes ensemble</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>model name</th>\n<th>error prediction for pseudo labeling</th>\n<th>use pseudo label</th>\n<th>act_fn for feedforward</th>\n<th>norm layer</th>\n<th>add norm and act for conv1d</th>\n<th>use contrafold(second bpps)</th>\n<th>connect attn_weight to bpps bias</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>exp064</td>\n<td>yes</td>\n<td>no</td>\n<td>SwiGLU</td>\n<td>Layernorm</td>\n<td>no</td>\n<td>no</td>\n<td>no</td>\n<td>0.12087</td>\n</tr>\n<tr>\n<td>exp070</td>\n<td>no</td>\n<td>exp064</td>\n<td>SwiGLU</td>\n<td>RMSnorm</td>\n<td>yes</td>\n<td>no</td>\n<td>no</td>\n<td>0.1199 / tiny: 0.12143</td>\n</tr>\n<tr>\n<td>exp071</td>\n<td>yes</td>\n<td>no</td>\n<td>GeGLU</td>\n<td>RMSnorm</td>\n<td>yes</td>\n<td>yes</td>\n<td>no</td>\n<td>0.12146</td>\n</tr>\n<tr>\n<td>exp072</td>\n<td>no</td>\n<td>exp064 + exp071</td>\n<td>GeGLU</td>\n<td>RMSnorm</td>\n<td>yes</td>\n<td>no</td>\n<td>yes</td>\n<td>0.11976</td>\n</tr>\n</tbody>\n</table>\n<h3>Training</h3>\n<ul>\n<li>simple kfold (k = 5)</li>\n<li>1st stage: First train with sn &gt; 0.5 (300epoch)</li>\n<li>2nd stage: Then finetune with sn &gt; 1.0 for a short number of epochs (15epoch)</li>\n<li>When training with train dataset only, output errors for each nucleotide in addition to reactivity</li>\n<li>For pseudo labels, use sn_pred&gt;0.75 for the 1st stage, sn_pred&gt;1.0 for the 2nd stage, and future=1 only<ul>\n<li>Train from scratch with the pseudo labels added</li></ul></li>\n<li>lr=1e-3, bs=256, AdamW(eps=1e-6), with warmup for 1st stage</li>\n</ul>\n<h3>Score for the single model</h3>\n<p>best model: exp072  </p>\n<ul>\n<li>CV (k-fold): 0.11976</li>\n<li>Public Score: 0.13681</li>\n<li>Private Score: 0.14124</li>\n</ul>\n<h2>1D Conv &amp; Residual BPP Attention</h2>\n<p>code: <a href=\"https://github.com/fuumin621/stanford-ribonanza-rna-folding-4th\" target=\"_blank\">https://github.com/fuumin621/stanford-ribonanza-rna-folding-4th</a></p>\n<p>Based on the 1D Conv + BPP Attention architecture proposed by <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189241\" target=\"_blank\">nyanp</a>, we made improvements.</p>\n<h3>Architecture</h3>\n<ol>\n<li>Sequence Embedding</li>\n<li>Conv 1d</li>\n<li>SE Residual &amp; Residual BPP Attention x 12 layers<ul>\n<li>Using BPP as attention weight</li>\n<li>Reduced the kernel size of 1D convolution as the layer depth increased</li></ul></li>\n<li>Bi-LSTM x 2 layers </li>\n<li>Linear</li>\n</ol>\n<h3>Input</h3>\n<ul>\n<li>Sequence</li>\n<li>BPP Matrix (by EternaFold)</li>\n</ul>\n<h3>Parameter</h3>\n<ul>\n<li>Drop out rate: 0.1</li>\n<li>n_dim: 256</li>\n<li>Kernel size: 9, 7, 5, 3 (decreases as the layer depth increases)</li>\n<li>Learning rate: 4e-3 with cosine scheduler</li>\n<li>Batch size: 64</li>\n</ul>\n<h3>Training</h3>\n<p>The training strategy essentially adopted the same approach as the above Modified RNAdegformer</p>\n<h3>Score for the single model</h3>\n<ul>\n<li>CV (k-fold): 0.12161</li>\n<li>Public Score: 0.13889</li>\n<li>Private Score: 0.1425</li>\n</ul>\n<h3>What Didn't Work</h3>\n<ul>\n<li>BPP packages other than EternaFold (contrafold_2)</li>\n<li>Distance Matrix</li>\n<li>Structure</li>\n<li>BPP Feature Engineering (max, sum, nb_count)</li>\n<li>Sample Weight by SN</li>\n<li>etc…</li>\n</ul>\n<h2>Transformer with BPP Attention Bias (@ren4yu's Part)</h2>\n<p>code: <a href=\"https://github.com/yu4u/kaggle-stanford-ribonanza-rna-folding-4th-place-solution\" target=\"_blank\">https://github.com/yu4u/kaggle-stanford-ribonanza-rna-folding-4th-place-solution</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F7eea07453fa8011a1435c012533a084d%2Ftransformer.png?generation=1702099940846948&amp;alt=media\"></p>\n<p>This model borrows the bpp attention bias idea from the RNAdegformer, but is closer to the original Transformer architecture.<br>\nInstead of using positional embedding, Conv1D was used in FFN to give relative positional information.</p>\n<h3>Training Procedure</h3>\n<ul>\n<li>Train model with AdamW LR=2e-3 to 2e-4, BS=128, use s/n filter threshold=0.5</li>\n<li>Finetune model with AdamW LR=2e-4 to 0, BS=256, use s/n filter threshold=1.0</li>\n</ul>\n<h3>Score for the single model</h3>\n<ul>\n<li>CV (k-fold): 0.12188</li>\n<li>Public Score: 0.13948</li>\n<li>Private Score: 0.14267</li>\n</ul>\n<h3>Does not work for me</h3>\n<ul>\n<li>Increasing dimension or number of layers</li>\n<li>UNet-like hierarchical architecture</li>\n</ul>",
      "rawMarkdown": "First of all, a big thank you to Kaggle staff and host for providing a fun competition.  \n[tattaka](https://www.kaggle.com/tattaka) and [monnu](https://www.kaggle.com/fuumin621) were able to complete their journey to Competitions Grandmaster with this result.\n\n## Summary\n* Various models ensemble\n  * Modified RNAdegformer proposed by Shujun\n  * 1D Conv & Residual BPP Attention\n  * Transformer with BPP Attention Bias\n* bp_matrix generated by eternafold (partially contrafold)\n  * I didn't realize until halfway through that it was being provided.\n* finetuning with higher s/n threshold\n* Pseudo Labeling using reactivity error prediction\n\n## Modified RNAdegformer\ncode: https://github.com/tattaka/stanford-ribonanza-rna-folding-public \n\n### Input\n* Sequence\n* BPP Matrix (by EternaFold and Contrafold)\n\n### Architecture\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fd89808656d85bd143d94055a543a5a55%2F2023-12-09%2018.25.55.png?generation=1702114143461138&alt=media)\nMade some changes to the [RNAdegformer](https://academic.oup.com/bib/article/24/1/bbac581/6986359)(https://github.com/Shujun-He/RNAdegformer) proposed by [Shujun](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189564), including the order of layers.\n* kernel_size = 7 except for the last transformer, which is 1\n* postnorm\n* GLU family activation\n* [ALiBi](https://arxiv.org/abs/2108.12409) positional encoding is applied separately head from bp_matrix\n* Other minor changes ensemble\n\n| model name | error prediction for pseudo labeling | use pseudo label | act_fn for feedforward | norm layer | add norm and act for conv1d | use contrafold(second bpps) | connect attn_weight to bpps bias | CV                     |\n| ---------- | ------------------------------------ | ---------------- | ---------------------- | ---------- | --------------------------- | --------------------------- | -------------------------------- | ---------------------- |\n| exp064     | yes                                  | no               | SwiGLU                 | Layernorm  | no                          | no                          | no                               | 0.12087                |\n| exp070     | no                                   | exp064           | SwiGLU                 | RMSnorm    | yes                         | no                          | no                               | 0.1199 / tiny: 0.12143 |\n| exp071     | yes                                  | no               | GeGLU                  | RMSnorm    | yes                         | yes                         | no                               | 0.12146                |\n| exp072     | no                                   | exp064 + exp071  | GeGLU                  | RMSnorm    | yes                         | no                          | yes                              | 0.11976                     |\n\n### Training\n* simple kfold (k = 5)\n* 1st stage: First train with sn > 0.5 (300epoch)\n* 2nd stage: Then finetune with sn > 1.0 for a short number of epochs (15epoch)\n* When training with train dataset only, output errors for each nucleotide in addition to reactivity\n* For pseudo labels, use sn_pred>0.75 for the 1st stage, sn_pred>1.0 for the 2nd stage, and future=1 only\n  * Train from scratch with the pseudo labels added\n* lr=1e-3, bs=256, AdamW(eps=1e-6), with warmup for 1st stage\n\n### Score for the single model\nbest model: exp072  \n* CV (k-fold): 0.11976\n* Public Score: 0.13681\n* Private Score: 0.14124\n\n## 1D Conv & Residual BPP Attention\ncode: https://github.com/fuumin621/stanford-ribonanza-rna-folding-4th\n\nBased on the 1D Conv + BPP Attention architecture proposed by [nyanp](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189241), we made improvements.\n\n### Architecture\n1. Sequence Embedding\n2. Conv 1d\n3. SE Residual & Residual BPP Attention x 12 layers\n   * Using BPP as attention weight\n   * Reduced the kernel size of 1D convolution as the layer depth increased\n4. Bi-LSTM x 2 layers \n5. Linear\n\n### Input\n* Sequence\n* BPP Matrix (by EternaFold)\n\n### Parameter\n* Drop out rate: 0.1\n* n_dim: 256\n* Kernel size: 9, 7, 5, 3 (decreases as the layer depth increases)\n* Learning rate: 4e-3 with cosine scheduler\n* Batch size: 64\n\n### Training\nThe training strategy essentially adopted the same approach as the above Modified RNAdegformer\n\n### Score for the single model\n* CV (k-fold): 0.12161\n* Public Score: 0.13889\n* Private Score: 0.1425\n\n### What Didn't Work\n* BPP packages other than EternaFold (contrafold_2)\n* Distance Matrix\n* Structure\n* BPP Feature Engineering (max, sum, nb_count)\n* Sample Weight by SN\n* etc...\n\n## Transformer with BPP Attention Bias (@ren4yu's Part)\n\ncode: https://github.com/yu4u/kaggle-stanford-ribonanza-rna-folding-4th-place-solution\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F7eea07453fa8011a1435c012533a084d%2Ftransformer.png?generation=1702099940846948&alt=media\" width=\"50%\">\n\nThis model borrows the bpp attention bias idea from the RNAdegformer, but is closer to the original Transformer architecture.\nInstead of using positional embedding, Conv1D was used in FFN to give relative positional information.\n### Training Procedure\n- Train model with AdamW LR=2e-3 to 2e-4, BS=128, use s/n filter threshold=0.5\n- Finetune model with AdamW LR=2e-4 to 0, BS=256, use s/n filter threshold=1.0\n### Score for the single model\n* CV (k-fold): 0.12188\n* Public Score: 0.13948\n* Private Score: 0.14267\n### Does not work for me\n- Increasing dimension or number of layers\n- UNet-like hierarchical architecture\n\n",
      "votes": 29
    },
    {
      "id": 2553449,
      "postDate": "2023-12-08T08:56:19.663Z",
      "content": "<p>Congratulations on achieving 4th position in this competition. Thanks for sharing details of your solution. </p>",
      "rawMarkdown": "Congratulations on achieving 4th position in this competition. Thanks for sharing details of your solution. "
    },
    {
      "id": 2553439,
      "postDate": "2023-12-08T08:46:48.940Z",
      "content": "<p>Thanks for sharing and congratulations on reaching the GM.</p>",
      "rawMarkdown": "Thanks for sharing and congratulations on reaching the GM."
    },
    {
      "id": 2553437,
      "postDate": "2023-12-08T08:45:44.220Z",
      "content": "<p>Congratulations on you and your team，may I ask how large is your machine memory, I tried to load BPP Matrix but it caused OOM error.</p>",
      "rawMarkdown": "Congratulations on you and your team，may I ask how large is your machine memory, I tried to load BPP Matrix but it caused OOM error.",
      "replies": [
        {
          "id": 2553441,
          "postDate": "2023-12-08T08:50:43.073Z",
          "content": "<p>In my experiment, A4000 (16GB) x 4, bs=64 for each GPU (bs=32 for exp072)</p>",
          "rawMarkdown": "In my experiment, A4000 (16GB) x 4, bs=64 for each GPU (bs=32 for exp072)",
          "votes": 1,
          "replies": [
            {
              "id": 2553472,
              "postDate": "2023-12-08T09:25:35.890Z",
              "content": "<p>Thanks for your reply.</p>",
              "rawMarkdown": "Thanks for your reply."
            }
          ]
        }
      ]
    },
    {
      "id": 2553466,
      "postDate": "2023-12-08T09:15:57.980Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2553672,
      "postDate": "2023-12-08T12:33:30.243Z",
      "content": "<p>Congrats, thanks for sharing.</p>",
      "rawMarkdown": "Congrats, thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 2553449,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-12-08T08:56:19.663000",
      "content": "<p>Congratulations on achieving 4th position in this competition. Thanks for sharing details of your solution. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2553439,
      "author_name": "Patrick Chan",
      "author_url": "",
      "post_date": "2023-12-08T08:46:48.940000",
      "content": "<p>Thanks for sharing and congratulations on reaching the GM.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2553437,
      "author_name": "Xiang Huang",
      "author_url": "",
      "post_date": "2023-12-08T08:45:44.220000",
      "content": "<p>Congratulations on you and your team，may I ask how large is your machine memory, I tried to load BPP Matrix but it caused OOM error.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2553441,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2023-12-08T08:50:43.073000",
          "content": "<p>In my experiment, A4000 (16GB) x 4, bs=64 for each GPU (bs=32 for exp072)</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2553472,
              "author_name": "Xiang Huang",
              "author_url": "",
              "post_date": "2023-12-08T09:25:35.890000",
              "content": "<p>Thanks for your reply.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2553466,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-08T09:15:57.980000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2553672,
      "author_name": "Antonio Félix",
      "author_url": "",
      "post_date": "2023-12-08T12:33:30.243000",
      "content": "<p>Congrats, thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2553394": "First of all, a big thank you to Kaggle staff and host for providing a fun competition.  \n[tattaka](https://www.kaggle.com/tattaka) and [monnu](https://www.kaggle.com/fuumin621) were able to complete their journey to Competitions Grandmaster with this result.\n\n## Summary\n* Various models ensemble\n  * Modified RNAdegformer proposed by Shujun\n  * 1D Conv & Residual BPP Attention\n  * Transformer with BPP Attention Bias\n* bp_matrix generated by eternafold (partially contrafold)\n  * I didn't realize until halfway through that it was being provided.\n* finetuning with higher s/n threshold\n* Pseudo Labeling using reactivity error prediction\n\n## Modified RNAdegformer\ncode: https://github.com/tattaka/stanford-ribonanza-rna-folding-public \n\n### Input\n* Sequence\n* BPP Matrix (by EternaFold and Contrafold)\n\n### Architecture\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fd89808656d85bd143d94055a543a5a55%2F2023-12-09%2018.25.55.png?generation=1702114143461138&alt=media)\nMade some changes to the [RNAdegformer](https://academic.oup.com/bib/article/24/1/bbac581/6986359)(https://github.com/Shujun-He/RNAdegformer) proposed by [Shujun](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189564), including the order of layers.\n* kernel_size = 7 except for the last transformer, which is 1\n* postnorm\n* GLU family activation\n* [ALiBi](https://arxiv.org/abs/2108.12409) positional encoding is applied separately head from bp_matrix\n* Other minor changes ensemble\n\n| model name | error prediction for pseudo labeling | use pseudo label | act_fn for feedforward | norm layer | add norm and act for conv1d | use contrafold(second bpps) | connect attn_weight to bpps bias | CV                     |\n| ---------- | ------------------------------------ | ---------------- | ---------------------- | ---------- | --------------------------- | --------------------------- | -------------------------------- | ---------------------- |\n| exp064     | yes                                  | no               | SwiGLU                 | Layernorm  | no                          | no                          | no                               | 0.12087                |\n| exp070     | no                                   | exp064           | SwiGLU                 | RMSnorm    | yes                         | no                          | no                               | 0.1199 / tiny: 0.12143 |\n| exp071     | yes                                  | no               | GeGLU                  | RMSnorm    | yes                         | yes                         | no                               | 0.12146                |\n| exp072     | no                                   | exp064 + exp071  | GeGLU                  | RMSnorm    | yes                         | no                          | yes                              | 0.11976                     |\n\n### Training\n* simple kfold (k = 5)\n* 1st stage: First train with sn > 0.5 (300epoch)\n* 2nd stage: Then finetune with sn > 1.0 for a short number of epochs (15epoch)\n* When training with train dataset only, output errors for each nucleotide in addition to reactivity\n* For pseudo labels, use sn_pred>0.75 for the 1st stage, sn_pred>1.0 for the 2nd stage, and future=1 only\n  * Train from scratch with the pseudo labels added\n* lr=1e-3, bs=256, AdamW(eps=1e-6), with warmup for 1st stage\n\n### Score for the single model\nbest model: exp072  \n* CV (k-fold): 0.11976\n* Public Score: 0.13681\n* Private Score: 0.14124\n\n## 1D Conv & Residual BPP Attention\ncode: https://github.com/fuumin621/stanford-ribonanza-rna-folding-4th\n\nBased on the 1D Conv + BPP Attention architecture proposed by [nyanp](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189241), we made improvements.\n\n### Architecture\n1. Sequence Embedding\n2. Conv 1d\n3. SE Residual & Residual BPP Attention x 12 layers\n   * Using BPP as attention weight\n   * Reduced the kernel size of 1D convolution as the layer depth increased\n4. Bi-LSTM x 2 layers \n5. Linear\n\n### Input\n* Sequence\n* BPP Matrix (by EternaFold)\n\n### Parameter\n* Drop out rate: 0.1\n* n_dim: 256\n* Kernel size: 9, 7, 5, 3 (decreases as the layer depth increases)\n* Learning rate: 4e-3 with cosine scheduler\n* Batch size: 64\n\n### Training\nThe training strategy essentially adopted the same approach as the above Modified RNAdegformer\n\n### Score for the single model\n* CV (k-fold): 0.12161\n* Public Score: 0.13889\n* Private Score: 0.1425\n\n### What Didn't Work\n* BPP packages other than EternaFold (contrafold_2)\n* Distance Matrix\n* Structure\n* BPP Feature Engineering (max, sum, nb_count)\n* Sample Weight by SN\n* etc...\n\n## Transformer with BPP Attention Bias (@ren4yu's Part)\n\ncode: https://github.com/yu4u/kaggle-stanford-ribonanza-rna-folding-4th-place-solution\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F7eea07453fa8011a1435c012533a084d%2Ftransformer.png?generation=1702099940846948&alt=media\" width=\"50%\">\n\nThis model borrows the bpp attention bias idea from the RNAdegformer, but is closer to the original Transformer architecture.\nInstead of using positional embedding, Conv1D was used in FFN to give relative positional information.\n### Training Procedure\n- Train model with AdamW LR=2e-3 to 2e-4, BS=128, use s/n filter threshold=0.5\n- Finetune model with AdamW LR=2e-4 to 0, BS=256, use s/n filter threshold=1.0\n### Score for the single model\n* CV (k-fold): 0.12188\n* Public Score: 0.13948\n* Private Score: 0.14267\n### Does not work for me\n- Increasing dimension or number of layers\n- UNet-like hierarchical architecture\n\n",
    "2553449": "Congratulations on achieving 4th position in this competition. Thanks for sharing details of your solution. ",
    "2553439": "Thanks for sharing and congratulations on reaching the GM.",
    "2553437": "Congratulations on you and your team，may I ask how large is your machine memory, I tried to load BPP Matrix but it caused OOM error.",
    "2553466": "",
    "2553672": "Congrats, thanks for sharing."
  }
}