{
  "id": 460287,
  "title": "36th solution : Cheap Transformers that can run on gaming laptops and does not require a BPP matrix",
  "url": "/competitions/stanford-ribonanza-rna-folding/writeups/qhapaq49-36th-solution-cheap-transformers-that-can",
  "author_name": "",
  "post_date": "2023-12-08T15:07:18.923Z",
  "votes": 19,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First of all, I would like to thank all the participants in the competition. My solution is a cheap transformer that does not require bpp-matrix and can learn and predict with poor computational resources named \"Cheap Transformers\"</p>\n<h1>Generalization to long sequences</h1>\n<p>My starting point is <a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb\" target=\"_blank\">Iafoss's amazing work</a> where the reactivities are predicted by simple transformers.</p>\n<p>It works very well in cross validation in random fold splitting but gives poor performance when train and validation data is split by length of sequence like following. </p>\n<pre><code>short_indices = df[df[]..() &lt; ].index.tolist()\nlong_indices = df[df[]..() &gt; ].index.tolist()\n</code></pre>\n<p>Therefore, I need another architecture to generalize to long sequences.</p>\n<h1>Cheap Transformer</h1>\n<p>This is an architecture of cheap transformer</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1233385%2Fef5a1e8baadc56507bee658eb624c633%2FScreenshot%20from%202023-12-08%2023-14-44.png?generation=1702046810306490&amp;alt=media\" alt=\"\">er. It just cut a certain length sequence from the original sequence.</p>\n<p>When regression, it cuts the sequence in multiple ways and ensemble the prediction results for each character. The ensemble is important. It improves cv-score at about 0.005.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1233385%2Fbbaf5eee4d3b85a8c3ce62c7bc6cf34f%2FScreenshot%20from%202023-12-08%2023-14-52.png?generation=1702046847620018&amp;alt=media\" alt=\"\"></p>\n<p>This architecture uses short sequences, reduces training time, and requires less GPU memory during execution. Therefore, it runs fast enough even with the GPU of a gaming laptop.</p>\n<h1>Final submission</h1>\n<p>I prepared multiple models with cut sequence lengths of 64, 96, and 128 and with various transformer's parameters and created an ensemble. Pseudo labeling was also performed, but it had little effect on LB. As you can see the ensemble contributed LB but its effect is not so big, therefore, using single model would be the most practival.</p>\n<h1>Scores</h1>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Iafoss's architecture</td>\n<td>0.15669</td>\n<td>0.20204</td>\n</tr>\n<tr>\n<td>Cheap transformers (len=64)</td>\n<td>0.15635</td>\n<td>0.15661</td>\n</tr>\n<tr>\n<td>Cheap transformers (len=96)</td>\n<td>0.15458</td>\n<td>0.15445</td>\n</tr>\n<tr>\n<td>Cheap transformers (len=128)</td>\n<td>0.15518</td>\n<td>0.15625</td>\n</tr>\n<tr>\n<td>Public best (ensemble with Iafoss's arch)</td>\n<td>0.14716</td>\n<td>0.18222</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>0.15271</td>\n<td>0.15346</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "2553860",
      "postDate": "12/08/2023 14:56:55",
      "content": "<p>First of all, I would like to thank all the participants in the competition. My solution is a cheap transformer that does not require bpp-matrix and can learn and predict with poor computational resources named \"Cheap Transformers\"</p>\n<h1>Generalization to long sequences</h1>\n<p>My starting point is <a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb\" target=\"_blank\">Iafoss's amazing work</a> where the reactivities are predicted by simple transformers.</p>\n<p>It works very well in cross validation in random fold splitting but gives poor performance when train and validation data is split by length of sequence like following. </p>\n<pre><code>short_indices = df[df[]..() &lt; ].index.tolist()\nlong_indices = df[df[]..() &gt; ].index.tolist()\n</code></pre>\n<p>Therefore, I need another architecture to generalize to long sequences.</p>\n<h1>Cheap Transformer</h1>\n<p>This is an architecture of cheap transformer</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1233385%2Fef5a1e8baadc56507bee658eb624c633%2FScreenshot%20from%202023-12-08%2023-14-44.png?generation=1702046810306490&amp;alt=media\" alt=\"\">er. It just cut a certain length sequence from the original sequence.</p>\n<p>When regression, it cuts the sequence in multiple ways and ensemble the prediction results for each character. The ensemble is important. It improves cv-score at about 0.005.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1233385%2Fbbaf5eee4d3b85a8c3ce62c7bc6cf34f%2FScreenshot%20from%202023-12-08%2023-14-52.png?generation=1702046847620018&amp;alt=media\" alt=\"\"></p>\n<p>This architecture uses short sequences, reduces training time, and requires less GPU memory during execution. Therefore, it runs fast enough even with the GPU of a gaming laptop.</p>\n<h1>Final submission</h1>\n<p>I prepared multiple models with cut sequence lengths of 64, 96, and 128 and with various transformer's parameters and created an ensemble. Pseudo labeling was also performed, but it had little effect on LB. As you can see the ensemble contributed LB but its effect is not so big, therefore, using single model would be the most practival.</p>\n<h1>Scores</h1>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Iafoss's architecture</td>\n<td>0.15669</td>\n<td>0.20204</td>\n</tr>\n<tr>\n<td>Cheap transformers (len=64)</td>\n<td>0.15635</td>\n<td>0.15661</td>\n</tr>\n<tr>\n<td>Cheap transformers (len=96)</td>\n<td>0.15458</td>\n<td>0.15445</td>\n</tr>\n<tr>\n<td>Cheap transformers (len=128)</td>\n<td>0.15518</td>\n<td>0.15625</td>\n</tr>\n<tr>\n<td>Public best (ensemble with Iafoss's arch)</td>\n<td>0.14716</td>\n<td>0.18222</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>0.15271</td>\n<td>0.15346</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "First of all, I would like to thank all the participants in the competition. My solution is a cheap transformer that does not require bpp-matrix and can learn and predict with poor computational resources named \"Cheap Transformers\"\n\n# Generalization to long sequences\nMy starting point is [Iafoss's amazing work](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb) where the reactivities are predicted by simple transformers.\n\nIt works very well in cross validation in random fold splitting but gives poor performance when train and validation data is split by length of sequence like following. \n\n```python\nshort_indices = df[df['sequence'].str.len() < 177].index.tolist()\nlong_indices = df[df['sequence'].str.len() > 177].index.tolist()\n```\n\nTherefore, I need another architecture to generalize to long sequences.\n\n# Cheap Transformer\n\nThis is an architecture of cheap transformer\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1233385%2Fef5a1e8baadc56507bee658eb624c633%2FScreenshot%20from%202023-12-08%2023-14-44.png?generation=1702046810306490&alt=media)er. It just cut a certain length sequence from the original sequence.\n\nWhen regression, it cuts the sequence in multiple ways and ensemble the prediction results for each character. The ensemble is important. It improves cv-score at about 0.005.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1233385%2Fbbaf5eee4d3b85a8c3ce62c7bc6cf34f%2FScreenshot%20from%202023-12-08%2023-14-52.png?generation=1702046847620018&alt=media)\n\nThis architecture uses short sequences, reduces training time, and requires less GPU memory during execution. Therefore, it runs fast enough even with the GPU of a gaming laptop.\n\n# Final submission\nI prepared multiple models with cut sequence lengths of 64, 96, and 128 and with various transformer's parameters and created an ensemble. Pseudo labeling was also performed, but it had little effect on LB. As you can see the ensemble contributed LB but its effect is not so big, therefore, using single model would be the most practival.\n\n\n\n# Scores\n| Method | Public | Private |\n| --- | --- | --- |\n| Iafoss's architecture | 0.15669 | 0.20204  |\n| Cheap transformers (len=64) | 0.15635 | 0.15661  |\n| Cheap transformers (len=96) | 0.15458 | 0.15445  |\n| Cheap transformers (len=128) | 0.15518 | 0.15625  |\n|Public best (ensemble with Iafoss's arch) | 0.14716 | 0.18222  |\n| ensemble | 0.15271 |  0.15346 |",
      "votes": null
    },
    {
      "id": "2553970",
      "postDate": "12/08/2023 17:00:40",
      "content": "<p>Very similar with the approach I discarted but that obtained best private scoring. With a small difference. I'm giving it a second chance while trying to remember the particular details of that submission. I will prepare the results and so we can compare them.</p>",
      "rawMarkdown": "Very similar with the approach I discarted but that obtained best private scoring. With a small difference. I'm giving it a second chance while trying to remember the particular details of that submission. I will prepare the results and so we can compare them.",
      "votes": null
    },
    {
      "id": "2554219",
      "postDate": "12/08/2023 23:12:28",
      "content": "<p>Thank you for your comment. As far as I know, the following point is important to use this architecture.</p>\n<h2>Using Eternafold to get structure</h2>\n<p>It improves CV score at approximately 0.01. I tried other libraries such as U-Fold, R-fold, Vienna RNA etc but none of them worked well.</p>\n<h2>Keeping sequence length 64-128</h2>\n<p>Too long/short length of sequence makes LB score worse. </p>",
      "rawMarkdown": "Thank you for your comment. As far as I know, the following point is important to use this architecture.\n\n## Using Eternafold to get structure\nIt improves CV score at approximately 0.01. I tried other libraries such as U-Fold, R-fold, Vienna RNA etc but none of them worked well.\n\n## Keeping sequence length 64-128\nToo long/short length of sequence makes LB score worse.",
      "votes": null
    },
    {
      "id": "2559485",
      "postDate": "12/12/2023 21:14:56",
      "content": "<p>It's really refreshing to see a well-placing solution that pays attention to hardware requirements. Great job!</p>",
      "rawMarkdown": "It's really refreshing to see a well-placing solution that pays attention to hardware requirements. Great job!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2553970,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "12/08/2023 17:00:40",
      "content": "<p>Very similar with the approach I discarted but that obtained best private scoring. With a small difference. I'm giving it a second chance while trying to remember the particular details of that submission. I will prepare the results and so we can compare them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2554219,
          "author_name": "qhapaq49",
          "author_url": "",
          "post_date": "12/08/2023 23:12:28",
          "content": "<p>Thank you for your comment. As far as I know, the following point is important to use this architecture.</p>\n<h2>Using Eternafold to get structure</h2>\n<p>It improves CV score at approximately 0.01. I tried other libraries such as U-Fold, R-fold, Vienna RNA etc but none of them worked well.</p>\n<h2>Keeping sequence length 64-128</h2>\n<p>Too long/short length of sequence makes LB score worse. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2559485,
      "author_name": "alberteinsten",
      "author_url": "",
      "post_date": "12/12/2023 21:14:56",
      "content": "<p>It's really refreshing to see a well-placing solution that pays attention to hardware requirements. Great job!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2553860": "First of all, I would like to thank all the participants in the competition. My solution is a cheap transformer that does not require bpp-matrix and can learn and predict with poor computational resources named \"Cheap Transformers\"\n\n# Generalization to long sequences\nMy starting point is [Iafoss's amazing work](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb) where the reactivities are predicted by simple transformers.\n\nIt works very well in cross validation in random fold splitting but gives poor performance when train and validation data is split by length of sequence like following. \n\n```python\nshort_indices = df[df['sequence'].str.len() < 177].index.tolist()\nlong_indices = df[df['sequence'].str.len() > 177].index.tolist()\n```\n\nTherefore, I need another architecture to generalize to long sequences.\n\n# Cheap Transformer\n\nThis is an architecture of cheap transformer\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1233385%2Fef5a1e8baadc56507bee658eb624c633%2FScreenshot%20from%202023-12-08%2023-14-44.png?generation=1702046810306490&alt=media)er. It just cut a certain length sequence from the original sequence.\n\nWhen regression, it cuts the sequence in multiple ways and ensemble the prediction results for each character. The ensemble is important. It improves cv-score at about 0.005.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1233385%2Fbbaf5eee4d3b85a8c3ce62c7bc6cf34f%2FScreenshot%20from%202023-12-08%2023-14-52.png?generation=1702046847620018&alt=media)\n\nThis architecture uses short sequences, reduces training time, and requires less GPU memory during execution. Therefore, it runs fast enough even with the GPU of a gaming laptop.\n\n# Final submission\nI prepared multiple models with cut sequence lengths of 64, 96, and 128 and with various transformer's parameters and created an ensemble. Pseudo labeling was also performed, but it had little effect on LB. As you can see the ensemble contributed LB but its effect is not so big, therefore, using single model would be the most practival.\n\n\n\n# Scores\n| Method | Public | Private |\n| --- | --- | --- |\n| Iafoss's architecture | 0.15669 | 0.20204  |\n| Cheap transformers (len=64) | 0.15635 | 0.15661  |\n| Cheap transformers (len=96) | 0.15458 | 0.15445  |\n| Cheap transformers (len=128) | 0.15518 | 0.15625  |\n|Public best (ensemble with Iafoss's arch) | 0.14716 | 0.18222  |\n| ensemble | 0.15271 |  0.15346 |",
    "2553970": "Very similar with the approach I discarted but that obtained best private scoring. With a small difference. I'm giving it a second chance while trying to remember the particular details of that submission. I will prepare the results and so we can compare them.",
    "2554219": "Thank you for your comment. As far as I know, the following point is important to use this architecture.\n\n## Using Eternafold to get structure\nIt improves CV score at approximately 0.01. I tried other libraries such as U-Fold, R-fold, Vienna RNA etc but none of them worked well.\n\n## Keeping sequence length 64-128\nToo long/short length of sequence makes LB score worse.",
    "2559485": "It's really refreshing to see a well-placing solution that pays attention to hardware requirements. Great job!"
  },
  "source": "meta"
}