{
  "id": 460285,
  "title": "21th solution: ESM2 + custom folding head",
  "url": "/competitions/stanford-ribonanza-rna-folding/writeups/cpmp-bo-21th-solution-esm2-custom-folding-head",
  "author_name": "",
  "post_date": "2023-12-08T14:45:46.766183800Z",
  "votes": 23,
  "comment_count": 9,
  "views": 0,
  "content": "<p><strong>TLDR</strong></p>\n<p>We entered the competition quite late, our first sub was from 10 days ago. Given the limited time, we decided to bet on a low development path and leverage the excellent ESM2 model family from META.</p>\n<p>We pretrained ESM2 backbones on competition data and RNA Central long sequences, then added custom folding heads and fine tuned on DMS and 2A3 labels.</p>\n<p>We honestly are a bit surprised by our end result, we expected worse when we started. </p>\n<p><strong>ESM2 Pretraining</strong></p>\n<p>We have some experience pretraining <a href=\"https://github.com/facebookresearch/esm\" target=\"_blank\">ESM2</a> protein language models, so we thought they would be a good fit for this competition. We pretrained three versions of ESM2 models (35M, 150M, 650M) on both train and test sequences using MLM (masked language modeling) loss. </p>\n<p>In particular, we used huggingface’s <a href=\"https://huggingface.co/docs/transformers/model_doc/esm#transformers.EsmForMaskedLM\" target=\"_blank\"><code>EsmForMaskedLM</code></a> model class and  <a href=\"https://huggingface.co/docs/transformers/main/main_classes/data_collator#transformers.DataCollatorForLanguageModeling\" target=\"_blank\"><code>DataCollatorForLanguageModeling(tokenizer, mlm=True)</code> </a> data collator which supports MLM loss.</p>\n<p>Since there are no train sequences and only 8000 test sequences longer than 208, we downloaded <a href=\"https://rnacentral.org/\" target=\"_blank\">RNA Central </a> data and selected 12 million samples with length between 115 and 457 (the range in this competition data) to complement the competition data.</p>\n<p>We pretrained 5-20 epochs depending on how many RNA central data used, for a total of 25M to 70M total samples seen.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2Feb2f9716c901ef7021c64e644a16d2a6%2Fesm2.png?generation=1702046716073336&amp;alt=media\" alt=\"\"></p>\n<p><strong>Folding Head</strong></p>\n<p>We adapted the model CPMP used in the Covid Vaccine competition three years ago to devise a folding head on top of ESM2 token embeddings. That model was presented <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189723\" target=\"_blank\">here</a>.</p>\n<p>The folding head takes as input both the ESM2 token embeddings and the bpp files. Its main component is an attention/convolution layer, repeated twice, then a linear classification head. </p>\n<p>The Attention/Convolution layers are similar to the transformer encoder structure: an attention layer followed by a convolution bloc with skip connection.<br>\nThe attention layer is a bpp attention layer followed by a bidirectional RNN (GRU or LSTM yields two model variants). Sure, GRU is not attention, but here it plays the same role as an attention layer where each node attends its two neighbors in the sequence.<br>\nThe bpp attention is a standard attention layer where attention weights are the bpp values.<br>\nThe convolution layer was designed after the Efficientnet convolution bloc. We then stack these as in this <a href=\"https://www.kaggle.com/thebigd8ta/open-vaccine-pytorch-v\" target=\"_blank\">notebook</a>.</p>\n<p>We bet on this RNN/Convolution architecture rather than a transformer for two reasons:</p>\n<ul>\n<li>We literally started working on it 5 days before competition end, and we did not have time to tune a new architecture</li>\n<li>we hoped that ESM2 attention would have already learned what a transformer head attention would learn.</li>\n</ul>\n<p><strong>Finetuning</strong></p>\n<p>Our best models are the ones with the custom folding head on top of the ESM2 backbone. We also trained simple models with just a linear head. The final ensemble is a mix of both, most being with a folding head.</p>\n<p>We used four  variants of training data:<br>\nQuick start<br>\nQuick start + SN &gt;= 0.3<br>\nQuick start + SN &gt;= 0.3, and SN as sample weight<br>\nQuick start + SN &gt;= 0.1, and SN as sample weight<br>\nSN is signal_to_noise clipped to (0, 1). We used it as sample noise in some variants. Model quality increased from top variant to bottom one.</p>\n<p>We used a 5 fold CV and submitted the average of each fold model predictions.</p>\n<p><strong>Scores</strong></p>\n<p>Our best single model is a 650M LSTM model with SN weight. CV = 0.12847, Public LB = 0.14308 (5 fold ensemble)<br>\nOur best ensemble is a 15 model ensemble (75 folds in total), CV = 0.12478, Public LB =  0.14174</p>\n<p><strong>What didn’t work</strong></p>\n<p>Pretrained models from <a href=\"https://github.com/ml4bio/RNA-FM\" target=\"_blank\">RNA-FM</a> were not as good as our ESM2 pretrained models.</p>\n<p>ESM2 can also output contact predictions. To do so it stacks all attention activations, then it applies a regression head.  We tried to use that to predict bpp values via an auxiliary loss. This improved the competition metric score a bit. But this comes at the expense of a 4x space and a 2x time increase, which prevented us from using it with 150M and 650M variants of ESM2.</p>\n<p>Bo &amp; CPMP</p>",
  "messages": [
    {
      "id": "2553840",
      "postDate": "12/08/2023 14:45:46",
      "content": "<p><strong>TLDR</strong></p>\n<p>We entered the competition quite late, our first sub was from 10 days ago. Given the limited time, we decided to bet on a low development path and leverage the excellent ESM2 model family from META.</p>\n<p>We pretrained ESM2 backbones on competition data and RNA Central long sequences, then added custom folding heads and fine tuned on DMS and 2A3 labels.</p>\n<p>We honestly are a bit surprised by our end result, we expected worse when we started. </p>\n<p><strong>ESM2 Pretraining</strong></p>\n<p>We have some experience pretraining <a href=\"https://github.com/facebookresearch/esm\" target=\"_blank\">ESM2</a> protein language models, so we thought they would be a good fit for this competition. We pretrained three versions of ESM2 models (35M, 150M, 650M) on both train and test sequences using MLM (masked language modeling) loss. </p>\n<p>In particular, we used huggingface’s <a href=\"https://huggingface.co/docs/transformers/model_doc/esm#transformers.EsmForMaskedLM\" target=\"_blank\"><code>EsmForMaskedLM</code></a> model class and  <a href=\"https://huggingface.co/docs/transformers/main/main_classes/data_collator#transformers.DataCollatorForLanguageModeling\" target=\"_blank\"><code>DataCollatorForLanguageModeling(tokenizer, mlm=True)</code> </a> data collator which supports MLM loss.</p>\n<p>Since there are no train sequences and only 8000 test sequences longer than 208, we downloaded <a href=\"https://rnacentral.org/\" target=\"_blank\">RNA Central </a> data and selected 12 million samples with length between 115 and 457 (the range in this competition data) to complement the competition data.</p>\n<p>We pretrained 5-20 epochs depending on how many RNA central data used, for a total of 25M to 70M total samples seen.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2Feb2f9716c901ef7021c64e644a16d2a6%2Fesm2.png?generation=1702046716073336&amp;alt=media\" alt=\"\"></p>\n<p><strong>Folding Head</strong></p>\n<p>We adapted the model CPMP used in the Covid Vaccine competition three years ago to devise a folding head on top of ESM2 token embeddings. That model was presented <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189723\" target=\"_blank\">here</a>.</p>\n<p>The folding head takes as input both the ESM2 token embeddings and the bpp files. Its main component is an attention/convolution layer, repeated twice, then a linear classification head. </p>\n<p>The Attention/Convolution layers are similar to the transformer encoder structure: an attention layer followed by a convolution bloc with skip connection.<br>\nThe attention layer is a bpp attention layer followed by a bidirectional RNN (GRU or LSTM yields two model variants). Sure, GRU is not attention, but here it plays the same role as an attention layer where each node attends its two neighbors in the sequence.<br>\nThe bpp attention is a standard attention layer where attention weights are the bpp values.<br>\nThe convolution layer was designed after the Efficientnet convolution bloc. We then stack these as in this <a href=\"https://www.kaggle.com/thebigd8ta/open-vaccine-pytorch-v\" target=\"_blank\">notebook</a>.</p>\n<p>We bet on this RNN/Convolution architecture rather than a transformer for two reasons:</p>\n<ul>\n<li>We literally started working on it 5 days before competition end, and we did not have time to tune a new architecture</li>\n<li>we hoped that ESM2 attention would have already learned what a transformer head attention would learn.</li>\n</ul>\n<p><strong>Finetuning</strong></p>\n<p>Our best models are the ones with the custom folding head on top of the ESM2 backbone. We also trained simple models with just a linear head. The final ensemble is a mix of both, most being with a folding head.</p>\n<p>We used four  variants of training data:<br>\nQuick start<br>\nQuick start + SN &gt;= 0.3<br>\nQuick start + SN &gt;= 0.3, and SN as sample weight<br>\nQuick start + SN &gt;= 0.1, and SN as sample weight<br>\nSN is signal_to_noise clipped to (0, 1). We used it as sample noise in some variants. Model quality increased from top variant to bottom one.</p>\n<p>We used a 5 fold CV and submitted the average of each fold model predictions.</p>\n<p><strong>Scores</strong></p>\n<p>Our best single model is a 650M LSTM model with SN weight. CV = 0.12847, Public LB = 0.14308 (5 fold ensemble)<br>\nOur best ensemble is a 15 model ensemble (75 folds in total), CV = 0.12478, Public LB =  0.14174</p>\n<p><strong>What didn’t work</strong></p>\n<p>Pretrained models from <a href=\"https://github.com/ml4bio/RNA-FM\" target=\"_blank\">RNA-FM</a> were not as good as our ESM2 pretrained models.</p>\n<p>ESM2 can also output contact predictions. To do so it stacks all attention activations, then it applies a regression head.  We tried to use that to predict bpp values via an auxiliary loss. This improved the competition metric score a bit. But this comes at the expense of a 4x space and a 2x time increase, which prevented us from using it with 150M and 650M variants of ESM2.</p>\n<p>Bo &amp; CPMP</p>",
      "rawMarkdown": "**TLDR**\n\nWe entered the competition quite late, our first sub was from 10 days ago. Given the limited time, we decided to bet on a low development path and leverage the excellent ESM2 model family from META.\n\nWe pretrained ESM2 backbones on competition data and RNA Central long sequences, then added custom folding heads and fine tuned on DMS and 2A3 labels.\n\nWe honestly are a bit surprised by our end result, we expected worse when we started. \n\n\n**ESM2 Pretraining**\n\nWe have some experience pretraining [ESM2](https://github.com/facebookresearch/esm) protein language models, so we thought they would be a good fit for this competition. We pretrained three versions of ESM2 models (35M, 150M, 650M) on both train and test sequences using MLM (masked language modeling) loss. \n\nIn particular, we used huggingface’s [`EsmForMaskedLM`](https://huggingface.co/docs/transformers/model_doc/esm#transformers.EsmForMaskedLM) model class and  [`DataCollatorForLanguageModeling(tokenizer, mlm=True)` ](https://huggingface.co/docs/transformers/main/main_classes/data_collator#transformers.DataCollatorForLanguageModeling) data collator which supports MLM loss.\n\nSince there are no train sequences and only 8000 test sequences longer than 208, we downloaded [RNA Central ](https://rnacentral.org/) data and selected 12 million samples with length between 115 and 457 (the range in this competition data) to complement the competition data.\n\nWe pretrained 5-20 epochs depending on how many RNA central data used, for a total of 25M to 70M total samples seen.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2Feb2f9716c901ef7021c64e644a16d2a6%2Fesm2.png?generation=1702046716073336&alt=media)\n\n**Folding Head**\n\nWe adapted the model CPMP used in the Covid Vaccine competition three years ago to devise a folding head on top of ESM2 token embeddings. That model was presented [here](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189723).\n\nThe folding head takes as input both the ESM2 token embeddings and the bpp files. Its main component is an attention/convolution layer, repeated twice, then a linear classification head. \n\nThe Attention/Convolution layers are similar to the transformer encoder structure: an attention layer followed by a convolution bloc with skip connection.\nThe attention layer is a bpp attention layer followed by a bidirectional RNN (GRU or LSTM yields two model variants). Sure, GRU is not attention, but here it plays the same role as an attention layer where each node attends its two neighbors in the sequence.\nThe bpp attention is a standard attention layer where attention weights are the bpp values.\nThe convolution layer was designed after the Efficientnet convolution bloc. We then stack these as in this [notebook](https://www.kaggle.com/thebigd8ta/open-vaccine-pytorch-v).\n\nWe bet on this RNN/Convolution architecture rather than a transformer for two reasons:\n- We literally started working on it 5 days before competition end, and we did not have time to tune a new architecture\n- we hoped that ESM2 attention would have already learned what a transformer head attention would learn.\n\n**Finetuning**\n\nOur best models are the ones with the custom folding head on top of the ESM2 backbone. We also trained simple models with just a linear head. The final ensemble is a mix of both, most being with a folding head.\n\nWe used four  variants of training data:\nQuick start\nQuick start + SN >= 0.3\nQuick start + SN >= 0.3, and SN as sample weight\nQuick start + SN >= 0.1, and SN as sample weight\nSN is signal_to_noise clipped to (0, 1). We used it as sample noise in some variants. Model quality increased from top variant to bottom one.\n\nWe used a 5 fold CV and submitted the average of each fold model predictions.\n\n**Scores**\n\nOur best single model is a 650M LSTM model with SN weight. CV = 0.12847, Public LB = 0.14308 (5 fold ensemble)\nOur best ensemble is a 15 model ensemble (75 folds in total), CV = 0.12478, Public LB =  0.14174\n\n**What didn’t work**\n\nPretrained models from [RNA-FM](https://github.com/ml4bio/RNA-FM) were not as good as our ESM2 pretrained models.\n\nESM2 can also output contact predictions. To do so it stacks all attention activations, then it applies a regression head.  We tried to use that to predict bpp values via an auxiliary loss. This improved the competition metric score a bit. But this comes at the expense of a 4x space and a 2x time increase, which prevented us from using it with 150M and 650M variants of ESM2.\n\nBo & CPMP",
      "votes": null
    },
    {
      "id": "2554204",
      "postDate": "12/08/2023 22:29:07",
      "content": "<p>Congrats! I also joined late and tried similar ideas, but wasn't so successful, lots of room to learn :) Do you know how much value you get from pre-training ESM with MLM? I also saw poor performance of RNA-FM and concluded MLM might not be helpful here, because we're dealing with a lot of synthetic sequences, as opposed to ESM that is trained on evolutionary designed proteins.</p>",
      "rawMarkdown": "Congrats! I also joined late and tried similar ideas, but wasn't so successful, lots of room to learn :) Do you know how much value you get from pre-training ESM with MLM? I also saw poor performance of RNA-FM and concluded MLM might not be helpful here, because we're dealing with a lot of synthetic sequences, as opposed to ESM that is trained on evolutionary designed proteins.",
      "votes": null
    },
    {
      "id": "2554223",
      "postDate": "12/08/2023 23:36:29",
      "content": "<p>Pretraining was key. We saw a performance improvement with more pretraining. I'm sorry I can't say more because I have not tried without a pretrained ESM2.</p>\n<p>Your score is not that bad!</p>",
      "rawMarkdown": "Pretraining was key. We saw a performance improvement with more pretraining. I'm sorry I can't say more because I have not tried without a pretrained ESM2.\n\nYour score is not that bad!",
      "votes": null
    },
    {
      "id": "2554905",
      "postDate": "12/09/2023 14:05:20",
      "content": "<p>Did you convert the RNA to proteins/AA?</p>",
      "rawMarkdown": "Did you convert the RNA to proteins/AA?",
      "votes": null
    },
    {
      "id": "2554955",
      "postDate": "12/09/2023 14:57:50",
      "content": "<p>no, we replaced the vocab.txt files used by the ESM tokenizer. We used this as <code>rna_vocab.txt</code>:</p>\n<pre><code>&lt;cls&gt;\n&lt;pad&gt;\n&lt;eos&gt;\n&lt;unk&gt;\nA\nC\nG\nU\n&lt;null_1&gt;\n&lt;mask&gt;\n</code></pre>\n<p>Then create the tokenizer with:</p>\n<p><code>tokenizer = EsmTokenizer('../rna_vocab.txt', model_max_length=context_length)</code></p>",
      "rawMarkdown": "no, we replaced the vocab.txt files used by the ESM tokenizer. We used this as `rna_vocab.txt`:\n\n```python\n<cls>\n<pad>\n<eos>\n<unk>\nA\nC\nG\nU\n<null_1>\n<mask>\n```\n\nThen create the tokenizer with:\n\n`tokenizer = EsmTokenizer('../rna_vocab.txt', model_max_length=context_length)`",
      "votes": null
    },
    {
      "id": "2554991",
      "postDate": "12/09/2023 15:41:15",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Did you try not using bpp? Curious how that would work</p>",
      "rawMarkdown": "cpmpml Did you try not using bpp? Curious how that would work",
      "votes": null
    },
    {
      "id": "2555708",
      "postDate": "12/10/2023 05:37:46",
      "content": "<p>Yes, we have submitted ESM2 alone. Best one has 0.14781 private / 0.14359 public. But it looks like an outlier, others have much worse scores. </p>\n<p><a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> did the ESM2 part, maybe he can say more.</p>",
      "rawMarkdown": "Yes, we have submitted ESM2 alone. Best one has 0.14781 private / 0.14359 public. But it looks like an outlier, others have much worse scores. \n\n@boliu0 did the ESM2 part, maybe he can say more.",
      "votes": null
    },
    {
      "id": "2556239",
      "postDate": "12/10/2023 15:01:31",
      "content": "<blockquote>\n  <p>Do you know how much value you get from pre-training ESM with MLM?</p>\n</blockquote>\n<p>A lot. In my very first baseline, I compared non-pretrained ESM 35M vs a pretrained ESM 35M (pre-trained on competition train set only, so it’s not the best pretrained version) on the downstream supervised task, the MAE is not close. 0.01-0.02 gap if I remember correctly. So I never tried non-pretrained ESM anymore.</p>",
      "rawMarkdown": "> Do you know how much value you get from pre-training ESM with MLM?\n\nA lot. In my very first baseline, I compared non-pretrained ESM 35M vs a pretrained ESM 35M (pre-trained on competition train set only, so it’s not the best pretrained version) on the downstream supervised task, the MAE is not close. 0.01-0.02 gap if I remember correctly. So I never tried non-pretrained ESM anymore.",
      "votes": null
    },
    {
      "id": "2556245",
      "postDate": "12/10/2023 15:07:19",
      "content": "<p>The one with 0.14781 private / 0.14359 public is not ESM2 alone actually. My incorrect naming of that sub confused you. </p>\n<p>Best ESM alone models:<br>\nESM 35M: public 0.15095, private 0.16162<br>\nESM 150M: public 0.15110, private 0.15836<br>\nESM 650M: public 0.14789, private 0.15331</p>\n<p>They are obviously not as good as bpp models, but they add diversity and boost cv scores, so we included them in the final ensemble.</p>",
      "rawMarkdown": "The one with 0.14781 private / 0.14359 public is not ESM2 alone actually. My incorrect naming of that sub confused you. \n\nBest ESM alone models:\nESM 35M: public 0.15095, private 0.16162\nESM 150M: public 0.15110, private 0.15836\nESM 650M: public 0.14789, private 0.15331\n\nThey are obviously not as good as bpp models, but they add diversity and boost cv scores, so we included them in the final ensemble.",
      "votes": null
    },
    {
      "id": "2559478",
      "postDate": "12/12/2023 21:10:30",
      "content": "<p>Wow, 21st in just 10 days is insane! Congratulations to you and your team!</p>",
      "rawMarkdown": "Wow, 21st in just 10 days is insane! Congratulations to you and your team!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2554204,
      "author_name": "thedrcat",
      "author_url": "",
      "post_date": "12/08/2023 22:29:07",
      "content": "<p>Congrats! I also joined late and tried similar ideas, but wasn't so successful, lots of room to learn :) Do you know how much value you get from pre-training ESM with MLM? I also saw poor performance of RNA-FM and concluded MLM might not be helpful here, because we're dealing with a lot of synthetic sequences, as opposed to ESM that is trained on evolutionary designed proteins.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2554223,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/08/2023 23:36:29",
          "content": "<p>Pretraining was key. We saw a performance improvement with more pretraining. I'm sorry I can't say more because I have not tried without a pretrained ESM2.</p>\n<p>Your score is not that bad!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2556239,
              "author_name": "boliu0",
              "author_url": "",
              "post_date": "12/10/2023 15:01:31",
              "content": "<blockquote>\n  <p>Do you know how much value you get from pre-training ESM with MLM?</p>\n</blockquote>\n<p>A lot. In my very first baseline, I compared non-pretrained ESM 35M vs a pretrained ESM 35M (pre-trained on competition train set only, so it’s not the best pretrained version) on the downstream supervised task, the MAE is not close. 0.01-0.02 gap if I remember correctly. So I never tried non-pretrained ESM anymore.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2554905,
      "author_name": "danofer",
      "author_url": "",
      "post_date": "12/09/2023 14:05:20",
      "content": "<p>Did you convert the RNA to proteins/AA?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2554955,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/09/2023 14:57:50",
          "content": "<p>no, we replaced the vocab.txt files used by the ESM tokenizer. We used this as <code>rna_vocab.txt</code>:</p>\n<pre><code>&lt;cls&gt;\n&lt;pad&gt;\n&lt;eos&gt;\n&lt;unk&gt;\nA\nC\nG\nU\n&lt;null_1&gt;\n&lt;mask&gt;\n</code></pre>\n<p>Then create the tokenizer with:</p>\n<p><code>tokenizer = EsmTokenizer('../rna_vocab.txt', model_max_length=context_length)</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2554991,
      "author_name": "shujun717",
      "author_url": "",
      "post_date": "12/09/2023 15:41:15",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Did you try not using bpp? Curious how that would work</p>",
      "votes": null,
      "replies": [
        {
          "id": 2555708,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/10/2023 05:37:46",
          "content": "<p>Yes, we have submitted ESM2 alone. Best one has 0.14781 private / 0.14359 public. But it looks like an outlier, others have much worse scores. </p>\n<p><a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> did the ESM2 part, maybe he can say more.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2556245,
              "author_name": "boliu0",
              "author_url": "",
              "post_date": "12/10/2023 15:07:19",
              "content": "<p>The one with 0.14781 private / 0.14359 public is not ESM2 alone actually. My incorrect naming of that sub confused you. </p>\n<p>Best ESM alone models:<br>\nESM 35M: public 0.15095, private 0.16162<br>\nESM 150M: public 0.15110, private 0.15836<br>\nESM 650M: public 0.14789, private 0.15331</p>\n<p>They are obviously not as good as bpp models, but they add diversity and boost cv scores, so we included them in the final ensemble.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2559478,
      "author_name": "alberteinsten",
      "author_url": "",
      "post_date": "12/12/2023 21:10:30",
      "content": "<p>Wow, 21st in just 10 days is insane! Congratulations to you and your team!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2553840": "**TLDR**\n\nWe entered the competition quite late, our first sub was from 10 days ago. Given the limited time, we decided to bet on a low development path and leverage the excellent ESM2 model family from META.\n\nWe pretrained ESM2 backbones on competition data and RNA Central long sequences, then added custom folding heads and fine tuned on DMS and 2A3 labels.\n\nWe honestly are a bit surprised by our end result, we expected worse when we started. \n\n\n**ESM2 Pretraining**\n\nWe have some experience pretraining [ESM2](https://github.com/facebookresearch/esm) protein language models, so we thought they would be a good fit for this competition. We pretrained three versions of ESM2 models (35M, 150M, 650M) on both train and test sequences using MLM (masked language modeling) loss. \n\nIn particular, we used huggingface’s [`EsmForMaskedLM`](https://huggingface.co/docs/transformers/model_doc/esm#transformers.EsmForMaskedLM) model class and  [`DataCollatorForLanguageModeling(tokenizer, mlm=True)` ](https://huggingface.co/docs/transformers/main/main_classes/data_collator#transformers.DataCollatorForLanguageModeling) data collator which supports MLM loss.\n\nSince there are no train sequences and only 8000 test sequences longer than 208, we downloaded [RNA Central ](https://rnacentral.org/) data and selected 12 million samples with length between 115 and 457 (the range in this competition data) to complement the competition data.\n\nWe pretrained 5-20 epochs depending on how many RNA central data used, for a total of 25M to 70M total samples seen.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2Feb2f9716c901ef7021c64e644a16d2a6%2Fesm2.png?generation=1702046716073336&alt=media)\n\n**Folding Head**\n\nWe adapted the model CPMP used in the Covid Vaccine competition three years ago to devise a folding head on top of ESM2 token embeddings. That model was presented [here](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189723).\n\nThe folding head takes as input both the ESM2 token embeddings and the bpp files. Its main component is an attention/convolution layer, repeated twice, then a linear classification head. \n\nThe Attention/Convolution layers are similar to the transformer encoder structure: an attention layer followed by a convolution bloc with skip connection.\nThe attention layer is a bpp attention layer followed by a bidirectional RNN (GRU or LSTM yields two model variants). Sure, GRU is not attention, but here it plays the same role as an attention layer where each node attends its two neighbors in the sequence.\nThe bpp attention is a standard attention layer where attention weights are the bpp values.\nThe convolution layer was designed after the Efficientnet convolution bloc. We then stack these as in this [notebook](https://www.kaggle.com/thebigd8ta/open-vaccine-pytorch-v).\n\nWe bet on this RNN/Convolution architecture rather than a transformer for two reasons:\n- We literally started working on it 5 days before competition end, and we did not have time to tune a new architecture\n- we hoped that ESM2 attention would have already learned what a transformer head attention would learn.\n\n**Finetuning**\n\nOur best models are the ones with the custom folding head on top of the ESM2 backbone. We also trained simple models with just a linear head. The final ensemble is a mix of both, most being with a folding head.\n\nWe used four  variants of training data:\nQuick start\nQuick start + SN >= 0.3\nQuick start + SN >= 0.3, and SN as sample weight\nQuick start + SN >= 0.1, and SN as sample weight\nSN is signal_to_noise clipped to (0, 1). We used it as sample noise in some variants. Model quality increased from top variant to bottom one.\n\nWe used a 5 fold CV and submitted the average of each fold model predictions.\n\n**Scores**\n\nOur best single model is a 650M LSTM model with SN weight. CV = 0.12847, Public LB = 0.14308 (5 fold ensemble)\nOur best ensemble is a 15 model ensemble (75 folds in total), CV = 0.12478, Public LB =  0.14174\n\n**What didn’t work**\n\nPretrained models from [RNA-FM](https://github.com/ml4bio/RNA-FM) were not as good as our ESM2 pretrained models.\n\nESM2 can also output contact predictions. To do so it stacks all attention activations, then it applies a regression head.  We tried to use that to predict bpp values via an auxiliary loss. This improved the competition metric score a bit. But this comes at the expense of a 4x space and a 2x time increase, which prevented us from using it with 150M and 650M variants of ESM2.\n\nBo & CPMP",
    "2554204": "Congrats! I also joined late and tried similar ideas, but wasn't so successful, lots of room to learn :) Do you know how much value you get from pre-training ESM with MLM? I also saw poor performance of RNA-FM and concluded MLM might not be helpful here, because we're dealing with a lot of synthetic sequences, as opposed to ESM that is trained on evolutionary designed proteins.",
    "2554223": "Pretraining was key. We saw a performance improvement with more pretraining. I'm sorry I can't say more because I have not tried without a pretrained ESM2.\n\nYour score is not that bad!",
    "2554905": "Did you convert the RNA to proteins/AA?",
    "2554955": "no, we replaced the vocab.txt files used by the ESM tokenizer. We used this as `rna_vocab.txt`:\n\n```python\n<cls>\n<pad>\n<eos>\n<unk>\nA\nC\nG\nU\n<null_1>\n<mask>\n```\n\nThen create the tokenizer with:\n\n`tokenizer = EsmTokenizer('../rna_vocab.txt', model_max_length=context_length)`",
    "2554991": "cpmpml Did you try not using bpp? Curious how that would work",
    "2555708": "Yes, we have submitted ESM2 alone. Best one has 0.14781 private / 0.14359 public. But it looks like an outlier, others have much worse scores. \n\n@boliu0 did the ESM2 part, maybe he can say more.",
    "2556239": "> Do you know how much value you get from pre-training ESM with MLM?\n\nA lot. In my very first baseline, I compared non-pretrained ESM 35M vs a pretrained ESM 35M (pre-trained on competition train set only, so it’s not the best pretrained version) on the downstream supervised task, the MAE is not close. 0.01-0.02 gap if I remember correctly. So I never tried non-pretrained ESM anymore.",
    "2556245": "The one with 0.14781 private / 0.14359 public is not ESM2 alone actually. My incorrect naming of that sub confused you. \n\nBest ESM alone models:\nESM 35M: public 0.15095, private 0.16162\nESM 150M: public 0.15110, private 0.15836\nESM 650M: public 0.14789, private 0.15331\n\nThey are obviously not as good as bpp models, but they add diversity and boost cv scores, so we included them in the final ensemble.",
    "2559478": "Wow, 21st in just 10 days is insane! Congratulations to you and your team!"
  },
  "source": "meta"
}