{
  "id": 460192,
  "title": "11th place solution: Transformer with MFE distance embeddings (with code)",
  "url": "/competitions/stanford-ribonanza-rna-folding/writeups/maxim-tsygankov-11th-place-solution-transformer-wi",
  "author_name": "",
  "post_date": "2023-12-08T06:35:24.887123700Z",
  "votes": 16,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I greatly appreciate the efforts of the hosts in organizing this competition. This was an unfamiliar domain filled with both challenges and fun. I wonder if the competition was worth it for organizers after all. Hopefully, the solutions shared by kagglers will prove beneficial to you and other researchers.</p>\n<p>Here is a quick summary of my solution:</p>\n<h2>Features</h2>\n<ul>\n<li><strong>Eternafold MFE and BPPs</strong>: I experimented with other packages (RNAstructure, Vienna, CONTRAfold, rnasoft, ipknot).  I found that only the features from Eternafold contributed significantly to my model.</li>\n<li><strong>CapR structure</strong>. This led to faster convergence of the model.</li>\n<li><strong>bpRNA structure</strong>. I am unsure of its contribution to the score.</li>\n</ul>\n<h2>Model architecture</h2>\n<p>I forked Huggingface's BERT and implemented several modifications:</p>\n<ul>\n<li><strong>Attention injection support</strong>: I added BPPs with 2d cnn on top to 8/16 of model's attention heads. The most effective variant involved a separate CNN layer for each attention head. However, this significantly increased training time with only a marginal improvement in score. I sacrificed a few score points in favor of faster experimentation. Most of my models in the final ensemble only use one 2d cnn layer per transformer block.</li>\n<li><strong>1d CNN layers between transformer layers</strong>: This addition was based on the premise that closely situated nucleotides influence each other, and a CNN layer could effectively extract these features.</li>\n<li><strong>Relative embeddings</strong>: I only used relative embeddings in my model. Initially, they seemed a safer choice due to the length differences between train/test sequences. More importantly, they appeared more intuitive in the context of RNA structure, as opposed to absolute embeddings. I used two distance types:<ul>\n<li>The shortest distance along ribose-phosphate chain.</li>\n<li>The shortest distance along the MFE structure.  </li>\n<li>To illustrate this point, consider the image below, which depicts an RNA molecule from <a href=\"https://www.kaggle.com/code/brainbowrna/rna-science-computational-environment\" target=\"_blank\">this notebook</a>. C and G are encircled and located at opposite ends of the RNA chain. Despite their positioning, they are actually paired with each other, indicating that the distance between them is merely 1 in terms of the RNA folding structure.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1420334%2Fbdcae78bf822ea1e1f00b4192ab4e784%2Frna_image_dists.png?generation=1702011163871853&amp;alt=media\" alt=\"\"></li></ul></li>\n</ul>\n<h2>Performance</h2>\n<p>The single model achieved a public score of approximately 0.1432. However, by averaging models with minor variations, I was able to achieve a score below 0.14.  </p>\n<p>The code is available at <a href=\"https://github.com/chubasik/stanford-ribonanza-rna-folding\" target=\"_blank\">https://github.com/chubasik/stanford-ribonanza-rna-folding</a></p>\n<p>P.S: I had to post it a second time for it to be marked as a write-up</p>",
  "messages": [
    {
      "id": "2553321",
      "postDate": "12/08/2023 06:35:24",
      "content": "<p>I greatly appreciate the efforts of the hosts in organizing this competition. This was an unfamiliar domain filled with both challenges and fun. I wonder if the competition was worth it for organizers after all. Hopefully, the solutions shared by kagglers will prove beneficial to you and other researchers.</p>\n<p>Here is a quick summary of my solution:</p>\n<h2>Features</h2>\n<ul>\n<li><strong>Eternafold MFE and BPPs</strong>: I experimented with other packages (RNAstructure, Vienna, CONTRAfold, rnasoft, ipknot).  I found that only the features from Eternafold contributed significantly to my model.</li>\n<li><strong>CapR structure</strong>. This led to faster convergence of the model.</li>\n<li><strong>bpRNA structure</strong>. I am unsure of its contribution to the score.</li>\n</ul>\n<h2>Model architecture</h2>\n<p>I forked Huggingface's BERT and implemented several modifications:</p>\n<ul>\n<li><strong>Attention injection support</strong>: I added BPPs with 2d cnn on top to 8/16 of model's attention heads. The most effective variant involved a separate CNN layer for each attention head. However, this significantly increased training time with only a marginal improvement in score. I sacrificed a few score points in favor of faster experimentation. Most of my models in the final ensemble only use one 2d cnn layer per transformer block.</li>\n<li><strong>1d CNN layers between transformer layers</strong>: This addition was based on the premise that closely situated nucleotides influence each other, and a CNN layer could effectively extract these features.</li>\n<li><strong>Relative embeddings</strong>: I only used relative embeddings in my model. Initially, they seemed a safer choice due to the length differences between train/test sequences. More importantly, they appeared more intuitive in the context of RNA structure, as opposed to absolute embeddings. I used two distance types:<ul>\n<li>The shortest distance along ribose-phosphate chain.</li>\n<li>The shortest distance along the MFE structure.  </li>\n<li>To illustrate this point, consider the image below, which depicts an RNA molecule from <a href=\"https://www.kaggle.com/code/brainbowrna/rna-science-computational-environment\" target=\"_blank\">this notebook</a>. C and G are encircled and located at opposite ends of the RNA chain. Despite their positioning, they are actually paired with each other, indicating that the distance between them is merely 1 in terms of the RNA folding structure.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1420334%2Fbdcae78bf822ea1e1f00b4192ab4e784%2Frna_image_dists.png?generation=1702011163871853&amp;alt=media\" alt=\"\"></li></ul></li>\n</ul>\n<h2>Performance</h2>\n<p>The single model achieved a public score of approximately 0.1432. However, by averaging models with minor variations, I was able to achieve a score below 0.14.  </p>\n<p>The code is available at <a href=\"https://github.com/chubasik/stanford-ribonanza-rna-folding\" target=\"_blank\">https://github.com/chubasik/stanford-ribonanza-rna-folding</a></p>\n<p>P.S: I had to post it a second time for it to be marked as a write-up</p>",
      "rawMarkdown": "I greatly appreciate the efforts of the hosts in organizing this competition. This was an unfamiliar domain filled with both challenges and fun. I wonder if the competition was worth it for organizers after all. Hopefully, the solutions shared by kagglers will prove beneficial to you and other researchers.\n\nHere is a quick summary of my solution:\n## Features\n* **Eternafold MFE and BPPs**: I experimented with other packages (RNAstructure, Vienna, CONTRAfold, rnasoft, ipknot).  I found that only the features from Eternafold contributed significantly to my model.\n* **CapR structure**. This led to faster convergence of the model.\n* **bpRNA structure**. I am unsure of its contribution to the score.\n\n## Model architecture\nI forked Huggingface's BERT and implemented several modifications:\n* **Attention injection support**: I added BPPs with 2d cnn on top to 8/16 of model's attention heads. The most effective variant involved a separate CNN layer for each attention head. However, this significantly increased training time with only a marginal improvement in score. I sacrificed a few score points in favor of faster experimentation. Most of my models in the final ensemble only use one 2d cnn layer per transformer block.\n* **1d CNN layers between transformer layers**: This addition was based on the premise that closely situated nucleotides influence each other, and a CNN layer could effectively extract these features.\n* **Relative embeddings**: I only used relative embeddings in my model. Initially, they seemed a safer choice due to the length differences between train/test sequences. More importantly, they appeared more intuitive in the context of RNA structure, as opposed to absolute embeddings. I used two distance types:\n  * The shortest distance along ribose-phosphate chain.\n  * The shortest distance along the MFE structure.  \n  * To illustrate this point, consider the image below, which depicts an RNA molecule from [this notebook](https://www.kaggle.com/code/brainbowrna/rna-science-computational-environment). C and G are encircled and located at opposite ends of the RNA chain. Despite their positioning, they are actually paired with each other, indicating that the distance between them is merely 1 in terms of the RNA folding structure.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1420334%2Fbdcae78bf822ea1e1f00b4192ab4e784%2Frna_image_dists.png?generation=1702011163871853&alt=media)\n\n## Performance\nThe single model achieved a public score of approximately 0.1432. However, by averaging models with minor variations, I was able to achieve a score below 0.14.  \n\nThe code is available at https://github.com/chubasik/stanford-ribonanza-rna-folding\n\nP.S: I had to post it a second time for it to be marked as a write-up",
      "votes": null
    },
    {
      "id": "2553470",
      "postDate": "12/08/2023 09:20:58",
      "content": "<p>Congratulations on being the topper in the competition. Also, thank you for sharing the method.</p>",
      "rawMarkdown": "Congratulations on being the topper in the competition. Also, thank you for sharing the method.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2553470,
      "author_name": "sarunpm",
      "author_url": "",
      "post_date": "12/08/2023 09:20:58",
      "content": "<p>Congratulations on being the topper in the competition. Also, thank you for sharing the method.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2553321": "I greatly appreciate the efforts of the hosts in organizing this competition. This was an unfamiliar domain filled with both challenges and fun. I wonder if the competition was worth it for organizers after all. Hopefully, the solutions shared by kagglers will prove beneficial to you and other researchers.\n\nHere is a quick summary of my solution:\n## Features\n* **Eternafold MFE and BPPs**: I experimented with other packages (RNAstructure, Vienna, CONTRAfold, rnasoft, ipknot).  I found that only the features from Eternafold contributed significantly to my model.\n* **CapR structure**. This led to faster convergence of the model.\n* **bpRNA structure**. I am unsure of its contribution to the score.\n\n## Model architecture\nI forked Huggingface's BERT and implemented several modifications:\n* **Attention injection support**: I added BPPs with 2d cnn on top to 8/16 of model's attention heads. The most effective variant involved a separate CNN layer for each attention head. However, this significantly increased training time with only a marginal improvement in score. I sacrificed a few score points in favor of faster experimentation. Most of my models in the final ensemble only use one 2d cnn layer per transformer block.\n* **1d CNN layers between transformer layers**: This addition was based on the premise that closely situated nucleotides influence each other, and a CNN layer could effectively extract these features.\n* **Relative embeddings**: I only used relative embeddings in my model. Initially, they seemed a safer choice due to the length differences between train/test sequences. More importantly, they appeared more intuitive in the context of RNA structure, as opposed to absolute embeddings. I used two distance types:\n  * The shortest distance along ribose-phosphate chain.\n  * The shortest distance along the MFE structure.  \n  * To illustrate this point, consider the image below, which depicts an RNA molecule from [this notebook](https://www.kaggle.com/code/brainbowrna/rna-science-computational-environment). C and G are encircled and located at opposite ends of the RNA chain. Despite their positioning, they are actually paired with each other, indicating that the distance between them is merely 1 in terms of the RNA folding structure.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1420334%2Fbdcae78bf822ea1e1f00b4192ab4e784%2Frna_image_dists.png?generation=1702011163871853&alt=media)\n\n## Performance\nThe single model achieved a public score of approximately 0.1432. However, by averaging models with minor variations, I was able to achieve a score below 0.14.  \n\nThe code is available at https://github.com/chubasik/stanford-ribonanza-rna-folding\n\nP.S: I had to post it a second time for it to be marked as a write-up",
    "2553470": "Congratulations on being the topper in the competition. Also, thank you for sharing the method."
  },
  "source": "meta"
}