{
  "id": 460222,
  "title": "8th place solution (KF Part)",
  "url": "/competitions/stanford-ribonanza-rna-folding/writeups/kazuki-2-dieter-8th-place-solution-kf-part",
  "author_name": "",
  "post_date": "2023-12-19T05:25:21.890Z",
  "votes": 27,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, I would like to thank the competition hosts for their exceptional organization and the challenging yet enriching environment they created. Their dedication to fostering a space for learning and innovation is deeply appreciated. I also express my profound gratitude to my teammates, <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>. Their collaboration, expertise, and unwavering commitment throughout this project were invaluable.</p>\n<p>Here I will try to summarize some of the main points of our solution.</p>\n<h1>Solutions from our temamates</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460956\" target=\"_blank\">onodera part</a></li>\n</ul>\n<h1>CV Strategy: GroupKFold</h1>\n<p>In the field of RNA secondary structure estimation, various studies including <a href=\"https://academic.oup.com/bib/article/24/4/bbad186/7179751?login=false\" target=\"_blank\">Sato et al., 2023,</a> have highlighted the inappropriateness of using random splits for evaluation.This stems from the concept of 'Families' in RNA experimental data - groups of RNA molecules sharing specific functions or structures, resulting in high structural similarity within the same family. For instance, <a href=\"https://academic.oup.com/nar/article/50/3/e14/6430845\" target=\"_blank\">Fu et al., 2022</a> noted that the <a href=\"https://academic.oup.com/view-large/figure/333767038/gkab1074fig3.jpg\" target=\"_blank\">E2EFold model was overestimated due to this issue</a>.</p>\n<p>To address this, we adopted clustering based on RNA sequence edit distances, using cluster IDs for GroupKFold. Additionally, we aggregated all the limited samples with seqlen=206 into fold=0. This strategy allowed for continuous evaluation of the model's performance on longer sequences.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2311404%2Fb39fa96937be0505cbe1b8480910784e%2Fsplit.png?generation=1702029388253122&amp;alt=media\" alt=\"\"></p>\n<h1>Datasets</h1>\n<p>As part of my models, I experimented with alternating training using both the RMDB dataset and the competition dataset. This approach was adopted with the anticipation that it would enhance the model's ability to handle longer RNA sequences, a critical aspect in RNA secondary structure prediction.</p>\n<h1>Features</h1>\n<p>We processed input sequences using the 'arnie' package, mainly utilizing eternafold, rnasoft, and rnastructure. The model inputs included embedding layers for structure and loop_type, and a Graph Neural Network (GNN) adjacency matrix for bpp (base pair probability).</p>\n<p>However, bpp alone, representing the probability of forming a pair, requires many layers when combined with CNNs to view the entire graph. As an alternative to using Transformers for a global view, we utilized “structure” (renamed to 'chunk' due to naming conflicts) and “segment”, as defined in <a href=\"https://academic.oup.com/nar/article/46/11/5381/4994207\" target=\"_blank\">bpRNA</a> [Danaee et al., 2018].</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2311404%2Fd2a22732d565f34536b3ad5c9af99220%2Fgraph_features.png?generation=1702027001721359&amp;alt=media\" alt=\"\"></p>\n<h1>Models</h1>\n<p>My architecture is inspired by <a href=\"https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784?login=false\" target=\"_blank\">LegNet</a> from <a href=\"https://www.kaggle.com/dmitrypenzar1996\" target=\"_blank\">@dmitrypenzar1996</a> and <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189241\" target=\"_blank\">OpenVaccine's 6th place solution</a> [nyanp]. It consists solely of CNN and GNN layers. The adjacency matrices for bpp, structure, chunk, and segment significantly differ, so I prepared independent CNN+GNN blocks for each type. Group Convolution and einsum enabled this without for loops.</p>\n<p>Following <a href=\"https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784?login=false\" target=\"_blank\">LegNet</a>, prediction head outputs a 100d vector instead of a 1d scalar, with weighted summation over bin values for the final output. This approach stabilized learning by controlling the output range in regression tasks.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2311404%2Feed37f263a4a1321dbd91ccaf915d66c%2Fmodel.png?generation=1702027231710653&amp;alt=media\" alt=\"\"></p>\n<h1>Loss</h1>\n<p>We opted for an MAE + MSE weighted by signal-to-noise ratio as my optimization function. MSE provides gradients similar to MAE in the early stages of training but approaches zero later, aiding in convergence.</p>\n<h1>Pseudo labels</h1>\n<p>We conducted pseudo label training using predictions on the test set created collaboratively with teammates <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>.</p>\n<h1>Result</h1>\n<table>\n<thead>\n<tr>\n<th>Name</th>\n<th>CV</th>\n<th>CV (seqlen=206)</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>single best (scratch training)</td>\n<td>0.1336</td>\n<td>0.1133</td>\n<td>0.14012</td>\n<td>0.14299</td>\n</tr>\n<tr>\n<td>single best (+pseudo label 1st)</td>\n<td>0.1313</td>\n<td>0.1152</td>\n<td>0.13828</td>\n<td>0.14222</td>\n</tr>\n<tr>\n<td>single best (+pseudo label 2nd)</td>\n<td>0.1306</td>\n<td>0.1221</td>\n<td>0.13739</td>\n<td>0.14186</td>\n</tr>\n<tr>\n<td>blending w/ all models</td>\n<td>0.127645</td>\n<td>0.109341</td>\n<td>0.13626</td>\n<td>0.14263</td>\n</tr>\n</tbody>\n</table>\n<h1>References</h1>\n<ul>\n<li>Fu, Laiyi, et al. \"<a href=\"https://academic.oup.com/nar/article/50/3/e14/6430845\" target=\"_blank\">UFold: fast and accurate RNA secondary structure prediction with deep learning.</a>\" Nucleic acids research 50.3 (2022): e14-e14.</li>\n<li>Sato, Kengo, and Michiaki Hamada. \"<a href=\"https://academic.oup.com/bib/article/24/4/bbad186/7179751?login=false\" target=\"_blank\">Recent trends in RNA informatics: a review of machine learning and deep learning for RNA secondary structure prediction and RNA drug discovery.</a>\" Briefings in Bioinformatics (2023): bbad186.</li>\n<li>Danaee, Padideh, et al. \"<a href=\"https://academic.oup.com/nar/article/46/11/5381/4994207\" target=\"_blank\">bpRNA: large-scale automated annotation and analysis of RNA secondary structure.</a>\"&nbsp;<em>Nucleic acids research</em>&nbsp;46.11 (2018): 5381-5394.</li>\n<li>Sato, Kengo, and Michiaki Hamada. \"<a href=\"https://academic.oup.com/bib/article/24/4/bbad186/7179751?login=false\" target=\"_blank\">Recent trends in RNA informatics: a review of machine learning and deep learning for RNA secondary structure prediction and RNA drug discovery.</a>\" Briefings in Bioinformatics (2023): bbad186.</li>\n<li>Penzar, Dmitry, et al. \"<a href=\"https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784?login=false\" target=\"_blank\">LegNet: a best-in-class deep learning model for short DNA regulatory regions.</a>\"&nbsp;<em>Bioinformatics</em>&nbsp;39.8 (2023): btad45</li>\n</ul>",
  "messages": [
    {
      "id": "2553502",
      "postDate": "12/08/2023 10:01:42",
      "content": "<p>First of all, I would like to thank the competition hosts for their exceptional organization and the challenging yet enriching environment they created. Their dedication to fostering a space for learning and innovation is deeply appreciated. I also express my profound gratitude to my teammates, <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>. Their collaboration, expertise, and unwavering commitment throughout this project were invaluable.</p>\n<p>Here I will try to summarize some of the main points of our solution.</p>\n<h1>Solutions from our temamates</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460956\" target=\"_blank\">onodera part</a></li>\n</ul>\n<h1>CV Strategy: GroupKFold</h1>\n<p>In the field of RNA secondary structure estimation, various studies including <a href=\"https://academic.oup.com/bib/article/24/4/bbad186/7179751?login=false\" target=\"_blank\">Sato et al., 2023,</a> have highlighted the inappropriateness of using random splits for evaluation.This stems from the concept of 'Families' in RNA experimental data - groups of RNA molecules sharing specific functions or structures, resulting in high structural similarity within the same family. For instance, <a href=\"https://academic.oup.com/nar/article/50/3/e14/6430845\" target=\"_blank\">Fu et al., 2022</a> noted that the <a href=\"https://academic.oup.com/view-large/figure/333767038/gkab1074fig3.jpg\" target=\"_blank\">E2EFold model was overestimated due to this issue</a>.</p>\n<p>To address this, we adopted clustering based on RNA sequence edit distances, using cluster IDs for GroupKFold. Additionally, we aggregated all the limited samples with seqlen=206 into fold=0. This strategy allowed for continuous evaluation of the model's performance on longer sequences.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2311404%2Fb39fa96937be0505cbe1b8480910784e%2Fsplit.png?generation=1702029388253122&amp;alt=media\" alt=\"\"></p>\n<h1>Datasets</h1>\n<p>As part of my models, I experimented with alternating training using both the RMDB dataset and the competition dataset. This approach was adopted with the anticipation that it would enhance the model's ability to handle longer RNA sequences, a critical aspect in RNA secondary structure prediction.</p>\n<h1>Features</h1>\n<p>We processed input sequences using the 'arnie' package, mainly utilizing eternafold, rnasoft, and rnastructure. The model inputs included embedding layers for structure and loop_type, and a Graph Neural Network (GNN) adjacency matrix for bpp (base pair probability).</p>\n<p>However, bpp alone, representing the probability of forming a pair, requires many layers when combined with CNNs to view the entire graph. As an alternative to using Transformers for a global view, we utilized “structure” (renamed to 'chunk' due to naming conflicts) and “segment”, as defined in <a href=\"https://academic.oup.com/nar/article/46/11/5381/4994207\" target=\"_blank\">bpRNA</a> [Danaee et al., 2018].</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2311404%2Fd2a22732d565f34536b3ad5c9af99220%2Fgraph_features.png?generation=1702027001721359&amp;alt=media\" alt=\"\"></p>\n<h1>Models</h1>\n<p>My architecture is inspired by <a href=\"https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784?login=false\" target=\"_blank\">LegNet</a> from <a href=\"https://www.kaggle.com/dmitrypenzar1996\" target=\"_blank\">@dmitrypenzar1996</a> and <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189241\" target=\"_blank\">OpenVaccine's 6th place solution</a> [nyanp]. It consists solely of CNN and GNN layers. The adjacency matrices for bpp, structure, chunk, and segment significantly differ, so I prepared independent CNN+GNN blocks for each type. Group Convolution and einsum enabled this without for loops.</p>\n<p>Following <a href=\"https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784?login=false\" target=\"_blank\">LegNet</a>, prediction head outputs a 100d vector instead of a 1d scalar, with weighted summation over bin values for the final output. This approach stabilized learning by controlling the output range in regression tasks.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2311404%2Feed37f263a4a1321dbd91ccaf915d66c%2Fmodel.png?generation=1702027231710653&amp;alt=media\" alt=\"\"></p>\n<h1>Loss</h1>\n<p>We opted for an MAE + MSE weighted by signal-to-noise ratio as my optimization function. MSE provides gradients similar to MAE in the early stages of training but approaches zero later, aiding in convergence.</p>\n<h1>Pseudo labels</h1>\n<p>We conducted pseudo label training using predictions on the test set created collaboratively with teammates <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>.</p>\n<h1>Result</h1>\n<table>\n<thead>\n<tr>\n<th>Name</th>\n<th>CV</th>\n<th>CV (seqlen=206)</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>single best (scratch training)</td>\n<td>0.1336</td>\n<td>0.1133</td>\n<td>0.14012</td>\n<td>0.14299</td>\n</tr>\n<tr>\n<td>single best (+pseudo label 1st)</td>\n<td>0.1313</td>\n<td>0.1152</td>\n<td>0.13828</td>\n<td>0.14222</td>\n</tr>\n<tr>\n<td>single best (+pseudo label 2nd)</td>\n<td>0.1306</td>\n<td>0.1221</td>\n<td>0.13739</td>\n<td>0.14186</td>\n</tr>\n<tr>\n<td>blending w/ all models</td>\n<td>0.127645</td>\n<td>0.109341</td>\n<td>0.13626</td>\n<td>0.14263</td>\n</tr>\n</tbody>\n</table>\n<h1>References</h1>\n<ul>\n<li>Fu, Laiyi, et al. \"<a href=\"https://academic.oup.com/nar/article/50/3/e14/6430845\" target=\"_blank\">UFold: fast and accurate RNA secondary structure prediction with deep learning.</a>\" Nucleic acids research 50.3 (2022): e14-e14.</li>\n<li>Sato, Kengo, and Michiaki Hamada. \"<a href=\"https://academic.oup.com/bib/article/24/4/bbad186/7179751?login=false\" target=\"_blank\">Recent trends in RNA informatics: a review of machine learning and deep learning for RNA secondary structure prediction and RNA drug discovery.</a>\" Briefings in Bioinformatics (2023): bbad186.</li>\n<li>Danaee, Padideh, et al. \"<a href=\"https://academic.oup.com/nar/article/46/11/5381/4994207\" target=\"_blank\">bpRNA: large-scale automated annotation and analysis of RNA secondary structure.</a>\"&nbsp;<em>Nucleic acids research</em>&nbsp;46.11 (2018): 5381-5394.</li>\n<li>Sato, Kengo, and Michiaki Hamada. \"<a href=\"https://academic.oup.com/bib/article/24/4/bbad186/7179751?login=false\" target=\"_blank\">Recent trends in RNA informatics: a review of machine learning and deep learning for RNA secondary structure prediction and RNA drug discovery.</a>\" Briefings in Bioinformatics (2023): bbad186.</li>\n<li>Penzar, Dmitry, et al. \"<a href=\"https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784?login=false\" target=\"_blank\">LegNet: a best-in-class deep learning model for short DNA regulatory regions.</a>\"&nbsp;<em>Bioinformatics</em>&nbsp;39.8 (2023): btad45</li>\n</ul>",
      "rawMarkdown": "First of all, I would like to thank the competition hosts for their exceptional organization and the challenging yet enriching environment they created. Their dedication to fostering a space for learning and innovation is deeply appreciated. I also express my profound gratitude to my teammates, @onodera and @christofhenkel. Their collaboration, expertise, and unwavering commitment throughout this project were invaluable.\n\nHere I will try to summarize some of the main points of our solution.\n\n# Solutions from our temamates\n\n- [onodera part](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460956)\n\n# CV Strategy: GroupKFold\n\nIn the field of RNA secondary structure estimation, various studies including [Sato et al., 2023,](https://academic.oup.com/bib/article/24/4/bbad186/7179751?login=false) have highlighted the inappropriateness of using random splits for evaluation.This stems from the concept of 'Families' in RNA experimental data - groups of RNA molecules sharing specific functions or structures, resulting in high structural similarity within the same family. For instance, [Fu et al., 2022](https://academic.oup.com/nar/article/50/3/e14/6430845) noted that the [E2EFold model was overestimated due to this issue](https://academic.oup.com/view-large/figure/333767038/gkab1074fig3.jpg).\n\nTo address this, we adopted clustering based on RNA sequence edit distances, using cluster IDs for GroupKFold. Additionally, we aggregated all the limited samples with seqlen=206 into fold=0. This strategy allowed for continuous evaluation of the model's performance on longer sequences.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2311404%2Fb39fa96937be0505cbe1b8480910784e%2Fsplit.png?generation=1702029388253122&alt=media)\n\n# Datasets\n\nAs part of my models, I experimented with alternating training using both the RMDB dataset and the competition dataset. This approach was adopted with the anticipation that it would enhance the model's ability to handle longer RNA sequences, a critical aspect in RNA secondary structure prediction.\n\n# Features\n\nWe processed input sequences using the 'arnie' package, mainly utilizing eternafold, rnasoft, and rnastructure. The model inputs included embedding layers for structure and loop_type, and a Graph Neural Network (GNN) adjacency matrix for bpp (base pair probability).\n\nHowever, bpp alone, representing the probability of forming a pair, requires many layers when combined with CNNs to view the entire graph. As an alternative to using Transformers for a global view, we utilized “structure” (renamed to 'chunk' due to naming conflicts) and “segment”, as defined in [bpRNA](https://academic.oup.com/nar/article/46/11/5381/4994207) [Danaee et al., 2018].\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2311404%2Fd2a22732d565f34536b3ad5c9af99220%2Fgraph_features.png?generation=1702027001721359&alt=media)\n\n# Models\n\nMy architecture is inspired by [LegNet](https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784?login=false) from @dmitrypenzar1996 and [OpenVaccine's 6th place solution](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189241) [nyanp]. It consists solely of CNN and GNN layers. The adjacency matrices for bpp, structure, chunk, and segment significantly differ, so I prepared independent CNN+GNN blocks for each type. Group Convolution and einsum enabled this without for loops.\n\nFollowing [LegNet](https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784?login=false), prediction head outputs a 100d vector instead of a 1d scalar, with weighted summation over bin values for the final output. This approach stabilized learning by controlling the output range in regression tasks.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2311404%2Feed37f263a4a1321dbd91ccaf915d66c%2Fmodel.png?generation=1702027231710653&alt=media)\n\n# Loss\n\nWe opted for an MAE + MSE weighted by signal-to-noise ratio as my optimization function. MSE provides gradients similar to MAE in the early stages of training but approaches zero later, aiding in convergence.\n\n# Pseudo labels\n\nWe conducted pseudo label training using predictions on the test set created collaboratively with teammates @onodera and @christofhenkel.\n\n# Result\n\n|Name|CV|CV (seqlen=206)|Public LB|Private LB|\n|:--:|:--:|:--:|\n|single best (scratch training)|0.1336|0.1133|0.14012|0.14299|\n|single best (+pseudo label 1st)|0.1313|0.1152|0.13828|0.14222|\n|single best (+pseudo label 2nd)|0.1306|0.1221|0.13739|0.14186|\n|blending w/ all models|0.127645|0.109341|0.13626|0.14263|\n\n# References\n\n- Fu, Laiyi, et al. \"[UFold: fast and accurate RNA secondary structure prediction with deep learning.](https://academic.oup.com/nar/article/50/3/e14/6430845)\" Nucleic acids research 50.3 (2022): e14-e14.\n- Sato, Kengo, and Michiaki Hamada. \"[Recent trends in RNA informatics: a review of machine learning and deep learning for RNA secondary structure prediction and RNA drug discovery.](https://academic.oup.com/bib/article/24/4/bbad186/7179751?login=false)\" Briefings in Bioinformatics (2023): bbad186.\n- Danaee, Padideh, et al. \"[bpRNA: large-scale automated annotation and analysis of RNA secondary structure.](https://academic.oup.com/nar/article/46/11/5381/4994207)\" *Nucleic acids research* 46.11 (2018): 5381-5394.\n- Sato, Kengo, and Michiaki Hamada. \"[Recent trends in RNA informatics: a review of machine learning and deep learning for RNA secondary structure prediction and RNA drug discovery.](https://academic.oup.com/bib/article/24/4/bbad186/7179751?login=false)\" Briefings in Bioinformatics (2023): bbad186.\n- Penzar, Dmitry, et al. \"[LegNet: a best-in-class deep learning model for short DNA regulatory regions.](https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784?login=false)\" *Bioinformatics* 39.8 (2023): btad45",
      "votes": null
    },
    {
      "id": "2553667",
      "postDate": "12/08/2023 12:26:07",
      "content": "<p>Oh, you made LegNet work for this task! Can't wait to see the code if you are going to publish it! </p>",
      "rawMarkdown": "Oh, you made LegNet work for this task! Can't wait to see the code if you are going to publish it!",
      "votes": null
    },
    {
      "id": "2556082",
      "postDate": "12/10/2023 13:02:16",
      "content": "<p>Thanks for your comment.<br>\nMy code is dirty so I don't have a plan to publish yet, but the main differences between my approach and LegNet are the presence of GNN layers and the objective function.<br>\nGNN layers based on secondary structures mentioned above enable the model to attend to long-range contacts.</p>",
      "rawMarkdown": "Thanks for your comment.\nMy code is dirty so I don't have a plan to publish yet, but the main differences between my approach and LegNet are the presence of GNN layers and the objective function.\nGNN layers based on secondary structures mentioned above enable the model to attend to long-range contacts.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2553667,
      "author_name": "dmitrypenzar1996",
      "author_url": "",
      "post_date": "12/08/2023 12:26:07",
      "content": "<p>Oh, you made LegNet work for this task! Can't wait to see the code if you are going to publish it! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2556082,
          "author_name": "kfujikawa",
          "author_url": "",
          "post_date": "12/10/2023 13:02:16",
          "content": "<p>Thanks for your comment.<br>\nMy code is dirty so I don't have a plan to publish yet, but the main differences between my approach and LegNet are the presence of GNN layers and the objective function.<br>\nGNN layers based on secondary structures mentioned above enable the model to attend to long-range contacts.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2553502": "First of all, I would like to thank the competition hosts for their exceptional organization and the challenging yet enriching environment they created. Their dedication to fostering a space for learning and innovation is deeply appreciated. I also express my profound gratitude to my teammates, @onodera and @christofhenkel. Their collaboration, expertise, and unwavering commitment throughout this project were invaluable.\n\nHere I will try to summarize some of the main points of our solution.\n\n# Solutions from our temamates\n\n- [onodera part](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460956)\n\n# CV Strategy: GroupKFold\n\nIn the field of RNA secondary structure estimation, various studies including [Sato et al., 2023,](https://academic.oup.com/bib/article/24/4/bbad186/7179751?login=false) have highlighted the inappropriateness of using random splits for evaluation.This stems from the concept of 'Families' in RNA experimental data - groups of RNA molecules sharing specific functions or structures, resulting in high structural similarity within the same family. For instance, [Fu et al., 2022](https://academic.oup.com/nar/article/50/3/e14/6430845) noted that the [E2EFold model was overestimated due to this issue](https://academic.oup.com/view-large/figure/333767038/gkab1074fig3.jpg).\n\nTo address this, we adopted clustering based on RNA sequence edit distances, using cluster IDs for GroupKFold. Additionally, we aggregated all the limited samples with seqlen=206 into fold=0. This strategy allowed for continuous evaluation of the model's performance on longer sequences.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2311404%2Fb39fa96937be0505cbe1b8480910784e%2Fsplit.png?generation=1702029388253122&alt=media)\n\n# Datasets\n\nAs part of my models, I experimented with alternating training using both the RMDB dataset and the competition dataset. This approach was adopted with the anticipation that it would enhance the model's ability to handle longer RNA sequences, a critical aspect in RNA secondary structure prediction.\n\n# Features\n\nWe processed input sequences using the 'arnie' package, mainly utilizing eternafold, rnasoft, and rnastructure. The model inputs included embedding layers for structure and loop_type, and a Graph Neural Network (GNN) adjacency matrix for bpp (base pair probability).\n\nHowever, bpp alone, representing the probability of forming a pair, requires many layers when combined with CNNs to view the entire graph. As an alternative to using Transformers for a global view, we utilized “structure” (renamed to 'chunk' due to naming conflicts) and “segment”, as defined in [bpRNA](https://academic.oup.com/nar/article/46/11/5381/4994207) [Danaee et al., 2018].\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2311404%2Fd2a22732d565f34536b3ad5c9af99220%2Fgraph_features.png?generation=1702027001721359&alt=media)\n\n# Models\n\nMy architecture is inspired by [LegNet](https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784?login=false) from @dmitrypenzar1996 and [OpenVaccine's 6th place solution](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189241) [nyanp]. It consists solely of CNN and GNN layers. The adjacency matrices for bpp, structure, chunk, and segment significantly differ, so I prepared independent CNN+GNN blocks for each type. Group Convolution and einsum enabled this without for loops.\n\nFollowing [LegNet](https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784?login=false), prediction head outputs a 100d vector instead of a 1d scalar, with weighted summation over bin values for the final output. This approach stabilized learning by controlling the output range in regression tasks.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2311404%2Feed37f263a4a1321dbd91ccaf915d66c%2Fmodel.png?generation=1702027231710653&alt=media)\n\n# Loss\n\nWe opted for an MAE + MSE weighted by signal-to-noise ratio as my optimization function. MSE provides gradients similar to MAE in the early stages of training but approaches zero later, aiding in convergence.\n\n# Pseudo labels\n\nWe conducted pseudo label training using predictions on the test set created collaboratively with teammates @onodera and @christofhenkel.\n\n# Result\n\n|Name|CV|CV (seqlen=206)|Public LB|Private LB|\n|:--:|:--:|:--:|\n|single best (scratch training)|0.1336|0.1133|0.14012|0.14299|\n|single best (+pseudo label 1st)|0.1313|0.1152|0.13828|0.14222|\n|single best (+pseudo label 2nd)|0.1306|0.1221|0.13739|0.14186|\n|blending w/ all models|0.127645|0.109341|0.13626|0.14263|\n\n# References\n\n- Fu, Laiyi, et al. \"[UFold: fast and accurate RNA secondary structure prediction with deep learning.](https://academic.oup.com/nar/article/50/3/e14/6430845)\" Nucleic acids research 50.3 (2022): e14-e14.\n- Sato, Kengo, and Michiaki Hamada. \"[Recent trends in RNA informatics: a review of machine learning and deep learning for RNA secondary structure prediction and RNA drug discovery.](https://academic.oup.com/bib/article/24/4/bbad186/7179751?login=false)\" Briefings in Bioinformatics (2023): bbad186.\n- Danaee, Padideh, et al. \"[bpRNA: large-scale automated annotation and analysis of RNA secondary structure.](https://academic.oup.com/nar/article/46/11/5381/4994207)\" *Nucleic acids research* 46.11 (2018): 5381-5394.\n- Sato, Kengo, and Michiaki Hamada. \"[Recent trends in RNA informatics: a review of machine learning and deep learning for RNA secondary structure prediction and RNA drug discovery.](https://academic.oup.com/bib/article/24/4/bbad186/7179751?login=false)\" Briefings in Bioinformatics (2023): bbad186.\n- Penzar, Dmitry, et al. \"[LegNet: a best-in-class deep learning model for short DNA regulatory regions.](https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784?login=false)\" *Bioinformatics* 39.8 (2023): btad45",
    "2553667": "Oh, you made LegNet work for this task! Can't wait to see the code if you are going to publish it!",
    "2556082": "Thanks for your comment.\nMy code is dirty so I don't have a plan to publish yet, but the main differences between my approach and LegNet are the presence of GNN layers and the objective function.\nGNN layers based on secondary structures mentioned above enable the model to attend to long-range contacts."
  },
  "source": "meta"
}