{
  "id": 460190,
  "title": "7th place solution",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460190",
  "author_name": "Iafoss",
  "post_date": "2023-12-08T06:33:17.327000",
  "votes": 44,
  "comment_count": 32,
  "views": 0,
  "content": "<h2>Summary</h2>\n<ul>\n<li>Transformer based solution</li>\n<li>Masked Conv1D instead of MLP</li>\n<li>bpp injection into attention matrix, bpp-based token mixing, matrix mixing, and dual stream setup (attention boosting) for incorporating bpp</li>\n</ul>\n<h2>Introduction</h2>\n<p>Our team would like to thank the organizers and Kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> and <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> for their incredible contribution toward our final result.</p>\n<h2>Details</h2>\n<h3>Data</h3>\n<p><strong>BPP</strong>: We generated additional bpp using vienna_2 (we also used SS output), contrafold_2, rnaformerv1, and rnafm. However, we saw nearly negligible improvement in comparison to using only bpp provided by organizers.<br>\n<a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/454397\" target=\"_blank\"><strong>EX data</strong></a>: We tried to do fine-tuning on this data since it is only the source that provides GT for sequence ends. We got 5-10 bps CV improvement from this procedure (Iafoss/DrHB) and got the best visualization for long-range dependencies. However, at the sequence ends the predictions tend to get values close to zero for unknown reasons. At private LB EX data did not give any boost.<br>\n<strong>CV split</strong>. slime used a random 4-fold split. DrHB and Iafoss used a split based on sequence similarity to avoid any possible leaks and also we excluded any train data overlapping with the test test from the train. The size of the val set is ~20k samples that passed SN criterion.</p>\n<h3>Model</h3>\n<p>At the beginning of the competition we quickly realized that bpp is very helpful for model performance. Initially, we tried to use it as an auxiliary output computed as the mean attention matrix, but convergence was quite slow (the typical run was 200-250 epochs). Then we build setups that also take bpp as an input. It drastically accelerated convergence, so that it took only several tens of epochs, while giving similar results.<br>\nWe considered several setups:<br>\n(1) <strong>bpp injection</strong> (Iafoss): add bpp (in the logit form) directly to the attention matrix (in case of multiple bpp, each bpp is added to the particular head). We gradually attenuate the injection towards zero in the last layers to give the mode the freedom to develop interactions missed in bpp. The critical part of the model was replacing MLP with a masked 1d x5 convolution that performs the mixing of neighboring tokens. The transformer has 384 width, a head size of 64, 24 depth, a droppath of 0.3, dropout of 0.1. We use rotary enc to define the relative position of tokens. Single bpp and 6 bpp setups were considered. The best single model got 0.14375 private and 0.14013 public LB and 0.14292 and 0.13973 with using pseudo labels.<br>\n(2) <strong>bpp mixing</strong> (DrHB): instead of injection of bpp into the attention matrix, we tried to add matrix multiplication modules performing mixing tokens based on bpp. The total structure of the model: 3 x [bpp eterna -&gt; conv -&gt;transformer -&gt; bpp rest mean -&gt;conv -&gt;transformer -&gt;ss- as adj -&gt; graph transformer ].<br>\n(3) <strong>matrix mixing</strong> (slime): trying to follow the previous competition solutions we added a learnable attention bias produced by a convolutional stream applied to bpp and ALIBI-like positional encoding. We used 2 bpp sources. The matrix mixer is represented by several conv layers with SE modules. The transformer consists of 12 matrix mixing + 12 regular transformer blocks with a width of 384. The best single model is 0.14509 private and 0.14066 public LB. This model used a different training pipeline and cannot be directly compared to others.<br>\n(4) <strong>dual stream</strong> (Iafoss). The drawback of matrix mixing is that the attention bias is updated independently from the transformer based on the input bpp. How about attention boosting? We perform a simple projection of the attention state followed by scaled tanh nonlinearity to limit accumulated values and stabilize training. This value is added to the Attention matrix and then the result is fed to the next transformer layer. This modification has improved the model performance giving 0.14296 at private and 0.13697 public LB as a single model. Unfortunately, we discovered this setup only a few days before the end of the competition, and didn't have a chance to run multiple trainings and PL setup, which we expect to give a further ~10 bps improvement of LB.</p>\n<p>The models are schematically depicted in the image below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fd537664beced5e0923d38a5cd8a612e3%2Fplots.png?generation=1702136589769473&amp;alt=media\" alt=\"\"></p>\n<h3>Training</h3>\n<p><strong>(Iafoss/DrHB)</strong> The loss is weighted based on the error, while we do not downselect data based on SN criterion: <code>w = 1/sqrt(1/6 + err.clip(100))</code>. Use AdamW, cosine annealing with warmup, lr=5e-4, wd=0.05, bs=16 (small bs was working better for the reason we could not identify). We used flip augmentation for both train and TTA, but the key thing here is using bpp computed for the correct order of nucleotides. It quite improves CV giving 10-15 bps boost. We also used bin based auxiliary loss which gave a slight improvement.</p>\n<p>We split the training portion of the data into 4 folds, train 4 models (40-48 epochs), fine-tune on external data + train data (5-6 epochs with x10 EX data oversampling), fine-tune on noise-free samples for 12 epochs. Then we corrected the provided train data weighting GT and PL based on the inverse error (assuming 0.15 error for PL). The provided train data has a large level of noise slowing down the convergence. Test data and the sequence ends were labeled solely based on PL. Then we train a PL model on the whole train (excluding val samples) + test data, fine-tune on EX data, and fine-tune on noise-free samples. This procedure improved CV by 10 bpp, and the weighted average of all models generated in the procedure gets further improvement.</p>\n<p><strong>(slime)</strong> <em>Pre-training</em>: First we performed MLM pre-training on a whole dataset for 5 epochs (40% of input tokens are replaced with mask tokens). We used cosine decay to zero with one epoch warmup and AdamW optimizer with base_lr= 5e-4 and wd=0.05. The model was trained to predict missing nucleotides with cross-entropy loss</p>\n<p><em>Fine-tuning</em>: During the fine-tuning stage we initialized our model with weights from MLM-pretraining, it gave a noticeable improvement to the final result [10 bps CV]. We fine-tuned our models for 50 epochs with batch_size=16 on samples where either SN(DMS) == 1 or SN(2A3) == 1, we masked the loss depending on SN of the given test [10 bps boost compared to filtering train dataset with SN(DMS) == 1 and SN(2a3) == 1). In addition, we weighted the samples based on their reactivity error provided by the ground-truth data as <code>loss *= torch.log(1.1 + snr) / 2</code>. Similar to pertaining, we used AdamW optimizer [lr=5e-4, wd=0.05] and cosine scheduler with lr decay to zero. Since in matrix mixing model masking was not considered in convolutions, this model was trained with length-matching batch sampling (samples of exactly the same length).</p>\n<h3>Best singe model end ensemble</h3>\n<p>Best single model (dual stream model): 0.14292 at private and 0.13711 public LB, which can take top 10 itself. This result would be further improved by 10 bps to ~0.1419 if if we had time to run our full PL pipeline.<br>\nOur final submission is a combination of ~20 models that got 0.14189 at private and 0.13604 at public LB.</p>\n<h3>Things didn't work</h3>\n<ul>\n<li>EX data was helpful at CV and public LB (5-10 bps boost), but not helpful at private LB</li>\n<li>Additional bpps gave only a negligible improvement in comparison to the use of single bpp provided by organizers</li>\n<li>MLM on 30M external RNA sequence dataset</li>\n<li>EMA, AWP, Floyd-warshall distance matrices </li>\n<li>2D Ushape models</li>\n</ul>",
  "messages": [
    {
      "id": 2553317,
      "postDate": "2023-12-08T06:33:17.327Z",
      "content": "<h2>Summary</h2>\n<ul>\n<li>Transformer based solution</li>\n<li>Masked Conv1D instead of MLP</li>\n<li>bpp injection into attention matrix, bpp-based token mixing, matrix mixing, and dual stream setup (attention boosting) for incorporating bpp</li>\n</ul>\n<h2>Introduction</h2>\n<p>Our team would like to thank the organizers and Kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> and <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> for their incredible contribution toward our final result.</p>\n<h2>Details</h2>\n<h3>Data</h3>\n<p><strong>BPP</strong>: We generated additional bpp using vienna_2 (we also used SS output), contrafold_2, rnaformerv1, and rnafm. However, we saw nearly negligible improvement in comparison to using only bpp provided by organizers.<br>\n<a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/454397\" target=\"_blank\"><strong>EX data</strong></a>: We tried to do fine-tuning on this data since it is only the source that provides GT for sequence ends. We got 5-10 bps CV improvement from this procedure (Iafoss/DrHB) and got the best visualization for long-range dependencies. However, at the sequence ends the predictions tend to get values close to zero for unknown reasons. At private LB EX data did not give any boost.<br>\n<strong>CV split</strong>. slime used a random 4-fold split. DrHB and Iafoss used a split based on sequence similarity to avoid any possible leaks and also we excluded any train data overlapping with the test test from the train. The size of the val set is ~20k samples that passed SN criterion.</p>\n<h3>Model</h3>\n<p>At the beginning of the competition we quickly realized that bpp is very helpful for model performance. Initially, we tried to use it as an auxiliary output computed as the mean attention matrix, but convergence was quite slow (the typical run was 200-250 epochs). Then we build setups that also take bpp as an input. It drastically accelerated convergence, so that it took only several tens of epochs, while giving similar results.<br>\nWe considered several setups:<br>\n(1) <strong>bpp injection</strong> (Iafoss): add bpp (in the logit form) directly to the attention matrix (in case of multiple bpp, each bpp is added to the particular head). We gradually attenuate the injection towards zero in the last layers to give the mode the freedom to develop interactions missed in bpp. The critical part of the model was replacing MLP with a masked 1d x5 convolution that performs the mixing of neighboring tokens. The transformer has 384 width, a head size of 64, 24 depth, a droppath of 0.3, dropout of 0.1. We use rotary enc to define the relative position of tokens. Single bpp and 6 bpp setups were considered. The best single model got 0.14375 private and 0.14013 public LB and 0.14292 and 0.13973 with using pseudo labels.<br>\n(2) <strong>bpp mixing</strong> (DrHB): instead of injection of bpp into the attention matrix, we tried to add matrix multiplication modules performing mixing tokens based on bpp. The total structure of the model: 3 x [bpp eterna -&gt; conv -&gt;transformer -&gt; bpp rest mean -&gt;conv -&gt;transformer -&gt;ss- as adj -&gt; graph transformer ].<br>\n(3) <strong>matrix mixing</strong> (slime): trying to follow the previous competition solutions we added a learnable attention bias produced by a convolutional stream applied to bpp and ALIBI-like positional encoding. We used 2 bpp sources. The matrix mixer is represented by several conv layers with SE modules. The transformer consists of 12 matrix mixing + 12 regular transformer blocks with a width of 384. The best single model is 0.14509 private and 0.14066 public LB. This model used a different training pipeline and cannot be directly compared to others.<br>\n(4) <strong>dual stream</strong> (Iafoss). The drawback of matrix mixing is that the attention bias is updated independently from the transformer based on the input bpp. How about attention boosting? We perform a simple projection of the attention state followed by scaled tanh nonlinearity to limit accumulated values and stabilize training. This value is added to the Attention matrix and then the result is fed to the next transformer layer. This modification has improved the model performance giving 0.14296 at private and 0.13697 public LB as a single model. Unfortunately, we discovered this setup only a few days before the end of the competition, and didn't have a chance to run multiple trainings and PL setup, which we expect to give a further ~10 bps improvement of LB.</p>\n<p>The models are schematically depicted in the image below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fd537664beced5e0923d38a5cd8a612e3%2Fplots.png?generation=1702136589769473&amp;alt=media\" alt=\"\"></p>\n<h3>Training</h3>\n<p><strong>(Iafoss/DrHB)</strong> The loss is weighted based on the error, while we do not downselect data based on SN criterion: <code>w = 1/sqrt(1/6 + err.clip(100))</code>. Use AdamW, cosine annealing with warmup, lr=5e-4, wd=0.05, bs=16 (small bs was working better for the reason we could not identify). We used flip augmentation for both train and TTA, but the key thing here is using bpp computed for the correct order of nucleotides. It quite improves CV giving 10-15 bps boost. We also used bin based auxiliary loss which gave a slight improvement.</p>\n<p>We split the training portion of the data into 4 folds, train 4 models (40-48 epochs), fine-tune on external data + train data (5-6 epochs with x10 EX data oversampling), fine-tune on noise-free samples for 12 epochs. Then we corrected the provided train data weighting GT and PL based on the inverse error (assuming 0.15 error for PL). The provided train data has a large level of noise slowing down the convergence. Test data and the sequence ends were labeled solely based on PL. Then we train a PL model on the whole train (excluding val samples) + test data, fine-tune on EX data, and fine-tune on noise-free samples. This procedure improved CV by 10 bpp, and the weighted average of all models generated in the procedure gets further improvement.</p>\n<p><strong>(slime)</strong> <em>Pre-training</em>: First we performed MLM pre-training on a whole dataset for 5 epochs (40% of input tokens are replaced with mask tokens). We used cosine decay to zero with one epoch warmup and AdamW optimizer with base_lr= 5e-4 and wd=0.05. The model was trained to predict missing nucleotides with cross-entropy loss</p>\n<p><em>Fine-tuning</em>: During the fine-tuning stage we initialized our model with weights from MLM-pretraining, it gave a noticeable improvement to the final result [10 bps CV]. We fine-tuned our models for 50 epochs with batch_size=16 on samples where either SN(DMS) == 1 or SN(2A3) == 1, we masked the loss depending on SN of the given test [10 bps boost compared to filtering train dataset with SN(DMS) == 1 and SN(2a3) == 1). In addition, we weighted the samples based on their reactivity error provided by the ground-truth data as <code>loss *= torch.log(1.1 + snr) / 2</code>. Similar to pertaining, we used AdamW optimizer [lr=5e-4, wd=0.05] and cosine scheduler with lr decay to zero. Since in matrix mixing model masking was not considered in convolutions, this model was trained with length-matching batch sampling (samples of exactly the same length).</p>\n<h3>Best singe model end ensemble</h3>\n<p>Best single model (dual stream model): 0.14292 at private and 0.13711 public LB, which can take top 10 itself. This result would be further improved by 10 bps to ~0.1419 if if we had time to run our full PL pipeline.<br>\nOur final submission is a combination of ~20 models that got 0.14189 at private and 0.13604 at public LB.</p>\n<h3>Things didn't work</h3>\n<ul>\n<li>EX data was helpful at CV and public LB (5-10 bps boost), but not helpful at private LB</li>\n<li>Additional bpps gave only a negligible improvement in comparison to the use of single bpp provided by organizers</li>\n<li>MLM on 30M external RNA sequence dataset</li>\n<li>EMA, AWP, Floyd-warshall distance matrices </li>\n<li>2D Ushape models</li>\n</ul>",
      "rawMarkdown": "## Summary\n- Transformer based solution\n- Masked Conv1D instead of MLP\n- bpp injection into attention matrix, bpp-based token mixing, matrix mixing, and dual stream setup (attention boosting) for incorporating bpp\n\n## Introduction\nOur team would like to thank the organizers and Kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates @drhabib and @martynoveduard for their incredible contribution toward our final result.\n\n## Details\n### Data\n**BPP**: We generated additional bpp using vienna_2 (we also used SS output), contrafold_2, rnaformerv1, and rnafm. However, we saw nearly negligible improvement in comparison to using only bpp provided by organizers.\n[**EX data**](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/454397): We tried to do fine-tuning on this data since it is only the source that provides GT for sequence ends. We got 5-10 bps CV improvement from this procedure (Iafoss/DrHB) and got the best visualization for long-range dependencies. However, at the sequence ends the predictions tend to get values close to zero for unknown reasons. At private LB EX data did not give any boost.\n**CV split**. slime used a random 4-fold split. DrHB and Iafoss used a split based on sequence similarity to avoid any possible leaks and also we excluded any train data overlapping with the test test from the train. The size of the val set is ~20k samples that passed SN criterion.\n\n### Model\nAt the beginning of the competition we quickly realized that bpp is very helpful for model performance. Initially, we tried to use it as an auxiliary output computed as the mean attention matrix, but convergence was quite slow (the typical run was 200-250 epochs). Then we build setups that also take bpp as an input. It drastically accelerated convergence, so that it took only several tens of epochs, while giving similar results.\nWe considered several setups:\n(1) **bpp injection** (Iafoss): add bpp (in the logit form) directly to the attention matrix (in case of multiple bpp, each bpp is added to the particular head). We gradually attenuate the injection towards zero in the last layers to give the mode the freedom to develop interactions missed in bpp. The critical part of the model was replacing MLP with a masked 1d x5 convolution that performs the mixing of neighboring tokens. The transformer has 384 width, a head size of 64, 24 depth, a droppath of 0.3, dropout of 0.1. We use rotary enc to define the relative position of tokens. Single bpp and 6 bpp setups were considered. The best single model got 0.14375 private and 0.14013 public LB and 0.14292 and 0.13973 with using pseudo labels.\n(2) **bpp mixing** (DrHB): instead of injection of bpp into the attention matrix, we tried to add matrix multiplication modules performing mixing tokens based on bpp. The total structure of the model: 3 x [bpp eterna -> conv ->transformer -> bpp rest mean ->conv ->transformer ->ss- as adj -> graph transformer ].\n(3) **matrix mixing** (slime): trying to follow the previous competition solutions we added a learnable attention bias produced by a convolutional stream applied to bpp and ALIBI-like positional encoding. We used 2 bpp sources. The matrix mixer is represented by several conv layers with SE modules. The transformer consists of 12 matrix mixing + 12 regular transformer blocks with a width of 384. The best single model is 0.14509 private and 0.14066 public LB. This model used a different training pipeline and cannot be directly compared to others.\n(4) **dual stream** (Iafoss). The drawback of matrix mixing is that the attention bias is updated independently from the transformer based on the input bpp. How about attention boosting? We perform a simple projection of the attention state followed by scaled tanh nonlinearity to limit accumulated values and stabilize training. This value is added to the Attention matrix and then the result is fed to the next transformer layer. This modification has improved the model performance giving 0.14296 at private and 0.13697 public LB as a single model. Unfortunately, we discovered this setup only a few days before the end of the competition, and didn't have a chance to run multiple trainings and PL setup, which we expect to give a further ~10 bps improvement of LB.\n\nThe models are schematically depicted in the image below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fd537664beced5e0923d38a5cd8a612e3%2Fplots.png?generation=1702136589769473&alt=media)\n\n### Training\n**(Iafoss/DrHB)** The loss is weighted based on the error, while we do not downselect data based on SN criterion: `w = 1/sqrt(1/6 + err.clip(100))`. Use AdamW, cosine annealing with warmup, lr=5e-4, wd=0.05, bs=16 (small bs was working better for the reason we could not identify). We used flip augmentation for both train and TTA, but the key thing here is using bpp computed for the correct order of nucleotides. It quite improves CV giving 10-15 bps boost. We also used bin based auxiliary loss which gave a slight improvement.\n\nWe split the training portion of the data into 4 folds, train 4 models (40-48 epochs), fine-tune on external data + train data (5-6 epochs with x10 EX data oversampling), fine-tune on noise-free samples for 12 epochs. Then we corrected the provided train data weighting GT and PL based on the inverse error (assuming 0.15 error for PL). The provided train data has a large level of noise slowing down the convergence. Test data and the sequence ends were labeled solely based on PL. Then we train a PL model on the whole train (excluding val samples) + test data, fine-tune on EX data, and fine-tune on noise-free samples. This procedure improved CV by 10 bpp, and the weighted average of all models generated in the procedure gets further improvement.\n\n**(slime)** *Pre-training*: First we performed MLM pre-training on a whole dataset for 5 epochs (40% of input tokens are replaced with mask tokens). We used cosine decay to zero with one epoch warmup and AdamW optimizer with base_lr= 5e-4 and wd=0.05. The model was trained to predict missing nucleotides with cross-entropy loss\n\n*Fine-tuning*: During the fine-tuning stage we initialized our model with weights from MLM-pretraining, it gave a noticeable improvement to the final result [10 bps CV]. We fine-tuned our models for 50 epochs with batch_size=16 on samples where either SN(DMS) == 1 or SN(2A3) == 1, we masked the loss depending on SN of the given test [10 bps boost compared to filtering train dataset with SN(DMS) == 1 and SN(2a3) == 1). In addition, we weighted the samples based on their reactivity error provided by the ground-truth data as `loss *= torch.log(1.1 + snr) / 2`. Similar to pertaining, we used AdamW optimizer [lr=5e-4, wd=0.05] and cosine scheduler with lr decay to zero. Since in matrix mixing model masking was not considered in convolutions, this model was trained with length-matching batch sampling (samples of exactly the same length).\n\n### Best singe model end ensemble\nBest single model (dual stream model): 0.14292 at private and 0.13711 public LB, which can take top 10 itself. This result would be further improved by 10 bps to ~0.1419 if if we had time to run our full PL pipeline.\nOur final submission is a combination of ~20 models that got 0.14189 at private and 0.13604 at public LB.\n\n### Things didn't work\n- EX data was helpful at CV and public LB (5-10 bps boost), but not helpful at private LB\n- Additional bpps gave only a negligible improvement in comparison to the use of single bpp provided by organizers\n- MLM on 30M external RNA sequence dataset\n- EMA, AWP, Floyd-warshall distance matrices \n- 2D Ushape models",
      "votes": 44
    },
    {
      "id": 2554198,
      "postDate": "2023-12-08T22:16:29.313Z",
      "content": "<p>One of the interesting things we tried during early experiments with merging bpp and distance matrices is to incorporate Floyd-Warshall distance matrices obtained from bpp graph, the motivation was to provide the model with more clues about distances from the given nucleotide to the paired nodes.</p>\n<p>We constructed the graph from the nucleotides in the following way:<br>\nFor the given sequence we extract hard-edges from eterna-bpp matrix, when base pair probability exceeds 0.5, to make the graph connected, we make edges between nodes (i, i+1) as well, after that we run Floyd-Warshall algorithm on this graph.</p>\n<p>We implemented it in C++, also we tried CUDA as well, but it gave comparable speed when running 32-threads CPU version, since graphs are relatively sparse</p>\n<p>left plot: resulting distance matrix, right plot: hard-clipped bpp matrix + hardcoded (i, i+1) edges<br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2F91af1798b6e88ce6ba77cf3d9604cab6%2Fimage%20(10).png?generation=1702073617798053&amp;alt=media\" alt=\"\"></p>\n<p>Unfortunately, it didn't give an improvement compared to just using bpp &amp; distance matrices with MatrixMixer layers, it seems that model learned this property by itself</p>",
      "rawMarkdown": "One of the interesting things we tried during early experiments with merging bpp and distance matrices is to incorporate Floyd-Warshall distance matrices obtained from bpp graph, the motivation was to provide the model with more clues about distances from the given nucleotide to the paired nodes.\n\nWe constructed the graph from the nucleotides in the following way:\nFor the given sequence we extract hard-edges from eterna-bpp matrix, when base pair probability exceeds 0.5, to make the graph connected, we make edges between nodes (i, i+1) as well, after that we run Floyd-Warshall algorithm on this graph.\n\nWe implemented it in C++, also we tried CUDA as well, but it gave comparable speed when running 32-threads CPU version, since graphs are relatively sparse\n\nleft plot: resulting distance matrix, right plot: hard-clipped bpp matrix + hardcoded (i, i+1) edges\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2F91af1798b6e88ce6ba77cf3d9604cab6%2Fimage%20(10).png?generation=1702073617798053&alt=media)\n\nUnfortunately, it didn't give an improvement compared to just using bpp & distance matrices with MatrixMixer layers, it seems that model learned this property by itself",
      "votes": 8
    },
    {
      "id": 2553670,
      "postDate": "2023-12-08T12:32:27.510Z",
      "content": "<p>Congrats on the 7th position!</p>",
      "rawMarkdown": "Congrats on the 7th position!",
      "votes": 3
    },
    {
      "id": 2558898,
      "postDate": "2023-12-12T12:27:58.767Z",
      "content": "<p>Hi there! <br>\nFirst congrats for 7th place. </p>\n<p>Seeing this amazing approach, my question is:  Is this intuition from any paper? If yes would you mind to share as i would like to read more?</p>",
      "rawMarkdown": "Hi there! \nFirst congrats for 7th place. \n\nSeeing this amazing approach, my question is:  Is this intuition from any paper? If yes would you mind to share as i would like to read more?",
      "votes": 2,
      "replies": [
        {
          "id": 2559156,
          "postDate": "2023-12-12T16:21:47.620Z",
          "content": "<p>Thanks. During the competition we tried some ideas from papers, like 2d Ushape networks, etc., but they did not work well. Our matrix mixer approach is motivated by <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189564\" target=\"_blank\">one of OpenVacine solutions</a> as I understand, but other models were developed based on our work on this challenge.</p>",
          "rawMarkdown": "Thanks. During the competition we tried some ideas from papers, like 2d Ushape networks, etc., but they did not work well. Our matrix mixer approach is motivated by [one of OpenVacine solutions](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189564) as I understand, but other models were developed based on our work on this challenge.",
          "votes": 2,
          "replies": [
            {
              "id": 2559910,
              "postDate": "2023-12-13T06:25:14.070Z",
              "content": "<p>Do you have any references for the masked1D conv? (thanks)</p>",
              "rawMarkdown": "Do you have any references for the masked1D conv? (thanks)",
              "votes": 1
            },
            {
              "id": 2560794,
              "postDate": "2023-12-14T02:31:57.823Z",
              "content": "<p>No, but it is pretty simple. Before each convolution in the network you must set all padded tokens to 0, so the output of the network is exactly the same regardless the padding you add. Otherwise you may start accumulating some states in the padding tokens, which is not the thing you really want to have. You may try to skip this masking by just having all sequences of the same length and zero padding, but at inference whenever you add padding you get an undefined behaviour.</p>",
              "rawMarkdown": "No, but it is pretty simple. Before each convolution in the network you must set all padded tokens to 0, so the output of the network is exactly the same regardless the padding you add. Otherwise you may start accumulating some states in the padding tokens, which is not the thing you really want to have. You may try to skip this masking by just having all sequences of the same length and zero padding, but at inference whenever you add padding you get an undefined behaviour.",
              "votes": 5
            },
            {
              "id": 2579451,
              "postDate": "2023-12-30T05:23:17.663Z",
              "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> do you have head to head comparison on how much this padded convolution helps? I have seen some unexpected behavior in my modeling which may be due to padding tokens accumulating</p>",
              "rawMarkdown": "@iafoss do you have head to head comparison on how much this padded convolution helps? I have seen some unexpected behavior in my modeling which may be due to padding tokens accumulating"
            },
            {
              "id": 2579469,
              "postDate": "2023-12-30T05:42:45.967Z",
              "content": "<p>I didn't do a careful analysis of that. I just tried to make make sure that I do not have any unexpected things in inference when batches may be composed of different length sequences. </p>",
              "rawMarkdown": "I didn't do a careful analysis of that. I just tried to make make sure that I do not have any unexpected things in inference when batches may be composed of different length sequences. "
            },
            {
              "id": 2579474,
              "postDate": "2023-12-30T05:52:44.820Z",
              "content": "<p>We tried both variants (with and without masked convolutions) and padding definitely helps</p>",
              "rawMarkdown": "We tried both variants (with and without masked convolutions) and padding definitely helps"
            },
            {
              "id": 2579479,
              "postDate": "2023-12-30T06:01:17.213Z",
              "content": "<p>I see thanks a lot!</p>",
              "rawMarkdown": "I see thanks a lot!"
            }
          ]
        }
      ]
    },
    {
      "id": 2555786,
      "postDate": "2023-12-10T07:47:24.397Z",
      "content": "<p>DUde that's cool</p>",
      "rawMarkdown": "DUde that's cool",
      "votes": 2
    },
    {
      "id": 2554720,
      "postDate": "2023-12-09T10:54:05.323Z",
      "content": "<p>Very nice details - thank you! <br>\nCongrats on 7th position.</p>\n<p>Naive question  what tool you used to create network images?</p>",
      "rawMarkdown": "Very nice details - thank you! \nCongrats on 7th position.\n\nNaive question  what tool you used to create network images?",
      "votes": 2,
      "replies": [
        {
          "id": 2554981,
          "postDate": "2023-12-09T15:32:34.423Z",
          "content": "<p>Thanks. I just used PowerPoint.</p>",
          "rawMarkdown": "Thanks. I just used PowerPoint.",
          "votes": 2
        },
        {
          "id": 2556495,
          "postDate": "2023-12-10T18:06:09.580Z",
          "content": "<blockquote>\n  <p>Very nice details - thank you! <br>\n  Congrats on 7th position.</p>\n  <p>Naive question  what tool you used to create network images?</p>\n</blockquote>",
          "rawMarkdown": "> Very nice details - thank you! \n> Congrats on 7th position.\n> \n> Naive question  what tool you used to create network images?\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2553361,
      "postDate": "2023-12-08T07:37:24.400Z",
      "content": "<p>Hi. Do you mind explaining how you weighted according to the loss?</p>",
      "rawMarkdown": "Hi. Do you mind explaining how you weighted according to the loss?",
      "votes": 2,
      "replies": [
        {
          "id": 2554163,
          "postDate": "2023-12-08T20:54:16.200Z",
          "content": "<p>I added these details.</p>",
          "rawMarkdown": "I added these details.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2553354,
      "postDate": "2023-12-08T07:19:10.840Z",
      "content": "<p>Hi<br>\nI do not understand the bpp, <br>\nwould mind give me more information about the bpp.<br>\nthank you very much. </p>",
      "rawMarkdown": "Hi\nI do not understand the bpp, \nwould mind give me more information about the bpp.\nthank you very much. ",
      "votes": 2,
      "replies": [
        {
          "id": 2554167,
          "postDate": "2023-12-08T20:57:16.770Z",
          "content": "<p>It is an abbreviation for base pairing probability, i.e. the probability that nucleotides interact in the secondary structure as I understand. It is provided by organizers for eternafold package, and we computed it with several more packages.</p>",
          "rawMarkdown": "It is an abbreviation for base pairing probability, i.e. the probability that nucleotides interact in the secondary structure as I understand. It is provided by organizers for eternafold package, and we computed it with several more packages.",
          "votes": 4,
          "replies": [
            {
              "id": 2554822,
              "postDate": "2023-12-09T12:41:49.660Z",
              "content": "<p>thank you very much,<br>\nfinally understand…<br>\nnice study.<br>\nありがとうございます。<br>\n非常感谢。。。</p>",
              "rawMarkdown": "thank you very much,\nfinally understand...\nnice study.\nありがとうございます。\n非常感谢。。。",
              "votes": 1
            },
            {
              "id": 2555232,
              "postDate": "2023-12-09T18:42:56.097Z",
              "content": "<p>thank you! I must admit your helps in other places is very helpful to me. </p>",
              "rawMarkdown": "thank you! I must admit your helps in other places is very helpful to me. ",
              "votes": 1
            },
            {
              "id": 2555234,
              "postDate": "2023-12-09T18:47:04.220Z",
              "content": "<p>Whether bpp was also available in test data?</p>",
              "rawMarkdown": "Whether bpp was also available in test data?",
              "votes": 1
            },
            {
              "id": 2555261,
              "postDate": "2023-12-09T19:16:50.703Z",
              "content": "<p>Yes, as far as I remember. Use bpp was the key thing in the competition, especially for generalization at the private LB. It appeared, that my very early models got 0.18-&gt;0.15 private LB improvement by using bpp auxiliary loss. As the next step we realized that adding bpp as an input, like injection into the attention matrix, for example, increases the convergence speed by ~x10 giving comparable performance to aux setup, and we just focused on that. Not sure, though, if the performance of our latest setups could be reproduced with bpp aux.</p>",
              "rawMarkdown": "Yes, as far as I remember. Use bpp was the key thing in the competition, especially for generalization at the private LB. It appeared, that my very early models got 0.18->0.15 private LB improvement by using bpp auxiliary loss. As the next step we realized that adding bpp as an input, like injection into the attention matrix, for example, increases the convergence speed by ~x10 giving comparable performance to aux setup, and we just focused on that. Not sure, though, if the performance of our latest setups could be reproduced with bpp aux.",
              "votes": 3
            },
            {
              "id": 2555378,
              "postDate": "2023-12-09T21:28:29.047Z",
              "content": "<p><a href=\"https://www.kaggle.com/liuyanfeng\" target=\"_blank\">@liuyanfeng</a> The base pair probabilities reflect the different conformations an RNA molecule can form. Some researchers visualize this behavior as RNA dancing. I explain more in this thread: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/446900\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/446900</a> </p>",
              "rawMarkdown": "@liuyanfeng The base pair probabilities reflect the different conformations an RNA molecule can form. Some researchers visualize this behavior as RNA dancing. I explain more in this thread: https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/446900 ",
              "votes": 1
            },
            {
              "id": 2555436,
              "postDate": "2023-12-09T22:13:22.087Z",
              "content": "<p>Very important information about auxilary loss, I think. One more argument that the final \"production\" model can use no bpp as input. It's totally OK to train such model longer given additional benefits at inference time </p>",
              "rawMarkdown": "Very important information about auxilary loss, I think. One more argument that the final \"production\" model can use no bpp as input. It's totally OK to train such model longer given additional benefits at inference time ",
              "votes": 1
            },
            {
              "id": 2555439,
              "postDate": "2023-12-09T22:15:17.627Z",
              "content": "<p>What kind of loss have you used on bpp matrix? BCE? or smth mae-like?</p>",
              "rawMarkdown": "What kind of loss have you used on bpp matrix? BCE? or smth mae-like?",
              "votes": 1
            },
            {
              "id": 2555454,
              "postDate": "2023-12-09T22:41:00.047Z",
              "content": "<p>BCE            </p>",
              "rawMarkdown": "BCE            ",
              "votes": 3
            },
            {
              "id": 2555765,
              "postDate": "2023-12-10T07:05:25.123Z",
              "content": "<p><a href=\"https://www.kaggle.com/dmitrypenzar1996\" target=\"_blank\">@dmitrypenzar1996</a> Just wanted to check what do you mean by <code>model can use no bpp as input</code>? It seems from <a href=\"https://www.kaggle.com/lafoss\" target=\"_blank\">@lafoss</a> that there is bpp for test sequences</p>",
              "rawMarkdown": "@dmitrypenzar1996 Just wanted to check what do you mean by ` model can use no bpp as input`? It seems from @lafoss that there is bpp for test sequences",
              "votes": 1
            },
            {
              "id": 2556206,
              "postDate": "2023-12-10T14:32:48.553Z",
              "content": "<p>Yes, but bpp are not perfect features -- 1) they tend to hide pseudoknots 2) you need to precalculate them before running a model</p>\n<p>While given the fact they are computed based on a quite limited number of parameters it seems reasonable to expect that model can at least learn that parameters given bpps as a target. </p>",
              "rawMarkdown": "Yes, but bpp are not perfect features -- 1) they tend to hide pseudoknots 2) you need to precalculate them before running a model\n\nWhile given the fact they are computed based on a quite limited number of parameters it seems reasonable to expect that model can at least learn that parameters given bpps as a target. ",
              "votes": 1
            },
            {
              "id": 2556394,
              "postDate": "2023-12-10T16:33:10.493Z",
              "rawMarkdown": "",
              "votes": 2,
              "isDeleted": true
            },
            {
              "id": 2556614,
              "postDate": "2023-12-10T20:48:15.890Z",
              "content": "<p>Yes, about tendency. As far as I remember sRoger plot had more visible pseudoknot although his model has lower performance.</p>",
              "rawMarkdown": "Yes, about tendency. As far as I remember sRoger plot had more visible pseudoknot although his model has lower performance.",
              "votes": 1
            },
            {
              "id": 2556615,
              "postDate": "2023-12-10T20:50:45.070Z",
              "content": "<p>There could be difference in behaviour (e.g hiding pseudoknots) and it will be difference in ease of use as no need in additional features </p>",
              "rawMarkdown": "There could be difference in behaviour (e.g hiding pseudoknots) and it will be difference in ease of use as no need in additional features ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2553319,
      "postDate": "2023-12-08T06:34:47.300Z",
      "content": "<p>Congratulations on achieving 7th position in this competition. Thanks for sharing the details of your solution. </p>",
      "rawMarkdown": "Congratulations on achieving 7th position in this competition. Thanks for sharing the details of your solution. ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2554198,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2023-12-08T22:16:29.313000",
      "content": "<p>One of the interesting things we tried during early experiments with merging bpp and distance matrices is to incorporate Floyd-Warshall distance matrices obtained from bpp graph, the motivation was to provide the model with more clues about distances from the given nucleotide to the paired nodes.</p>\n<p>We constructed the graph from the nucleotides in the following way:<br>\nFor the given sequence we extract hard-edges from eterna-bpp matrix, when base pair probability exceeds 0.5, to make the graph connected, we make edges between nodes (i, i+1) as well, after that we run Floyd-Warshall algorithm on this graph.</p>\n<p>We implemented it in C++, also we tried CUDA as well, but it gave comparable speed when running 32-threads CPU version, since graphs are relatively sparse</p>\n<p>left plot: resulting distance matrix, right plot: hard-clipped bpp matrix + hardcoded (i, i+1) edges<br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2F91af1798b6e88ce6ba77cf3d9604cab6%2Fimage%20(10).png?generation=1702073617798053&amp;alt=media\" alt=\"\"></p>\n<p>Unfortunately, it didn't give an improvement compared to just using bpp &amp; distance matrices with MatrixMixer layers, it seems that model learned this property by itself</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 2553670,
      "author_name": "Antonio Félix",
      "author_url": "",
      "post_date": "2023-12-08T12:32:27.510000",
      "content": "<p>Congrats on the 7th position!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2558898,
      "author_name": "Diogo Santiago",
      "author_url": "",
      "post_date": "2023-12-12T12:27:58.767000",
      "content": "<p>Hi there! <br>\nFirst congrats for 7th place. </p>\n<p>Seeing this amazing approach, my question is:  Is this intuition from any paper? If yes would you mind to share as i would like to read more?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2559156,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-12-12T16:21:47.620000",
          "content": "<p>Thanks. During the competition we tried some ideas from papers, like 2d Ushape networks, etc., but they did not work well. Our matrix mixer approach is motivated by <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189564\" target=\"_blank\">one of OpenVacine solutions</a> as I understand, but other models were developed based on our work on this challenge.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2559910,
              "author_name": "FdotRK",
              "author_url": "",
              "post_date": "2023-12-13T06:25:14.070000",
              "content": "<p>Do you have any references for the masked1D conv? (thanks)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2560794,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-12-14T02:31:57.823000",
              "content": "<p>No, but it is pretty simple. Before each convolution in the network you must set all padded tokens to 0, so the output of the network is exactly the same regardless the padding you add. Otherwise you may start accumulating some states in the padding tokens, which is not the thing you really want to have. You may try to skip this masking by just having all sequences of the same length and zero padding, but at inference whenever you add padding you get an undefined behaviour.</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2579451,
              "author_name": "Shujun",
              "author_url": "",
              "post_date": "2023-12-30T05:23:17.663000",
              "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> do you have head to head comparison on how much this padded convolution helps? I have seen some unexpected behavior in my modeling which may be due to padding tokens accumulating</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2579469,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-12-30T05:42:45.967000",
              "content": "<p>I didn't do a careful analysis of that. I just tried to make make sure that I do not have any unexpected things in inference when batches may be composed of different length sequences. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2579474,
              "author_name": "Penzar Dmitry",
              "author_url": "",
              "post_date": "2023-12-30T05:52:44.820000",
              "content": "<p>We tried both variants (with and without masked convolutions) and padding definitely helps</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2579479,
              "author_name": "Shujun",
              "author_url": "",
              "post_date": "2023-12-30T06:01:17.213000",
              "content": "<p>I see thanks a lot!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2555786,
      "author_name": "Luficer G",
      "author_url": "",
      "post_date": "2023-12-10T07:47:24.397000",
      "content": "<p>DUde that's cool</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2554720,
      "author_name": "FdotRK",
      "author_url": "",
      "post_date": "2023-12-09T10:54:05.323000",
      "content": "<p>Very nice details - thank you! <br>\nCongrats on 7th position.</p>\n<p>Naive question  what tool you used to create network images?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2554981,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-12-09T15:32:34.423000",
          "content": "<p>Thanks. I just used PowerPoint.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2556495,
          "author_name": "Rupes kumar",
          "author_url": "",
          "post_date": "2023-12-10T18:06:09.580000",
          "content": "<blockquote>\n  <p>Very nice details - thank you! <br>\n  Congrats on 7th position.</p>\n  <p>Naive question  what tool you used to create network images?</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2553361,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2023-12-08T07:37:24.400000",
      "content": "<p>Hi. Do you mind explaining how you weighted according to the loss?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2554163,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-12-08T20:54:16.200000",
          "content": "<p>I added these details.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2553354,
      "author_name": "swordsman",
      "author_url": "",
      "post_date": "2023-12-08T07:19:10.840000",
      "content": "<p>Hi<br>\nI do not understand the bpp, <br>\nwould mind give me more information about the bpp.<br>\nthank you very much. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2554167,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-12-08T20:57:16.770000",
          "content": "<p>It is an abbreviation for base pairing probability, i.e. the probability that nucleotides interact in the secondary structure as I understand. It is provided by organizers for eternafold package, and we computed it with several more packages.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2554822,
              "author_name": "swordsman",
              "author_url": "",
              "post_date": "2023-12-09T12:41:49.660000",
              "content": "<p>thank you very much,<br>\nfinally understand…<br>\nnice study.<br>\nありがとうございます。<br>\n非常感谢。。。</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2555232,
              "author_name": "FdotRK",
              "author_url": "",
              "post_date": "2023-12-09T18:42:56.097000",
              "content": "<p>thank you! I must admit your helps in other places is very helpful to me. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2555234,
              "author_name": "FdotRK",
              "author_url": "",
              "post_date": "2023-12-09T18:47:04.220000",
              "content": "<p>Whether bpp was also available in test data?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2555261,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-12-09T19:16:50.703000",
              "content": "<p>Yes, as far as I remember. Use bpp was the key thing in the competition, especially for generalization at the private LB. It appeared, that my very early models got 0.18-&gt;0.15 private LB improvement by using bpp auxiliary loss. As the next step we realized that adding bpp as an input, like injection into the attention matrix, for example, increases the convergence speed by ~x10 giving comparable performance to aux setup, and we just focused on that. Not sure, though, if the performance of our latest setups could be reproduced with bpp aux.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2555378,
              "author_name": "DigitalEmbrace",
              "author_url": "",
              "post_date": "2023-12-09T21:28:29.047000",
              "content": "<p><a href=\"https://www.kaggle.com/liuyanfeng\" target=\"_blank\">@liuyanfeng</a> The base pair probabilities reflect the different conformations an RNA molecule can form. Some researchers visualize this behavior as RNA dancing. I explain more in this thread: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/446900\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/446900</a> </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2555436,
              "author_name": "Penzar Dmitry",
              "author_url": "",
              "post_date": "2023-12-09T22:13:22.087000",
              "content": "<p>Very important information about auxilary loss, I think. One more argument that the final \"production\" model can use no bpp as input. It's totally OK to train such model longer given additional benefits at inference time </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2555439,
              "author_name": "Penzar Dmitry",
              "author_url": "",
              "post_date": "2023-12-09T22:15:17.627000",
              "content": "<p>What kind of loss have you used on bpp matrix? BCE? or smth mae-like?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2555454,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-12-09T22:41:00.047000",
              "content": "<p>BCE            </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2555765,
              "author_name": "FdotRK",
              "author_url": "",
              "post_date": "2023-12-10T07:05:25.123000",
              "content": "<p><a href=\"https://www.kaggle.com/dmitrypenzar1996\" target=\"_blank\">@dmitrypenzar1996</a> Just wanted to check what do you mean by <code>model can use no bpp as input</code>? It seems from <a href=\"https://www.kaggle.com/lafoss\" target=\"_blank\">@lafoss</a> that there is bpp for test sequences</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2556206,
              "author_name": "Penzar Dmitry",
              "author_url": "",
              "post_date": "2023-12-10T14:32:48.553000",
              "content": "<p>Yes, but bpp are not perfect features -- 1) they tend to hide pseudoknots 2) you need to precalculate them before running a model</p>\n<p>While given the fact they are computed based on a quite limited number of parameters it seems reasonable to expect that model can at least learn that parameters given bpps as a target. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2556394,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-12-10T16:33:10.493000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2556614,
              "author_name": "Penzar Dmitry",
              "author_url": "",
              "post_date": "2023-12-10T20:48:15.890000",
              "content": "<p>Yes, about tendency. As far as I remember sRoger plot had more visible pseudoknot although his model has lower performance.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2556615,
              "author_name": "Penzar Dmitry",
              "author_url": "",
              "post_date": "2023-12-10T20:50:45.070000",
              "content": "<p>There could be difference in behaviour (e.g hiding pseudoknots) and it will be difference in ease of use as no need in additional features </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2553319,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-12-08T06:34:47.300000",
      "content": "<p>Congratulations on achieving 7th position in this competition. Thanks for sharing the details of your solution. </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2553317": "## Summary\n- Transformer based solution\n- Masked Conv1D instead of MLP\n- bpp injection into attention matrix, bpp-based token mixing, matrix mixing, and dual stream setup (attention boosting) for incorporating bpp\n\n## Introduction\nOur team would like to thank the organizers and Kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates @drhabib and @martynoveduard for their incredible contribution toward our final result.\n\n## Details\n### Data\n**BPP**: We generated additional bpp using vienna_2 (we also used SS output), contrafold_2, rnaformerv1, and rnafm. However, we saw nearly negligible improvement in comparison to using only bpp provided by organizers.\n[**EX data**](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/454397): We tried to do fine-tuning on this data since it is only the source that provides GT for sequence ends. We got 5-10 bps CV improvement from this procedure (Iafoss/DrHB) and got the best visualization for long-range dependencies. However, at the sequence ends the predictions tend to get values close to zero for unknown reasons. At private LB EX data did not give any boost.\n**CV split**. slime used a random 4-fold split. DrHB and Iafoss used a split based on sequence similarity to avoid any possible leaks and also we excluded any train data overlapping with the test test from the train. The size of the val set is ~20k samples that passed SN criterion.\n\n### Model\nAt the beginning of the competition we quickly realized that bpp is very helpful for model performance. Initially, we tried to use it as an auxiliary output computed as the mean attention matrix, but convergence was quite slow (the typical run was 200-250 epochs). Then we build setups that also take bpp as an input. It drastically accelerated convergence, so that it took only several tens of epochs, while giving similar results.\nWe considered several setups:\n(1) **bpp injection** (Iafoss): add bpp (in the logit form) directly to the attention matrix (in case of multiple bpp, each bpp is added to the particular head). We gradually attenuate the injection towards zero in the last layers to give the mode the freedom to develop interactions missed in bpp. The critical part of the model was replacing MLP with a masked 1d x5 convolution that performs the mixing of neighboring tokens. The transformer has 384 width, a head size of 64, 24 depth, a droppath of 0.3, dropout of 0.1. We use rotary enc to define the relative position of tokens. Single bpp and 6 bpp setups were considered. The best single model got 0.14375 private and 0.14013 public LB and 0.14292 and 0.13973 with using pseudo labels.\n(2) **bpp mixing** (DrHB): instead of injection of bpp into the attention matrix, we tried to add matrix multiplication modules performing mixing tokens based on bpp. The total structure of the model: 3 x [bpp eterna -> conv ->transformer -> bpp rest mean ->conv ->transformer ->ss- as adj -> graph transformer ].\n(3) **matrix mixing** (slime): trying to follow the previous competition solutions we added a learnable attention bias produced by a convolutional stream applied to bpp and ALIBI-like positional encoding. We used 2 bpp sources. The matrix mixer is represented by several conv layers with SE modules. The transformer consists of 12 matrix mixing + 12 regular transformer blocks with a width of 384. The best single model is 0.14509 private and 0.14066 public LB. This model used a different training pipeline and cannot be directly compared to others.\n(4) **dual stream** (Iafoss). The drawback of matrix mixing is that the attention bias is updated independently from the transformer based on the input bpp. How about attention boosting? We perform a simple projection of the attention state followed by scaled tanh nonlinearity to limit accumulated values and stabilize training. This value is added to the Attention matrix and then the result is fed to the next transformer layer. This modification has improved the model performance giving 0.14296 at private and 0.13697 public LB as a single model. Unfortunately, we discovered this setup only a few days before the end of the competition, and didn't have a chance to run multiple trainings and PL setup, which we expect to give a further ~10 bps improvement of LB.\n\nThe models are schematically depicted in the image below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fd537664beced5e0923d38a5cd8a612e3%2Fplots.png?generation=1702136589769473&alt=media)\n\n### Training\n**(Iafoss/DrHB)** The loss is weighted based on the error, while we do not downselect data based on SN criterion: `w = 1/sqrt(1/6 + err.clip(100))`. Use AdamW, cosine annealing with warmup, lr=5e-4, wd=0.05, bs=16 (small bs was working better for the reason we could not identify). We used flip augmentation for both train and TTA, but the key thing here is using bpp computed for the correct order of nucleotides. It quite improves CV giving 10-15 bps boost. We also used bin based auxiliary loss which gave a slight improvement.\n\nWe split the training portion of the data into 4 folds, train 4 models (40-48 epochs), fine-tune on external data + train data (5-6 epochs with x10 EX data oversampling), fine-tune on noise-free samples for 12 epochs. Then we corrected the provided train data weighting GT and PL based on the inverse error (assuming 0.15 error for PL). The provided train data has a large level of noise slowing down the convergence. Test data and the sequence ends were labeled solely based on PL. Then we train a PL model on the whole train (excluding val samples) + test data, fine-tune on EX data, and fine-tune on noise-free samples. This procedure improved CV by 10 bpp, and the weighted average of all models generated in the procedure gets further improvement.\n\n**(slime)** *Pre-training*: First we performed MLM pre-training on a whole dataset for 5 epochs (40% of input tokens are replaced with mask tokens). We used cosine decay to zero with one epoch warmup and AdamW optimizer with base_lr= 5e-4 and wd=0.05. The model was trained to predict missing nucleotides with cross-entropy loss\n\n*Fine-tuning*: During the fine-tuning stage we initialized our model with weights from MLM-pretraining, it gave a noticeable improvement to the final result [10 bps CV]. We fine-tuned our models for 50 epochs with batch_size=16 on samples where either SN(DMS) == 1 or SN(2A3) == 1, we masked the loss depending on SN of the given test [10 bps boost compared to filtering train dataset with SN(DMS) == 1 and SN(2a3) == 1). In addition, we weighted the samples based on their reactivity error provided by the ground-truth data as `loss *= torch.log(1.1 + snr) / 2`. Similar to pertaining, we used AdamW optimizer [lr=5e-4, wd=0.05] and cosine scheduler with lr decay to zero. Since in matrix mixing model masking was not considered in convolutions, this model was trained with length-matching batch sampling (samples of exactly the same length).\n\n### Best singe model end ensemble\nBest single model (dual stream model): 0.14292 at private and 0.13711 public LB, which can take top 10 itself. This result would be further improved by 10 bps to ~0.1419 if if we had time to run our full PL pipeline.\nOur final submission is a combination of ~20 models that got 0.14189 at private and 0.13604 at public LB.\n\n### Things didn't work\n- EX data was helpful at CV and public LB (5-10 bps boost), but not helpful at private LB\n- Additional bpps gave only a negligible improvement in comparison to the use of single bpp provided by organizers\n- MLM on 30M external RNA sequence dataset\n- EMA, AWP, Floyd-warshall distance matrices \n- 2D Ushape models",
    "2554198": "One of the interesting things we tried during early experiments with merging bpp and distance matrices is to incorporate Floyd-Warshall distance matrices obtained from bpp graph, the motivation was to provide the model with more clues about distances from the given nucleotide to the paired nodes.\n\nWe constructed the graph from the nucleotides in the following way:\nFor the given sequence we extract hard-edges from eterna-bpp matrix, when base pair probability exceeds 0.5, to make the graph connected, we make edges between nodes (i, i+1) as well, after that we run Floyd-Warshall algorithm on this graph.\n\nWe implemented it in C++, also we tried CUDA as well, but it gave comparable speed when running 32-threads CPU version, since graphs are relatively sparse\n\nleft plot: resulting distance matrix, right plot: hard-clipped bpp matrix + hardcoded (i, i+1) edges\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2F91af1798b6e88ce6ba77cf3d9604cab6%2Fimage%20(10).png?generation=1702073617798053&alt=media)\n\nUnfortunately, it didn't give an improvement compared to just using bpp & distance matrices with MatrixMixer layers, it seems that model learned this property by itself",
    "2553670": "Congrats on the 7th position!",
    "2558898": "Hi there! \nFirst congrats for 7th place. \n\nSeeing this amazing approach, my question is:  Is this intuition from any paper? If yes would you mind to share as i would like to read more?",
    "2555786": "DUde that's cool",
    "2554720": "Very nice details - thank you! \nCongrats on 7th position.\n\nNaive question  what tool you used to create network images?",
    "2553361": "Hi. Do you mind explaining how you weighted according to the loss?",
    "2553354": "Hi\nI do not understand the bpp, \nwould mind give me more information about the bpp.\nthank you very much. ",
    "2553319": "Congratulations on achieving 7th position in this competition. Thanks for sharing the details of your solution. "
  }
}