{
  "id": 460121,
  "title": "[1st place solution] Transformer model with Dynamic positional encoding + CNN for BPPM features",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460121",
  "author_name": "vyaltsevvaleriy",
  "post_date": "2023-12-08T00:33:45.895000",
  "votes": 147,
  "comment_count": 58,
  "views": 0,
  "content": "<p>The Ribonanza RNA Folding competition has been an amazing opportunity and we are really glad to have participated in it. Kudos to the hosts for their substantial effort in keeping the high organizational level and giving the contestants plenty of valuable insight and support. We would also like to thank the community for the fruitful public discussions and coding initiatives. Congratulations to everyone involved! 🎉</p>\n<h1>Code</h1>\n<p>Open source code is available on <a href=\"https://github.com/autosome-ru/vigg_ribonanza/\" target=\"_blank\">github</a>.</p>\n<h1>TLDR</h1>\n<p>Our solution is based on transformer architecture. We predicted base pair probability matrix(BPPM) for each RNA sequence using EternaFold and added convolutional blocks to our architecture to process BPPM features. For some models we added Squeeze-and-Excitation layer in convolutional blocks Also we add these features with attention values before softmax operation in Self-Attention block. To allow better generalization for longer input we implemented Dynamic Positional Bias. Then we ensembled models with slight differences in architecture and training process.</p>\n<h1>Data Preprocessing</h1>\n<p>At first, for each sequence in train and test datasets we calculated a base pair probability matrix (BPPM) using EternaFold. During training and inference phases we passed to our model the RNA sequences encoded with a learnable embedding layer. Each nucleotide was considered a token and special <strong>&lt; start &gt;</strong> and <strong>&lt; end &gt;</strong> tokens were placed at both ends as well. To provide the model with the information about whether the sequence comes from a “clean” subset of training dataset (which SN_filter values show - 1 corresponds to “clean” sequences with high signal-to-noise ratio and 0 is otherwise) we encoded SN_filter values with a learnable embedding layer and added corresponding  embeddings to sequence embeddings. The embedding dimension was chosen to be of size <strong>192</strong>. BPPMs were padded with zeros at their margins to account for adding <strong>&lt; start &gt;</strong> and <strong>&lt; end &gt;</strong> tokens. The figure below summarizes the data preprocessing part.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F482c2d20fce57fa513021909a515e98b%2Fpreprocessing.png?generation=1701994647204619&amp;alt=media\" alt=\"\"></p>\n<h1>Model</h1>\n<p>The model idea is partially inspired by <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> solution to the Open Vaccine challenge, <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189564\" target=\"_blank\">link</a></p>\n<p>The model takes a sequence of tokens and BPPM as input and outputs DMS_MaP and 2A3_MaP reactivities for each input nucleotide. Its architecture comprises 12 consecutive Transformer Encoder Layers and an output projection linear layer. Each Transformer block takes a sequence of tokens and BPPM features from the previous layer and outputs updated feature maps as shown below. The Transformer Encoder block adopts common transformer encoder architecture, except we modified the Self-Attention block to ensure interaction between BPPM features and sequence features.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F0c70879344b972d64ba13945e163a126%2Fmodel_general_view.png?generation=1701994737102906&amp;alt=media\" alt=\"\"></p>\n<h2>Self-Attention Block</h2>\n<p>In the Self-Attention (SA) block we implemented an attention mechanism with the following modifications: after attention values for each head are calculated, we add BPPM features updated by the ‘Convolutional block’, which outputs BPPM features with the number of channels corresponding to the number of heads in the SA block. We set the number of heads in SA and corresponding channels of BPPM features to be 6. Thus, the hidden dimension size of Q, K, V matrices is 32. The overall structure of the SA block is shown below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F5a7bf6e6f893938ced8bdd2e9a1fb08c%2Fselfattention.png?generation=1701994715778085&amp;alt=media\" alt=\"\"></p>\n<h2>Dynamic Positional Bias</h2>\n<p>The sequence length is distributed differently in train and test datasets (test sequences are generally longer). To allow for the better generalization of longer inputs we implemented a positional encoding to be added to attention values. We found Dynamic Positional Bias to be of more use compared to other relative positional encoding methods we tried, such as xPos (rotary positional encoding) and ALiBi. Dynamic Positional Bias calculates for each head a relative positional bias map, which is learnable and depends on sequence length. Relative positional bias doesn’t allow to leverage distance from start and end of sequence, so tokens <strong>&lt; start &gt;</strong> and <strong>&lt; end &gt;</strong> were added to fix that.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F8d8468983d7a7d057dcf74d7511b3ac8%2Fdynpos%20(1).png?generation=1701994757603986&amp;alt=media\" alt=\"\"></p>\n<h2>Convolutional Block and SE block</h2>\n<p>The models in the final ensemble come in two versions that are slightly different in the structure of the Convolutional block. The basic Convolutional block consists of 2D convolutional layer, batchnorm layer, activation and learnable gamma parameters that scale the output feature channels, whereas the modified version of this block (SE-Convolutional block) also contains Squeeze-and-Excitation layer, hence the name. The SE layer applies input-dependent rescaling of values along the channels as shown below. Thus, the only difference between models in the ensemble is the presence of the SE layer.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F4b8f15f5aa950ba64ecbd55df123cad4%2Fconv_se.png?generation=1701994778649467&amp;alt=media\" alt=\"\"></p>\n<h1>Training Process</h1>\n<p>The training process has been adapted from the <a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb/notebook\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/IAFOSS\" target=\"_blank\">@IAFOSS</a><br>\nWe used a one-cycle learning rate schedule (pct_start=0.05, lr_max=2.5e-3,) coupled with AdamW optimizer (wd=0.05), and batch size set to 128.<br>\nThe number of epochs for model training was determined from the dataset size.<br>\nFor final models (trained almost on the whole dataset) we used 270 epochs. <br>\nIn each epoch, 1791 batches were processed (we kept this number due to historical reasons), elements of each batch were sampled from dataset with the weight = 0.5 * torch.clamp_min(torch.log(sn + 1.01),0.01)</p>\n<p>We have also found that training model with a simple SGD optimizer for ~15 epochs of 500 batches each additionally improves model performance (the number of epochs varies so we used a small validation set to determine the exact number of epochs)</p>\n<p>Additionally, one model was trained to predict 2A3 given DMS, RNA sequence and BPPM (dms-to-2a3 model).</p>\n<h1>Inference</h1>\n<p>At the inference phase for all inputs we set the SN filter value to be 1 as if they come from a “clean” dataset.</p>\n<h1>Ensembling</h1>\n<p>We ensembled 15 models with SE-Convolutional block, 10 models with plain Convolutional block, 2 with plain Convolutional block trained on split by sequence lengths (one of these two models accepts bracket features).  We took the average of their predictions. Then we predicted 2a3 reactivities based on averaged dms reactivities using dms-to-2a3 model and added these predictions to averaged 2a3 reactivities in the following way: (27/28)<em>averaged_2a3 + (1/28)</em>predicted_2a3.</p>\n<h1>Other splits</h1>\n<p>The clear problem with a simple KFold split is the high sequence similarity across the training dataset and the fact that test sequences are very distinct from the training data. This might hamper the model development, because the increase in quality on the validation set could be due to overfitting rather than actual improvement.</p>\n<p>We have calculated a hamming distance matrix for all the sequences present in the training set and performed a modified DBSCAN clustering procedure with distance threshold set at 0.2. In the following picture we show the clusters mapped to their respective cluster identifier (a cluster ID was assigned as a number of the smallest sequence within that cluster in the train dataset).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F55dc21823386d878fb212d681cae593e%2Fcluster.jpg?generation=1701994798203967&amp;alt=media\" alt=\"\"></p>\n<p>We tested several splits based on sequence identity (the easiest one is just to split data into folds without shuffling) and found out that while the model final validation performance degrades as we choose more and more stricter distance threshold, its relative value behave the same way as for a simple KFold split. So, we decided to use simple KFold split than training final models </p>\n<p>Also, we have conducted a test of the model performance on length-based split. For that, we trained a model on the short sequences (length &lt; 206) and validated it on sequences of size = 206. The behavior of the validation metric was slightly noisier but still highly correlated with a metric for a simple KFold split.</p>\n<h1>Public data leakage</h1>\n<p>Approximately 13% of public test sequences are identical to the ones present in the train dataset (by sequence). To avoid selecting a model that is memorizing more of these sequences rather than learning RNA-related stuff, we zeroed out the predictions for these sequences in most of our submissions, while sometimes sending non-zeroed out submissions in order to compare our performance to other participants.</p>\n<h1>Features we also tried</h1>\n<h2>capR</h2>\n<p>We don’t have conclusive results for that feature. It seems like it is not beneficial for the model on average but sometimes it resulted in a better model and sometimes – in worse. We decided to not use this feature.</p>\n<h2>Brackets</h2>\n<p>We have tried to use brackets generated by EternaFold, ContraFold, ViennaRNA, etc., as well as programs for pseudoknots prediction (like IPknot). However, after adding EternaFold BPPMs to the model, adding other features yields no significant increase in model performance. For some models we used brackets just to augment model</p>\n<h2>Different BPPMs</h2>\n<p>All other BPPMs (ContraFold, ViennaRNA, RNAsoft, RNAstructure) result in a suboptimal model.<br>\nAveraging BPPMs doesn’t result in better performance.<br>\nBPPMs from RFold yield the same quality as ViennaRNA BPPMs.<br>\nThe RNA-FM model produces both per-nucleotide embeddings and BPPM-like matrix. Still, those don’t help the model at all, and using them results only in a slight increase in model performance then compared to sequence only model</p>\n<h2>SQUARNA matrix</h2>\n<p>SQUARNA (<a href=\"https://github.com/febos/SQUARNA\" target=\"_blank\">github</a>) outputs a matrix different from BPPM, but can be used in the same manner. Unfortunately, this feature also didn’t yield any additional performance increase.</p>\n<h1>What we also tried</h1>\n<h2>Fully-Convolutional architecture</h2>\n<p>The initial reason we decided to take part in the competition was to test our model LegNet <a href=\"https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784\" target=\"_blank\">link</a>, which shows SOTA results on DNA sequences and worked well on some RNA-related tasks (not yet published).<br>\nUnfortunately, any modifications of this architecture resulted in subpar performance when compared with properly tuned transformer models. This can be explained by the fact that predicting RNA secondary structure requires attending to long-range contacts. The transformer architecture suits better for such cases.</p>\n<h2>Subsetting data</h2>\n<p>Subsetting data (filtering by different thresholds on SN ratio) resulted in a performance boost for all models. However, this technique was superseded by weight sampling, which proved itself to be more effective.</p>\n<h2>Fine-tuning on public datasets</h2>\n<p>We tried to fine-tune our model on a public dataset, gathered by the organizers (<a href=\"https://www.kaggle.com/datasets/rhijudas/rmdb-rna-mapping-database-2023-data/data\" target=\"_blank\">link</a>) by training model to predict the results of the public experiments (excluding ones with a small number of samples) and for Ribonanza data simultaneously. Unfortunately, this also gave no boost to the model performance.</p>\n<h2>Using 3D data</h2>\n<p>We tried to use data about predicted 3D structure of 100k sequences from the train dataset but gave up on that once we had visually analyzed them:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F2d054d1679176e0664130f6c46ccb0ae%2Ftrashfold.jpg?generation=1701994817511350&amp;alt=media\" alt=\"\"></p>\n<h2>Absolute positional embedding</h2>\n<p>Using absolute positional embedding leads to unsolvable issues when generalizing upon longer sequences.</p>\n<h2>Relative positional embedding</h2>\n<p>The simplest approach is to augment absolute positional encoding during the training phase to shift randomly from 0 to (Lmax - seqlen) position. This indeed solves the issue with extrapolation, but works worse than other methods.<br>\nRotational positional embedding, unfortunately, doesn’t help the model to generalize on larger lengths.<br>\nALiBi positional embedding solves the issue with extrapolation but even after keeping it only for a part of heads (as suggested in <a href=\"https://github.com/lucidrains/x-transformers\" target=\"_blank\">https://github.com/lucidrains/x-transformers</a>) still behaves worse than dynamic positional bias.</p>\n<h2>Augmentation</h2>\n<p>First, we tried to use reverse augmentation. This can be done in three ways:<br>\nreversing sequence before any modifications,<br>\nreversing sequence before padding, but after adding  and  tokens,<br>\nreversing sequence after padding.</p>\n<p>The first two ways yield no gain for all variants of models we tested. Yet, the third one (upon coupling with additional finetuning) gave us a good result for a model with xpos positional encoding. The resulting single-model performance was 0.13937 on the public leaderboard. Unfortunately, xpos shows rather poor performance on long sequences so we abstained from using this model in the final submission.</p>\n<p>We have also tried shift augmentation and different sequence padding approaches. This didn’t improve our model performance as well.</p>\n<h2>Sliding window</h2>\n<p>One of the possible ways to generalize for larger sequences is to predict reactivities using the sliding window. However, the idea is somewhat wrong in a biological sense, and it results in performance degradation when testing on train dataset sequences.</p>\n<h2>Pseudolabelling</h2>\n<p>Once we obtained ensembles of best-performing models we tried to use them to pseudolabel test dataset and use predictions with the highest confidence to train new models. While this indeed results in a better single-performing model, adding such a model to ensemble doesn’t improve ensemble performance.</p>\n<h2>Changing loss</h2>\n<p>Instead of filtering sequences with low SN we tried to mask positions with high reactivity error as it was done by <a href=\"https://www.kaggle.com/nullrecurrent\" target=\"_blank\">@nullrecurrent</a> in Open Vaccine challenge (<a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620\" target=\"_blank\">link</a>)<br>\nThis resulted in poor performance.<br>\nWe tried to weight loss for each sequence by its SN – it didn’t result in any improvement</p>\n<h1>Tools</h1>\n<p><a href=\"https://github.com/DasLab/arnie/tree/master\" target=\"_blank\">arnie</a><br>\n<a href=\"https://github.com/eternagame/eternafold\" target=\"_blank\">EternaFold</a></p>\n<h1>Links</h1>\n<h2>Squeeze-and-excitation block</h2>\n<p>Hu, J., Shen, L., &amp; Sun, G. (2018). Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7132-7141).</p>\n<h2>Dynamic positional bias</h2>\n<p>Wang, W., Chen, W., Qiu, Q., Chen, L., Wu, B., Lin, B., … &amp; Liu, W. (2023). Crossformer++: A versatile vision transformer hinging on cross-scale attention. arXiv preprint arXiv:2303.06908.</p>",
  "messages": [
    {
      "id": 2553003,
      "postDate": "2023-12-08T00:33:45.897Z",
      "content": "<p>The Ribonanza RNA Folding competition has been an amazing opportunity and we are really glad to have participated in it. Kudos to the hosts for their substantial effort in keeping the high organizational level and giving the contestants plenty of valuable insight and support. We would also like to thank the community for the fruitful public discussions and coding initiatives. Congratulations to everyone involved! 🎉</p>\n<h1>Code</h1>\n<p>Open source code is available on <a href=\"https://github.com/autosome-ru/vigg_ribonanza/\" target=\"_blank\">github</a>.</p>\n<h1>TLDR</h1>\n<p>Our solution is based on transformer architecture. We predicted base pair probability matrix(BPPM) for each RNA sequence using EternaFold and added convolutional blocks to our architecture to process BPPM features. For some models we added Squeeze-and-Excitation layer in convolutional blocks Also we add these features with attention values before softmax operation in Self-Attention block. To allow better generalization for longer input we implemented Dynamic Positional Bias. Then we ensembled models with slight differences in architecture and training process.</p>\n<h1>Data Preprocessing</h1>\n<p>At first, for each sequence in train and test datasets we calculated a base pair probability matrix (BPPM) using EternaFold. During training and inference phases we passed to our model the RNA sequences encoded with a learnable embedding layer. Each nucleotide was considered a token and special <strong>&lt; start &gt;</strong> and <strong>&lt; end &gt;</strong> tokens were placed at both ends as well. To provide the model with the information about whether the sequence comes from a “clean” subset of training dataset (which SN_filter values show - 1 corresponds to “clean” sequences with high signal-to-noise ratio and 0 is otherwise) we encoded SN_filter values with a learnable embedding layer and added corresponding  embeddings to sequence embeddings. The embedding dimension was chosen to be of size <strong>192</strong>. BPPMs were padded with zeros at their margins to account for adding <strong>&lt; start &gt;</strong> and <strong>&lt; end &gt;</strong> tokens. The figure below summarizes the data preprocessing part.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F482c2d20fce57fa513021909a515e98b%2Fpreprocessing.png?generation=1701994647204619&amp;alt=media\" alt=\"\"></p>\n<h1>Model</h1>\n<p>The model idea is partially inspired by <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> solution to the Open Vaccine challenge, <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189564\" target=\"_blank\">link</a></p>\n<p>The model takes a sequence of tokens and BPPM as input and outputs DMS_MaP and 2A3_MaP reactivities for each input nucleotide. Its architecture comprises 12 consecutive Transformer Encoder Layers and an output projection linear layer. Each Transformer block takes a sequence of tokens and BPPM features from the previous layer and outputs updated feature maps as shown below. The Transformer Encoder block adopts common transformer encoder architecture, except we modified the Self-Attention block to ensure interaction between BPPM features and sequence features.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F0c70879344b972d64ba13945e163a126%2Fmodel_general_view.png?generation=1701994737102906&amp;alt=media\" alt=\"\"></p>\n<h2>Self-Attention Block</h2>\n<p>In the Self-Attention (SA) block we implemented an attention mechanism with the following modifications: after attention values for each head are calculated, we add BPPM features updated by the ‘Convolutional block’, which outputs BPPM features with the number of channels corresponding to the number of heads in the SA block. We set the number of heads in SA and corresponding channels of BPPM features to be 6. Thus, the hidden dimension size of Q, K, V matrices is 32. The overall structure of the SA block is shown below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F5a7bf6e6f893938ced8bdd2e9a1fb08c%2Fselfattention.png?generation=1701994715778085&amp;alt=media\" alt=\"\"></p>\n<h2>Dynamic Positional Bias</h2>\n<p>The sequence length is distributed differently in train and test datasets (test sequences are generally longer). To allow for the better generalization of longer inputs we implemented a positional encoding to be added to attention values. We found Dynamic Positional Bias to be of more use compared to other relative positional encoding methods we tried, such as xPos (rotary positional encoding) and ALiBi. Dynamic Positional Bias calculates for each head a relative positional bias map, which is learnable and depends on sequence length. Relative positional bias doesn’t allow to leverage distance from start and end of sequence, so tokens <strong>&lt; start &gt;</strong> and <strong>&lt; end &gt;</strong> were added to fix that.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F8d8468983d7a7d057dcf74d7511b3ac8%2Fdynpos%20(1).png?generation=1701994757603986&amp;alt=media\" alt=\"\"></p>\n<h2>Convolutional Block and SE block</h2>\n<p>The models in the final ensemble come in two versions that are slightly different in the structure of the Convolutional block. The basic Convolutional block consists of 2D convolutional layer, batchnorm layer, activation and learnable gamma parameters that scale the output feature channels, whereas the modified version of this block (SE-Convolutional block) also contains Squeeze-and-Excitation layer, hence the name. The SE layer applies input-dependent rescaling of values along the channels as shown below. Thus, the only difference between models in the ensemble is the presence of the SE layer.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F4b8f15f5aa950ba64ecbd55df123cad4%2Fconv_se.png?generation=1701994778649467&amp;alt=media\" alt=\"\"></p>\n<h1>Training Process</h1>\n<p>The training process has been adapted from the <a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb/notebook\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/IAFOSS\" target=\"_blank\">@IAFOSS</a><br>\nWe used a one-cycle learning rate schedule (pct_start=0.05, lr_max=2.5e-3,) coupled with AdamW optimizer (wd=0.05), and batch size set to 128.<br>\nThe number of epochs for model training was determined from the dataset size.<br>\nFor final models (trained almost on the whole dataset) we used 270 epochs. <br>\nIn each epoch, 1791 batches were processed (we kept this number due to historical reasons), elements of each batch were sampled from dataset with the weight = 0.5 * torch.clamp_min(torch.log(sn + 1.01),0.01)</p>\n<p>We have also found that training model with a simple SGD optimizer for ~15 epochs of 500 batches each additionally improves model performance (the number of epochs varies so we used a small validation set to determine the exact number of epochs)</p>\n<p>Additionally, one model was trained to predict 2A3 given DMS, RNA sequence and BPPM (dms-to-2a3 model).</p>\n<h1>Inference</h1>\n<p>At the inference phase for all inputs we set the SN filter value to be 1 as if they come from a “clean” dataset.</p>\n<h1>Ensembling</h1>\n<p>We ensembled 15 models with SE-Convolutional block, 10 models with plain Convolutional block, 2 with plain Convolutional block trained on split by sequence lengths (one of these two models accepts bracket features).  We took the average of their predictions. Then we predicted 2a3 reactivities based on averaged dms reactivities using dms-to-2a3 model and added these predictions to averaged 2a3 reactivities in the following way: (27/28)<em>averaged_2a3 + (1/28)</em>predicted_2a3.</p>\n<h1>Other splits</h1>\n<p>The clear problem with a simple KFold split is the high sequence similarity across the training dataset and the fact that test sequences are very distinct from the training data. This might hamper the model development, because the increase in quality on the validation set could be due to overfitting rather than actual improvement.</p>\n<p>We have calculated a hamming distance matrix for all the sequences present in the training set and performed a modified DBSCAN clustering procedure with distance threshold set at 0.2. In the following picture we show the clusters mapped to their respective cluster identifier (a cluster ID was assigned as a number of the smallest sequence within that cluster in the train dataset).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F55dc21823386d878fb212d681cae593e%2Fcluster.jpg?generation=1701994798203967&amp;alt=media\" alt=\"\"></p>\n<p>We tested several splits based on sequence identity (the easiest one is just to split data into folds without shuffling) and found out that while the model final validation performance degrades as we choose more and more stricter distance threshold, its relative value behave the same way as for a simple KFold split. So, we decided to use simple KFold split than training final models </p>\n<p>Also, we have conducted a test of the model performance on length-based split. For that, we trained a model on the short sequences (length &lt; 206) and validated it on sequences of size = 206. The behavior of the validation metric was slightly noisier but still highly correlated with a metric for a simple KFold split.</p>\n<h1>Public data leakage</h1>\n<p>Approximately 13% of public test sequences are identical to the ones present in the train dataset (by sequence). To avoid selecting a model that is memorizing more of these sequences rather than learning RNA-related stuff, we zeroed out the predictions for these sequences in most of our submissions, while sometimes sending non-zeroed out submissions in order to compare our performance to other participants.</p>\n<h1>Features we also tried</h1>\n<h2>capR</h2>\n<p>We don’t have conclusive results for that feature. It seems like it is not beneficial for the model on average but sometimes it resulted in a better model and sometimes – in worse. We decided to not use this feature.</p>\n<h2>Brackets</h2>\n<p>We have tried to use brackets generated by EternaFold, ContraFold, ViennaRNA, etc., as well as programs for pseudoknots prediction (like IPknot). However, after adding EternaFold BPPMs to the model, adding other features yields no significant increase in model performance. For some models we used brackets just to augment model</p>\n<h2>Different BPPMs</h2>\n<p>All other BPPMs (ContraFold, ViennaRNA, RNAsoft, RNAstructure) result in a suboptimal model.<br>\nAveraging BPPMs doesn’t result in better performance.<br>\nBPPMs from RFold yield the same quality as ViennaRNA BPPMs.<br>\nThe RNA-FM model produces both per-nucleotide embeddings and BPPM-like matrix. Still, those don’t help the model at all, and using them results only in a slight increase in model performance then compared to sequence only model</p>\n<h2>SQUARNA matrix</h2>\n<p>SQUARNA (<a href=\"https://github.com/febos/SQUARNA\" target=\"_blank\">github</a>) outputs a matrix different from BPPM, but can be used in the same manner. Unfortunately, this feature also didn’t yield any additional performance increase.</p>\n<h1>What we also tried</h1>\n<h2>Fully-Convolutional architecture</h2>\n<p>The initial reason we decided to take part in the competition was to test our model LegNet <a href=\"https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784\" target=\"_blank\">link</a>, which shows SOTA results on DNA sequences and worked well on some RNA-related tasks (not yet published).<br>\nUnfortunately, any modifications of this architecture resulted in subpar performance when compared with properly tuned transformer models. This can be explained by the fact that predicting RNA secondary structure requires attending to long-range contacts. The transformer architecture suits better for such cases.</p>\n<h2>Subsetting data</h2>\n<p>Subsetting data (filtering by different thresholds on SN ratio) resulted in a performance boost for all models. However, this technique was superseded by weight sampling, which proved itself to be more effective.</p>\n<h2>Fine-tuning on public datasets</h2>\n<p>We tried to fine-tune our model on a public dataset, gathered by the organizers (<a href=\"https://www.kaggle.com/datasets/rhijudas/rmdb-rna-mapping-database-2023-data/data\" target=\"_blank\">link</a>) by training model to predict the results of the public experiments (excluding ones with a small number of samples) and for Ribonanza data simultaneously. Unfortunately, this also gave no boost to the model performance.</p>\n<h2>Using 3D data</h2>\n<p>We tried to use data about predicted 3D structure of 100k sequences from the train dataset but gave up on that once we had visually analyzed them:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F2d054d1679176e0664130f6c46ccb0ae%2Ftrashfold.jpg?generation=1701994817511350&amp;alt=media\" alt=\"\"></p>\n<h2>Absolute positional embedding</h2>\n<p>Using absolute positional embedding leads to unsolvable issues when generalizing upon longer sequences.</p>\n<h2>Relative positional embedding</h2>\n<p>The simplest approach is to augment absolute positional encoding during the training phase to shift randomly from 0 to (Lmax - seqlen) position. This indeed solves the issue with extrapolation, but works worse than other methods.<br>\nRotational positional embedding, unfortunately, doesn’t help the model to generalize on larger lengths.<br>\nALiBi positional embedding solves the issue with extrapolation but even after keeping it only for a part of heads (as suggested in <a href=\"https://github.com/lucidrains/x-transformers\" target=\"_blank\">https://github.com/lucidrains/x-transformers</a>) still behaves worse than dynamic positional bias.</p>\n<h2>Augmentation</h2>\n<p>First, we tried to use reverse augmentation. This can be done in three ways:<br>\nreversing sequence before any modifications,<br>\nreversing sequence before padding, but after adding  and  tokens,<br>\nreversing sequence after padding.</p>\n<p>The first two ways yield no gain for all variants of models we tested. Yet, the third one (upon coupling with additional finetuning) gave us a good result for a model with xpos positional encoding. The resulting single-model performance was 0.13937 on the public leaderboard. Unfortunately, xpos shows rather poor performance on long sequences so we abstained from using this model in the final submission.</p>\n<p>We have also tried shift augmentation and different sequence padding approaches. This didn’t improve our model performance as well.</p>\n<h2>Sliding window</h2>\n<p>One of the possible ways to generalize for larger sequences is to predict reactivities using the sliding window. However, the idea is somewhat wrong in a biological sense, and it results in performance degradation when testing on train dataset sequences.</p>\n<h2>Pseudolabelling</h2>\n<p>Once we obtained ensembles of best-performing models we tried to use them to pseudolabel test dataset and use predictions with the highest confidence to train new models. While this indeed results in a better single-performing model, adding such a model to ensemble doesn’t improve ensemble performance.</p>\n<h2>Changing loss</h2>\n<p>Instead of filtering sequences with low SN we tried to mask positions with high reactivity error as it was done by <a href=\"https://www.kaggle.com/nullrecurrent\" target=\"_blank\">@nullrecurrent</a> in Open Vaccine challenge (<a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620\" target=\"_blank\">link</a>)<br>\nThis resulted in poor performance.<br>\nWe tried to weight loss for each sequence by its SN – it didn’t result in any improvement</p>\n<h1>Tools</h1>\n<p><a href=\"https://github.com/DasLab/arnie/tree/master\" target=\"_blank\">arnie</a><br>\n<a href=\"https://github.com/eternagame/eternafold\" target=\"_blank\">EternaFold</a></p>\n<h1>Links</h1>\n<h2>Squeeze-and-excitation block</h2>\n<p>Hu, J., Shen, L., &amp; Sun, G. (2018). Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7132-7141).</p>\n<h2>Dynamic positional bias</h2>\n<p>Wang, W., Chen, W., Qiu, Q., Chen, L., Wu, B., Lin, B., … &amp; Liu, W. (2023). Crossformer++: A versatile vision transformer hinging on cross-scale attention. arXiv preprint arXiv:2303.06908.</p>",
      "rawMarkdown": "The Ribonanza RNA Folding competition has been an amazing opportunity and we are really glad to have participated in it. Kudos to the hosts for their substantial effort in keeping the high organizational level and giving the contestants plenty of valuable insight and support. We would also like to thank the community for the fruitful public discussions and coding initiatives. Congratulations to everyone involved! 🎉\n\n# Code\nOpen source code is available on [github](https://github.com/autosome-ru/vigg_ribonanza/).\n\n# TLDR\n\nOur solution is based on transformer architecture. We predicted base pair probability matrix(BPPM) for each RNA sequence using EternaFold and added convolutional blocks to our architecture to process BPPM features. For some models we added Squeeze-and-Excitation layer in convolutional blocks Also we add these features with attention values before softmax operation in Self-Attention block. To allow better generalization for longer input we implemented Dynamic Positional Bias. Then we ensembled models with slight differences in architecture and training process.\n\n# Data Preprocessing\n\nAt first, for each sequence in train and test datasets we calculated a base pair probability matrix (BPPM) using EternaFold. During training and inference phases we passed to our model the RNA sequences encoded with a learnable embedding layer. Each nucleotide was considered a token and special **< start >** and **< end >** tokens were placed at both ends as well. To provide the model with the information about whether the sequence comes from a “clean” subset of training dataset (which SN_filter values show - 1 corresponds to “clean” sequences with high signal-to-noise ratio and 0 is otherwise) we encoded SN_filter values with a learnable embedding layer and added corresponding  embeddings to sequence embeddings. The embedding dimension was chosen to be of size **192**. BPPMs were padded with zeros at their margins to account for adding **< start >** and **< end >** tokens. The figure below summarizes the data preprocessing part.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F482c2d20fce57fa513021909a515e98b%2Fpreprocessing.png?generation=1701994647204619&alt=media)\n\n# Model\n\nThe model idea is partially inspired by @shujun717 solution to the Open Vaccine challenge, [link](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189564)\n\nThe model takes a sequence of tokens and BPPM as input and outputs DMS_MaP and 2A3_MaP reactivities for each input nucleotide. Its architecture comprises 12 consecutive Transformer Encoder Layers and an output projection linear layer. Each Transformer block takes a sequence of tokens and BPPM features from the previous layer and outputs updated feature maps as shown below. The Transformer Encoder block adopts common transformer encoder architecture, except we modified the Self-Attention block to ensure interaction between BPPM features and sequence features.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F0c70879344b972d64ba13945e163a126%2Fmodel_general_view.png?generation=1701994737102906&alt=media)\n\n## Self-Attention Block\n\nIn the Self-Attention (SA) block we implemented an attention mechanism with the following modifications: after attention values for each head are calculated, we add BPPM features updated by the ‘Convolutional block’, which outputs BPPM features with the number of channels corresponding to the number of heads in the SA block. We set the number of heads in SA and corresponding channels of BPPM features to be 6. Thus, the hidden dimension size of Q, K, V matrices is 32. The overall structure of the SA block is shown below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F5a7bf6e6f893938ced8bdd2e9a1fb08c%2Fselfattention.png?generation=1701994715778085&alt=media)\n\n## Dynamic Positional Bias\n\nThe sequence length is distributed differently in train and test datasets (test sequences are generally longer). To allow for the better generalization of longer inputs we implemented a positional encoding to be added to attention values. We found Dynamic Positional Bias to be of more use compared to other relative positional encoding methods we tried, such as xPos (rotary positional encoding) and ALiBi. Dynamic Positional Bias calculates for each head a relative positional bias map, which is learnable and depends on sequence length. Relative positional bias doesn’t allow to leverage distance from start and end of sequence, so tokens **< start >** and **< end >** were added to fix that.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F8d8468983d7a7d057dcf74d7511b3ac8%2Fdynpos%20(1).png?generation=1701994757603986&alt=media)\n\n## Convolutional Block and SE block\n\nThe models in the final ensemble come in two versions that are slightly different in the structure of the Convolutional block. The basic Convolutional block consists of 2D convolutional layer, batchnorm layer, activation and learnable gamma parameters that scale the output feature channels, whereas the modified version of this block (SE-Convolutional block) also contains Squeeze-and-Excitation layer, hence the name. The SE layer applies input-dependent rescaling of values along the channels as shown below. Thus, the only difference between models in the ensemble is the presence of the SE layer.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F4b8f15f5aa950ba64ecbd55df123cad4%2Fconv_se.png?generation=1701994778649467&alt=media)\n\n# Training Process\n\nThe training process has been adapted from the [notebook](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb/notebook) by @IAFOSS\nWe used a one-cycle learning rate schedule (pct_start=0.05, lr_max=2.5e-3,) coupled with AdamW optimizer (wd=0.05), and batch size set to 128.\nThe number of epochs for model training was determined from the dataset size.\nFor final models (trained almost on the whole dataset) we used 270 epochs. \nIn each epoch, 1791 batches were processed (we kept this number due to historical reasons), elements of each batch were sampled from dataset with the weight = 0.5 * torch.clamp_min(torch.log(sn + 1.01),0.01)\n\nWe have also found that training model with a simple SGD optimizer for ~15 epochs of 500 batches each additionally improves model performance (the number of epochs varies so we used a small validation set to determine the exact number of epochs)\n\nAdditionally, one model was trained to predict 2A3 given DMS, RNA sequence and BPPM (dms-to-2a3 model).\n\n# Inference\nAt the inference phase for all inputs we set the SN filter value to be 1 as if they come from a “clean” dataset.\n\n# Ensembling\n\nWe ensembled 15 models with SE-Convolutional block, 10 models with plain Convolutional block, 2 with plain Convolutional block trained on split by sequence lengths (one of these two models accepts bracket features).  We took the average of their predictions. Then we predicted 2a3 reactivities based on averaged dms reactivities using dms-to-2a3 model and added these predictions to averaged 2a3 reactivities in the following way: (27/28)*averaged_2a3 + (1/28)*predicted_2a3.\n\n# Other splits \n\nThe clear problem with a simple KFold split is the high sequence similarity across the training dataset and the fact that test sequences are very distinct from the training data. This might hamper the model development, because the increase in quality on the validation set could be due to overfitting rather than actual improvement.\n\nWe have calculated a hamming distance matrix for all the sequences present in the training set and performed a modified DBSCAN clustering procedure with distance threshold set at 0.2. In the following picture we show the clusters mapped to their respective cluster identifier (a cluster ID was assigned as a number of the smallest sequence within that cluster in the train dataset).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F55dc21823386d878fb212d681cae593e%2Fcluster.jpg?generation=1701994798203967&alt=media)\n\nWe tested several splits based on sequence identity (the easiest one is just to split data into folds without shuffling) and found out that while the model final validation performance degrades as we choose more and more stricter distance threshold, its relative value behave the same way as for a simple KFold split. So, we decided to use simple KFold split than training final models \n\nAlso, we have conducted a test of the model performance on length-based split. For that, we trained a model on the short sequences (length < 206) and validated it on sequences of size = 206. The behavior of the validation metric was slightly noisier but still highly correlated with a metric for a simple KFold split.\n\n# Public data leakage \n\nApproximately 13% of public test sequences are identical to the ones present in the train dataset (by sequence). To avoid selecting a model that is memorizing more of these sequences rather than learning RNA-related stuff, we zeroed out the predictions for these sequences in most of our submissions, while sometimes sending non-zeroed out submissions in order to compare our performance to other participants.\n\n# Features we also tried\n\n## capR\n\nWe don’t have conclusive results for that feature. It seems like it is not beneficial for the model on average but sometimes it resulted in a better model and sometimes – in worse. We decided to not use this feature.\n\n## Brackets\n\nWe have tried to use brackets generated by EternaFold, ContraFold, ViennaRNA, etc., as well as programs for pseudoknots prediction (like IPknot). However, after adding EternaFold BPPMs to the model, adding other features yields no significant increase in model performance. For some models we used brackets just to augment model\n\n## Different BPPMs\n\nAll other BPPMs (ContraFold, ViennaRNA, RNAsoft, RNAstructure) result in a suboptimal model.\nAveraging BPPMs doesn’t result in better performance.\nBPPMs from RFold yield the same quality as ViennaRNA BPPMs.\nThe RNA-FM model produces both per-nucleotide embeddings and BPPM-like matrix. Still, those don’t help the model at all, and using them results only in a slight increase in model performance then compared to sequence only model\n \n## SQUARNA matrix\n\nSQUARNA ([github](https://github.com/febos/SQUARNA)) outputs a matrix different from BPPM, but can be used in the same manner. Unfortunately, this feature also didn’t yield any additional performance increase.\n\n# What we also tried\n\n## Fully-Convolutional architecture\n\nThe initial reason we decided to take part in the competition was to test our model LegNet [link](https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784), which shows SOTA results on DNA sequences and worked well on some RNA-related tasks (not yet published).\nUnfortunately, any modifications of this architecture resulted in subpar performance when compared with properly tuned transformer models. This can be explained by the fact that predicting RNA secondary structure requires attending to long-range contacts. The transformer architecture suits better for such cases.\n\n## Subsetting data \n\nSubsetting data (filtering by different thresholds on SN ratio) resulted in a performance boost for all models. However, this technique was superseded by weight sampling, which proved itself to be more effective.\n\n## Fine-tuning on public datasets \n\nWe tried to fine-tune our model on a public dataset, gathered by the organizers ([link](https://www.kaggle.com/datasets/rhijudas/rmdb-rna-mapping-database-2023-data/data)) by training model to predict the results of the public experiments (excluding ones with a small number of samples) and for Ribonanza data simultaneously. Unfortunately, this also gave no boost to the model performance.\n\n## Using 3D data \n\nWe tried to use data about predicted 3D structure of 100k sequences from the train dataset but gave up on that once we had visually analyzed them:\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F2d054d1679176e0664130f6c46ccb0ae%2Ftrashfold.jpg?generation=1701994817511350&alt=media)\n\n## Absolute positional embedding\n\nUsing absolute positional embedding leads to unsolvable issues when generalizing upon longer sequences.\n\n## Relative positional embedding \n\nThe simplest approach is to augment absolute positional encoding during the training phase to shift randomly from 0 to (Lmax - seqlen) position. This indeed solves the issue with extrapolation, but works worse than other methods.\nRotational positional embedding, unfortunately, doesn’t help the model to generalize on larger lengths.\nALiBi positional embedding solves the issue with extrapolation but even after keeping it only for a part of heads (as suggested in https://github.com/lucidrains/x-transformers) still behaves worse than dynamic positional bias.\n\n## Augmentation \n\nFirst, we tried to use reverse augmentation. This can be done in three ways:\nreversing sequence before any modifications,\nreversing sequence before padding, but after adding <start> and <end> tokens,\nreversing sequence after padding.\n\nThe first two ways yield no gain for all variants of models we tested. Yet, the third one (upon coupling with additional finetuning) gave us a good result for a model with xpos positional encoding. The resulting single-model performance was 0.13937 on the public leaderboard. Unfortunately, xpos shows rather poor performance on long sequences so we abstained from using this model in the final submission.\n\nWe have also tried shift augmentation and different sequence padding approaches. This didn’t improve our model performance as well.\n\n## Sliding window\n\nOne of the possible ways to generalize for larger sequences is to predict reactivities using the sliding window. However, the idea is somewhat wrong in a biological sense, and it results in performance degradation when testing on train dataset sequences.\n\n## Pseudolabelling\n\nOnce we obtained ensembles of best-performing models we tried to use them to pseudolabel test dataset and use predictions with the highest confidence to train new models. While this indeed results in a better single-performing model, adding such a model to ensemble doesn’t improve ensemble performance.\n\n## Changing loss\n\nInstead of filtering sequences with low SN we tried to mask positions with high reactivity error as it was done by @nullrecurrent in Open Vaccine challenge ([link](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620))\nThis resulted in poor performance.\nWe tried to weight loss for each sequence by its SN – it didn’t result in any improvement\n\n# Tools\n\n[arnie](https://github.com/DasLab/arnie/tree/master)\n[EternaFold](https://github.com/eternagame/eternafold)\n\n# Links\n\n## Squeeze-and-excitation block\nHu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7132-7141).\n\n## Dynamic positional bias\nWang, W., Chen, W., Qiu, Q., Chen, L., Wu, B., Lin, B., ... & Liu, W. (2023). Crossformer++: A versatile vision transformer hinging on cross-scale attention. arXiv preprint arXiv:2303.06908.\n\n\n",
      "votes": 147
    },
    {
      "id": 2553074,
      "postDate": "2023-12-08T01:56:43.433Z",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/vyaltsevvaleriy\" target=\"_blank\">@vyaltsevvaleriy</a> 🎉🎉🎉 This is a very, very great method. May I ask what software was used to draw this picture? This is so beautiful</p>",
      "rawMarkdown": "Congratulations! @vyaltsevvaleriy 🎉🎉🎉 This is a very, very great method. May I ask what software was used to draw this picture? This is so beautiful",
      "votes": 9,
      "replies": [
        {
          "id": 2553451,
          "postDate": "2023-12-08T08:56:55.440Z",
          "content": "<p>Haha, I also want to know</p>",
          "rawMarkdown": "Haha, I also want to know",
          "votes": 2
        },
        {
          "id": 2553696,
          "postDate": "2023-12-08T12:56:38.877Z",
          "content": "<p>Hi! Thank you for your appreciation! I used Inkscape - it is a vector graphics editor allowing to flexibly implement any ideas.</p>",
          "rawMarkdown": "Hi! Thank you for your appreciation! I used Inkscape - it is a vector graphics editor allowing to flexibly implement any ideas.",
          "votes": 10,
          "replies": [
            {
              "id": 2553829,
              "postDate": "2023-12-08T14:32:42.767Z",
              "content": "<p>Thank you very much! vyaltsevvaleriy🥰</p>",
              "rawMarkdown": "Thank you very much! vyaltsevvaleriy🥰",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2553004,
      "postDate": "2023-12-08T00:36:09.530Z",
      "content": "<p>Wonderful work and presentation! Congratulations! Would you mind providing more improvement details (improvement points) with each operation?</p>",
      "rawMarkdown": "Wonderful work and presentation! Congratulations! Would you mind providing more improvement details (improvement points) with each operation?",
      "votes": 5,
      "replies": [
        {
          "id": 2553022,
          "postDate": "2023-12-08T00:52:28.640Z",
          "content": "<p>Yes, sure, I hope we will be able to explain this observation - <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/458478\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/458478</a></p>",
          "rawMarkdown": "Yes, sure, I hope we will be able to explain this observation - https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/458478",
          "votes": 4
        }
      ]
    },
    {
      "id": 2553008,
      "postDate": "2023-12-08T00:39:09.843Z",
      "content": "<p>Nice! Actually my model has also been published here: <a href=\"https://academic.oup.com/bib/article/24/1/bbac581/6986359\" target=\"_blank\">https://academic.oup.com/bib/article/24/1/bbac581/6986359</a>. Congrats!</p>",
      "rawMarkdown": "Nice! Actually my model has also been published here: https://academic.oup.com/bib/article/24/1/bbac581/6986359. Congrats!",
      "votes": 6,
      "replies": [
        {
          "id": 2553023,
          "postDate": "2023-12-08T00:53:19.323Z",
          "content": "<p>We failed to google your publication initially, thanks!) </p>",
          "rawMarkdown": "We failed to google your publication initially, thanks!) ",
          "votes": 1,
          "replies": [
            {
              "id": 2553026,
              "postDate": "2023-12-08T00:58:28.277Z",
              "content": "<p>No problem! I actually didn't share my model because we realized models using bpp are a bit conservative on pseudo-knots and often miss them, so I didn't want to influence competitors to use bpp by sharing my paper, but I guess it ended up happening anyway haha</p>",
              "rawMarkdown": "No problem! I actually didn't share my model because we realized models using bpp are a bit conservative on pseudo-knots and often miss them, so I didn't want to influence competitors to use bpp by sharing my paper, but I guess it ended up happening anyway haha",
              "votes": 6
            }
          ]
        }
      ]
    },
    {
      "id": 2553032,
      "postDate": "2023-12-08T01:07:23.680Z",
      "content": "<p>Great solution. I’m very surprised by how similar it is to mine. congrats on winning!</p>",
      "rawMarkdown": "Great solution. I’m very surprised by how similar it is to mine. congrats on winning!",
      "votes": 3
    },
    {
      "id": 2553018,
      "postDate": "2023-12-08T00:50:21.310Z",
      "content": "<p>Congratulations, your solution is so stable!</p>",
      "rawMarkdown": "Congratulations, your solution is so stable!",
      "votes": 3
    },
    {
      "id": 2553173,
      "postDate": "2023-12-08T03:56:25.133Z",
      "content": "<p>Congratulations on topping the leaderboard. Thanks for sharing the details of your solution with colorful diagrams and charts. </p>",
      "rawMarkdown": "Congratulations on topping the leaderboard. Thanks for sharing the details of your solution with colorful diagrams and charts. ",
      "votes": 4
    },
    {
      "id": 2575203,
      "postDate": "2023-12-26T16:52:02.887Z",
      "content": "<p>Wow! Amazing… and congratulations</p>",
      "rawMarkdown": "Wow! Amazing... and congratulations",
      "votes": 1
    },
    {
      "id": 2566435,
      "postDate": "2023-12-18T19:47:51.347Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations\n",
      "votes": 1
    },
    {
      "id": 2558443,
      "postDate": "2023-12-12T07:01:45.770Z",
      "content": "<p>Interesting read, congrats!</p>",
      "rawMarkdown": "Interesting read, congrats!",
      "votes": 1
    },
    {
      "id": 2557661,
      "postDate": "2023-12-11T15:46:04.120Z",
      "content": "<p>congratulations, this is really nice!</p>",
      "rawMarkdown": "congratulations, this is really nice!",
      "votes": 1
    },
    {
      "id": 2557425,
      "postDate": "2023-12-11T13:33:40.570Z",
      "content": "<p>For Dynamic Positional Bias: is there a relative length limit? </p>",
      "rawMarkdown": "For Dynamic Positional Bias: is there a relative length limit? \n",
      "votes": 1,
      "replies": [
        {
          "id": 2557913,
          "postDate": "2023-12-11T18:06:37.723Z",
          "content": "<p>We did not impose any explicit relative distance constraint; however, this layer may perform differently at relative distances greater than those in the training data set.</p>",
          "rawMarkdown": "We did not impose any explicit relative distance constraint; however, this layer may perform differently at relative distances greater than those in the training data set.",
          "votes": 1,
          "replies": [
            {
              "id": 2557965,
              "postDate": "2023-12-11T19:08:54.850Z",
              "content": "<p>interesting as this may lead to unexpected behavior at long distances. surprised you didn't cap it</p>",
              "rawMarkdown": "interesting as this may lead to unexpected behavior at long distances. surprised you didn't cap it"
            },
            {
              "id": 2558064,
              "postDate": "2023-12-11T20:55:49.773Z",
              "content": "<p>We followed the realization from <a href=\"https://github.com/bob80333/investigating_extrapolation\" target=\"_blank\">https://github.com/bob80333/investigating_extrapolation</a> and lucidrains x-transformers, where clipping is not done. And it's shows a good (<a href=\"https://github.com/bob80333/investigating_extrapolation/blob/master/experiments.txt\" target=\"_blank\">https://github.com/bob80333/investigating_extrapolation/blob/master/experiments.txt</a>) extrapolation performance</p>\n<p>If you look at Swin-Transformer realization - they also do no capping (although they do some scaling for distances). Also they use log-distances but in already mentioned experiments log-distances shown to behave worse than linear</p>\n<p>We have plans to test AlphaFold positional embedding used by the 3st place. This embedding clips positions at 32 but we are not sure if transferring the choice from protein modelling is the best one. </p>",
              "rawMarkdown": "We followed the realization from https://github.com/bob80333/investigating_extrapolation and lucidrains x-transformers, where clipping is not done. And it's shows a good (https://github.com/bob80333/investigating_extrapolation/blob/master/experiments.txt) extrapolation performance\n\nIf you look at Swin-Transformer realization - they also do no capping (although they do some scaling for distances). Also they use log-distances but in already mentioned experiments log-distances shown to behave worse than linear\n \nWe have plans to test AlphaFold positional embedding used by the 3st place. This embedding clips positions at 32 but we are not sure if transferring the choice from protein modelling is the best one. \n\n\n"
            }
          ]
        }
      ]
    },
    {
      "id": 2555740,
      "postDate": "2023-12-10T06:21:00.163Z",
      "content": "<p>Congratulations this is truly amazing . great idea and implementation </p>",
      "rawMarkdown": "Congratulations this is truly amazing . great idea and implementation ",
      "votes": 1
    },
    {
      "id": 2555205,
      "postDate": "2023-12-09T18:17:58.893Z",
      "content": "<p>Wow, congratulations on the win! Your write up looks so professional too. It’s not easy to maintain first place from public to private, and especially not in this competition!</p>",
      "rawMarkdown": "Wow, congratulations on the win! Your write up looks so professional too. It’s not easy to maintain first place from public to private, and especially not in this competition!",
      "votes": 1
    },
    {
      "id": 2554352,
      "postDate": "2023-12-09T04:47:15.327Z",
      "content": "<p>It is truly amazing</p>",
      "rawMarkdown": "It is truly amazing",
      "votes": 1
    },
    {
      "id": 2553452,
      "postDate": "2023-12-08T08:56:58.873Z",
      "content": "<p>Congratulations🎉 <a href=\"https://www.kaggle.com/vyaltsevvaleriy\" target=\"_blank\">@vyaltsevvaleriy</a> on your victory in the competition! Your work was not only impressive but also visually beautiful. I'm curious, what software did you use to create those visuals?</p>",
      "rawMarkdown": "Congratulations🎉 @vyaltsevvaleriy on your victory in the competition! Your work was not only impressive but also visually beautiful. I'm curious, what software did you use to create those visuals?",
      "votes": 1,
      "replies": [
        {
          "id": 2553711,
          "postDate": "2023-12-08T13:08:15.627Z",
          "content": "<p>Hi! I'm very glad to hear that! I used a vector graphics editor called Inkscape.</p>",
          "rawMarkdown": "Hi! I'm very glad to hear that! I used a vector graphics editor called Inkscape.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2553407,
      "postDate": "2023-12-08T08:23:42.537Z",
      "content": "<p>Congratulations well done. Nice explanations and charts 👏</p>",
      "rawMarkdown": "Congratulations well done. Nice explanations and charts 👏",
      "votes": 1
    },
    {
      "id": 2553326,
      "postDate": "2023-12-08T06:39:13.563Z",
      "content": "<p>A few assorted thoughts that came to mind that could be interesting to explore/think about</p>\n<ul>\n<li>I wonder if using hamming distance could be limiting, and whether it would be better to use a method that better accounts for \"shifted\" subsequences (eg, I would expect AAAACCCCAAAAGGGG and CCCCAAAAGGGGAAAA to both have the same characteristics and ultimately fold into the same structure). Though I'm not sure how prevalent that kind of thing actually winds up being in the training data.</li>\n<li>Per <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460121#2553008\" target=\"_blank\">Shijun's comment</a> about concerns about ability to capture pseudoknots in BPPs, I wonder if an Eternafold (or more generally, Contrafold) model that fully predicts pseudoknots would lead to better performance. Though if NUPACK's implementation is anything to go by, it would be awfully computationally expensive.</li>\n<li>That being said as far as performance is concerned (even under the current situation), it would be interesting to see how well this model would perform using LinearPartition with Eternafold's parameters rather than Eternafold itself to speed up that step.</li>\n</ul>",
      "rawMarkdown": "A few assorted thoughts that came to mind that could be interesting to explore/think about\n- I wonder if using hamming distance could be limiting, and whether it would be better to use a method that better accounts for \"shifted\" subsequences (eg, I would expect AAAACCCCAAAAGGGG and CCCCAAAAGGGGAAAA to both have the same characteristics and ultimately fold into the same structure). Though I'm not sure how prevalent that kind of thing actually winds up being in the training data.\n- Per [Shijun's comment](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460121#2553008) about concerns about ability to capture pseudoknots in BPPs, I wonder if an Eternafold (or more generally, Contrafold) model that fully predicts pseudoknots would lead to better performance. Though if NUPACK's implementation is anything to go by, it would be awfully computationally expensive.\n- That being said as far as performance is concerned (even under the current situation), it would be interesting to see how well this model would perform using LinearPartition with Eternafold's parameters rather than Eternafold itself to speed up that step.",
      "votes": 1,
      "replies": [
        {
          "id": 2553424,
          "postDate": "2023-12-08T08:40:25.430Z",
          "content": "<p>I think it would still be possible to use NUPACK. Just apply it only for some sequences. For example, we could select these sequences using active learning.</p>",
          "rawMarkdown": "I think it would still be possible to use NUPACK. Just apply it only for some sequences. For example, we could select these sequences using active learning.",
          "replies": [
            {
              "id": 2553841,
              "postDate": "2023-12-08T14:47:00.250Z",
              "content": "<p>If you're doing that, you would still presumably need to run nupack during inference though. For our use case on Eterna, even ~155 base strands were a problem in both runtime and memory usage <em>just for MFE prediction</em> (especially on lower-end hardware we target), not even BPPs (which has worse computational complexity). The time/memory complexity as it gets longer is pretty bad, so trying to use it on strands that are, say, a few thousand bases long (think SARS-CoV-2 spike protein mRNA or the larger parts of ribosomal RNA) is not viable (even presumably for higher-end hardware).</p>",
              "rawMarkdown": "If you're doing that, you would still presumably need to run nupack during inference though. For our use case on Eterna, even ~155 base strands were a problem in both runtime and memory usage _just for MFE prediction_ (especially on lower-end hardware we target), not even BPPs (which has worse computational complexity). The time/memory complexity as it gets longer is pretty bad, so trying to use it on strands that are, say, a few thousand bases long (think SARS-CoV-2 spike protein mRNA or the larger parts of ribosomal RNA) is not viable (even presumably for higher-end hardware).",
              "votes": 1
            },
            {
              "id": 2553894,
              "postDate": "2023-12-08T15:33:53.010Z",
              "content": "<p>Exactly, I apologize, I didn't think about it in the middle of the night. But it can still be used, we just need to slightly change the approach. In theory, there is nothing stopping the Transformer from learning the \"real\" BPP in the attention layers simply from the sequence, if we train it on a large amount of data. We have limited data, but we could use distillation. First, create an accurate model on sequences and BPP, and then distill the knowledge from it into a model based only on sequences. In other domains, this usually leads to a significant improvement in metrics, and I believe this task will not be an exception.</p>",
              "rawMarkdown": "Exactly, I apologize, I didn't think about it in the middle of the night. But it can still be used, we just need to slightly change the approach. In theory, there is nothing stopping the Transformer from learning the \"real\" BPP in the attention layers simply from the sequence, if we train it on a large amount of data. We have limited data, but we could use distillation. First, create an accurate model on sequences and BPP, and then distill the knowledge from it into a model based only on sequences. In other domains, this usually leads to a significant improvement in metrics, and I believe this task will not be an exception.",
              "votes": 2
            },
            {
              "id": 2553897,
              "postDate": "2023-12-08T15:35:35.060Z",
              "content": "<p>I'd love to see that approach!</p>",
              "rawMarkdown": "I'd love to see that approach!"
            },
            {
              "id": 2554335,
              "postDate": "2023-12-09T04:17:33.890Z",
              "content": "<p>Yeap, we thought about that. But it seems <a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a> has a solution that is able to achieve that without distillation. Lets wait.</p>\n<p>It seems, there are multiple ways to integrate bpps knowledge in neural network (maybe, just forcing model to predict instead of providing model with bpps). Or somehow integrate the data used by eternafold during deducing parameters for their bpp-predicting model (it's not so complex so definitely can be learned by the network itself)</p>\n<p>Unfortunately, we have no time during the competition to test this approaches. Actually, we tested so many things the post still lacking 10-20% of them and we will try to add it as soon as we will have time for that. E.g, we tried to allow the model to modify bpps given attention matricies. This lead to faster convergence (in number of epochs), but resulted in the same model performance. </p>\n<p>Anyway we a not going to stop at this, maybe the further model development will be a part of <a href=\"https://www.kaggle.com/vyaltsevvaleriy\" target=\"_blank\">@vyaltsevvaleriy</a> diploma project. Especially given the fact organizers are going to publish raw data -- a whole room for testing and additional improvements.</p>",
              "rawMarkdown": "Yeap, we thought about that. But it seems @dankrstev has a solution that is able to achieve that without distillation. Lets wait.\n\nIt seems, there are multiple ways to integrate bpps knowledge in neural network (maybe, just forcing model to predict instead of providing model with bpps). Or somehow integrate the data used by eternafold during deducing parameters for their bpp-predicting model (it's not so complex so definitely can be learned by the network itself)\n\nUnfortunately, we have no time during the competition to test this approaches. Actually, we tested so many things the post still lacking 10-20% of them and we will try to add it as soon as we will have time for that. E.g, we tried to allow the model to modify bpps given attention matricies. This lead to faster convergence (in number of epochs), but resulted in the same model performance. \n\nAnyway we a not going to stop at this, maybe the further model development will be a part of @vyaltsevvaleriy diploma project. Especially given the fact organizers are going to publish raw data -- a whole room for testing and additional improvements.",
              "votes": 1
            },
            {
              "id": 2555084,
              "postDate": "2023-12-09T17:02:32.180Z",
              "content": "<p>Yes, I think you can just embed information about BPP into the transformer. Although, I would take a Graph Transformer for such a case, it could be taught to simply predict BPP for edges, it seems more intuitive.<br>\nIn addition, there is a NodeFormer for graphs, where, with the same global attention, the complexity is reduced to O(N).<br>\nThere were thoughts of crossing NodeFormer with BPP, but we joined the competition only 2 weeks before the end, so I did not try.</p>",
              "rawMarkdown": "Yes, I think you can just embed information about BPP into the transformer. Although, I would take a Graph Transformer for such a case, it could be taught to simply predict BPP for edges, it seems more intuitive.\nIn addition, there is a NodeFormer for graphs, where, with the same global attention, the complexity is reduced to O(N).\nThere were thoughts of crossing NodeFormer with BPP, but we joined the competition only 2 weeks before the end, so I did not try."
            }
          ]
        },
        {
          "id": 2554867,
          "postDate": "2023-12-09T13:27:05.037Z",
          "content": "<p><a href=\"https://www.kaggle.com/jonathanromano\" target=\"_blank\">@jonathanromano</a> These bpps are all generated by LinearPartition with Eternafold's parameters actually, though locally it doesn't seem that much faster. </p>",
          "rawMarkdown": "@jonathanromano These bpps are all generated by LinearPartition with Eternafold's parameters actually, though locally it doesn't seem that much faster. ",
          "votes": 2,
          "replies": [
            {
              "id": 2555431,
              "postDate": "2023-12-09T22:10:06.383Z",
              "content": "<p>Interesting - admittedly I haven't gotten the chance to really sit down and play with it much. I assume linearpartition is being run with default parameters? IIRC depending on how you set beam size it changes how close to the original complexity it is (changing how much of the solution space gets thrown out)</p>",
              "rawMarkdown": "Interesting - admittedly I haven't gotten the chance to really sit down and play with it much. I assume linearpartition is being run with default parameters? IIRC depending on how you set beam size it changes how close to the original complexity it is (changing how much of the solution space gets thrown out)",
              "votes": 1
            },
            {
              "id": 2555441,
              "postDate": "2023-12-09T22:20:44.210Z",
              "content": "<p>Yes. I used the default beam size which is kind of big (100), so for 200 long sequences, not much faster</p>",
              "rawMarkdown": "Yes. I used the default beam size which is kind of big (100), so for 200 long sequences, not much faster"
            },
            {
              "id": 2555446,
              "postDate": "2023-12-09T22:28:31.573Z",
              "content": "<p>Oh yeah, we are working on relatively small sequences - this would matter more if we were looking at longer sequences (going into the thousands of bases), which we do have use cases for</p>",
              "rawMarkdown": "Oh yeah, we are working on relatively small sequences - this would matter more if we were looking at longer sequences (going into the thousands of bases), which we do have use cases for"
            }
          ]
        }
      ]
    },
    {
      "id": 2553246,
      "postDate": "2023-12-08T05:36:49.340Z",
      "content": "<p>Not to rush on as you said you would release the code. <br>\nI am trying to understand the split of dataset and would want to understand:</p>\n<ul>\n<li>how (what entities) hamming distance was computed. </li>\n<li>what does clustering by dbscan did to help split <br>\n(thanks)</li>\n</ul>",
      "rawMarkdown": "Not to rush on as you said you would release the code. \nI am trying to understand the split of dataset and would want to understand:\n- how (what entities) hamming distance was computed. \n- what does clustering by dbscan did to help split \n(thanks)",
      "votes": 1
    },
    {
      "id": 2553234,
      "postDate": "2023-12-08T05:28:04.527Z",
      "content": "<p>Congratulations on winning. Excellent explanation of the methods used.</p>",
      "rawMarkdown": "Congratulations on winning. Excellent explanation of the methods used.",
      "votes": 1
    },
    {
      "id": 2553208,
      "postDate": "2023-12-08T04:46:37.470Z",
      "content": "<p>Congrats! Definitely see a lot of similarities, but also some differences.<br>\nFor us relative worked better than dynamic positional bias and we were able to utilize pseudolabeling for a sizable boost.</p>",
      "rawMarkdown": "Congrats! Definitely see a lot of similarities, but also some differences.\nFor us relative worked better than dynamic positional bias and we were able to utilize pseudolabeling for a sizable boost.",
      "votes": 1
    },
    {
      "id": 2553016,
      "postDate": "2023-12-08T00:47:19.840Z",
      "content": "<p>Congratulations! Would you consider sharing the code? </p>",
      "rawMarkdown": "Congratulations! Would you consider sharing the code? ",
      "votes": 1,
      "replies": [
        {
          "id": 2553019,
          "postDate": "2023-12-08T00:50:48.200Z",
          "content": "<p>Yes, sure. We need some time to wrap it up</p>",
          "rawMarkdown": "Yes, sure. We need some time to wrap it up",
          "votes": 6
        }
      ]
    },
    {
      "id": 2553055,
      "postDate": "2023-12-08T01:45:17.623Z",
      "content": "<p>Congratulations!!<br>\nWonderful work and great to learn.</p>",
      "rawMarkdown": "Congratulations!!\nWonderful work and great to learn.",
      "votes": 2
    },
    {
      "id": 2553053,
      "postDate": "2023-12-08T01:42:07.857Z",
      "content": "<p>Thanks for the incredibly detailed write up and congratulations!</p>",
      "rawMarkdown": "Thanks for the incredibly detailed write up and congratulations!",
      "votes": 2
    },
    {
      "id": 2553036,
      "postDate": "2023-12-08T01:14:17.540Z",
      "content": "<p>Congratulations!<br>\nReally a wonderful work. learn a lot!</p>",
      "rawMarkdown": "Congratulations!\nReally a wonderful work. learn a lot!\n",
      "votes": 2
    },
    {
      "id": 2553034,
      "postDate": "2023-12-08T01:11:54.620Z",
      "content": "<p>Congratulations on your first place! Our solution is very similar to yours! Great solutuion!</p>",
      "rawMarkdown": "Congratulations on your first place! Our solution is very similar to yours! Great solutuion!",
      "votes": 2
    },
    {
      "id": 2553020,
      "postDate": "2023-12-08T00:51:20.427Z",
      "content": "<p>Congrats!</p>\n<p>Thanks for the great explanation.</p>",
      "rawMarkdown": "Congrats!\n\nThanks for the great explanation.",
      "votes": 2
    },
    {
      "id": 2602045,
      "postDate": "2024-01-14T20:58:22.943Z",
      "content": "<p><a href=\"https://www.kaggle.com/vyaltsevvaleriy\" target=\"_blank\">@vyaltsevvaleriy</a> <a href=\"https://www.kaggle.com/dmitrypenzar1996\" target=\"_blank\">@dmitrypenzar1996</a> I read your codes, and I have a question. <br>\nYou prepared</p>\n<ol>\n<li>EternaFold BPP</li>\n<li>EternaFold MFE</li>\n<li>ipknot from 2</li>\n</ol>\n<p>But you actually used only EternaFold BPP? Because this option doesn't use <strong>--brackets</strong>.<br>\n<code>python3 python_scripts/train_uni_adjnet_se.py --bpp_path /projects_nvme/deepbeer/eterna/ --train_path /projects/deepbeer/ribonanza/train_data/train_data.parquet  --out_path  outmodel_dir --device 0 --num_workers 20 --wd 0.05 --epoch 270 --lr_max 5e-3 --pct_start 0.05 --batch_cnt 1791 --sgd_lr 5e-5 --sgd_epochs 25 --sgd_batch_cnt 500 --sgd_wd 0.05 --fold 0 --nfolds 1000 --pos_embedding dyn --adj_ks 3 --seed 42 --use_se</code></p>",
      "rawMarkdown": "@vyaltsevvaleriy @dmitrypenzar1996 I read your codes, and I have a question. \nYou prepared\n1. EternaFold BPP\n2. EternaFold MFE\n3. ipknot from 2\n\nBut you actually used only EternaFold BPP? Because this option doesn't use **--brackets**.\n`python3 python_scripts/train_uni_adjnet_se.py --bpp_path /projects_nvme/deepbeer/eterna/ --train_path /projects/deepbeer/ribonanza/train_data/train_data.parquet  --out_path  outmodel_dir --device 0 --num_workers 20 --wd 0.05 --epoch 270 --lr_max 5e-3 --pct_start 0.05 --batch_cnt 1791 --sgd_lr 5e-5 --sgd_epochs 25 --sgd_batch_cnt 500 --sgd_wd 0.05 --fold 0 --nfolds 1000 --pos_embedding dyn --adj_ks 3 --seed 42 --use_se`",
      "replies": [
        {
          "id": 2602971,
          "postDate": "2024-01-15T13:37:36.570Z",
          "content": "<p>We used 2 and 3 for one model in our ensemble. In general, this features are completely useless, but they were used in final submission so we have to include them to git</p>",
          "rawMarkdown": "We used 2 and 3 for one model in our ensemble. In general, this features are completely useless, but they were used in final submission so we have to include them to git",
          "votes": 1,
          "replies": [
            {
              "id": 2603110,
              "postDate": "2024-01-15T15:25:48.123Z",
              "content": "<blockquote>\n  <p>this features are completely useless</p>\n</blockquote>\n<p>Unbelievable! I used bpp, structure, loop_type, chunk and segment.<br>\nYour model is so much simpler than I thought.</p>\n<p>Did you tune the number of 12 transformer layers and 192 embedding?</p>",
              "rawMarkdown": ">this features are completely useless\n\nUnbelievable! I used bpp, structure, loop_type, chunk and segment.\nYour model is so much simpler than I thought.\n\nDid you tune the number of 12 transformer layers and 192 embedding?"
            },
            {
              "id": 2618348,
              "postDate": "2024-01-24T17:20:59.567Z",
              "content": "<p>We tried but there was no significan improvement.</p>",
              "rawMarkdown": "We tried but there was no significan improvement."
            }
          ]
        }
      ]
    },
    {
      "id": 2557147,
      "postDate": "2023-12-11T08:46:28.860Z",
      "content": "<p>May you share your CV-LB correlation?</p>",
      "rawMarkdown": "May you share your CV-LB correlation?"
    },
    {
      "id": 2554943,
      "postDate": "2023-12-09T14:49:15.010Z",
      "content": "<p>nice. congratulations</p>",
      "rawMarkdown": "nice. congratulations"
    },
    {
      "id": 2553780,
      "postDate": "2023-12-08T14:06:06.373Z",
      "content": "<p>Was any data pre-processing done on reactivity values, wherein the reactivity values are adjusted based on Signal to Noise and reactivity error?</p>",
      "rawMarkdown": "Was any data pre-processing done on reactivity values, wherein the reactivity values are adjusted based on Signal to Noise and reactivity error?",
      "replies": [
        {
          "id": 2554322,
          "postDate": "2023-12-09T04:00:24.640Z",
          "content": "<p>No, only preprocessing done by the organizers. We tried to use additional losses based on reactivities errors but they haven't resulted in increased performance</p>",
          "rawMarkdown": "No, only preprocessing done by the organizers. We tried to use additional losses based on reactivities errors but they haven't resulted in increased performance"
        }
      ]
    },
    {
      "id": 2553759,
      "postDate": "2023-12-08T13:52:35.540Z",
      "content": "<p>congratulations !</p>\n<p>predicting reactivity of DMS given 2A3 and vice-versa is smart.<br>\nIs there some relation between DMS and 2A3 ? What was the intuition behind this thought?<br>\nIf you can share that will be great.</p>",
      "rawMarkdown": "congratulations !\n\npredicting reactivity of DMS given 2A3 and vice-versa is smart.\nIs there some relation between DMS and 2A3 ? What was the intuition behind this thought?\nIf you can share that will be great.\n",
      "replies": [
        {
          "id": 2554325,
          "postDate": "2023-12-09T04:02:22.970Z",
          "content": "<p>Both dms and 2a3 are reactivities and somehow show the RNA secondary structure. So, if we know one of them it should be easier to predict the second. We are really good at predicting dms values while 2a3 seems to be much more noisier and harder to predict. So predicting dms and then using it to predict 2a3 works (while 2a3-&gt;dms conversion doens't)</p>",
          "rawMarkdown": "Both dms and 2a3 are reactivities and somehow show the RNA secondary structure. So, if we know one of them it should be easier to predict the second. We are really good at predicting dms values while 2a3 seems to be much more noisier and harder to predict. So predicting dms and then using it to predict 2a3 works (while 2a3->dms conversion doens't)",
          "votes": 2
        }
      ]
    },
    {
      "id": 2553051,
      "postDate": "2023-12-08T01:38:00.170Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2553074,
      "author_name": "BarryZhou",
      "author_url": "",
      "post_date": "2023-12-08T01:56:43.433000",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/vyaltsevvaleriy\" target=\"_blank\">@vyaltsevvaleriy</a> 🎉🎉🎉 This is a very, very great method. May I ask what software was used to draw this picture? This is so beautiful</p>",
      "votes": 9,
      "replies": [
        {
          "id": 2553451,
          "author_name": "try",
          "author_url": "",
          "post_date": "2023-12-08T08:56:55.440000",
          "content": "<p>Haha, I also want to know</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2553696,
          "author_name": "vyaltsevvaleriy",
          "author_url": "",
          "post_date": "2023-12-08T12:56:38.877000",
          "content": "<p>Hi! Thank you for your appreciation! I used Inkscape - it is a vector graphics editor allowing to flexibly implement any ideas.</p>",
          "votes": 10,
          "replies": [
            {
              "id": 2553829,
              "author_name": "BarryZhou",
              "author_url": "",
              "post_date": "2023-12-08T14:32:42.767000",
              "content": "<p>Thank you very much! vyaltsevvaleriy🥰</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2553004,
      "author_name": "medicine-wave",
      "author_url": "",
      "post_date": "2023-12-08T00:36:09.530000",
      "content": "<p>Wonderful work and presentation! Congratulations! Would you mind providing more improvement details (improvement points) with each operation?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2553022,
          "author_name": "Penzar Dmitry",
          "author_url": "",
          "post_date": "2023-12-08T00:52:28.640000",
          "content": "<p>Yes, sure, I hope we will be able to explain this observation - <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/458478\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/458478</a></p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2553008,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2023-12-08T00:39:09.843000",
      "content": "<p>Nice! Actually my model has also been published here: <a href=\"https://academic.oup.com/bib/article/24/1/bbac581/6986359\" target=\"_blank\">https://academic.oup.com/bib/article/24/1/bbac581/6986359</a>. Congrats!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2553023,
          "author_name": "Penzar Dmitry",
          "author_url": "",
          "post_date": "2023-12-08T00:53:19.323000",
          "content": "<p>We failed to google your publication initially, thanks!) </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2553026,
              "author_name": "Shujun",
              "author_url": "",
              "post_date": "2023-12-08T00:58:28.277000",
              "content": "<p>No problem! I actually didn't share my model because we realized models using bpp are a bit conservative on pseudo-knots and often miss them, so I didn't want to influence competitors to use bpp by sharing my paper, but I guess it ended up happening anyway haha</p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2553032,
      "author_name": "hoyso48",
      "author_url": "",
      "post_date": "2023-12-08T01:07:23.680000",
      "content": "<p>Great solution. I’m very surprised by how similar it is to mine. congrats on winning!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2553018,
      "author_name": "Xiang Huang",
      "author_url": "",
      "post_date": "2023-12-08T00:50:21.310000",
      "content": "<p>Congratulations, your solution is so stable!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2553173,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-12-08T03:56:25.133000",
      "content": "<p>Congratulations on topping the leaderboard. Thanks for sharing the details of your solution with colorful diagrams and charts. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2575203,
      "author_name": "Shivansh",
      "author_url": "",
      "post_date": "2023-12-26T16:52:02.887000",
      "content": "<p>Wow! Amazing… and congratulations</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2566435,
      "author_name": "Md Nazrul Islam",
      "author_url": "",
      "post_date": "2023-12-18T19:47:51.347000",
      "content": "<p>Congratulations</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2558443,
      "author_name": "Borna Jamali",
      "author_url": "",
      "post_date": "2023-12-12T07:01:45.770000",
      "content": "<p>Interesting read, congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2557661,
      "author_name": "Zafir Khalid",
      "author_url": "",
      "post_date": "2023-12-11T15:46:04.120000",
      "content": "<p>congratulations, this is really nice!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2557425,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2023-12-11T13:33:40.570000",
      "content": "<p>For Dynamic Positional Bias: is there a relative length limit? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2557913,
          "author_name": "vyaltsevvaleriy",
          "author_url": "",
          "post_date": "2023-12-11T18:06:37.723000",
          "content": "<p>We did not impose any explicit relative distance constraint; however, this layer may perform differently at relative distances greater than those in the training data set.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2557965,
              "author_name": "Shujun",
              "author_url": "",
              "post_date": "2023-12-11T19:08:54.850000",
              "content": "<p>interesting as this may lead to unexpected behavior at long distances. surprised you didn't cap it</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2558064,
              "author_name": "Penzar Dmitry",
              "author_url": "",
              "post_date": "2023-12-11T20:55:49.773000",
              "content": "<p>We followed the realization from <a href=\"https://github.com/bob80333/investigating_extrapolation\" target=\"_blank\">https://github.com/bob80333/investigating_extrapolation</a> and lucidrains x-transformers, where clipping is not done. And it's shows a good (<a href=\"https://github.com/bob80333/investigating_extrapolation/blob/master/experiments.txt\" target=\"_blank\">https://github.com/bob80333/investigating_extrapolation/blob/master/experiments.txt</a>) extrapolation performance</p>\n<p>If you look at Swin-Transformer realization - they also do no capping (although they do some scaling for distances). Also they use log-distances but in already mentioned experiments log-distances shown to behave worse than linear</p>\n<p>We have plans to test AlphaFold positional embedding used by the 3st place. This embedding clips positions at 32 but we are not sure if transferring the choice from protein modelling is the best one. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2555740,
      "author_name": "mhdaw",
      "author_url": "",
      "post_date": "2023-12-10T06:21:00.163000",
      "content": "<p>Congratulations this is truly amazing . great idea and implementation </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2555205,
      "author_name": "Abhishek Shah",
      "author_url": "",
      "post_date": "2023-12-09T18:17:58.893000",
      "content": "<p>Wow, congratulations on the win! Your write up looks so professional too. It’s not easy to maintain first place from public to private, and especially not in this competition!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2554352,
      "author_name": "Izam Mohammed",
      "author_url": "",
      "post_date": "2023-12-09T04:47:15.327000",
      "content": "<p>It is truly amazing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2553452,
      "author_name": "Choekyel",
      "author_url": "",
      "post_date": "2023-12-08T08:56:58.873000",
      "content": "<p>Congratulations🎉 <a href=\"https://www.kaggle.com/vyaltsevvaleriy\" target=\"_blank\">@vyaltsevvaleriy</a> on your victory in the competition! Your work was not only impressive but also visually beautiful. I'm curious, what software did you use to create those visuals?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2553711,
          "author_name": "vyaltsevvaleriy",
          "author_url": "",
          "post_date": "2023-12-08T13:08:15.627000",
          "content": "<p>Hi! I'm very glad to hear that! I used a vector graphics editor called Inkscape.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2553407,
      "author_name": "B@rro_07",
      "author_url": "",
      "post_date": "2023-12-08T08:23:42.537000",
      "content": "<p>Congratulations well done. Nice explanations and charts 👏</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2553326,
      "author_name": "Jonathan Romano",
      "author_url": "",
      "post_date": "2023-12-08T06:39:13.563000",
      "content": "<p>A few assorted thoughts that came to mind that could be interesting to explore/think about</p>\n<ul>\n<li>I wonder if using hamming distance could be limiting, and whether it would be better to use a method that better accounts for \"shifted\" subsequences (eg, I would expect AAAACCCCAAAAGGGG and CCCCAAAAGGGGAAAA to both have the same characteristics and ultimately fold into the same structure). Though I'm not sure how prevalent that kind of thing actually winds up being in the training data.</li>\n<li>Per <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460121#2553008\" target=\"_blank\">Shijun's comment</a> about concerns about ability to capture pseudoknots in BPPs, I wonder if an Eternafold (or more generally, Contrafold) model that fully predicts pseudoknots would lead to better performance. Though if NUPACK's implementation is anything to go by, it would be awfully computationally expensive.</li>\n<li>That being said as far as performance is concerned (even under the current situation), it would be interesting to see how well this model would perform using LinearPartition with Eternafold's parameters rather than Eternafold itself to speed up that step.</li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 2553424,
          "author_name": "Aleksey Trepetsky",
          "author_url": "",
          "post_date": "2023-12-08T08:40:25.430000",
          "content": "<p>I think it would still be possible to use NUPACK. Just apply it only for some sequences. For example, we could select these sequences using active learning.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2553841,
              "author_name": "Jonathan Romano",
              "author_url": "",
              "post_date": "2023-12-08T14:47:00.250000",
              "content": "<p>If you're doing that, you would still presumably need to run nupack during inference though. For our use case on Eterna, even ~155 base strands were a problem in both runtime and memory usage <em>just for MFE prediction</em> (especially on lower-end hardware we target), not even BPPs (which has worse computational complexity). The time/memory complexity as it gets longer is pretty bad, so trying to use it on strands that are, say, a few thousand bases long (think SARS-CoV-2 spike protein mRNA or the larger parts of ribosomal RNA) is not viable (even presumably for higher-end hardware).</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2553894,
              "author_name": "Aleksey Trepetsky",
              "author_url": "",
              "post_date": "2023-12-08T15:33:53.010000",
              "content": "<p>Exactly, I apologize, I didn't think about it in the middle of the night. But it can still be used, we just need to slightly change the approach. In theory, there is nothing stopping the Transformer from learning the \"real\" BPP in the attention layers simply from the sequence, if we train it on a large amount of data. We have limited data, but we could use distillation. First, create an accurate model on sequences and BPP, and then distill the knowledge from it into a model based only on sequences. In other domains, this usually leads to a significant improvement in metrics, and I believe this task will not be an exception.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2553897,
              "author_name": "Jonathan Romano",
              "author_url": "",
              "post_date": "2023-12-08T15:35:35.060000",
              "content": "<p>I'd love to see that approach!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2554335,
              "author_name": "Penzar Dmitry",
              "author_url": "",
              "post_date": "2023-12-09T04:17:33.890000",
              "content": "<p>Yeap, we thought about that. But it seems <a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a> has a solution that is able to achieve that without distillation. Lets wait.</p>\n<p>It seems, there are multiple ways to integrate bpps knowledge in neural network (maybe, just forcing model to predict instead of providing model with bpps). Or somehow integrate the data used by eternafold during deducing parameters for their bpp-predicting model (it's not so complex so definitely can be learned by the network itself)</p>\n<p>Unfortunately, we have no time during the competition to test this approaches. Actually, we tested so many things the post still lacking 10-20% of them and we will try to add it as soon as we will have time for that. E.g, we tried to allow the model to modify bpps given attention matricies. This lead to faster convergence (in number of epochs), but resulted in the same model performance. </p>\n<p>Anyway we a not going to stop at this, maybe the further model development will be a part of <a href=\"https://www.kaggle.com/vyaltsevvaleriy\" target=\"_blank\">@vyaltsevvaleriy</a> diploma project. Especially given the fact organizers are going to publish raw data -- a whole room for testing and additional improvements.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2555084,
              "author_name": "Aleksey Trepetsky",
              "author_url": "",
              "post_date": "2023-12-09T17:02:32.180000",
              "content": "<p>Yes, I think you can just embed information about BPP into the transformer. Although, I would take a Graph Transformer for such a case, it could be taught to simply predict BPP for edges, it seems more intuitive.<br>\nIn addition, there is a NodeFormer for graphs, where, with the same global attention, the complexity is reduced to O(N).<br>\nThere were thoughts of crossing NodeFormer with BPP, but we joined the competition only 2 weeks before the end, so I did not try.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2554867,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2023-12-09T13:27:05.037000",
          "content": "<p><a href=\"https://www.kaggle.com/jonathanromano\" target=\"_blank\">@jonathanromano</a> These bpps are all generated by LinearPartition with Eternafold's parameters actually, though locally it doesn't seem that much faster. </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2555431,
              "author_name": "Jonathan Romano",
              "author_url": "",
              "post_date": "2023-12-09T22:10:06.383000",
              "content": "<p>Interesting - admittedly I haven't gotten the chance to really sit down and play with it much. I assume linearpartition is being run with default parameters? IIRC depending on how you set beam size it changes how close to the original complexity it is (changing how much of the solution space gets thrown out)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2555441,
              "author_name": "Shujun",
              "author_url": "",
              "post_date": "2023-12-09T22:20:44.210000",
              "content": "<p>Yes. I used the default beam size which is kind of big (100), so for 200 long sequences, not much faster</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2555446,
              "author_name": "Jonathan Romano",
              "author_url": "",
              "post_date": "2023-12-09T22:28:31.573000",
              "content": "<p>Oh yeah, we are working on relatively small sequences - this would matter more if we were looking at longer sequences (going into the thousands of bases), which we do have use cases for</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2553246,
      "author_name": "FdotRK",
      "author_url": "",
      "post_date": "2023-12-08T05:36:49.340000",
      "content": "<p>Not to rush on as you said you would release the code. <br>\nI am trying to understand the split of dataset and would want to understand:</p>\n<ul>\n<li>how (what entities) hamming distance was computed. </li>\n<li>what does clustering by dbscan did to help split <br>\n(thanks)</li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2553234,
      "author_name": "Sarun P M",
      "author_url": "",
      "post_date": "2023-12-08T05:28:04.527000",
      "content": "<p>Congratulations on winning. Excellent explanation of the methods used.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2553208,
      "author_name": "sroger",
      "author_url": "",
      "post_date": "2023-12-08T04:46:37.470000",
      "content": "<p>Congrats! Definitely see a lot of similarities, but also some differences.<br>\nFor us relative worked better than dynamic positional bias and we were able to utilize pseudolabeling for a sizable boost.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2553016,
      "author_name": "Swikwislkdjc",
      "author_url": "",
      "post_date": "2023-12-08T00:47:19.840000",
      "content": "<p>Congratulations! Would you consider sharing the code? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2553019,
          "author_name": "Penzar Dmitry",
          "author_url": "",
          "post_date": "2023-12-08T00:50:48.200000",
          "content": "<p>Yes, sure. We need some time to wrap it up</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 2553055,
      "author_name": "FdotRK",
      "author_url": "",
      "post_date": "2023-12-08T01:45:17.623000",
      "content": "<p>Congratulations!!<br>\nWonderful work and great to learn.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2553053,
      "author_name": "Javier Martín",
      "author_url": "",
      "post_date": "2023-12-08T01:42:07.857000",
      "content": "<p>Thanks for the incredibly detailed write up and congratulations!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2553036,
      "author_name": "Horikita Saku",
      "author_url": "",
      "post_date": "2023-12-08T01:14:17.540000",
      "content": "<p>Congratulations!<br>\nReally a wonderful work. learn a lot!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2553034,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "2023-12-08T01:11:54.620000",
      "content": "<p>Congratulations on your first place! Our solution is very similar to yours! Great solutuion!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2553020,
      "author_name": "Antonio Félix",
      "author_url": "",
      "post_date": "2023-12-08T00:51:20.427000",
      "content": "<p>Congrats!</p>\n<p>Thanks for the great explanation.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2602045,
      "author_name": "ONODERA",
      "author_url": "",
      "post_date": "2024-01-14T20:58:22.943000",
      "content": "<p><a href=\"https://www.kaggle.com/vyaltsevvaleriy\" target=\"_blank\">@vyaltsevvaleriy</a> <a href=\"https://www.kaggle.com/dmitrypenzar1996\" target=\"_blank\">@dmitrypenzar1996</a> I read your codes, and I have a question. <br>\nYou prepared</p>\n<ol>\n<li>EternaFold BPP</li>\n<li>EternaFold MFE</li>\n<li>ipknot from 2</li>\n</ol>\n<p>But you actually used only EternaFold BPP? Because this option doesn't use <strong>--brackets</strong>.<br>\n<code>python3 python_scripts/train_uni_adjnet_se.py --bpp_path /projects_nvme/deepbeer/eterna/ --train_path /projects/deepbeer/ribonanza/train_data/train_data.parquet  --out_path  outmodel_dir --device 0 --num_workers 20 --wd 0.05 --epoch 270 --lr_max 5e-3 --pct_start 0.05 --batch_cnt 1791 --sgd_lr 5e-5 --sgd_epochs 25 --sgd_batch_cnt 500 --sgd_wd 0.05 --fold 0 --nfolds 1000 --pos_embedding dyn --adj_ks 3 --seed 42 --use_se</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2602971,
          "author_name": "Penzar Dmitry",
          "author_url": "",
          "post_date": "2024-01-15T13:37:36.570000",
          "content": "<p>We used 2 and 3 for one model in our ensemble. In general, this features are completely useless, but they were used in final submission so we have to include them to git</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2603110,
              "author_name": "ONODERA",
              "author_url": "",
              "post_date": "2024-01-15T15:25:48.123000",
              "content": "<blockquote>\n  <p>this features are completely useless</p>\n</blockquote>\n<p>Unbelievable! I used bpp, structure, loop_type, chunk and segment.<br>\nYour model is so much simpler than I thought.</p>\n<p>Did you tune the number of 12 transformer layers and 192 embedding?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2618348,
              "author_name": "Penzar Dmitry",
              "author_url": "",
              "post_date": "2024-01-24T17:20:59.567000",
              "content": "<p>We tried but there was no significan improvement.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2557147,
      "author_name": "Correlation",
      "author_url": "",
      "post_date": "2023-12-11T08:46:28.860000",
      "content": "<p>May you share your CV-LB correlation?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2554943,
      "author_name": "Durlav",
      "author_url": "",
      "post_date": "2023-12-09T14:49:15.010000",
      "content": "<p>nice. congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2553780,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-08T14:06:06.373000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2554322,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-12-09T04:00:24.640000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2553759,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-08T13:52:35.540000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2554325,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-12-09T04:02:22.970000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2553051,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-08T01:38:00.170000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2553003": "The Ribonanza RNA Folding competition has been an amazing opportunity and we are really glad to have participated in it. Kudos to the hosts for their substantial effort in keeping the high organizational level and giving the contestants plenty of valuable insight and support. We would also like to thank the community for the fruitful public discussions and coding initiatives. Congratulations to everyone involved! 🎉\n\n# Code\nOpen source code is available on [github](https://github.com/autosome-ru/vigg_ribonanza/).\n\n# TLDR\n\nOur solution is based on transformer architecture. We predicted base pair probability matrix(BPPM) for each RNA sequence using EternaFold and added convolutional blocks to our architecture to process BPPM features. For some models we added Squeeze-and-Excitation layer in convolutional blocks Also we add these features with attention values before softmax operation in Self-Attention block. To allow better generalization for longer input we implemented Dynamic Positional Bias. Then we ensembled models with slight differences in architecture and training process.\n\n# Data Preprocessing\n\nAt first, for each sequence in train and test datasets we calculated a base pair probability matrix (BPPM) using EternaFold. During training and inference phases we passed to our model the RNA sequences encoded with a learnable embedding layer. Each nucleotide was considered a token and special **< start >** and **< end >** tokens were placed at both ends as well. To provide the model with the information about whether the sequence comes from a “clean” subset of training dataset (which SN_filter values show - 1 corresponds to “clean” sequences with high signal-to-noise ratio and 0 is otherwise) we encoded SN_filter values with a learnable embedding layer and added corresponding  embeddings to sequence embeddings. The embedding dimension was chosen to be of size **192**. BPPMs were padded with zeros at their margins to account for adding **< start >** and **< end >** tokens. The figure below summarizes the data preprocessing part.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F482c2d20fce57fa513021909a515e98b%2Fpreprocessing.png?generation=1701994647204619&alt=media)\n\n# Model\n\nThe model idea is partially inspired by @shujun717 solution to the Open Vaccine challenge, [link](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189564)\n\nThe model takes a sequence of tokens and BPPM as input and outputs DMS_MaP and 2A3_MaP reactivities for each input nucleotide. Its architecture comprises 12 consecutive Transformer Encoder Layers and an output projection linear layer. Each Transformer block takes a sequence of tokens and BPPM features from the previous layer and outputs updated feature maps as shown below. The Transformer Encoder block adopts common transformer encoder architecture, except we modified the Self-Attention block to ensure interaction between BPPM features and sequence features.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F0c70879344b972d64ba13945e163a126%2Fmodel_general_view.png?generation=1701994737102906&alt=media)\n\n## Self-Attention Block\n\nIn the Self-Attention (SA) block we implemented an attention mechanism with the following modifications: after attention values for each head are calculated, we add BPPM features updated by the ‘Convolutional block’, which outputs BPPM features with the number of channels corresponding to the number of heads in the SA block. We set the number of heads in SA and corresponding channels of BPPM features to be 6. Thus, the hidden dimension size of Q, K, V matrices is 32. The overall structure of the SA block is shown below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F5a7bf6e6f893938ced8bdd2e9a1fb08c%2Fselfattention.png?generation=1701994715778085&alt=media)\n\n## Dynamic Positional Bias\n\nThe sequence length is distributed differently in train and test datasets (test sequences are generally longer). To allow for the better generalization of longer inputs we implemented a positional encoding to be added to attention values. We found Dynamic Positional Bias to be of more use compared to other relative positional encoding methods we tried, such as xPos (rotary positional encoding) and ALiBi. Dynamic Positional Bias calculates for each head a relative positional bias map, which is learnable and depends on sequence length. Relative positional bias doesn’t allow to leverage distance from start and end of sequence, so tokens **< start >** and **< end >** were added to fix that.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F8d8468983d7a7d057dcf74d7511b3ac8%2Fdynpos%20(1).png?generation=1701994757603986&alt=media)\n\n## Convolutional Block and SE block\n\nThe models in the final ensemble come in two versions that are slightly different in the structure of the Convolutional block. The basic Convolutional block consists of 2D convolutional layer, batchnorm layer, activation and learnable gamma parameters that scale the output feature channels, whereas the modified version of this block (SE-Convolutional block) also contains Squeeze-and-Excitation layer, hence the name. The SE layer applies input-dependent rescaling of values along the channels as shown below. Thus, the only difference between models in the ensemble is the presence of the SE layer.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F4b8f15f5aa950ba64ecbd55df123cad4%2Fconv_se.png?generation=1701994778649467&alt=media)\n\n# Training Process\n\nThe training process has been adapted from the [notebook](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb/notebook) by @IAFOSS\nWe used a one-cycle learning rate schedule (pct_start=0.05, lr_max=2.5e-3,) coupled with AdamW optimizer (wd=0.05), and batch size set to 128.\nThe number of epochs for model training was determined from the dataset size.\nFor final models (trained almost on the whole dataset) we used 270 epochs. \nIn each epoch, 1791 batches were processed (we kept this number due to historical reasons), elements of each batch were sampled from dataset with the weight = 0.5 * torch.clamp_min(torch.log(sn + 1.01),0.01)\n\nWe have also found that training model with a simple SGD optimizer for ~15 epochs of 500 batches each additionally improves model performance (the number of epochs varies so we used a small validation set to determine the exact number of epochs)\n\nAdditionally, one model was trained to predict 2A3 given DMS, RNA sequence and BPPM (dms-to-2a3 model).\n\n# Inference\nAt the inference phase for all inputs we set the SN filter value to be 1 as if they come from a “clean” dataset.\n\n# Ensembling\n\nWe ensembled 15 models with SE-Convolutional block, 10 models with plain Convolutional block, 2 with plain Convolutional block trained on split by sequence lengths (one of these two models accepts bracket features).  We took the average of their predictions. Then we predicted 2a3 reactivities based on averaged dms reactivities using dms-to-2a3 model and added these predictions to averaged 2a3 reactivities in the following way: (27/28)*averaged_2a3 + (1/28)*predicted_2a3.\n\n# Other splits \n\nThe clear problem with a simple KFold split is the high sequence similarity across the training dataset and the fact that test sequences are very distinct from the training data. This might hamper the model development, because the increase in quality on the validation set could be due to overfitting rather than actual improvement.\n\nWe have calculated a hamming distance matrix for all the sequences present in the training set and performed a modified DBSCAN clustering procedure with distance threshold set at 0.2. In the following picture we show the clusters mapped to their respective cluster identifier (a cluster ID was assigned as a number of the smallest sequence within that cluster in the train dataset).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F55dc21823386d878fb212d681cae593e%2Fcluster.jpg?generation=1701994798203967&alt=media)\n\nWe tested several splits based on sequence identity (the easiest one is just to split data into folds without shuffling) and found out that while the model final validation performance degrades as we choose more and more stricter distance threshold, its relative value behave the same way as for a simple KFold split. So, we decided to use simple KFold split than training final models \n\nAlso, we have conducted a test of the model performance on length-based split. For that, we trained a model on the short sequences (length < 206) and validated it on sequences of size = 206. The behavior of the validation metric was slightly noisier but still highly correlated with a metric for a simple KFold split.\n\n# Public data leakage \n\nApproximately 13% of public test sequences are identical to the ones present in the train dataset (by sequence). To avoid selecting a model that is memorizing more of these sequences rather than learning RNA-related stuff, we zeroed out the predictions for these sequences in most of our submissions, while sometimes sending non-zeroed out submissions in order to compare our performance to other participants.\n\n# Features we also tried\n\n## capR\n\nWe don’t have conclusive results for that feature. It seems like it is not beneficial for the model on average but sometimes it resulted in a better model and sometimes – in worse. We decided to not use this feature.\n\n## Brackets\n\nWe have tried to use brackets generated by EternaFold, ContraFold, ViennaRNA, etc., as well as programs for pseudoknots prediction (like IPknot). However, after adding EternaFold BPPMs to the model, adding other features yields no significant increase in model performance. For some models we used brackets just to augment model\n\n## Different BPPMs\n\nAll other BPPMs (ContraFold, ViennaRNA, RNAsoft, RNAstructure) result in a suboptimal model.\nAveraging BPPMs doesn’t result in better performance.\nBPPMs from RFold yield the same quality as ViennaRNA BPPMs.\nThe RNA-FM model produces both per-nucleotide embeddings and BPPM-like matrix. Still, those don’t help the model at all, and using them results only in a slight increase in model performance then compared to sequence only model\n \n## SQUARNA matrix\n\nSQUARNA ([github](https://github.com/febos/SQUARNA)) outputs a matrix different from BPPM, but can be used in the same manner. Unfortunately, this feature also didn’t yield any additional performance increase.\n\n# What we also tried\n\n## Fully-Convolutional architecture\n\nThe initial reason we decided to take part in the competition was to test our model LegNet [link](https://academic.oup.com/bioinformatics/article/39/8/btad457/7230784), which shows SOTA results on DNA sequences and worked well on some RNA-related tasks (not yet published).\nUnfortunately, any modifications of this architecture resulted in subpar performance when compared with properly tuned transformer models. This can be explained by the fact that predicting RNA secondary structure requires attending to long-range contacts. The transformer architecture suits better for such cases.\n\n## Subsetting data \n\nSubsetting data (filtering by different thresholds on SN ratio) resulted in a performance boost for all models. However, this technique was superseded by weight sampling, which proved itself to be more effective.\n\n## Fine-tuning on public datasets \n\nWe tried to fine-tune our model on a public dataset, gathered by the organizers ([link](https://www.kaggle.com/datasets/rhijudas/rmdb-rna-mapping-database-2023-data/data)) by training model to predict the results of the public experiments (excluding ones with a small number of samples) and for Ribonanza data simultaneously. Unfortunately, this also gave no boost to the model performance.\n\n## Using 3D data \n\nWe tried to use data about predicted 3D structure of 100k sequences from the train dataset but gave up on that once we had visually analyzed them:\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16725112%2F2d054d1679176e0664130f6c46ccb0ae%2Ftrashfold.jpg?generation=1701994817511350&alt=media)\n\n## Absolute positional embedding\n\nUsing absolute positional embedding leads to unsolvable issues when generalizing upon longer sequences.\n\n## Relative positional embedding \n\nThe simplest approach is to augment absolute positional encoding during the training phase to shift randomly from 0 to (Lmax - seqlen) position. This indeed solves the issue with extrapolation, but works worse than other methods.\nRotational positional embedding, unfortunately, doesn’t help the model to generalize on larger lengths.\nALiBi positional embedding solves the issue with extrapolation but even after keeping it only for a part of heads (as suggested in https://github.com/lucidrains/x-transformers) still behaves worse than dynamic positional bias.\n\n## Augmentation \n\nFirst, we tried to use reverse augmentation. This can be done in three ways:\nreversing sequence before any modifications,\nreversing sequence before padding, but after adding <start> and <end> tokens,\nreversing sequence after padding.\n\nThe first two ways yield no gain for all variants of models we tested. Yet, the third one (upon coupling with additional finetuning) gave us a good result for a model with xpos positional encoding. The resulting single-model performance was 0.13937 on the public leaderboard. Unfortunately, xpos shows rather poor performance on long sequences so we abstained from using this model in the final submission.\n\nWe have also tried shift augmentation and different sequence padding approaches. This didn’t improve our model performance as well.\n\n## Sliding window\n\nOne of the possible ways to generalize for larger sequences is to predict reactivities using the sliding window. However, the idea is somewhat wrong in a biological sense, and it results in performance degradation when testing on train dataset sequences.\n\n## Pseudolabelling\n\nOnce we obtained ensembles of best-performing models we tried to use them to pseudolabel test dataset and use predictions with the highest confidence to train new models. While this indeed results in a better single-performing model, adding such a model to ensemble doesn’t improve ensemble performance.\n\n## Changing loss\n\nInstead of filtering sequences with low SN we tried to mask positions with high reactivity error as it was done by @nullrecurrent in Open Vaccine challenge ([link](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620))\nThis resulted in poor performance.\nWe tried to weight loss for each sequence by its SN – it didn’t result in any improvement\n\n# Tools\n\n[arnie](https://github.com/DasLab/arnie/tree/master)\n[EternaFold](https://github.com/eternagame/eternafold)\n\n# Links\n\n## Squeeze-and-excitation block\nHu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7132-7141).\n\n## Dynamic positional bias\nWang, W., Chen, W., Qiu, Q., Chen, L., Wu, B., Lin, B., ... & Liu, W. (2023). Crossformer++: A versatile vision transformer hinging on cross-scale attention. arXiv preprint arXiv:2303.06908.\n\n\n",
    "2553074": "Congratulations! @vyaltsevvaleriy 🎉🎉🎉 This is a very, very great method. May I ask what software was used to draw this picture? This is so beautiful",
    "2553004": "Wonderful work and presentation! Congratulations! Would you mind providing more improvement details (improvement points) with each operation?",
    "2553008": "Nice! Actually my model has also been published here: https://academic.oup.com/bib/article/24/1/bbac581/6986359. Congrats!",
    "2553032": "Great solution. I’m very surprised by how similar it is to mine. congrats on winning!",
    "2553018": "Congratulations, your solution is so stable!",
    "2553173": "Congratulations on topping the leaderboard. Thanks for sharing the details of your solution with colorful diagrams and charts. ",
    "2575203": "Wow! Amazing... and congratulations",
    "2566435": "Congratulations\n",
    "2558443": "Interesting read, congrats!",
    "2557661": "congratulations, this is really nice!",
    "2557425": "For Dynamic Positional Bias: is there a relative length limit? \n",
    "2555740": "Congratulations this is truly amazing . great idea and implementation ",
    "2555205": "Wow, congratulations on the win! Your write up looks so professional too. It’s not easy to maintain first place from public to private, and especially not in this competition!",
    "2554352": "It is truly amazing",
    "2553452": "Congratulations🎉 @vyaltsevvaleriy on your victory in the competition! Your work was not only impressive but also visually beautiful. I'm curious, what software did you use to create those visuals?",
    "2553407": "Congratulations well done. Nice explanations and charts 👏",
    "2553326": "A few assorted thoughts that came to mind that could be interesting to explore/think about\n- I wonder if using hamming distance could be limiting, and whether it would be better to use a method that better accounts for \"shifted\" subsequences (eg, I would expect AAAACCCCAAAAGGGG and CCCCAAAAGGGGAAAA to both have the same characteristics and ultimately fold into the same structure). Though I'm not sure how prevalent that kind of thing actually winds up being in the training data.\n- Per [Shijun's comment](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460121#2553008) about concerns about ability to capture pseudoknots in BPPs, I wonder if an Eternafold (or more generally, Contrafold) model that fully predicts pseudoknots would lead to better performance. Though if NUPACK's implementation is anything to go by, it would be awfully computationally expensive.\n- That being said as far as performance is concerned (even under the current situation), it would be interesting to see how well this model would perform using LinearPartition with Eternafold's parameters rather than Eternafold itself to speed up that step.",
    "2553246": "Not to rush on as you said you would release the code. \nI am trying to understand the split of dataset and would want to understand:\n- how (what entities) hamming distance was computed. \n- what does clustering by dbscan did to help split \n(thanks)",
    "2553234": "Congratulations on winning. Excellent explanation of the methods used.",
    "2553208": "Congrats! Definitely see a lot of similarities, but also some differences.\nFor us relative worked better than dynamic positional bias and we were able to utilize pseudolabeling for a sizable boost.",
    "2553016": "Congratulations! Would you consider sharing the code? ",
    "2553055": "Congratulations!!\nWonderful work and great to learn.",
    "2553053": "Thanks for the incredibly detailed write up and congratulations!",
    "2553036": "Congratulations!\nReally a wonderful work. learn a lot!\n",
    "2553034": "Congratulations on your first place! Our solution is very similar to yours! Great solutuion!",
    "2553020": "Congrats!\n\nThanks for the great explanation.",
    "2602045": "@vyaltsevvaleriy @dmitrypenzar1996 I read your codes, and I have a question. \nYou prepared\n1. EternaFold BPP\n2. EternaFold MFE\n3. ipknot from 2\n\nBut you actually used only EternaFold BPP? Because this option doesn't use **--brackets**.\n`python3 python_scripts/train_uni_adjnet_se.py --bpp_path /projects_nvme/deepbeer/eterna/ --train_path /projects/deepbeer/ribonanza/train_data/train_data.parquet  --out_path  outmodel_dir --device 0 --num_workers 20 --wd 0.05 --epoch 270 --lr_max 5e-3 --pct_start 0.05 --batch_cnt 1791 --sgd_lr 5e-5 --sgd_epochs 25 --sgd_batch_cnt 500 --sgd_wd 0.05 --fold 0 --nfolds 1000 --pos_embedding dyn --adj_ks 3 --seed 42 --use_se`",
    "2557147": "May you share your CV-LB correlation?",
    "2554943": "nice. congratulations",
    "2553780": "Was any data pre-processing done on reactivity values, wherein the reactivity values are adjusted based on Signal to Noise and reactivity error?",
    "2553759": "congratulations !\n\npredicting reactivity of DMS given 2A3 and vice-versa is smart.\nIs there some relation between DMS and 2A3 ? What was the intuition behind this thought?\nIf you can share that will be great.\n",
    "2553051": ""
  }
}