{
  "id": 463352,
  "title": "10th Place Solution for the Stanford Ribonanza RNA Folding Competition",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/463352",
  "author_name": "greySnow",
  "post_date": "2023-12-24T19:17:14.325000",
  "votes": 13,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Initially, I had not intended to publish my solution since I used an interesting method for the loss (that I did not see in other solutions, although I might have missed it), and since I did not finish within the price range, I wanted to keep it to myself for some other future compatible competition. However, since apparently there are plans for a paper by the hosts, I eventually opted to publish it.</p>\n<h1>Context section</h1>\n<p><a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding\" target=\"_blank\">Business context</a>.<br>\n<a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">Data context</a>.</p>\n<h1>Overview of the Approach</h1>\n<p>I used a 1Dconv+transformer model. The model does not include positional encoding since the conv layers take care of positional relationship ‘automatically,’ allowing for straightforward generalization for longer sequences. I based my model on a previous work of <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> in the ISLR competition; see <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">HERE</a> and <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406978\" target=\"_blank\">HERE</a>. BTW, congrats for 2nd place. You know, Hoyeol is probably the person from whose work I learned the most. This solution is already the second gold I get, thanks to the knowledge I got from his work. So, this is a special thank you. <br>\nHoiso’s original model was causal in time; I changed it to be symmetric to both sides by employing suitable paddings and maskings. <br>\nI trained two kinds of models; model_1 was trained only on the sequences, and model_2 included CapR and bpp sum&amp;max per nucleotide.<br>\nThings that helped me (I did not perform a proper ablation study, so I will not give exact numbers most of the time):</p>\n<ol>\n<li>Clipping the predictions with a sigmoid (regular clipping also helps, but a sigmoid helps even more). It also made the convergence smoother.</li>\n<li>One-hot encoding of the nucleotides instead of embeddings (this is mainly for the seq-only model. When including more features, one-hot is a given)</li>\n<li>CapR</li>\n<li>BPP sum&amp;max<br>\nI also found the injection of bpps values to the attention helpful, even more so with 2D conv layers on the bpps in various schemes. HOWEVER…I was very very suspicious of bpps. I feared that they would be much worse on the private test set since it consists of new sequences, and I thought that even if it was good for the public set, maybe the BPPs models were developed on similar sequences, but for the private set, it would fail. So I chose only to include sum&amp;max per nucleotide, thinking that these ‘averaged’ values would be less prone to overfitting. Also, tbh, I did not have enough time at the end to do proper research and ensemble of the more complicated models due to trying too many ideas in too little time, lol.</li>\n<li>Weighted loss function:<br>\nLet's start from the end.<br>\nloss(reactivity, reactivity_error, pred)=MAE(reactivity, pred)*potentials  <br>\nwith:  <br>\npotentials = (1-2*dist.cdf(reactivity-|reactivity-pred|)  <br>\nand:  <br>\ndist = normal_distribution(mu=reactivity, sigma= reactivity_error)<br>\nNow, to explain why.<br>\nI thought a lot about the loss function. Since we were given the errors, I searched for an appropriate way to incorporate uncertainty in the measurements into the loss. I was a bit surprised when I could not find something useful- of course, I maybe did not dig deep enough, but at the very least, it seems that a lot of digging is required. I found several things about incorporating uncertainties into the predictions (i.e., predicting the uncertainties of the predictions), but I wanted the opposite.<br>\nNow, the obvious thing to do is directly weigh the loss by the error- to give some examples, Hoiso (2nd place) used log1p error, and I saw suggestions like 1/(1+err), etc. However, I needed more than this: it punished high and low errors too much. To understand why, let's talk for a second about the meaning of error- generally, it means that instead of knowing the exact value, we can only give an approximation for the true value, with a distribution of probability for it to be any number, and the probability given by normal distribution with sigma = error. When we say that reactivity=0 ± 1, we give a probability of 0.68 for the exact value to be between -1 to 1 (i.e., within one standard deviation). <br>\nLets assume that we have reactivity1 =0 ± 1 , reactivity2 =0 ± 3 , pred1 =pred2 =100 . If we normalize, for example, by 1/(1+err), and with MAE, then loss1 =50, loss2 = 25. So the first predictions contribute to the loss twice as much as the second one, even though, if we look at the distribution of probabilities- the first one has 0.68 to be in [-1,1] and the second one has 0.68 to be in [-3,3] (and 0.95 to be in [-6,6]) so for a prediction so far away from most of the distribution, I expect both predictions to contribute to the loss about the same. Of course, we clip the predictions between 0,1, making the analysis more complicated and depending on the exact loss and clipping each person applies. <br>\nOn the other hand, assume we have reactivity1=0 ± 0.01, rectivity2=0 ± 0.1, pred1=pred2=0.05, then we get loss1=0.0495, loss2=0.0454. This time, they are almost the same- even though we know that reactivity1 has a probability of 0.68 to be in [-0.01, 0.01], which is far from the prediction relatively to the prediction for reactivity2, which lies inside one standard deviation of the ground truth distribution, so I expect it to contribute much less to the loss. <br>\nOne way to address this issue is to perturb the reactivity values according to their probability distribution. However, this leads to unstabilized training since many values have a high error, and moreover, they are perturbed outside the range of [0,1], leading to more complications.<br>\nI wanted to properly incorporate our knowledge about the probability distribution of the values into the loss function. At this point, I took some inspiration from the shell theorem (check it <a href=\"https://en.wikipedia.org/wiki/Shell_theorem\" target=\"_blank\">on Wikipedia</a>), and by considering the probability distribution as a form of ‘potential,’ I finished with the above formula. <br>\nThis loss helped my convergence. I don’t remember precisely how much, maybe about 0.001? Not by a huge amount. But it certainly helped. It also made it even smoother. I think this method allows the inclusion of high error values while squeezing as much information as possible from them. Also, I believe this is better than trying to pseudo-label high-error values since noisy data is still data with real information inside. We just need to squeeze this information out…<br>\nBut wait, there is more. What if I suspect the errors we were given themselves are inaccurate?<br>\nNo problem. Just multiply the given errors by some factor, add some constant errors, and assemble many models with randomly chosen factors.<br>\nAt this point, I remembered the method of label smoothing for logistic regression, and, well, this is basically a form of label smoothing for regression. Once I framed it as label smoothing, it was easy to find similar ideas, e.g. <a href=\"https://proceedings.mlr.press/v80/imani18a/imani18a.pdf\" target=\"_blank\">this paper</a>. Although it seems like they overcomplicated things there a bit.<br>\nNote that this works the same for other losses, e.g., for MSE, multiply MSE by the same 'potentials,' etc.</li>\n</ol>\n<h1>Details of the submission</h1>\n<ol>\n<li><p>Ensembling (obviously):<br>\nHere is a nice plot of the power of ensembling for the sequence-only models (model_1)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F81ab8378d36c6bc5a08c64da32642008%2Fensembling_1.png?generation=1703442429686814&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Ensembling models with bpp+CapR with sequences-only models<br>\nI trusted my sequences-only models much more than the +bpp&amp;CapR ones, so I submitted one ensemble of pure sequences-only models. I planned the other ensemble to be only +bpp&amp;CapR models. But then, on the very last day, just when I had the two ensembles ready, I looked again at the pictures of the results for R1138v1 predictions. See more details <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653\" target=\"_blank\">here</a>. Basically, we wanted the models to replicate the following image:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F506500c54df757dd1b8046843f9a25e8%2FR1138v1.png?generation=1703442477381564&amp;alt=media\" alt=\"\"><br>\nHere is the picture of my sequence-only models ensemble:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F6a3fece98bcbd1a2deb22b7b5dfa52cc%2FR1138v1%20model_1.png?generation=1703442512137941&amp;alt=media\" alt=\"\"></p>\n<p>It’s a bit hard to see, but it recover the central line, especially in the 2A3 picture. However, the sidelines are missing. On the other hand, for the +bpp&amp;CapR models ensemble:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F8cd376a57c59e86e7bcb09917ef4e6e6%2FR1138v1%20model_2.png?generation=1703442537753029&amp;alt=media\" alt=\"\"><br>\nHere, it’s the opposite. I have visible traces of the sidelines but almost nothing for the central one (maybe one very vague point).<br>\nThese images bring an immediate idea…more ensembling…(this was not obvious at first since the BPP&amp;CapR models/ensemble had a much higher validation/public score than the sequences-only ones)<br>\nI had only three submissions left. I tried the following weights: 6:2, 5:3, and 4:4 in favor of a sequence-only ensemble. 4:4 had a better score but I was too suspisious of bpp and CapR, so I chose 5:3 weights which had a similar score to the original only bpp+CapR ensemble but was expected to be more robust.<br>\nAfter the competition ended, I tried other weights. Here is the complete analysis:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2Fa86e7996fb22dca69e22fe486ff72282%2Fensembling_2.png?generation=1703442561705283&amp;alt=media\" alt=\"\"><br>\nEven if I chose the best weight, I would still end outside the price range. So I don’t feel too terrible about missing here.  I could end up in 8th place, though, if I chose the 4:4 weights on the last day. Keep in mind that my solution does not include the edge information of the BPPs, as opposed to most (all?) other top solutions. Maybe, except the 3rd place, he got an extremely strong sequences-only solution, much better than mine.</p></li>\n</ol>\n<h1>More things I tried but did not included in the final solution or didn’t work out (possibly also due to lack of time):</h1>\n<p>Using the external data source as another head for the loss, using the external data source to train a model, predicting on the train set, and use the predictions as features, different predictions for bpps (eterna, contra, etc.) did not contribute too much or at all. Structures (i.e. (.) notation) also were not especially useful. I had plans for the 3D features but not enough time…I tried other schemes for the loss, for example, minimizing log loss (sigmoid) (i.e., framing it as logistic regression since the values are between [0,1]) or directly minimizing the ‘potentials’ (see #5. Weighted loss functions)- did not work. I tried various schemes for direct injection of bpp to the attention with or without 2D conv, and it did help (at least +0.001 in validation, probably would be more if I submitted), but as I said before, I did not trusted the bpps and the public LB enough. </p>\n<h1>Validation scheme</h1>\n<p>During experimentations, I trained on the middle reactivities of the length 177 sequences (between nucleotides #36 and #116, since we were told that the private set might include data on the nucleotides on the edges). I validated on the length 206 sequences and the start/end of 177 length sequences (up to the #36 nucleotide and from the #116 one). For submission, I trained on sequences from all sizes, including the edges reactivities, and validated each model on a random part of the sequences I did not train on for the specific model.</p>\n<h1>Main points to take from my solution</h1>\n<p>My loss function and ensembling with sequences-only models might have a chance to improve the scores of the top models/ensembles a bit. However, I would not give it too high a chance since their scores are much above mine, and they probably captured most of what there was to capture already. Even so, I think that when we have more data (private test is 1M new sequences, and there are talks about producing ten times more data), bpps and CapR methods would be less and less useful while squeezing more data from the measurements (as I tried to do with my loss function) would become more useful than now. So I still hope that my work will be of some help.</p>\n<h1>Sources</h1>\n<p>I based my model on Hoyso's solution to the ISLR competition; see <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">HERE</a> and <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406978\" target=\"_blank\">HERE</a>.&nbsp;<br>\nI used <a href=\"https://www.kaggle.com/code/ratthachat/preprocessing-deep-learning-input-from-rna-string\" target=\"_blank\">ratthachat's notebook</a> to calculate the CapR values.<br>\nCheck my <a href=\"https://github.com/shlomoron/Stanford-Ribonanza-RNA-Folding-10th-place-solution\" target=\"_blank\">GitHub</a> for data preparation and training code.</p>",
  "messages": [
    {
      "id": 2573176,
      "postDate": "2023-12-24T19:17:14.327Z",
      "content": "<p>Initially, I had not intended to publish my solution since I used an interesting method for the loss (that I did not see in other solutions, although I might have missed it), and since I did not finish within the price range, I wanted to keep it to myself for some other future compatible competition. However, since apparently there are plans for a paper by the hosts, I eventually opted to publish it.</p>\n<h1>Context section</h1>\n<p><a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding\" target=\"_blank\">Business context</a>.<br>\n<a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">Data context</a>.</p>\n<h1>Overview of the Approach</h1>\n<p>I used a 1Dconv+transformer model. The model does not include positional encoding since the conv layers take care of positional relationship ‘automatically,’ allowing for straightforward generalization for longer sequences. I based my model on a previous work of <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> in the ISLR competition; see <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">HERE</a> and <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406978\" target=\"_blank\">HERE</a>. BTW, congrats for 2nd place. You know, Hoyeol is probably the person from whose work I learned the most. This solution is already the second gold I get, thanks to the knowledge I got from his work. So, this is a special thank you. <br>\nHoiso’s original model was causal in time; I changed it to be symmetric to both sides by employing suitable paddings and maskings. <br>\nI trained two kinds of models; model_1 was trained only on the sequences, and model_2 included CapR and bpp sum&amp;max per nucleotide.<br>\nThings that helped me (I did not perform a proper ablation study, so I will not give exact numbers most of the time):</p>\n<ol>\n<li>Clipping the predictions with a sigmoid (regular clipping also helps, but a sigmoid helps even more). It also made the convergence smoother.</li>\n<li>One-hot encoding of the nucleotides instead of embeddings (this is mainly for the seq-only model. When including more features, one-hot is a given)</li>\n<li>CapR</li>\n<li>BPP sum&amp;max<br>\nI also found the injection of bpps values to the attention helpful, even more so with 2D conv layers on the bpps in various schemes. HOWEVER…I was very very suspicious of bpps. I feared that they would be much worse on the private test set since it consists of new sequences, and I thought that even if it was good for the public set, maybe the BPPs models were developed on similar sequences, but for the private set, it would fail. So I chose only to include sum&amp;max per nucleotide, thinking that these ‘averaged’ values would be less prone to overfitting. Also, tbh, I did not have enough time at the end to do proper research and ensemble of the more complicated models due to trying too many ideas in too little time, lol.</li>\n<li>Weighted loss function:<br>\nLet's start from the end.<br>\nloss(reactivity, reactivity_error, pred)=MAE(reactivity, pred)*potentials  <br>\nwith:  <br>\npotentials = (1-2*dist.cdf(reactivity-|reactivity-pred|)  <br>\nand:  <br>\ndist = normal_distribution(mu=reactivity, sigma= reactivity_error)<br>\nNow, to explain why.<br>\nI thought a lot about the loss function. Since we were given the errors, I searched for an appropriate way to incorporate uncertainty in the measurements into the loss. I was a bit surprised when I could not find something useful- of course, I maybe did not dig deep enough, but at the very least, it seems that a lot of digging is required. I found several things about incorporating uncertainties into the predictions (i.e., predicting the uncertainties of the predictions), but I wanted the opposite.<br>\nNow, the obvious thing to do is directly weigh the loss by the error- to give some examples, Hoiso (2nd place) used log1p error, and I saw suggestions like 1/(1+err), etc. However, I needed more than this: it punished high and low errors too much. To understand why, let's talk for a second about the meaning of error- generally, it means that instead of knowing the exact value, we can only give an approximation for the true value, with a distribution of probability for it to be any number, and the probability given by normal distribution with sigma = error. When we say that reactivity=0 ± 1, we give a probability of 0.68 for the exact value to be between -1 to 1 (i.e., within one standard deviation). <br>\nLets assume that we have reactivity1 =0 ± 1 , reactivity2 =0 ± 3 , pred1 =pred2 =100 . If we normalize, for example, by 1/(1+err), and with MAE, then loss1 =50, loss2 = 25. So the first predictions contribute to the loss twice as much as the second one, even though, if we look at the distribution of probabilities- the first one has 0.68 to be in [-1,1] and the second one has 0.68 to be in [-3,3] (and 0.95 to be in [-6,6]) so for a prediction so far away from most of the distribution, I expect both predictions to contribute to the loss about the same. Of course, we clip the predictions between 0,1, making the analysis more complicated and depending on the exact loss and clipping each person applies. <br>\nOn the other hand, assume we have reactivity1=0 ± 0.01, rectivity2=0 ± 0.1, pred1=pred2=0.05, then we get loss1=0.0495, loss2=0.0454. This time, they are almost the same- even though we know that reactivity1 has a probability of 0.68 to be in [-0.01, 0.01], which is far from the prediction relatively to the prediction for reactivity2, which lies inside one standard deviation of the ground truth distribution, so I expect it to contribute much less to the loss. <br>\nOne way to address this issue is to perturb the reactivity values according to their probability distribution. However, this leads to unstabilized training since many values have a high error, and moreover, they are perturbed outside the range of [0,1], leading to more complications.<br>\nI wanted to properly incorporate our knowledge about the probability distribution of the values into the loss function. At this point, I took some inspiration from the shell theorem (check it <a href=\"https://en.wikipedia.org/wiki/Shell_theorem\" target=\"_blank\">on Wikipedia</a>), and by considering the probability distribution as a form of ‘potential,’ I finished with the above formula. <br>\nThis loss helped my convergence. I don’t remember precisely how much, maybe about 0.001? Not by a huge amount. But it certainly helped. It also made it even smoother. I think this method allows the inclusion of high error values while squeezing as much information as possible from them. Also, I believe this is better than trying to pseudo-label high-error values since noisy data is still data with real information inside. We just need to squeeze this information out…<br>\nBut wait, there is more. What if I suspect the errors we were given themselves are inaccurate?<br>\nNo problem. Just multiply the given errors by some factor, add some constant errors, and assemble many models with randomly chosen factors.<br>\nAt this point, I remembered the method of label smoothing for logistic regression, and, well, this is basically a form of label smoothing for regression. Once I framed it as label smoothing, it was easy to find similar ideas, e.g. <a href=\"https://proceedings.mlr.press/v80/imani18a/imani18a.pdf\" target=\"_blank\">this paper</a>. Although it seems like they overcomplicated things there a bit.<br>\nNote that this works the same for other losses, e.g., for MSE, multiply MSE by the same 'potentials,' etc.</li>\n</ol>\n<h1>Details of the submission</h1>\n<ol>\n<li><p>Ensembling (obviously):<br>\nHere is a nice plot of the power of ensembling for the sequence-only models (model_1)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F81ab8378d36c6bc5a08c64da32642008%2Fensembling_1.png?generation=1703442429686814&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Ensembling models with bpp+CapR with sequences-only models<br>\nI trusted my sequences-only models much more than the +bpp&amp;CapR ones, so I submitted one ensemble of pure sequences-only models. I planned the other ensemble to be only +bpp&amp;CapR models. But then, on the very last day, just when I had the two ensembles ready, I looked again at the pictures of the results for R1138v1 predictions. See more details <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653\" target=\"_blank\">here</a>. Basically, we wanted the models to replicate the following image:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F506500c54df757dd1b8046843f9a25e8%2FR1138v1.png?generation=1703442477381564&amp;alt=media\" alt=\"\"><br>\nHere is the picture of my sequence-only models ensemble:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F6a3fece98bcbd1a2deb22b7b5dfa52cc%2FR1138v1%20model_1.png?generation=1703442512137941&amp;alt=media\" alt=\"\"></p>\n<p>It’s a bit hard to see, but it recover the central line, especially in the 2A3 picture. However, the sidelines are missing. On the other hand, for the +bpp&amp;CapR models ensemble:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F8cd376a57c59e86e7bcb09917ef4e6e6%2FR1138v1%20model_2.png?generation=1703442537753029&amp;alt=media\" alt=\"\"><br>\nHere, it’s the opposite. I have visible traces of the sidelines but almost nothing for the central one (maybe one very vague point).<br>\nThese images bring an immediate idea…more ensembling…(this was not obvious at first since the BPP&amp;CapR models/ensemble had a much higher validation/public score than the sequences-only ones)<br>\nI had only three submissions left. I tried the following weights: 6:2, 5:3, and 4:4 in favor of a sequence-only ensemble. 4:4 had a better score but I was too suspisious of bpp and CapR, so I chose 5:3 weights which had a similar score to the original only bpp+CapR ensemble but was expected to be more robust.<br>\nAfter the competition ended, I tried other weights. Here is the complete analysis:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2Fa86e7996fb22dca69e22fe486ff72282%2Fensembling_2.png?generation=1703442561705283&amp;alt=media\" alt=\"\"><br>\nEven if I chose the best weight, I would still end outside the price range. So I don’t feel too terrible about missing here.  I could end up in 8th place, though, if I chose the 4:4 weights on the last day. Keep in mind that my solution does not include the edge information of the BPPs, as opposed to most (all?) other top solutions. Maybe, except the 3rd place, he got an extremely strong sequences-only solution, much better than mine.</p></li>\n</ol>\n<h1>More things I tried but did not included in the final solution or didn’t work out (possibly also due to lack of time):</h1>\n<p>Using the external data source as another head for the loss, using the external data source to train a model, predicting on the train set, and use the predictions as features, different predictions for bpps (eterna, contra, etc.) did not contribute too much or at all. Structures (i.e. (.) notation) also were not especially useful. I had plans for the 3D features but not enough time…I tried other schemes for the loss, for example, minimizing log loss (sigmoid) (i.e., framing it as logistic regression since the values are between [0,1]) or directly minimizing the ‘potentials’ (see #5. Weighted loss functions)- did not work. I tried various schemes for direct injection of bpp to the attention with or without 2D conv, and it did help (at least +0.001 in validation, probably would be more if I submitted), but as I said before, I did not trusted the bpps and the public LB enough. </p>\n<h1>Validation scheme</h1>\n<p>During experimentations, I trained on the middle reactivities of the length 177 sequences (between nucleotides #36 and #116, since we were told that the private set might include data on the nucleotides on the edges). I validated on the length 206 sequences and the start/end of 177 length sequences (up to the #36 nucleotide and from the #116 one). For submission, I trained on sequences from all sizes, including the edges reactivities, and validated each model on a random part of the sequences I did not train on for the specific model.</p>\n<h1>Main points to take from my solution</h1>\n<p>My loss function and ensembling with sequences-only models might have a chance to improve the scores of the top models/ensembles a bit. However, I would not give it too high a chance since their scores are much above mine, and they probably captured most of what there was to capture already. Even so, I think that when we have more data (private test is 1M new sequences, and there are talks about producing ten times more data), bpps and CapR methods would be less and less useful while squeezing more data from the measurements (as I tried to do with my loss function) would become more useful than now. So I still hope that my work will be of some help.</p>\n<h1>Sources</h1>\n<p>I based my model on Hoyso's solution to the ISLR competition; see <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">HERE</a> and <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406978\" target=\"_blank\">HERE</a>.&nbsp;<br>\nI used <a href=\"https://www.kaggle.com/code/ratthachat/preprocessing-deep-learning-input-from-rna-string\" target=\"_blank\">ratthachat's notebook</a> to calculate the CapR values.<br>\nCheck my <a href=\"https://github.com/shlomoron/Stanford-Ribonanza-RNA-Folding-10th-place-solution\" target=\"_blank\">GitHub</a> for data preparation and training code.</p>",
      "rawMarkdown": "Initially, I had not intended to publish my solution since I used an interesting method for the loss (that I did not see in other solutions, although I might have missed it), and since I did not finish within the price range, I wanted to keep it to myself for some other future compatible competition. However, since apparently there are plans for a paper by the hosts, I eventually opted to publish it.\n\n# Context section\n[Business context](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding).\n[Data context](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data).\n\n# Overview of the Approach\nI used a 1Dconv+transformer model. The model does not include positional encoding since the conv layers take care of positional relationship ‘automatically,’ allowing for straightforward generalization for longer sequences. I based my model on a previous work of @hoyso48 in the ISLR competition; see [HERE](https://www.kaggle.com/competitions/asl-signs/discussion/406684) and [HERE](https://www.kaggle.com/competitions/asl-signs/discussion/406978). BTW, congrats for 2nd place. You know, Hoyeol is probably the person from whose work I learned the most. This solution is already the second gold I get, thanks to the knowledge I got from his work. So, this is a special thank you. \nHoiso’s original model was causal in time; I changed it to be symmetric to both sides by employing suitable paddings and maskings. \nI trained two kinds of models; model_1 was trained only on the sequences, and model_2 included CapR and bpp sum&max per nucleotide.\nThings that helped me (I did not perform a proper ablation study, so I will not give exact numbers most of the time):\n1. Clipping the predictions with a sigmoid (regular clipping also helps, but a sigmoid helps even more). It also made the convergence smoother.\n2. One-hot encoding of the nucleotides instead of embeddings (this is mainly for the seq-only model. When including more features, one-hot is a given)\n3. CapR\n4. BPP sum&max\nI also found the injection of bpps values to the attention helpful, even more so with 2D conv layers on the bpps in various schemes. HOWEVER...I was very very suspicious of bpps. I feared that they would be much worse on the private test set since it consists of new sequences, and I thought that even if it was good for the public set, maybe the BPPs models were developed on similar sequences, but for the private set, it would fail. So I chose only to include sum&max per nucleotide, thinking that these ‘averaged’ values would be less prone to overfitting. Also, tbh, I did not have enough time at the end to do proper research and ensemble of the more complicated models due to trying too many ideas in too little time, lol.\n5. Weighted loss function:\nLet's start from the end.\nloss(reactivity, reactivity_error, pred)=MAE(reactivity, pred)\\*potentials  \nwith:  \npotentials = (1-2\\*dist.cdf(reactivity-|reactivity-pred|)  \nand:  \ndist = normal_distribution(mu=reactivity, sigma= reactivity_error)\nNow, to explain why.\nI thought a lot about the loss function. Since we were given the errors, I searched for an appropriate way to incorporate uncertainty in the measurements into the loss. I was a bit surprised when I could not find something useful- of course, I maybe did not dig deep enough, but at the very least, it seems that a lot of digging is required. I found several things about incorporating uncertainties into the predictions (i.e., predicting the uncertainties of the predictions), but I wanted the opposite.\nNow, the obvious thing to do is directly weigh the loss by the error- to give some examples, Hoiso (2nd place) used log1p error, and I saw suggestions like 1/(1+err), etc. However, I needed more than this: it punished high and low errors too much. To understand why, let's talk for a second about the meaning of error- generally, it means that instead of knowing the exact value, we can only give an approximation for the true value, with a distribution of probability for it to be any number, and the probability given by normal distribution with sigma = error. When we say that reactivity=0 ± 1, we give a probability of 0.68 for the exact value to be between -1 to 1 (i.e., within one standard deviation). \nLets assume that we have reactivity1 =0 ± 1 , reactivity2 =0 ± 3 , pred1 =pred2 =100 . If we normalize, for example, by 1/(1+err), and with MAE, then loss1 =50, loss2 = 25. So the first predictions contribute to the loss twice as much as the second one, even though, if we look at the distribution of probabilities- the first one has 0.68 to be in [-1,1] and the second one has 0.68 to be in [-3,3] (and 0.95 to be in [-6,6]) so for a prediction so far away from most of the distribution, I expect both predictions to contribute to the loss about the same. Of course, we clip the predictions between 0,1, making the analysis more complicated and depending on the exact loss and clipping each person applies. \nOn the other hand, assume we have reactivity1=0 ± 0.01, rectivity2=0 ± 0.1, pred1=pred2=0.05, then we get loss1=0.0495, loss2=0.0454. This time, they are almost the same- even though we know that reactivity1 has a probability of 0.68 to be in [-0.01, 0.01], which is far from the prediction relatively to the prediction for reactivity2, which lies inside one standard deviation of the ground truth distribution, so I expect it to contribute much less to the loss. \nOne way to address this issue is to perturb the reactivity values according to their probability distribution. However, this leads to unstabilized training since many values have a high error, and moreover, they are perturbed outside the range of [0,1], leading to more complications.\nI wanted to properly incorporate our knowledge about the probability distribution of the values into the loss function. At this point, I took some inspiration from the shell theorem (check it [on Wikipedia](https://en.wikipedia.org/wiki/Shell_theorem)), and by considering the probability distribution as a form of ‘potential,’ I finished with the above formula. \nThis loss helped my convergence. I don’t remember precisely how much, maybe about 0.001? Not by a huge amount. But it certainly helped. It also made it even smoother. I think this method allows the inclusion of high error values while squeezing as much information as possible from them. Also, I believe this is better than trying to pseudo-label high-error values since noisy data is still data with real information inside. We just need to squeeze this information out…\nBut wait, there is more. What if I suspect the errors we were given themselves are inaccurate?\nNo problem. Just multiply the given errors by some factor, add some constant errors, and assemble many models with randomly chosen factors.\nAt this point, I remembered the method of label smoothing for logistic regression, and, well, this is basically a form of label smoothing for regression. Once I framed it as label smoothing, it was easy to find similar ideas, e.g. [this paper](https://proceedings.mlr.press/v80/imani18a/imani18a.pdf). Although it seems like they overcomplicated things there a bit.\nNote that this works the same for other losses, e.g., for MSE, multiply MSE by the same 'potentials,' etc.\n\n# Details of the submission\n1. Ensembling (obviously):\nHere is a nice plot of the power of ensembling for the sequence-only models (model_1)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F81ab8378d36c6bc5a08c64da32642008%2Fensembling_1.png?generation=1703442429686814&alt=media)\n\n2. Ensembling models with bpp+CapR with sequences-only models\nI trusted my sequences-only models much more than the +bpp&CapR ones, so I submitted one ensemble of pure sequences-only models. I planned the other ensemble to be only +bpp&CapR models. But then, on the very last day, just when I had the two ensembles ready, I looked again at the pictures of the results for R1138v1 predictions. See more details [here](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653). Basically, we wanted the models to replicate the following image:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F506500c54df757dd1b8046843f9a25e8%2FR1138v1.png?generation=1703442477381564&alt=media)\nHere is the picture of my sequence-only models ensemble:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F6a3fece98bcbd1a2deb22b7b5dfa52cc%2FR1138v1%20model_1.png?generation=1703442512137941&alt=media)\n\n It’s a bit hard to see, but it recover the central line, especially in the 2A3 picture. However, the sidelines are missing. On the other hand, for the +bpp&CapR models ensemble:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F8cd376a57c59e86e7bcb09917ef4e6e6%2FR1138v1%20model_2.png?generation=1703442537753029&alt=media)\nHere, it’s the opposite. I have visible traces of the sidelines but almost nothing for the central one (maybe one very vague point).\nThese images bring an immediate idea...more ensembling…(this was not obvious at first since the BPP&CapR models/ensemble had a much higher validation/public score than the sequences-only ones)\nI had only three submissions left. I tried the following weights: 6:2, 5:3, and 4:4 in favor of a sequence-only ensemble. 4:4 had a better score but I was too suspisious of bpp and CapR, so I chose 5:3 weights which had a similar score to the original only bpp+CapR ensemble but was expected to be more robust.\nAfter the competition ended, I tried other weights. Here is the complete analysis:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2Fa86e7996fb22dca69e22fe486ff72282%2Fensembling_2.png?generation=1703442561705283&alt=media)\nEven if I chose the best weight, I would still end outside the price range. So I don’t feel too terrible about missing here.  I could end up in 8th place, though, if I chose the 4:4 weights on the last day. Keep in mind that my solution does not include the edge information of the BPPs, as opposed to most (all?) other top solutions. Maybe, except the 3rd place, he got an extremely strong sequences-only solution, much better than mine.\n# More things I tried but did not included in the final solution or didn’t work out (possibly also due to lack of time):\nUsing the external data source as another head for the loss, using the external data source to train a model, predicting on the train set, and use the predictions as features, different predictions for bpps (eterna, contra, etc.) did not contribute too much or at all. Structures (i.e. (.) notation) also were not especially useful. I had plans for the 3D features but not enough time...I tried other schemes for the loss, for example, minimizing log loss (sigmoid) (i.e., framing it as logistic regression since the values are between [0,1]) or directly minimizing the ‘potentials’ (see #5. Weighted loss functions)- did not work. I tried various schemes for direct injection of bpp to the attention with or without 2D conv, and it did help (at least +0.001 in validation, probably would be more if I submitted), but as I said before, I did not trusted the bpps and the public LB enough. \n\n#Validation scheme\nDuring experimentations, I trained on the middle reactivities of the length 177 sequences (between nucleotides #36 and #116, since we were told that the private set might include data on the nucleotides on the edges). I validated on the length 206 sequences and the start/end of 177 length sequences (up to the #36 nucleotide and from the #116 one). For submission, I trained on sequences from all sizes, including the edges reactivities, and validated each model on a random part of the sequences I did not train on for the specific model.\n\n# Main points to take from my solution\nMy loss function and ensembling with sequences-only models might have a chance to improve the scores of the top models/ensembles a bit. However, I would not give it too high a chance since their scores are much above mine, and they probably captured most of what there was to capture already. Even so, I think that when we have more data (private test is 1M new sequences, and there are talks about producing ten times more data), bpps and CapR methods would be less and less useful while squeezing more data from the measurements (as I tried to do with my loss function) would become more useful than now. So I still hope that my work will be of some help.\n\n# Sources\nI based my model on Hoyso's solution to the ISLR competition; see [HERE](https://www.kaggle.com/competitions/asl-signs/discussion/406684) and [HERE](https://www.kaggle.com/competitions/asl-signs/discussion/406978). \nI used [ratthachat's notebook](https://www.kaggle.com/code/ratthachat/preprocessing-deep-learning-input-from-rna-string) to calculate the CapR values.\nCheck my [GitHub](https://github.com/shlomoron/Stanford-Ribonanza-RNA-Folding-10th-place-solution) for data preparation and training code.",
      "votes": 13
    },
    {
      "id": 2575030,
      "postDate": "2023-12-26T13:30:39.913Z",
      "content": "<p>Really nice post. And link to the paper. We used histogram loss for DREAM-2022 challenge and hadn't known its name was histogram loss and it has some theoretical guarantees. Definitely a thing to try in ribonanza setup</p>",
      "rawMarkdown": "Really nice post. And link to the paper. We used histogram loss for DREAM-2022 challenge and hadn't known its name was histogram loss and it has some theoretical guarantees. Definitely a thing to try in ribonanza setup",
      "votes": 1,
      "replies": [
        {
          "id": 2575098,
          "postDate": "2023-12-26T15:01:53.470Z",
          "content": "<p>Thank you. It can be interesting if you try this loss on your model. However, your score is already so high, so it's hard to tell if it would contribute anything with the current data. Maybe with some tinkering.</p>",
          "rawMarkdown": "Thank you. It can be interesting if you try this loss on your model. However, your score is already so high, so it's hard to tell if it would contribute anything with the current data. Maybe with some tinkering."
        }
      ]
    },
    {
      "id": 2947470,
      "postDate": "2024-08-05T07:57:45.370Z",
      "content": "<p>impressive</p>",
      "rawMarkdown": "impressive",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2575030,
      "author_name": "Penzar Dmitry",
      "author_url": "",
      "post_date": "2023-12-26T13:30:39.913000",
      "content": "<p>Really nice post. And link to the paper. We used histogram loss for DREAM-2022 challenge and hadn't known its name was histogram loss and it has some theoretical guarantees. Definitely a thing to try in ribonanza setup</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2575098,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2023-12-26T15:01:53.470000",
          "content": "<p>Thank you. It can be interesting if you try this loss on your model. However, your score is already so high, so it's hard to tell if it would contribute anything with the current data. Maybe with some tinkering.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2947470,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-08-05T07:57:45.370000",
      "content": "<p>impressive</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2573176": "Initially, I had not intended to publish my solution since I used an interesting method for the loss (that I did not see in other solutions, although I might have missed it), and since I did not finish within the price range, I wanted to keep it to myself for some other future compatible competition. However, since apparently there are plans for a paper by the hosts, I eventually opted to publish it.\n\n# Context section\n[Business context](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding).\n[Data context](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data).\n\n# Overview of the Approach\nI used a 1Dconv+transformer model. The model does not include positional encoding since the conv layers take care of positional relationship ‘automatically,’ allowing for straightforward generalization for longer sequences. I based my model on a previous work of @hoyso48 in the ISLR competition; see [HERE](https://www.kaggle.com/competitions/asl-signs/discussion/406684) and [HERE](https://www.kaggle.com/competitions/asl-signs/discussion/406978). BTW, congrats for 2nd place. You know, Hoyeol is probably the person from whose work I learned the most. This solution is already the second gold I get, thanks to the knowledge I got from his work. So, this is a special thank you. \nHoiso’s original model was causal in time; I changed it to be symmetric to both sides by employing suitable paddings and maskings. \nI trained two kinds of models; model_1 was trained only on the sequences, and model_2 included CapR and bpp sum&max per nucleotide.\nThings that helped me (I did not perform a proper ablation study, so I will not give exact numbers most of the time):\n1. Clipping the predictions with a sigmoid (regular clipping also helps, but a sigmoid helps even more). It also made the convergence smoother.\n2. One-hot encoding of the nucleotides instead of embeddings (this is mainly for the seq-only model. When including more features, one-hot is a given)\n3. CapR\n4. BPP sum&max\nI also found the injection of bpps values to the attention helpful, even more so with 2D conv layers on the bpps in various schemes. HOWEVER...I was very very suspicious of bpps. I feared that they would be much worse on the private test set since it consists of new sequences, and I thought that even if it was good for the public set, maybe the BPPs models were developed on similar sequences, but for the private set, it would fail. So I chose only to include sum&max per nucleotide, thinking that these ‘averaged’ values would be less prone to overfitting. Also, tbh, I did not have enough time at the end to do proper research and ensemble of the more complicated models due to trying too many ideas in too little time, lol.\n5. Weighted loss function:\nLet's start from the end.\nloss(reactivity, reactivity_error, pred)=MAE(reactivity, pred)\\*potentials  \nwith:  \npotentials = (1-2\\*dist.cdf(reactivity-|reactivity-pred|)  \nand:  \ndist = normal_distribution(mu=reactivity, sigma= reactivity_error)\nNow, to explain why.\nI thought a lot about the loss function. Since we were given the errors, I searched for an appropriate way to incorporate uncertainty in the measurements into the loss. I was a bit surprised when I could not find something useful- of course, I maybe did not dig deep enough, but at the very least, it seems that a lot of digging is required. I found several things about incorporating uncertainties into the predictions (i.e., predicting the uncertainties of the predictions), but I wanted the opposite.\nNow, the obvious thing to do is directly weigh the loss by the error- to give some examples, Hoiso (2nd place) used log1p error, and I saw suggestions like 1/(1+err), etc. However, I needed more than this: it punished high and low errors too much. To understand why, let's talk for a second about the meaning of error- generally, it means that instead of knowing the exact value, we can only give an approximation for the true value, with a distribution of probability for it to be any number, and the probability given by normal distribution with sigma = error. When we say that reactivity=0 ± 1, we give a probability of 0.68 for the exact value to be between -1 to 1 (i.e., within one standard deviation). \nLets assume that we have reactivity1 =0 ± 1 , reactivity2 =0 ± 3 , pred1 =pred2 =100 . If we normalize, for example, by 1/(1+err), and with MAE, then loss1 =50, loss2 = 25. So the first predictions contribute to the loss twice as much as the second one, even though, if we look at the distribution of probabilities- the first one has 0.68 to be in [-1,1] and the second one has 0.68 to be in [-3,3] (and 0.95 to be in [-6,6]) so for a prediction so far away from most of the distribution, I expect both predictions to contribute to the loss about the same. Of course, we clip the predictions between 0,1, making the analysis more complicated and depending on the exact loss and clipping each person applies. \nOn the other hand, assume we have reactivity1=0 ± 0.01, rectivity2=0 ± 0.1, pred1=pred2=0.05, then we get loss1=0.0495, loss2=0.0454. This time, they are almost the same- even though we know that reactivity1 has a probability of 0.68 to be in [-0.01, 0.01], which is far from the prediction relatively to the prediction for reactivity2, which lies inside one standard deviation of the ground truth distribution, so I expect it to contribute much less to the loss. \nOne way to address this issue is to perturb the reactivity values according to their probability distribution. However, this leads to unstabilized training since many values have a high error, and moreover, they are perturbed outside the range of [0,1], leading to more complications.\nI wanted to properly incorporate our knowledge about the probability distribution of the values into the loss function. At this point, I took some inspiration from the shell theorem (check it [on Wikipedia](https://en.wikipedia.org/wiki/Shell_theorem)), and by considering the probability distribution as a form of ‘potential,’ I finished with the above formula. \nThis loss helped my convergence. I don’t remember precisely how much, maybe about 0.001? Not by a huge amount. But it certainly helped. It also made it even smoother. I think this method allows the inclusion of high error values while squeezing as much information as possible from them. Also, I believe this is better than trying to pseudo-label high-error values since noisy data is still data with real information inside. We just need to squeeze this information out…\nBut wait, there is more. What if I suspect the errors we were given themselves are inaccurate?\nNo problem. Just multiply the given errors by some factor, add some constant errors, and assemble many models with randomly chosen factors.\nAt this point, I remembered the method of label smoothing for logistic regression, and, well, this is basically a form of label smoothing for regression. Once I framed it as label smoothing, it was easy to find similar ideas, e.g. [this paper](https://proceedings.mlr.press/v80/imani18a/imani18a.pdf). Although it seems like they overcomplicated things there a bit.\nNote that this works the same for other losses, e.g., for MSE, multiply MSE by the same 'potentials,' etc.\n\n# Details of the submission\n1. Ensembling (obviously):\nHere is a nice plot of the power of ensembling for the sequence-only models (model_1)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F81ab8378d36c6bc5a08c64da32642008%2Fensembling_1.png?generation=1703442429686814&alt=media)\n\n2. Ensembling models with bpp+CapR with sequences-only models\nI trusted my sequences-only models much more than the +bpp&CapR ones, so I submitted one ensemble of pure sequences-only models. I planned the other ensemble to be only +bpp&CapR models. But then, on the very last day, just when I had the two ensembles ready, I looked again at the pictures of the results for R1138v1 predictions. See more details [here](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653). Basically, we wanted the models to replicate the following image:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F506500c54df757dd1b8046843f9a25e8%2FR1138v1.png?generation=1703442477381564&alt=media)\nHere is the picture of my sequence-only models ensemble:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F6a3fece98bcbd1a2deb22b7b5dfa52cc%2FR1138v1%20model_1.png?generation=1703442512137941&alt=media)\n\n It’s a bit hard to see, but it recover the central line, especially in the 2A3 picture. However, the sidelines are missing. On the other hand, for the +bpp&CapR models ensemble:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F8cd376a57c59e86e7bcb09917ef4e6e6%2FR1138v1%20model_2.png?generation=1703442537753029&alt=media)\nHere, it’s the opposite. I have visible traces of the sidelines but almost nothing for the central one (maybe one very vague point).\nThese images bring an immediate idea...more ensembling…(this was not obvious at first since the BPP&CapR models/ensemble had a much higher validation/public score than the sequences-only ones)\nI had only three submissions left. I tried the following weights: 6:2, 5:3, and 4:4 in favor of a sequence-only ensemble. 4:4 had a better score but I was too suspisious of bpp and CapR, so I chose 5:3 weights which had a similar score to the original only bpp+CapR ensemble but was expected to be more robust.\nAfter the competition ended, I tried other weights. Here is the complete analysis:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2Fa86e7996fb22dca69e22fe486ff72282%2Fensembling_2.png?generation=1703442561705283&alt=media)\nEven if I chose the best weight, I would still end outside the price range. So I don’t feel too terrible about missing here.  I could end up in 8th place, though, if I chose the 4:4 weights on the last day. Keep in mind that my solution does not include the edge information of the BPPs, as opposed to most (all?) other top solutions. Maybe, except the 3rd place, he got an extremely strong sequences-only solution, much better than mine.\n# More things I tried but did not included in the final solution or didn’t work out (possibly also due to lack of time):\nUsing the external data source as another head for the loss, using the external data source to train a model, predicting on the train set, and use the predictions as features, different predictions for bpps (eterna, contra, etc.) did not contribute too much or at all. Structures (i.e. (.) notation) also were not especially useful. I had plans for the 3D features but not enough time...I tried other schemes for the loss, for example, minimizing log loss (sigmoid) (i.e., framing it as logistic regression since the values are between [0,1]) or directly minimizing the ‘potentials’ (see #5. Weighted loss functions)- did not work. I tried various schemes for direct injection of bpp to the attention with or without 2D conv, and it did help (at least +0.001 in validation, probably would be more if I submitted), but as I said before, I did not trusted the bpps and the public LB enough. \n\n#Validation scheme\nDuring experimentations, I trained on the middle reactivities of the length 177 sequences (between nucleotides #36 and #116, since we were told that the private set might include data on the nucleotides on the edges). I validated on the length 206 sequences and the start/end of 177 length sequences (up to the #36 nucleotide and from the #116 one). For submission, I trained on sequences from all sizes, including the edges reactivities, and validated each model on a random part of the sequences I did not train on for the specific model.\n\n# Main points to take from my solution\nMy loss function and ensembling with sequences-only models might have a chance to improve the scores of the top models/ensembles a bit. However, I would not give it too high a chance since their scores are much above mine, and they probably captured most of what there was to capture already. Even so, I think that when we have more data (private test is 1M new sequences, and there are talks about producing ten times more data), bpps and CapR methods would be less and less useful while squeezing more data from the measurements (as I tried to do with my loss function) would become more useful than now. So I still hope that my work will be of some help.\n\n# Sources\nI based my model on Hoyso's solution to the ISLR competition; see [HERE](https://www.kaggle.com/competitions/asl-signs/discussion/406684) and [HERE](https://www.kaggle.com/competitions/asl-signs/discussion/406978). \nI used [ratthachat's notebook](https://www.kaggle.com/code/ratthachat/preprocessing-deep-learning-input-from-rna-string) to calculate the CapR values.\nCheck my [GitHub](https://github.com/shlomoron/Stanford-Ribonanza-RNA-Folding-10th-place-solution) for data preparation and training code.",
    "2575030": "Really nice post. And link to the paper. We used histogram loss for DREAM-2022 challenge and hadn't known its name was histogram loss and it has some theoretical guarantees. Definitely a thing to try in ribonanza setup",
    "2947470": "impressive"
  }
}