{
  "id": 462850,
  "title": "Help us train RibonanzaNet, a single model that integrates all of our efforts! [NEEDED BY DEC 24 2023: Labels for train sequences]",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/462850",
  "author_name": "Rhiju Das",
  "post_date": "2023-12-22T00:26:27.170000",
  "votes": 5,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> and I are working to prepare our models for use by the broader research community.   </p>\n<p>For the broader community, it would be wonderful to have a single model that might integrate and perhaps outperform any previous single model in Ribonanza. </p>\n<p>As part of this effort, it would be helpful to have labels of your models for the train data sequences. </p>\n<p>Yes, that seems weird, but as you know, most of the experimental train data are noisy, and we expect that model predictions will potentially be more accurate than the data themselves!</p>\n<p>We're making the train data sequences available as <code>train_sequences.csv</code> file here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/rhijudas/ribonanza-train-sequences/data\" target=\"_blank\">https://www.kaggle.com/datasets/rhijudas/ribonanza-train-sequences/data</a></p>\n<p>To join us in this research effort, would you be willing to re-run your notebooks to generate sample submissions for these 806,573 sequences, upload as a public data set, and provide link in this thread?  </p>\n<p>If running a large ensemble of models is painful, having the outputs of one or two of the best models you developed would still be helpful.  </p>\n<p>It would be most useful to have the predictions in the next few days, by <em>Dec. 24, 2023</em>, so we can train over the holidays when we have priority access to compute. 😉 </p>\n<p>Gratefully, for the hosts,<br>\nRhiju</p>",
  "messages": [
    {
      "id": 2570145,
      "postDate": "2023-12-22T00:26:27.170Z",
      "content": "<p>Hi everyone!</p>\n<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> and I are working to prepare our models for use by the broader research community.   </p>\n<p>For the broader community, it would be wonderful to have a single model that might integrate and perhaps outperform any previous single model in Ribonanza. </p>\n<p>As part of this effort, it would be helpful to have labels of your models for the train data sequences. </p>\n<p>Yes, that seems weird, but as you know, most of the experimental train data are noisy, and we expect that model predictions will potentially be more accurate than the data themselves!</p>\n<p>We're making the train data sequences available as <code>train_sequences.csv</code> file here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/rhijudas/ribonanza-train-sequences/data\" target=\"_blank\">https://www.kaggle.com/datasets/rhijudas/ribonanza-train-sequences/data</a></p>\n<p>To join us in this research effort, would you be willing to re-run your notebooks to generate sample submissions for these 806,573 sequences, upload as a public data set, and provide link in this thread?  </p>\n<p>If running a large ensemble of models is painful, having the outputs of one or two of the best models you developed would still be helpful.  </p>\n<p>It would be most useful to have the predictions in the next few days, by <em>Dec. 24, 2023</em>, so we can train over the holidays when we have priority access to compute. 😉 </p>\n<p>Gratefully, for the hosts,<br>\nRhiju</p>",
      "rawMarkdown": "Hi everyone!\n\n@shujun717 and I are working to prepare our models for use by the broader research community.   \n\nFor the broader community, it would be wonderful to have a single model that might integrate and perhaps outperform any previous single model in Ribonanza. \n\nAs part of this effort, it would be helpful to have labels of your models for the train data sequences. \n\nYes, that seems weird, but as you know, most of the experimental train data are noisy, and we expect that model predictions will potentially be more accurate than the data themselves!\n\nWe're making the train data sequences available as `train_sequences.csv` file here:\n\nhttps://www.kaggle.com/datasets/rhijudas/ribonanza-train-sequences/data\n\nTo join us in this research effort, would you be willing to re-run your notebooks to generate sample submissions for these 806,573 sequences, upload as a public data set, and provide link in this thread?  \n\nIf running a large ensemble of models is painful, having the outputs of one or two of the best models you developed would still be helpful.  \n\nIt would be most useful to have the predictions in the next few days, by *Dec. 24, 2023*, so we can train over the holidays when we have priority access to compute. 😉 \n\nGratefully, for the hosts,\nRhiju\n",
      "votes": 5
    },
    {
      "id": 2571292,
      "postDate": "2023-12-23T05:07:28.077Z",
      "content": "<p>I uploaded predictions for my blend <a href=\"https://www.kaggle.com/datasets/iafoss/stanford-ribonanza-rna-folding-train-predictions\" target=\"_blank\">here</a>. It should be ~0.142 at private LB. Following the request from the post, it is not OOF predictions. So, these predictions may not be that beneficial since the model tries to memorize training data, while a better filtering could be achieved with OOF.</p>",
      "rawMarkdown": "I uploaded predictions for my blend [here](https://www.kaggle.com/datasets/iafoss/stanford-ribonanza-rna-folding-train-predictions). It should be ~0.142 at private LB. Following the request from the post, it is not OOF predictions. So, these predictions may not be that beneficial since the model tries to memorize training data, while a better filtering could be achieved with OOF.",
      "votes": 1,
      "replies": [
        {
          "id": 2571795,
          "postDate": "2023-12-23T15:33:08.367Z",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> did you use data with low signal to noise (&lt;1) for training? If not, your model predictions on those would be very useful as extra pl data.  And even if you did use them in some capacity, your predictions might be still be more accurate than the original noisy labels</p>",
          "rawMarkdown": "@iafoss did you use data with low signal to noise (<1) for training? If not, your model predictions on those would be very useful as extra pl data.  And even if you did use them in some capacity, your predictions might be still be more accurate than the original noisy labels",
          "replies": [
            {
              "id": 2571800,
              "postDate": "2023-12-23T15:38:17.607Z",
              "content": "<p>I did but with weighed loss. </p>",
              "rawMarkdown": "I did but with weighed loss. "
            },
            {
              "id": 2571831,
              "postDate": "2023-12-23T16:36:44.833Z",
              "content": "<p>By the way, we've calculated the correlation between reactivities predicted for train dataset and real reactivities. <br>\nFor clean data the correlation is 0.73 and 0.76 for  DMS and 2A3 respectively, <br>\nfor all data - 0.41 and 0.28. <br>\nSo, it seems that our model wasn't good at memorizing data and we've got about the same correlations for your dataset.   </p>",
              "rawMarkdown": "By the way, we've calculated the correlation between reactivities predicted for train dataset and real reactivities. \nFor clean data the correlation is 0.73 and 0.76 for  DMS and 2A3 respectively, \nfor all data - 0.41 and 0.28. \nSo, it seems that our model wasn't good at memorizing data and we've got about the same correlations for your dataset.   "
            },
            {
              "id": 2571832,
              "postDate": "2023-12-23T16:38:35.307Z",
              "content": "<p>Sounds good, thanks for checking.</p>",
              "rawMarkdown": "Sounds good, thanks for checking."
            },
            {
              "id": 2571923,
              "postDate": "2023-12-23T18:03:37.490Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/dmitrypenzar1996\" target=\"_blank\">@dmitrypenzar1996</a> . I think it might be because everyone is using MAE loss which has gradients of +1/-1 at each position so the model does not end up memorizing labels as much as MSE loss whose gradients are proportional to distance to target</p>",
              "rawMarkdown": "Thanks @dmitrypenzar1996 . I think it might be because everyone is using MAE loss which has gradients of +1/-1 at each position so the model does not end up memorizing labels as much as MSE loss whose gradients are proportional to distance to target"
            }
          ]
        }
      ]
    },
    {
      "id": 2573360,
      "postDate": "2023-12-24T23:19:53.510Z",
      "content": "<p>Train Labels predicted with the submission models are uploaded <a href=\"https://www.kaggle.com/datasets/dankrstev/stanford-ribonanza-rna-folding-train-labels\" target=\"_blank\">here.</a></p>\n<p>Sorry for the delay, hopefully I uploaded them on time.</p>",
      "rawMarkdown": "Train Labels predicted with the submission models are uploaded [here.](https://www.kaggle.com/datasets/dankrstev/stanford-ribonanza-rna-folding-train-labels)\n\nSorry for the delay, hopefully I uploaded them on time.",
      "replies": [
        {
          "id": 2574389,
          "postDate": "2023-12-25T22:00:15.957Z",
          "content": "<p>i can't see this. did you make it public?</p>",
          "rawMarkdown": "i can't see this. did you make it public?",
          "replies": [
            {
              "id": 2574711,
              "postDate": "2023-12-26T07:18:05.103Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2572641,
      "postDate": "2023-12-24T12:32:18.617Z",
      "content": "<p>I uploaded mine <a href=\"https://www.kaggle.com/datasets/hoyso48/ribonanza-train-sequences-predictions\" target=\"_blank\">here</a>. Hope it's not too late.</p>\n<p>For the predictions I used the 4 identical models trained on the whole train dataset each with different seed, which scored 0.137 on public LB and 0.140 on private LB.</p>",
      "rawMarkdown": "I uploaded mine [here](https://www.kaggle.com/datasets/hoyso48/ribonanza-train-sequences-predictions). Hope it's not too late.\n\nFor the predictions I used the 4 identical models trained on the whole train dataset each with different seed, which scored 0.137 on public LB and 0.140 on private LB."
    },
    {
      "id": 2571948,
      "postDate": "2023-12-23T18:31:19.990Z",
      "content": "<p>As you requested, here are my predictions. My final ensemble had two kinds of models: the first one was trained only on the sequences, and the second one included CapR and bpp sum&amp;max across the columns. An ensemble of about forty models with equal weights to each kind of model achieved 0.14066/0.14238 public/private LB. (which would place me in 8th place. My final score is lower since I was too suspicious of CapR/bpp and chose 5:3 weights in favor of only train data models with 0.14133/0.14304 public/private LB). Here, I give a blend of 4 models from each kind. As others have pointed out, making OOF predictions for PL would probably be better. Also, I suspect using the original data would be better- noisy data is still data. Since we know the error, it's not wrong data that we need to 'fix,' so it should have more information than PL and give better results if handled correctly. I will write more about it in my write-up. I hope to finish it soon.<br>\n<a href=\"https://www.kaggle.com/datasets/shlomoron/srrf-model-1-train-preds\" target=\"_blank\">model 1 (only train)</a><br>\n<a href=\"https://www.kaggle.com/datasets/shlomoron/srrf-model-2-train-preds\" target=\"_blank\">model 2 (+CapR, bpp sum&amp;max)</a></p>",
      "rawMarkdown": "As you requested, here are my predictions. My final ensemble had two kinds of models: the first one was trained only on the sequences, and the second one included CapR and bpp sum&max across the columns. An ensemble of about forty models with equal weights to each kind of model achieved 0.14066/0.14238 public/private LB. (which would place me in 8th place. My final score is lower since I was too suspicious of CapR/bpp and chose 5:3 weights in favor of only train data models with 0.14133/0.14304 public/private LB). Here, I give a blend of 4 models from each kind. As others have pointed out, making OOF predictions for PL would probably be better. Also, I suspect using the original data would be better- noisy data is still data. Since we know the error, it's not wrong data that we need to 'fix,' so it should have more information than PL and give better results if handled correctly. I will write more about it in my write-up. I hope to finish it soon.\n[model 1 (only train)](https://www.kaggle.com/datasets/shlomoron/srrf-model-1-train-preds)\n[model 2 (+CapR, bpp sum&max)](https://www.kaggle.com/datasets/shlomoron/srrf-model-2-train-preds)"
    },
    {
      "id": 2571876,
      "postDate": "2023-12-23T17:20:08.840Z",
      "content": "<p>Uploaded our predictions <a href=\"https://www.kaggle.com/datasets/bacterio/ribonanza-train-dataset-predictions\" target=\"_blank\">here</a>. FWIW, this is a plot of <code>plt.hist(np.abs(preds_not_nan - targs_not_nan).clip(0, 0.25), bins=100, log=True)</code>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2Fd9164276feafc0f13b4dc43f5164cfb1%2Ftrain_hist.png?generation=1703351640892991&amp;alt=media\" alt=\"\"></p>\n<p>50% of the predictions have an absolute error &lt;0.0457. Other percentile stats of <code>np.abs(preds_not_nan - targs_not_nan)</code>:</p>\n<pre><code>  : .\n : .\n : .\n : .\n : .\n</code></pre>\n<p>Both <code>preds_not_nan</code> and <code>targs_not_nan</code> were clipped to 0, 1 prior to computing the above stats. Hope this makes sense and works for you.</p>",
      "rawMarkdown": "Uploaded our predictions [here](https://www.kaggle.com/datasets/bacterio/ribonanza-train-dataset-predictions). FWIW, this is a plot of `plt.hist(np.abs(preds_not_nan - targs_not_nan).clip(0, 0.25), bins=100, log=True)`:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2Fd9164276feafc0f13b4dc43f5164cfb1%2Ftrain_hist.png?generation=1703351640892991&alt=media)\n\n50% of the predictions have an absolute error <0.0457. Other percentile stats of `np.abs(preds_not_nan - targs_not_nan)`:\n\n```\nPercentile  5: 0.00056\nPercentile 25: 0.00283\nPercentile 50: 0.04570\nPercentile 75: 0.24829\nPercentile 95: 0.82649\n```\n\nBoth `preds_not_nan` and `targs_not_nan` were clipped to 0, 1 prior to computing the above stats. Hope this makes sense and works for you."
    },
    {
      "id": 2571755,
      "postDate": "2023-12-23T14:37:54.380Z",
      "content": "<p>I think something is wrong with your sample_submission_TRAIN.csv. It has a length of 27079811 instead of the expected 142513888.</p>",
      "rawMarkdown": "I think something is wrong with your sample_submission_TRAIN.csv. It has a length of 27079811 instead of the expected 142513888."
    },
    {
      "id": 2570971,
      "postDate": "2023-12-22T18:32:21.663Z",
      "content": "<p>Two questions, 1) how far down the private LB do you consider predictions to be useful? And 2) are models allowed to predict sequences in their training sets or would you prefer OOF predictions?</p>",
      "rawMarkdown": "Two questions, 1) how far down the private LB do you consider predictions to be useful? And 2) are models allowed to predict sequences in their training sets or would you prefer OOF predictions?",
      "replies": [
        {
          "id": 2570976,
          "postDate": "2023-12-22T18:37:34.180Z",
          "content": "<p>Prediction within training sets would be useful -- we're particularly curious about cases where models might have learned deviate from the (noisy) train labels in order to maximize generalization. </p>\n<p>To enable comparison across models, the best sequences for you to focus on would actually would be the ones in the competition train set, compiled here: <a href=\"https://www.kaggle.com/datasets/rhijudas/ribonanza-train-sequences/data\" target=\"_blank\">https://www.kaggle.com/datasets/rhijudas/ribonanza-train-sequences/data</a> </p>\n<p>We'd be particularly interested in models that are in top 20 on private LB, as those are better than our internal host model (listed as <code>RNAdegformer</code> in private LB: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/leaderboard)\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/leaderboard)</a>.</p>",
          "rawMarkdown": "Prediction within training sets would be useful -- we're particularly curious about cases where models might have learned deviate from the (noisy) train labels in order to maximize generalization. \n\nTo enable comparison across models, the best sequences for you to focus on would actually would be the ones in the competition train set, compiled here: https://www.kaggle.com/datasets/rhijudas/ribonanza-train-sequences/data \n\nWe'd be particularly interested in models that are in top 20 on private LB, as those are better than our internal host model (listed as `RNAdegformer` in private LB: https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/leaderboard).\n\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 2570861,
      "postDate": "2023-12-22T16:58:43.370Z",
      "content": "<p>Hi, could you clarity the deadline for publishing predictions? </p>\n<p>We are currently updating our approach with Squeezeformer used by other top-winners and doing some additional ablation study. We want to make predictions with this updated model and ensemble of such models (given single model  about 0.141 private score that should result in much better predictions). This step would require two weeks I think. Do we have them?</p>",
      "rawMarkdown": "Hi, could you clarity the deadline for publishing predictions? \n\nWe are currently updating our approach with Squeezeformer used by other top-winners and doing some additional ablation study. We want to make predictions with this updated model and ensemble of such models (given single model  about 0.141 private score that should result in much better predictions). This step would require two weeks I think. Do we have them?",
      "replies": [
        {
          "id": 2570927,
          "postDate": "2023-12-22T17:51:16.827Z",
          "content": "<p>We'd love labels in the next couple dats [I'm editing above to clarify]. If you can get us labels for one or two of your actual competition models soon that would be helpful. (We could take a stab at re-training with post-competition models in the new year.)</p>",
          "rawMarkdown": "We'd love labels in the next couple dats [I'm editing above to clarify]. If you can get us labels for one or two of your actual competition models soon that would be helpful. (We could take a stab at re-training with post-competition models in the new year.)",
          "votes": 1,
          "replies": [
            {
              "id": 2571871,
              "postDate": "2023-12-23T17:14:12.593Z",
              "content": "<p>Hi! <a href=\"https://www.kaggle.com/datasets/vyaltsevvaleriy/stanford-ribonanza-rna-folding-train-predictions/data\" target=\"_blank\">Here</a> we uploaded reactivities for train sequences predicted with our model. For convenience, predictions are provided in format of a submition and in a format of original train dataset.</p>",
              "rawMarkdown": "Hi! [Here](https://www.kaggle.com/datasets/vyaltsevvaleriy/stanford-ribonanza-rna-folding-train-predictions/data) we uploaded reactivities for train sequences predicted with our model. For convenience, predictions are provided in format of a submition and in a format of original train dataset.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2571125,
      "postDate": "2023-12-22T22:11:27.847Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2571292,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2023-12-23T05:07:28.077000",
      "content": "<p>I uploaded predictions for my blend <a href=\"https://www.kaggle.com/datasets/iafoss/stanford-ribonanza-rna-folding-train-predictions\" target=\"_blank\">here</a>. It should be ~0.142 at private LB. Following the request from the post, it is not OOF predictions. So, these predictions may not be that beneficial since the model tries to memorize training data, while a better filtering could be achieved with OOF.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2571795,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2023-12-23T15:33:08.367000",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> did you use data with low signal to noise (&lt;1) for training? If not, your model predictions on those would be very useful as extra pl data.  And even if you did use them in some capacity, your predictions might be still be more accurate than the original noisy labels</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2571800,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-12-23T15:38:17.607000",
              "content": "<p>I did but with weighed loss. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2571831,
              "author_name": "Penzar Dmitry",
              "author_url": "",
              "post_date": "2023-12-23T16:36:44.833000",
              "content": "<p>By the way, we've calculated the correlation between reactivities predicted for train dataset and real reactivities. <br>\nFor clean data the correlation is 0.73 and 0.76 for  DMS and 2A3 respectively, <br>\nfor all data - 0.41 and 0.28. <br>\nSo, it seems that our model wasn't good at memorizing data and we've got about the same correlations for your dataset.   </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2571832,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-12-23T16:38:35.307000",
              "content": "<p>Sounds good, thanks for checking.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2571923,
              "author_name": "Shujun",
              "author_url": "",
              "post_date": "2023-12-23T18:03:37.490000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/dmitrypenzar1996\" target=\"_blank\">@dmitrypenzar1996</a> . I think it might be because everyone is using MAE loss which has gradients of +1/-1 at each position so the model does not end up memorizing labels as much as MSE loss whose gradients are proportional to distance to target</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2573360,
      "author_name": "dan4o",
      "author_url": "",
      "post_date": "2023-12-24T23:19:53.510000",
      "content": "<p>Train Labels predicted with the submission models are uploaded <a href=\"https://www.kaggle.com/datasets/dankrstev/stanford-ribonanza-rna-folding-train-labels\" target=\"_blank\">here.</a></p>\n<p>Sorry for the delay, hopefully I uploaded them on time.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2574389,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2023-12-25T22:00:15.957000",
          "content": "<p>i can't see this. did you make it public?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2574711,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-12-26T07:18:05.103000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2572641,
      "author_name": "hoyso48",
      "author_url": "",
      "post_date": "2023-12-24T12:32:18.617000",
      "content": "<p>I uploaded mine <a href=\"https://www.kaggle.com/datasets/hoyso48/ribonanza-train-sequences-predictions\" target=\"_blank\">here</a>. Hope it's not too late.</p>\n<p>For the predictions I used the 4 identical models trained on the whole train dataset each with different seed, which scored 0.137 on public LB and 0.140 on private LB.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2571948,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2023-12-23T18:31:19.990000",
      "content": "<p>As you requested, here are my predictions. My final ensemble had two kinds of models: the first one was trained only on the sequences, and the second one included CapR and bpp sum&amp;max across the columns. An ensemble of about forty models with equal weights to each kind of model achieved 0.14066/0.14238 public/private LB. (which would place me in 8th place. My final score is lower since I was too suspicious of CapR/bpp and chose 5:3 weights in favor of only train data models with 0.14133/0.14304 public/private LB). Here, I give a blend of 4 models from each kind. As others have pointed out, making OOF predictions for PL would probably be better. Also, I suspect using the original data would be better- noisy data is still data. Since we know the error, it's not wrong data that we need to 'fix,' so it should have more information than PL and give better results if handled correctly. I will write more about it in my write-up. I hope to finish it soon.<br>\n<a href=\"https://www.kaggle.com/datasets/shlomoron/srrf-model-1-train-preds\" target=\"_blank\">model 1 (only train)</a><br>\n<a href=\"https://www.kaggle.com/datasets/shlomoron/srrf-model-2-train-preds\" target=\"_blank\">model 2 (+CapR, bpp sum&amp;max)</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2571876,
      "author_name": "Javier Martín",
      "author_url": "",
      "post_date": "2023-12-23T17:20:08.840000",
      "content": "<p>Uploaded our predictions <a href=\"https://www.kaggle.com/datasets/bacterio/ribonanza-train-dataset-predictions\" target=\"_blank\">here</a>. FWIW, this is a plot of <code>plt.hist(np.abs(preds_not_nan - targs_not_nan).clip(0, 0.25), bins=100, log=True)</code>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2Fd9164276feafc0f13b4dc43f5164cfb1%2Ftrain_hist.png?generation=1703351640892991&amp;alt=media\" alt=\"\"></p>\n<p>50% of the predictions have an absolute error &lt;0.0457. Other percentile stats of <code>np.abs(preds_not_nan - targs_not_nan)</code>:</p>\n<pre><code>  : .\n : .\n : .\n : .\n : .\n</code></pre>\n<p>Both <code>preds_not_nan</code> and <code>targs_not_nan</code> were clipped to 0, 1 prior to computing the above stats. Hope this makes sense and works for you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2571755,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2023-12-23T14:37:54.380000",
      "content": "<p>I think something is wrong with your sample_submission_TRAIN.csv. It has a length of 27079811 instead of the expected 142513888.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2570971,
      "author_name": "Javier Martín",
      "author_url": "",
      "post_date": "2023-12-22T18:32:21.663000",
      "content": "<p>Two questions, 1) how far down the private LB do you consider predictions to be useful? And 2) are models allowed to predict sequences in their training sets or would you prefer OOF predictions?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2570976,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2023-12-22T18:37:34.180000",
          "content": "<p>Prediction within training sets would be useful -- we're particularly curious about cases where models might have learned deviate from the (noisy) train labels in order to maximize generalization. </p>\n<p>To enable comparison across models, the best sequences for you to focus on would actually would be the ones in the competition train set, compiled here: <a href=\"https://www.kaggle.com/datasets/rhijudas/ribonanza-train-sequences/data\" target=\"_blank\">https://www.kaggle.com/datasets/rhijudas/ribonanza-train-sequences/data</a> </p>\n<p>We'd be particularly interested in models that are in top 20 on private LB, as those are better than our internal host model (listed as <code>RNAdegformer</code> in private LB: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/leaderboard)\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/leaderboard)</a>.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2570861,
      "author_name": "Penzar Dmitry",
      "author_url": "",
      "post_date": "2023-12-22T16:58:43.370000",
      "content": "<p>Hi, could you clarity the deadline for publishing predictions? </p>\n<p>We are currently updating our approach with Squeezeformer used by other top-winners and doing some additional ablation study. We want to make predictions with this updated model and ensemble of such models (given single model  about 0.141 private score that should result in much better predictions). This step would require two weeks I think. Do we have them?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2570927,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2023-12-22T17:51:16.827000",
          "content": "<p>We'd love labels in the next couple dats [I'm editing above to clarify]. If you can get us labels for one or two of your actual competition models soon that would be helpful. (We could take a stab at re-training with post-competition models in the new year.)</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2571871,
              "author_name": "vyaltsevvaleriy",
              "author_url": "",
              "post_date": "2023-12-23T17:14:12.593000",
              "content": "<p>Hi! <a href=\"https://www.kaggle.com/datasets/vyaltsevvaleriy/stanford-ribonanza-rna-folding-train-predictions/data\" target=\"_blank\">Here</a> we uploaded reactivities for train sequences predicted with our model. For convenience, predictions are provided in format of a submition and in a format of original train dataset.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2571125,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-22T22:11:27.847000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2570145": "Hi everyone!\n\n@shujun717 and I are working to prepare our models for use by the broader research community.   \n\nFor the broader community, it would be wonderful to have a single model that might integrate and perhaps outperform any previous single model in Ribonanza. \n\nAs part of this effort, it would be helpful to have labels of your models for the train data sequences. \n\nYes, that seems weird, but as you know, most of the experimental train data are noisy, and we expect that model predictions will potentially be more accurate than the data themselves!\n\nWe're making the train data sequences available as `train_sequences.csv` file here:\n\nhttps://www.kaggle.com/datasets/rhijudas/ribonanza-train-sequences/data\n\nTo join us in this research effort, would you be willing to re-run your notebooks to generate sample submissions for these 806,573 sequences, upload as a public data set, and provide link in this thread?  \n\nIf running a large ensemble of models is painful, having the outputs of one or two of the best models you developed would still be helpful.  \n\nIt would be most useful to have the predictions in the next few days, by *Dec. 24, 2023*, so we can train over the holidays when we have priority access to compute. 😉 \n\nGratefully, for the hosts,\nRhiju\n",
    "2571292": "I uploaded predictions for my blend [here](https://www.kaggle.com/datasets/iafoss/stanford-ribonanza-rna-folding-train-predictions). It should be ~0.142 at private LB. Following the request from the post, it is not OOF predictions. So, these predictions may not be that beneficial since the model tries to memorize training data, while a better filtering could be achieved with OOF.",
    "2573360": "Train Labels predicted with the submission models are uploaded [here.](https://www.kaggle.com/datasets/dankrstev/stanford-ribonanza-rna-folding-train-labels)\n\nSorry for the delay, hopefully I uploaded them on time.",
    "2572641": "I uploaded mine [here](https://www.kaggle.com/datasets/hoyso48/ribonanza-train-sequences-predictions). Hope it's not too late.\n\nFor the predictions I used the 4 identical models trained on the whole train dataset each with different seed, which scored 0.137 on public LB and 0.140 on private LB.",
    "2571948": "As you requested, here are my predictions. My final ensemble had two kinds of models: the first one was trained only on the sequences, and the second one included CapR and bpp sum&max across the columns. An ensemble of about forty models with equal weights to each kind of model achieved 0.14066/0.14238 public/private LB. (which would place me in 8th place. My final score is lower since I was too suspicious of CapR/bpp and chose 5:3 weights in favor of only train data models with 0.14133/0.14304 public/private LB). Here, I give a blend of 4 models from each kind. As others have pointed out, making OOF predictions for PL would probably be better. Also, I suspect using the original data would be better- noisy data is still data. Since we know the error, it's not wrong data that we need to 'fix,' so it should have more information than PL and give better results if handled correctly. I will write more about it in my write-up. I hope to finish it soon.\n[model 1 (only train)](https://www.kaggle.com/datasets/shlomoron/srrf-model-1-train-preds)\n[model 2 (+CapR, bpp sum&max)](https://www.kaggle.com/datasets/shlomoron/srrf-model-2-train-preds)",
    "2571876": "Uploaded our predictions [here](https://www.kaggle.com/datasets/bacterio/ribonanza-train-dataset-predictions). FWIW, this is a plot of `plt.hist(np.abs(preds_not_nan - targs_not_nan).clip(0, 0.25), bins=100, log=True)`:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2Fd9164276feafc0f13b4dc43f5164cfb1%2Ftrain_hist.png?generation=1703351640892991&alt=media)\n\n50% of the predictions have an absolute error <0.0457. Other percentile stats of `np.abs(preds_not_nan - targs_not_nan)`:\n\n```\nPercentile  5: 0.00056\nPercentile 25: 0.00283\nPercentile 50: 0.04570\nPercentile 75: 0.24829\nPercentile 95: 0.82649\n```\n\nBoth `preds_not_nan` and `targs_not_nan` were clipped to 0, 1 prior to computing the above stats. Hope this makes sense and works for you.",
    "2571755": "I think something is wrong with your sample_submission_TRAIN.csv. It has a length of 27079811 instead of the expected 142513888.",
    "2570971": "Two questions, 1) how far down the private LB do you consider predictions to be useful? And 2) are models allowed to predict sequences in their training sets or would you prefer OOF predictions?",
    "2570861": "Hi, could you clarity the deadline for publishing predictions? \n\nWe are currently updating our approach with Squeezeformer used by other top-winners and doing some additional ablation study. We want to make predictions with this updated model and ensemble of such models (given single model  about 0.141 private score that should result in much better predictions). This step would require two weeks I think. Do we have them?",
    "2571125": ""
  }
}