{
  "id": 455715,
  "title": "What are your expectations of shakeup?",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/455715",
  "author_name": "",
  "post_date": "2023-11-16T04:08:38.746826200Z",
  "votes": 9,
  "comment_count": 12,
  "views": 0,
  "content": "<p>In the private test set, number of sequences with length 206 is 1kk, number of sequences with length {307, 457} is 8k, meaning the ratio between these two numbers is 125 [or around 70~ if the count number of nucleiotids for which we have to make the predictions], I wonder if these longer sequences are going to have any impact at all, assuming we make predictions for these sequences that are \"somewhat good\"</p>",
  "messages": [
    {
      "id": "2526801",
      "postDate": "11/16/2023 04:08:38",
      "content": "<p>In the private test set, number of sequences with length 206 is 1kk, number of sequences with length {307, 457} is 8k, meaning the ratio between these two numbers is 125 [or around 70~ if the count number of nucleiotids for which we have to make the predictions], I wonder if these longer sequences are going to have any impact at all, assuming we make predictions for these sequences that are \"somewhat good\"</p>",
      "rawMarkdown": "In the private test set, number of sequences with length 206 is 1kk, number of sequences with length {307, 457} is 8k, meaning the ratio between these two numbers is 125 [or around 70~ if the count number of nucleiotids for which we have to make the predictions], I wonder if these longer sequences are going to have any impact at all, assuming we make predictions for these sequences that are \"somewhat good\"",
      "votes": null
    },
    {
      "id": "2527284",
      "postDate": "11/16/2023 12:19:41",
      "content": "<p>Hi,</p>\n<p>I wonder about this too, public lb have around 12% data leakage, private lb is around 98% out-of-distribution from the training set.</p>\n<p>Hard to know how many people is optimizing models validating on the set with data leakage.</p>",
      "rawMarkdown": "Hi,\n\nI wonder about this too, public lb have around 12% data leakage, private lb is around 98% out-of-distribution from the training set.\n\nHard to know how many people is optimizing models validating on the set with data leakage.",
      "votes": null
    },
    {
      "id": "2527303",
      "postDate": "11/16/2023 12:38:34",
      "content": "<p>I tried to validate the model on the available 2.3k L=206 sequences [and train on sequences with L&lt;206], and the metric fluctuated a lot, but the trend behaved as expected </p>",
      "rawMarkdown": "I tried to validate the model on the available 2.3k L=206 sequences [and train on sequences with L<206], and the metric fluctuated a lot, but the trend behaved as expected",
      "votes": null
    },
    {
      "id": "2527559",
      "postDate": "11/16/2023 16:07:54",
      "content": "<p>Hi, I'm a bit confused regarding the terms private and public lb here. By private leaderboard you mean the test sequences we can download, or just some of them with certain future label? If so then what is the public lb? :)</p>",
      "rawMarkdown": "Hi, I'm a bit confused regarding the terms private and public lb here. By private leaderboard you mean the test sequences we can download, or just some of them with certain future label? If so then what is the public lb? :)",
      "votes": null
    },
    {
      "id": "2527571",
      "postDate": "11/16/2023 16:24:02",
      "content": "<p>Hi,</p>\n<p>The organizers have 1,118,513 RNA sequences, 311,935 of those sequences are the public leaderboard, when you submit to the competition those sequences are used to evaluate the submission, when the competition ends, the submissions will be evaluate on 1,031,888 RNA sequences that are different from the public leaderboard, this is the private leaderboard.</p>\n<p>You can find more information in the section \"Additional Notes\" here: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data</a></p>\n<p>Edit -------------------------------------------------</p>\n<p>Check test set size vs private leaderboard + public leaderboard</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F6bfe6eadc8c30229b5de5588f308545c%2FScreenshot%20from%202023-11-16%2017-31-58.png?generation=1700152433191752&amp;alt=media\" alt=\"\"></p>\n<p>At submission time the organizers evaluate both the public and the private set, but only shows the public one.</p>",
      "rawMarkdown": "Hi,\n\nThe organizers have 1,118,513 RNA sequences, 311,935 of those sequences are the public leaderboard, when you submit to the competition those sequences are used to evaluate the submission, when the competition ends, the submissions will be evaluate on 1,031,888 RNA sequences that are different from the public leaderboard, this is the private leaderboard.\n\nYou can find more information in the section \"Additional Notes\" here: [https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data](url)\n\nEdit -------------------------------------------------\n\nCheck test set size vs private leaderboard + public leaderboard\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F6bfe6eadc8c30229b5de5588f308545c%2FScreenshot%20from%202023-11-16%2017-31-58.png?generation=1700152433191752&alt=media)\n\nAt submission time the organizers evaluate both the public and the private set, but only shows the public one.",
      "votes": null
    },
    {
      "id": "2527646",
      "postDate": "11/16/2023 17:31:03",
      "content": "<p>Now it's much clearer for me, thx! So we do not know exactly which sequences in the test set make up the public LB, but roughly (most of the ones with zeroes), right? Based on the data description, some of the zero labelled ones are also part of the private LB.</p>",
      "rawMarkdown": "Now it's much clearer for me, thx! So we do not know exactly which sequences in the test set make up the public LB, but roughly (most of the ones with zeroes), right? Based on the data description, some of the zero labelled ones are also part of the private LB.",
      "votes": null
    },
    {
      "id": "2527656",
      "postDate": "11/16/2023 17:43:22",
      "content": "<p>The private leaderboard is mostly sequences with a length greater than 206 (98%).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F660aa08b4e9c5593495731d1acc7fb62%2FScreenshot%20from%202023-11-16%2018-39-50.png?generation=1700156543065719&amp;alt=media\" alt=\"\"></p>\n<p>That will give you the sequences that are in the public leaderboard, plus a few of the private lb.</p>",
      "rawMarkdown": "The private leaderboard is mostly sequences with a length greater than 206 (98%).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F660aa08b4e9c5593495731d1acc7fb62%2FScreenshot%20from%202023-11-16%2018-39-50.png?generation=1700156543065719&alt=media)\n\nThat will give you the sequences that are in the public leaderboard, plus a few of the private lb.",
      "votes": null
    },
    {
      "id": "2528502",
      "postDate": "11/17/2023 12:33:08",
      "content": "<p>Same. I think the main problem will be not with out-of-distribution length - you are absolutely correct that sequences of length 300+ have small overall weight. But with the fact, that in train dataset we have relatively similar sequences and test dataset have ones very different from them. So, there may be many interactions model fails to predict and we don't observe that now. </p>",
      "rawMarkdown": "Same. I think the main problem will be not with out-of-distribution length - you are absolutely correct that sequences of length 300+ have small overall weight. But with the fact, that in train dataset we have relatively similar sequences and test dataset have ones very different from them. So, there may be many interactions model fails to predict and we don't observe that now.",
      "votes": null
    },
    {
      "id": "2549382",
      "postDate": "12/05/2023 07:53:28",
      "content": "<p>I think the top20 will shakeup. But maybe I missed some magic. Good lucky.</p>",
      "rawMarkdown": "I think the top20 will shakeup. But maybe I missed some magic. Good lucky.",
      "votes": null
    },
    {
      "id": "2552083",
      "postDate": "12/07/2023 07:11:54",
      "content": "<p>We wasn't able to get any split which would result in performance negatively correlated with a simple k-fold split. </p>\n<p>I 100% sure there will be a huge gap between public and private scores due to 13% data leak in public data (it's easy to check, actually). </p>\n<p>We can't predict shakeup due to private sequences completely different from the training ones or due to experimental noise (e.g batch effect). So for me, as in the old joke, the probability of shakeup is 0.5 - it will take place or not)</p>",
      "rawMarkdown": "We wasn't able to get any split which would result in performance negatively correlated with a simple k-fold split. \n\nI 100% sure there will be a huge gap between public and private scores due to 13% data leak in public data (it's easy to check, actually). \n\nWe can't predict shakeup due to private sequences completely different from the training ones or due to experimental noise (e.g batch effect). So for me, as in the old joke, the probability of shakeup is 0.5 - it will take place or not)",
      "votes": null
    },
    {
      "id": "2552620",
      "postDate": "12/07/2023 16:03:57",
      "content": "<p>Shake up has never been grateful with me. I'm 100% sure I have 0% chance to be in the top 0.000000001% 🤣</p>",
      "rawMarkdown": "Shake up has never been grateful with me. I'm 100% sure I have 0% chance to be in the top 0.000000001% 🤣",
      "votes": null
    },
    {
      "id": "2552946",
      "postDate": "12/07/2023 21:55:24",
      "content": "<p>Fun fact, with what I thought a \"cleaver\" way to split (and evaluate) the data i could not break the 0.141<br>\nFunnier fact, i still think it is the best way to evaluate the test set… we will see… lucky us there is more than one sub :)</p>",
      "rawMarkdown": "Fun fact, with what I thought a \"cleaver\" way to split (and evaluate) the data i could not break the 0.141\nFunnier fact, i still think it is the best way to evaluate the test set... we will see... lucky us there is more than one sub :)",
      "votes": null
    },
    {
      "id": "2552949",
      "postDate": "12/07/2023 22:04:16",
      "content": "<p>+1 for clever split <br>\nwe also have clever split … let see what will happen </p>",
      "rawMarkdown": "1 for clever split \nwe also have clever split ... let see what will happen",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2527284,
      "author_name": "enriquezaf",
      "author_url": "",
      "post_date": "11/16/2023 12:19:41",
      "content": "<p>Hi,</p>\n<p>I wonder about this too, public lb have around 12% data leakage, private lb is around 98% out-of-distribution from the training set.</p>\n<p>Hard to know how many people is optimizing models validating on the set with data leakage.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2527559,
          "author_name": "michauhl",
          "author_url": "",
          "post_date": "11/16/2023 16:07:54",
          "content": "<p>Hi, I'm a bit confused regarding the terms private and public lb here. By private leaderboard you mean the test sequences we can download, or just some of them with certain future label? If so then what is the public lb? :)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2527571,
              "author_name": "enriquezaf",
              "author_url": "",
              "post_date": "11/16/2023 16:24:02",
              "content": "<p>Hi,</p>\n<p>The organizers have 1,118,513 RNA sequences, 311,935 of those sequences are the public leaderboard, when you submit to the competition those sequences are used to evaluate the submission, when the competition ends, the submissions will be evaluate on 1,031,888 RNA sequences that are different from the public leaderboard, this is the private leaderboard.</p>\n<p>You can find more information in the section \"Additional Notes\" here: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data</a></p>\n<p>Edit -------------------------------------------------</p>\n<p>Check test set size vs private leaderboard + public leaderboard</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F6bfe6eadc8c30229b5de5588f308545c%2FScreenshot%20from%202023-11-16%2017-31-58.png?generation=1700152433191752&amp;alt=media\" alt=\"\"></p>\n<p>At submission time the organizers evaluate both the public and the private set, but only shows the public one.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2527646,
                  "author_name": "michauhl",
                  "author_url": "",
                  "post_date": "11/16/2023 17:31:03",
                  "content": "<p>Now it's much clearer for me, thx! So we do not know exactly which sequences in the test set make up the public LB, but roughly (most of the ones with zeroes), right? Based on the data description, some of the zero labelled ones are also part of the private LB.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2527656,
                      "author_name": "enriquezaf",
                      "author_url": "",
                      "post_date": "11/16/2023 17:43:22",
                      "content": "<p>The private leaderboard is mostly sequences with a length greater than 206 (98%).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F660aa08b4e9c5593495731d1acc7fb62%2FScreenshot%20from%202023-11-16%2018-39-50.png?generation=1700156543065719&amp;alt=media\" alt=\"\"></p>\n<p>That will give you the sequences that are in the public leaderboard, plus a few of the private lb.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2527303,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "11/16/2023 12:38:34",
      "content": "<p>I tried to validate the model on the available 2.3k L=206 sequences [and train on sequences with L&lt;206], and the metric fluctuated a lot, but the trend behaved as expected </p>",
      "votes": null,
      "replies": [
        {
          "id": 2528502,
          "author_name": "dmitrypenzar1996",
          "author_url": "",
          "post_date": "11/17/2023 12:33:08",
          "content": "<p>Same. I think the main problem will be not with out-of-distribution length - you are absolutely correct that sequences of length 300+ have small overall weight. But with the fact, that in train dataset we have relatively similar sequences and test dataset have ones very different from them. So, there may be many interactions model fails to predict and we don't observe that now. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2549382,
      "author_name": "daishu",
      "author_url": "",
      "post_date": "12/05/2023 07:53:28",
      "content": "<p>I think the top20 will shakeup. But maybe I missed some magic. Good lucky.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2552946,
          "author_name": "callmeb",
          "author_url": "",
          "post_date": "12/07/2023 21:55:24",
          "content": "<p>Fun fact, with what I thought a \"cleaver\" way to split (and evaluate) the data i could not break the 0.141<br>\nFunnier fact, i still think it is the best way to evaluate the test set… we will see… lucky us there is more than one sub :)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2552949,
              "author_name": "drhabib",
              "author_url": "",
              "post_date": "12/07/2023 22:04:16",
              "content": "<p>+1 for clever split <br>\nwe also have clever split … let see what will happen </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2552083,
      "author_name": "dmitrypenzar1996",
      "author_url": "",
      "post_date": "12/07/2023 07:11:54",
      "content": "<p>We wasn't able to get any split which would result in performance negatively correlated with a simple k-fold split. </p>\n<p>I 100% sure there will be a huge gap between public and private scores due to 13% data leak in public data (it's easy to check, actually). </p>\n<p>We can't predict shakeup due to private sequences completely different from the training ones or due to experimental noise (e.g batch effect). So for me, as in the old joke, the probability of shakeup is 0.5 - it will take place or not)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2552620,
      "author_name": "callmeb",
      "author_url": "",
      "post_date": "12/07/2023 16:03:57",
      "content": "<p>Shake up has never been grateful with me. I'm 100% sure I have 0% chance to be in the top 0.000000001% 🤣</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2526801": "In the private test set, number of sequences with length 206 is 1kk, number of sequences with length {307, 457} is 8k, meaning the ratio between these two numbers is 125 [or around 70~ if the count number of nucleiotids for which we have to make the predictions], I wonder if these longer sequences are going to have any impact at all, assuming we make predictions for these sequences that are \"somewhat good\"",
    "2527284": "Hi,\n\nI wonder about this too, public lb have around 12% data leakage, private lb is around 98% out-of-distribution from the training set.\n\nHard to know how many people is optimizing models validating on the set with data leakage.",
    "2527303": "I tried to validate the model on the available 2.3k L=206 sequences [and train on sequences with L<206], and the metric fluctuated a lot, but the trend behaved as expected",
    "2527559": "Hi, I'm a bit confused regarding the terms private and public lb here. By private leaderboard you mean the test sequences we can download, or just some of them with certain future label? If so then what is the public lb? :)",
    "2527571": "Hi,\n\nThe organizers have 1,118,513 RNA sequences, 311,935 of those sequences are the public leaderboard, when you submit to the competition those sequences are used to evaluate the submission, when the competition ends, the submissions will be evaluate on 1,031,888 RNA sequences that are different from the public leaderboard, this is the private leaderboard.\n\nYou can find more information in the section \"Additional Notes\" here: [https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data](url)\n\nEdit -------------------------------------------------\n\nCheck test set size vs private leaderboard + public leaderboard\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F6bfe6eadc8c30229b5de5588f308545c%2FScreenshot%20from%202023-11-16%2017-31-58.png?generation=1700152433191752&alt=media)\n\nAt submission time the organizers evaluate both the public and the private set, but only shows the public one.",
    "2527646": "Now it's much clearer for me, thx! So we do not know exactly which sequences in the test set make up the public LB, but roughly (most of the ones with zeroes), right? Based on the data description, some of the zero labelled ones are also part of the private LB.",
    "2527656": "The private leaderboard is mostly sequences with a length greater than 206 (98%).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F660aa08b4e9c5593495731d1acc7fb62%2FScreenshot%20from%202023-11-16%2018-39-50.png?generation=1700156543065719&alt=media)\n\nThat will give you the sequences that are in the public leaderboard, plus a few of the private lb.",
    "2528502": "Same. I think the main problem will be not with out-of-distribution length - you are absolutely correct that sequences of length 300+ have small overall weight. But with the fact, that in train dataset we have relatively similar sequences and test dataset have ones very different from them. So, there may be many interactions model fails to predict and we don't observe that now.",
    "2549382": "I think the top20 will shakeup. But maybe I missed some magic. Good lucky.",
    "2552083": "We wasn't able to get any split which would result in performance negatively correlated with a simple k-fold split. \n\nI 100% sure there will be a huge gap between public and private scores due to 13% data leak in public data (it's easy to check, actually). \n\nWe can't predict shakeup due to private sequences completely different from the training ones or due to experimental noise (e.g batch effect). So for me, as in the old joke, the probability of shakeup is 0.5 - it will take place or not)",
    "2552620": "Shake up has never been grateful with me. I'm 100% sure I have 0% chance to be in the top 0.000000001% 🤣",
    "2552946": "Fun fact, with what I thought a \"cleaver\" way to split (and evaluate) the data i could not break the 0.141\nFunnier fact, i still think it is the best way to evaluate the test set... we will see... lucky us there is more than one sub :)",
    "2552949": "1 for clever split \nwe also have clever split ... let see what will happen"
  },
  "source": "meta"
}