{
  "id": 450692,
  "title": "[Mistake Resolved]",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/450692",
  "author_name": "",
  "post_date": "2023-10-25T09:43:11.208896500Z",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p></p>\n<p></p>\n<p>Edit:<br>\n<a href=\"https://www.kaggle.com/Crimson206\" target=\"_blank\">@Crimson206</a>: public LB consists of only length 177 seqs</p>\n<p>Edit 2:<br>\nThere was indeed a mistake on my end and the drastic divergence has been resolved</p>",
  "messages": [
    {
      "id": "2498410",
      "postDate": "10/25/2023 09:43:11",
      "content": "<p></p>\n<p></p>\n<p>Edit:<br>\n<a href=\"https://www.kaggle.com/Crimson206\" target=\"_blank\">@Crimson206</a>: public LB consists of only length 177 seqs</p>\n<p>Edit 2:<br>\nThere was indeed a mistake on my end and the drastic divergence has been resolved</p>",
      "rawMarkdown": "~~I've moved away from using absolute positional embeddings since the private test set is exclusively of sequences longer than train. While I'm able to get CV close between absolute and relative embeddings, LB seems to no longer correlate with CV using relative embeddings. (Which is very strange considering their seq length distributions should be similar if not the same.)~~\n\n~~I'm wondering if others are experiencing this as well or if it's a mistake on my end, cheers.~~\n\nEdit:\n@Crimson206: public LB consists of only length 177 seqs\n\nEdit 2:\nThere was indeed a mistake on my end and the drastic divergence has been resolved",
      "votes": null
    },
    {
      "id": "2498417",
      "postDate": "10/25/2023 09:49:32",
      "content": "<p>I didn't fully understand your post. Can you explain more?<br>\nAnd, it seems that, you didn't notice that only 177 long sequences of test data are considered as public LB data.</p>",
      "rawMarkdown": "I didn't fully understand your post. Can you explain more?\nAnd, it seems that, you didn't notice that only 177 long sequences of test data are considered as public LB data.",
      "votes": null
    },
    {
      "id": "2498434",
      "postDate": "10/25/2023 10:00:05",
      "content": "<p>If you want to know the robustness of your model with longer unseen sequences,<br>\nI recommend you to train your model only with 115 sequences, and then evaluate its performance on other sequences.<br>\nWhen you test it with two setups above, you will be able to choose the better one.</p>",
      "rawMarkdown": "If you want to know the robustness of your model with longer unseen sequences,\nI recommend you to train your model only with 115 sequences, and then evaluate its performance on other sequences.\nWhen you test it with two setups above, you will be able to choose the better one.",
      "votes": null
    },
    {
      "id": "2498466",
      "postDate": "10/25/2023 10:19:37",
      "content": "<p>I must have missed that, could you link me to the post?</p>",
      "rawMarkdown": "I must have missed that, could you link me to the post?",
      "votes": null
    },
    {
      "id": "2498509",
      "postDate": "10/25/2023 10:44:27",
      "content": "<pre><code>test = pd.read_csv()\npublic_test = test[test[]==]\npublic_test[] = public_test[].apply()\npublic_test[].unique()\n</code></pre>\n<p>I quickly wrote for you. If you run the lines, you will see only 177.</p>",
      "rawMarkdown": "```python\ntest = pd.read_csv(\"/kaggle/input/stanford-ribonanza-rna-folding/test_sequences.csv\")\npublic_test = test[test[\"future\"]==0]\npublic_test[\"seq_len\"] = public_test[\"sequence\"].apply(len)\npublic_test[\"seq_len\"].unique()\n```\n\nI quickly wrote for you. If you run the lines, you will see only 177.",
      "votes": null
    },
    {
      "id": "2498521",
      "postDate": "10/25/2023 10:57:44",
      "content": "<p>Did not think to look there, thanks for the knowledge</p>",
      "rawMarkdown": "Did not think to look there, thanks for the knowledge",
      "votes": null
    },
    {
      "id": "2499465",
      "postDate": "10/26/2023 04:10:58",
      "content": "<p>Hi, you can have a look at this topic -- <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/440702\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/440702</a>. <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> posted a nice picture of correlation between CV and leaderboard. I can add that he was lucky enough - in case of our lab there was more discrepancy between our CV and LB scores. But a trend is the same</p>",
      "rawMarkdown": "Hi, you can have a look at this topic -- https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/440702. @martynoveduard posted a nice picture of correlation between CV and leaderboard. I can add that he was lucky enough - in case of our lab there was more discrepancy between our CV and LB scores. But a trend is the same",
      "votes": null
    },
    {
      "id": "2499569",
      "postDate": "10/26/2023 05:55:00",
      "content": "<p>There is probably a mistake on my end then, thanks for confirming.</p>",
      "rawMarkdown": "There is probably a mistake on my end then, thanks for confirming.",
      "votes": null
    },
    {
      "id": "2500318",
      "postDate": "10/26/2023 15:24:03",
      "content": "<p>I think it's not a mistake, it's more like that different features result in different correlation between lb and validation.  </p>",
      "rawMarkdown": "I think it's not a mistake, it's more like that different features result in different correlation between lb and validation.",
      "votes": null
    },
    {
      "id": "2500423",
      "postDate": "10/26/2023 16:48:11",
      "content": "<p>Gotcha, I think I have some ideas thanks</p>",
      "rawMarkdown": "Gotcha, I think I have some ideas thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2498417,
      "author_name": "",
      "author_url": "",
      "post_date": "10/25/2023 09:49:32",
      "content": "<p>I didn't fully understand your post. Can you explain more?<br>\nAnd, it seems that, you didn't notice that only 177 long sequences of test data are considered as public LB data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2498466,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "10/25/2023 10:19:37",
          "content": "<p>I must have missed that, could you link me to the post?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2498509,
              "author_name": "",
              "author_url": "",
              "post_date": "10/25/2023 10:44:27",
              "content": "<pre><code>test = pd.read_csv()\npublic_test = test[test[]==]\npublic_test[] = public_test[].apply()\npublic_test[].unique()\n</code></pre>\n<p>I quickly wrote for you. If you run the lines, you will see only 177.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2498521,
                  "author_name": "sroger",
                  "author_url": "",
                  "post_date": "10/25/2023 10:57:44",
                  "content": "<p>Did not think to look there, thanks for the knowledge</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2498434,
      "author_name": "",
      "author_url": "",
      "post_date": "10/25/2023 10:00:05",
      "content": "<p>If you want to know the robustness of your model with longer unseen sequences,<br>\nI recommend you to train your model only with 115 sequences, and then evaluate its performance on other sequences.<br>\nWhen you test it with two setups above, you will be able to choose the better one.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2499465,
      "author_name": "dmitrypenzar1996",
      "author_url": "",
      "post_date": "10/26/2023 04:10:58",
      "content": "<p>Hi, you can have a look at this topic -- <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/440702\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/440702</a>. <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> posted a nice picture of correlation between CV and leaderboard. I can add that he was lucky enough - in case of our lab there was more discrepancy between our CV and LB scores. But a trend is the same</p>",
      "votes": null,
      "replies": [
        {
          "id": 2499569,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "10/26/2023 05:55:00",
          "content": "<p>There is probably a mistake on my end then, thanks for confirming.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2500318,
              "author_name": "dmitrypenzar1996",
              "author_url": "",
              "post_date": "10/26/2023 15:24:03",
              "content": "<p>I think it's not a mistake, it's more like that different features result in different correlation between lb and validation.  </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2500423,
                  "author_name": "sroger",
                  "author_url": "",
                  "post_date": "10/26/2023 16:48:11",
                  "content": "<p>Gotcha, I think I have some ideas thanks</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2498410": "~~I've moved away from using absolute positional embeddings since the private test set is exclusively of sequences longer than train. While I'm able to get CV close between absolute and relative embeddings, LB seems to no longer correlate with CV using relative embeddings. (Which is very strange considering their seq length distributions should be similar if not the same.)~~\n\n~~I'm wondering if others are experiencing this as well or if it's a mistake on my end, cheers.~~\n\nEdit:\n@Crimson206: public LB consists of only length 177 seqs\n\nEdit 2:\nThere was indeed a mistake on my end and the drastic divergence has been resolved",
    "2498417": "I didn't fully understand your post. Can you explain more?\nAnd, it seems that, you didn't notice that only 177 long sequences of test data are considered as public LB data.",
    "2498434": "If you want to know the robustness of your model with longer unseen sequences,\nI recommend you to train your model only with 115 sequences, and then evaluate its performance on other sequences.\nWhen you test it with two setups above, you will be able to choose the better one.",
    "2498466": "I must have missed that, could you link me to the post?",
    "2498509": "```python\ntest = pd.read_csv(\"/kaggle/input/stanford-ribonanza-rna-folding/test_sequences.csv\")\npublic_test = test[test[\"future\"]==0]\npublic_test[\"seq_len\"] = public_test[\"sequence\"].apply(len)\npublic_test[\"seq_len\"].unique()\n```\n\nI quickly wrote for you. If you run the lines, you will see only 177.",
    "2498521": "Did not think to look there, thanks for the knowledge",
    "2499465": "Hi, you can have a look at this topic -- https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/440702. @martynoveduard posted a nice picture of correlation between CV and leaderboard. I can add that he was lucky enough - in case of our lab there was more discrepancy between our CV and LB scores. But a trend is the same",
    "2499569": "There is probably a mistake on my end then, thanks for confirming.",
    "2500318": "I think it's not a mistake, it's more like that different features result in different correlation between lb and validation.",
    "2500423": "Gotcha, I think I have some ideas thanks"
  },
  "source": "meta"
}