{
  "id": 569666,
  "title": "LB probing results in Stanford RNA 3D Folding",
  "url": "/competitions/stanford-rna-3d-folding/discussion/569666",
  "author_name": "",
  "post_date": "2025-03-23T10:26:16.534016300Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi, Kagglers!</p>\n<p>In <a href=\"https://www.kaggle.com/code/tomooinubushi/lb-probing-for-sr3f\" target=\"_blank\">this notebook</a>, I tested some basic assumptions about the test dataset for this competition.<br>\nI found that the following assumptions are all TRUE:</p>\n<ul>\n<li>Max sequence length: 800–1000 (4298 in the training dataset)</li>\n<li>Min sequence length: 50–100 (3 in the training dataset)</li>\n<li>Mean sequence length: 300–400 (162 in the training dataset)</li>\n<li>Ratio of sequences over 400: 0.45–0.5 (0.05 in the training dataset)</li>\n<li>Mean ratio of A: 0.28–0.32 (0.23 in the training dataset)</li>\n<li>Mean ratio of U: 0.20–0.24 (0.21 in the training dataset)</li>\n<li>Mean ratio of C: 0.20–0.24 (0.25 in the training dataset)</li>\n<li>Mean ratio of G: 0.24–0.28 (0.29 in the training dataset)</li>\n</ul>\n<p>Please note that these results are based on the dataset of current training phase. Maybe we can test them again after April 23rd???</p>\n<p>If I made any mistakes, I’d appreciate any corrections!<br>\nI’d also love to hear about any other assumptions or hypotheses you might have about the test dataset.</p>\n<p><strong>[Update in 2025/04/29]</strong><br>\n I resubmitted <a href=\"https://www.kaggle.com/code/tomooinubushi/lb-probing-for-sr3f?scriptVersionId=229184980\" target=\"_blank\">the version 29 of LB probing notebook</a> for the new test data, and confirmed that the results are the same.</p>",
  "messages": [
    {
      "id": "3157341",
      "postDate": "03/23/2025 10:26:16",
      "content": "<p>Hi, Kagglers!</p>\n<p>In <a href=\"https://www.kaggle.com/code/tomooinubushi/lb-probing-for-sr3f\" target=\"_blank\">this notebook</a>, I tested some basic assumptions about the test dataset for this competition.<br>\nI found that the following assumptions are all TRUE:</p>\n<ul>\n<li>Max sequence length: 800–1000 (4298 in the training dataset)</li>\n<li>Min sequence length: 50–100 (3 in the training dataset)</li>\n<li>Mean sequence length: 300–400 (162 in the training dataset)</li>\n<li>Ratio of sequences over 400: 0.45–0.5 (0.05 in the training dataset)</li>\n<li>Mean ratio of A: 0.28–0.32 (0.23 in the training dataset)</li>\n<li>Mean ratio of U: 0.20–0.24 (0.21 in the training dataset)</li>\n<li>Mean ratio of C: 0.20–0.24 (0.25 in the training dataset)</li>\n<li>Mean ratio of G: 0.24–0.28 (0.29 in the training dataset)</li>\n</ul>\n<p>Please note that these results are based on the dataset of current training phase. Maybe we can test them again after April 23rd???</p>\n<p>If I made any mistakes, I’d appreciate any corrections!<br>\nI’d also love to hear about any other assumptions or hypotheses you might have about the test dataset.</p>\n<p><strong>[Update in 2025/04/29]</strong><br>\n I resubmitted <a href=\"https://www.kaggle.com/code/tomooinubushi/lb-probing-for-sr3f?scriptVersionId=229184980\" target=\"_blank\">the version 29 of LB probing notebook</a> for the new test data, and confirmed that the results are the same.</p>",
      "rawMarkdown": "Hi, Kagglers!\n\nIn [this notebook](https://www.kaggle.com/code/tomooinubushi/lb-probing-for-sr3f), I tested some basic assumptions about the test dataset for this competition.\nI found that the following assumptions are all TRUE:\n\n- Max sequence length: 800–1000 (4298 in the training dataset)\n- Min sequence length: 50–100 (3 in the training dataset)\n- Mean sequence length: 300–400 (162 in the training dataset)\n- Ratio of sequences over 400: 0.45–0.5 (0.05 in the training dataset)\n- Mean ratio of A: 0.28–0.32 (0.23 in the training dataset)\n- Mean ratio of U: 0.20–0.24 (0.21 in the training dataset)\n- Mean ratio of C: 0.20–0.24 (0.25 in the training dataset)\n- Mean ratio of G: 0.24–0.28 (0.29 in the training dataset)\n\nPlease note that these results are based on the dataset of current training phase. Maybe we can test them again after April 23rd???\n\nIf I made any mistakes, I’d appreciate any corrections!\nI’d also love to hear about any other assumptions or hypotheses you might have about the test dataset.\n\n\n**[Update in 2025/04/29]**\n I resubmitted [the version 29 of LB probing notebook](https://www.kaggle.com/code/tomooinubushi/lb-probing-for-sr3f?scriptVersionId=229184980) for the new test data, and confirmed that the results are the same.",
      "votes": null
    },
    {
      "id": "3189044",
      "postDate": "04/28/2025 16:59:54",
      "content": "<p><a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> have you run this for the new test data?</p>",
      "rawMarkdown": "tomooinubushi have you run this for the new test data?",
      "votes": null
    },
    {
      "id": "3189262",
      "postDate": "04/29/2025 03:17:58",
      "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> Thank you for your comment.<br>\nI resubmitted <a href=\"https://www.kaggle.com/code/tomooinubushi/lb-probing-for-sr3f?scriptVersionId=229184980\" target=\"_blank\">the version 29 of LB probing notebook</a> for the new test data, and confirmed that the results are the same.</p>",
      "rawMarkdown": "alejopaullier Thank you for your comment.\nI resubmitted [the version 29 of LB probing notebook](https://www.kaggle.com/code/tomooinubushi/lb-probing-for-sr3f?scriptVersionId=229184980) for the new test data, and confirmed that the results are the same.",
      "votes": null
    },
    {
      "id": "3189624",
      "postDate": "04/29/2025 13:40:50",
      "content": "<p>Thanks! That's weird unless the new test data resembles quite well the old!</p>",
      "rawMarkdown": "Thanks! That's weird unless the new test data resembles quite well the old!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3189044,
      "author_name": "alejopaullier",
      "author_url": "",
      "post_date": "04/28/2025 16:59:54",
      "content": "<p><a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> have you run this for the new test data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3189262,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "04/29/2025 03:17:58",
          "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> Thank you for your comment.<br>\nI resubmitted <a href=\"https://www.kaggle.com/code/tomooinubushi/lb-probing-for-sr3f?scriptVersionId=229184980\" target=\"_blank\">the version 29 of LB probing notebook</a> for the new test data, and confirmed that the results are the same.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3189624,
              "author_name": "alejopaullier",
              "author_url": "",
              "post_date": "04/29/2025 13:40:50",
              "content": "<p>Thanks! That's weird unless the new test data resembles quite well the old!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3157341": "Hi, Kagglers!\n\nIn [this notebook](https://www.kaggle.com/code/tomooinubushi/lb-probing-for-sr3f), I tested some basic assumptions about the test dataset for this competition.\nI found that the following assumptions are all TRUE:\n\n- Max sequence length: 800–1000 (4298 in the training dataset)\n- Min sequence length: 50–100 (3 in the training dataset)\n- Mean sequence length: 300–400 (162 in the training dataset)\n- Ratio of sequences over 400: 0.45–0.5 (0.05 in the training dataset)\n- Mean ratio of A: 0.28–0.32 (0.23 in the training dataset)\n- Mean ratio of U: 0.20–0.24 (0.21 in the training dataset)\n- Mean ratio of C: 0.20–0.24 (0.25 in the training dataset)\n- Mean ratio of G: 0.24–0.28 (0.29 in the training dataset)\n\nPlease note that these results are based on the dataset of current training phase. Maybe we can test them again after April 23rd???\n\nIf I made any mistakes, I’d appreciate any corrections!\nI’d also love to hear about any other assumptions or hypotheses you might have about the test dataset.\n\n\n**[Update in 2025/04/29]**\n I resubmitted [the version 29 of LB probing notebook](https://www.kaggle.com/code/tomooinubushi/lb-probing-for-sr3f?scriptVersionId=229184980) for the new test data, and confirmed that the results are the same.",
    "3189044": "tomooinubushi have you run this for the new test data?",
    "3189262": "alejopaullier Thank you for your comment.\nI resubmitted [the version 29 of LB probing notebook](https://www.kaggle.com/code/tomooinubushi/lb-probing-for-sr3f?scriptVersionId=229184980) for the new test data, and confirmed that the results are the same.",
    "3189624": "Thanks! That's weird unless the new test data resembles quite well the old!"
  },
  "source": "meta"
}