{
  "id": 571121,
  "title": "Why scoring failed for my notebook",
  "url": "/competitions/stanford-rna-3d-folding/discussion/571121",
  "author_name": "",
  "post_date": "2025-04-01T13:29:55.398473900Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi, I have made my notebook public. I would really appreciate if anyone can help figure out why scoring always fails. I did many tests and the output csv file looks fine. Here is the notebook: <a href=\"https://www.kaggle.com/code/smliu1997/rhofold\" target=\"_blank\">https://www.kaggle.com/code/smliu1997/rhofold</a></p>",
  "messages": [
    {
      "id": "3167368",
      "postDate": "04/01/2025 13:29:55",
      "content": "<p>Hi, I have made my notebook public. I would really appreciate if anyone can help figure out why scoring always fails. I did many tests and the output csv file looks fine. Here is the notebook: <a href=\"https://www.kaggle.com/code/smliu1997/rhofold\" target=\"_blank\">https://www.kaggle.com/code/smliu1997/rhofold</a></p>",
      "rawMarkdown": "Hi, I have made my notebook public. I would really appreciate if anyone can help figure out why scoring always fails. I did many tests and the output csv file looks fine. Here is the notebook: https://www.kaggle.com/code/smliu1997/rhofold",
      "votes": null
    },
    {
      "id": "3167635",
      "postDate": "04/01/2025 18:00:41",
      "content": "<p>I believe the issue arises because specific IDs such as R1126, R1136, and R1138 are being skipped. Since these IDs do not exist in the private data evaluation, an error occurs. To avoid this, it may be better to specify skipping based on sequence length instead.</p>",
      "rawMarkdown": "I believe the issue arises because specific IDs such as R1126, R1136, and R1138 are being skipped. Since these IDs do not exist in the private data evaluation, an error occurs. To avoid this, it may be better to specify skipping based on sequence length instead.",
      "votes": null
    },
    {
      "id": "3167648",
      "postDate": "04/01/2025 18:09:49",
      "content": "<p>Thank you for the reply! Though I did not use rhofold for these 3 sequences, I did include them in the output csv file with zeros as coordinates. You can see in the notebook, I ensure the columns and ID orders are consistent with the sample submission csv file. </p>",
      "rawMarkdown": "Thank you for the reply! Though I did not use rhofold for these 3 sequences, I did include them in the output csv file with zeros as coordinates. You can see in the notebook, I ensure the columns and ID orders are consistent with the sample submission csv file.",
      "votes": null
    },
    {
      "id": "3167843",
      "postDate": "04/01/2025 21:44:17",
      "content": "<p>It’s a bit unclear, but I think there are 25 hidden data points in addition to the 12 sample data points. When I submit the notebook, those hidden data points might also be used for prediction. In the case of long sequences, they are not skipped, which likely causes a memory error and results in a failure.</p>",
      "rawMarkdown": "It’s a bit unclear, but I think there are 25 hidden data points in addition to the 12 sample data points. When I submit the notebook, those hidden data points might also be used for prediction. In the case of long sequences, they are not skipped, which likely causes a memory error and results in a failure.",
      "votes": null
    },
    {
      "id": "3167851",
      "postDate": "04/01/2025 21:51:09",
      "content": "<p>This sounds reasonable, I will modify the code to test. I tried also just produce the output csv file as the sample submission csv file and the scoring works. Based on your explanation, I guess maybe during scoring there is another set of test sequence and sample submission file. </p>",
      "rawMarkdown": "This sounds reasonable, I will modify the code to test. I tried also just produce the output csv file as the sample submission csv file and the scoring works. Based on your explanation, I guess maybe during scoring there is another set of test sequence and sample submission file.",
      "votes": null
    },
    {
      "id": "3168132",
      "postDate": "04/02/2025 06:25:45",
      "content": "<p>How about removing other files except submission.csv? I also had fails and solve the problems by removing files.</p>",
      "rawMarkdown": "How about removing other files except submission.csv? I also had fails and solve the problems by removing files.",
      "votes": null
    },
    {
      "id": "3168606",
      "postDate": "04/02/2025 15:43:42",
      "content": "<p>I think Shun Kuraishi's comment is correct. I fixed the issue following his advice. </p>",
      "rawMarkdown": "I think Shun Kuraishi's comment is correct. I fixed the issue following his advice.",
      "votes": null
    },
    {
      "id": "3176327",
      "postDate": "04/11/2025 08:49:24",
      "content": "<p>I have the same problem…</p>",
      "rawMarkdown": "I have the same problem...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3167635,
      "author_name": "shunkuraishi",
      "author_url": "",
      "post_date": "04/01/2025 18:00:41",
      "content": "<p>I believe the issue arises because specific IDs such as R1126, R1136, and R1138 are being skipped. Since these IDs do not exist in the private data evaluation, an error occurs. To avoid this, it may be better to specify skipping based on sequence length instead.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3167648,
          "author_name": "smliu1997",
          "author_url": "",
          "post_date": "04/01/2025 18:09:49",
          "content": "<p>Thank you for the reply! Though I did not use rhofold for these 3 sequences, I did include them in the output csv file with zeros as coordinates. You can see in the notebook, I ensure the columns and ID orders are consistent with the sample submission csv file. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3167843,
              "author_name": "shunkuraishi",
              "author_url": "",
              "post_date": "04/01/2025 21:44:17",
              "content": "<p>It’s a bit unclear, but I think there are 25 hidden data points in addition to the 12 sample data points. When I submit the notebook, those hidden data points might also be used for prediction. In the case of long sequences, they are not skipped, which likely causes a memory error and results in a failure.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3167851,
                  "author_name": "smliu1997",
                  "author_url": "",
                  "post_date": "04/01/2025 21:51:09",
                  "content": "<p>This sounds reasonable, I will modify the code to test. I tried also just produce the output csv file as the sample submission csv file and the scoring works. Based on your explanation, I guess maybe during scoring there is another set of test sequence and sample submission file. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3168132,
      "author_name": "youhanlee",
      "author_url": "",
      "post_date": "04/02/2025 06:25:45",
      "content": "<p>How about removing other files except submission.csv? I also had fails and solve the problems by removing files.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3168606,
          "author_name": "smliu1997",
          "author_url": "",
          "post_date": "04/02/2025 15:43:42",
          "content": "<p>I think Shun Kuraishi's comment is correct. I fixed the issue following his advice. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3176327,
      "author_name": "sofiyarachkevych",
      "author_url": "",
      "post_date": "04/11/2025 08:49:24",
      "content": "<p>I have the same problem…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3167368": "Hi, I have made my notebook public. I would really appreciate if anyone can help figure out why scoring always fails. I did many tests and the output csv file looks fine. Here is the notebook: https://www.kaggle.com/code/smliu1997/rhofold",
    "3167635": "I believe the issue arises because specific IDs such as R1126, R1136, and R1138 are being skipped. Since these IDs do not exist in the private data evaluation, an error occurs. To avoid this, it may be better to specify skipping based on sequence length instead.",
    "3167648": "Thank you for the reply! Though I did not use rhofold for these 3 sequences, I did include them in the output csv file with zeros as coordinates. You can see in the notebook, I ensure the columns and ID orders are consistent with the sample submission csv file.",
    "3167843": "It’s a bit unclear, but I think there are 25 hidden data points in addition to the 12 sample data points. When I submit the notebook, those hidden data points might also be used for prediction. In the case of long sequences, they are not skipped, which likely causes a memory error and results in a failure.",
    "3167851": "This sounds reasonable, I will modify the code to test. I tried also just produce the output csv file as the sample submission csv file and the scoring works. Based on your explanation, I guess maybe during scoring there is another set of test sequence and sample submission file.",
    "3168132": "How about removing other files except submission.csv? I also had fails and solve the problems by removing files.",
    "3168606": "I think Shun Kuraishi's comment is correct. I fixed the issue following his advice.",
    "3176327": "I have the same problem..."
  },
  "source": "meta"
}