{
  "id": 578396,
  "title": "Submission issues",
  "url": "/competitions/stanford-rna-3d-folding/discussion/578396",
  "author_name": "",
  "post_date": "2025-05-10T14:32:25.582581900Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<table>\n<thead>\n<tr>\n<th>LB score filling longer sequences with zeros</th>\n<th></th>\n<th>LB score crop and padding</th>\n<th></th>\n<th>casp15 filling with zeros</th>\n<th></th>\n<th>casp15 crop and padding</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.454</td>\n<td></td>\n<td>0.452</td>\n<td></td>\n<td>0.371</td>\n<td></td>\n<td>0.40</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>On the local CASP15 test set using the same evaluation method, zero-filling coordinates for longer sequences led to poor results, while cropping and padding those sequences produced better performance.</li>\n<li>However, on the Kaggle leaderboard, the opposite was observed — zero-filling coordinates yielded higher scores, whereas the cropping and padding approach resulted in lower performance.<br>\n<a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <br>\nCould you provide insight into how the evaluation is performed?</li>\n</ul>",
  "messages": [
    {
      "id": "3199154",
      "postDate": "05/10/2025 14:32:25",
      "content": "<table>\n<thead>\n<tr>\n<th>LB score filling longer sequences with zeros</th>\n<th></th>\n<th>LB score crop and padding</th>\n<th></th>\n<th>casp15 filling with zeros</th>\n<th></th>\n<th>casp15 crop and padding</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.454</td>\n<td></td>\n<td>0.452</td>\n<td></td>\n<td>0.371</td>\n<td></td>\n<td>0.40</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>On the local CASP15 test set using the same evaluation method, zero-filling coordinates for longer sequences led to poor results, while cropping and padding those sequences produced better performance.</li>\n<li>However, on the Kaggle leaderboard, the opposite was observed — zero-filling coordinates yielded higher scores, whereas the cropping and padding approach resulted in lower performance.<br>\n<a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <br>\nCould you provide insight into how the evaluation is performed?</li>\n</ul>",
      "rawMarkdown": "| LB score filling longer sequences with zeros |  | LB score crop and padding| | casp15 filling with zeros | | casp15 crop and padding|\n|---| --- | ---| --- | --- |\n| 0.454 |  | 0.452 | | 0.371| | 0.40|\n\n\n- On the local CASP15 test set using the same evaluation method, zero-filling coordinates for longer sequences led to poor results, while cropping and padding those sequences produced better performance.\n- However, on the Kaggle leaderboard, the opposite was observed — zero-filling coordinates yielded higher scores, whereas the cropping and padding approach resulted in lower performance.\n\n\n@rhijudas @inversion \nCould you provide insight into how the evaluation is performed?",
      "votes": null
    },
    {
      "id": "3199290",
      "postDate": "05/10/2025 18:39:14",
      "content": "<p>Thanks for posting! The results are indeed curious. Can you tell us more - or post a link to a public notebook - on the cropping/padding strategy that you’re exploring?</p>",
      "rawMarkdown": "Thanks for posting! The results are indeed curious. Can you tell us more - or post a link to a public notebook - on the cropping/padding strategy that you’re exploring?",
      "votes": null
    },
    {
      "id": "3199550",
      "postDate": "05/11/2025 05:50:35",
      "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>  <br>\n<code>cropped_seq = seq[:700],\n remaining_len = len(seq) - 700,\nPredict = cropped_sequence,\nfull_coord = [torch.tensor(predict),torch.zeros(remanining_len)] \n</code><br>\nThis is how I do it, I believe this should give better scoring rather than filling with zeros </p>",
      "rawMarkdown": "rhijudas  \n`cropped_seq = seq[:700],\n remaining_len = len(seq) - 700,\nPredict = cropped_sequence,\nfull_coord = [torch.tensor(predict),torch.zeros(remanining_len)] \n`\nThis is how I do it, I believe this should give better scoring rather than filling with zeros",
      "votes": null
    },
    {
      "id": "3199890",
      "postDate": "05/11/2025 16:22:35",
      "content": "<p>Thanks for the explanation. Have you tried to zero-center the prediction for the crop?</p>",
      "rawMarkdown": "Thanks for the explanation. Have you tried to zero-center the prediction for the crop?",
      "votes": null
    },
    {
      "id": "3200079",
      "postDate": "05/12/2025 02:45:57",
      "content": "<p>yes i did it </p>",
      "rawMarkdown": "yes i did it",
      "votes": null
    },
    {
      "id": "3200403",
      "postDate": "05/12/2025 14:24:50",
      "content": "<p>OK, then this response to 0.0's seems like a puzzling property of the Tm-score evaluation metric. </p>\n<p>My guess is that it's better to not include any 0.0's in the prediction -- try to get reasonable models for the entire targets, and avoid padding. </p>\n<p>Let us know what you find!</p>",
      "rawMarkdown": "OK, then this response to 0.0's seems like a puzzling property of the Tm-score evaluation metric. \n\nMy guess is that it's better to not include any 0.0's in the prediction -- try to get reasonable models for the entire targets, and avoid padding. \n\nLet us know what you find!",
      "votes": null
    },
    {
      "id": "3200408",
      "postDate": "05/12/2025 14:34:48",
      "content": "<p>sure, i will try and let u know</p>",
      "rawMarkdown": "sure, i will try and let u know",
      "votes": null
    },
    {
      "id": "3200437",
      "postDate": "05/12/2025 15:12:36",
      "content": "<p>there's a minor difference between zero-filled tail vs the centroid-filled tail</p>\n<pre><code>Full-length TM-score:             \nZero-fill tail (&gt;) TM-score:   \nCentroid-fill tail (&gt;) TM-score: \n</code></pre>\n<p>notebook is here: <a href=\"https://www.kaggle.com/code/jaejohn/rna-3d-folding-evaluation?scriptVersionId=239292902\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rna-3d-folding-evaluation?scriptVersionId=239292902</a></p>",
      "rawMarkdown": "there's a minor difference between zero-filled tail vs the centroid-filled tail\n\n```python\nFull-length TM-score:             0.1531\nZero-fill tail (>700) TM-score:   0.1519\nCentroid-fill tail (>700) TM-score: 0.1516\n```\n\nnotebook is here: https://www.kaggle.com/code/jaejohn/rna-3d-folding-evaluation?scriptVersionId=239292902",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3199290,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "05/10/2025 18:39:14",
      "content": "<p>Thanks for posting! The results are indeed curious. Can you tell us more - or post a link to a public notebook - on the cropping/padding strategy that you’re exploring?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3199550,
          "author_name": "arunodhayan",
          "author_url": "",
          "post_date": "05/11/2025 05:50:35",
          "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>  <br>\n<code>cropped_seq = seq[:700],\n remaining_len = len(seq) - 700,\nPredict = cropped_sequence,\nfull_coord = [torch.tensor(predict),torch.zeros(remanining_len)] \n</code><br>\nThis is how I do it, I believe this should give better scoring rather than filling with zeros </p>",
          "votes": null,
          "replies": [
            {
              "id": 3199890,
              "author_name": "rhijudas",
              "author_url": "",
              "post_date": "05/11/2025 16:22:35",
              "content": "<p>Thanks for the explanation. Have you tried to zero-center the prediction for the crop?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3200079,
                  "author_name": "arunodhayan",
                  "author_url": "",
                  "post_date": "05/12/2025 02:45:57",
                  "content": "<p>yes i did it </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3200403,
                      "author_name": "rhijudas",
                      "author_url": "",
                      "post_date": "05/12/2025 14:24:50",
                      "content": "<p>OK, then this response to 0.0's seems like a puzzling property of the Tm-score evaluation metric. </p>\n<p>My guess is that it's better to not include any 0.0's in the prediction -- try to get reasonable models for the entire targets, and avoid padding. </p>\n<p>Let us know what you find!</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3200408,
                          "author_name": "arunodhayan",
                          "author_url": "",
                          "post_date": "05/12/2025 14:34:48",
                          "content": "<p>sure, i will try and let u know</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 3200437,
              "author_name": "jaejohn",
              "author_url": "",
              "post_date": "05/12/2025 15:12:36",
              "content": "<p>there's a minor difference between zero-filled tail vs the centroid-filled tail</p>\n<pre><code>Full-length TM-score:             \nZero-fill tail (&gt;) TM-score:   \nCentroid-fill tail (&gt;) TM-score: \n</code></pre>\n<p>notebook is here: <a href=\"https://www.kaggle.com/code/jaejohn/rna-3d-folding-evaluation?scriptVersionId=239292902\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rna-3d-folding-evaluation?scriptVersionId=239292902</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3199154": "| LB score filling longer sequences with zeros |  | LB score crop and padding| | casp15 filling with zeros | | casp15 crop and padding|\n|---| --- | ---| --- | --- |\n| 0.454 |  | 0.452 | | 0.371| | 0.40|\n\n\n- On the local CASP15 test set using the same evaluation method, zero-filling coordinates for longer sequences led to poor results, while cropping and padding those sequences produced better performance.\n- However, on the Kaggle leaderboard, the opposite was observed — zero-filling coordinates yielded higher scores, whereas the cropping and padding approach resulted in lower performance.\n\n\n@rhijudas @inversion \nCould you provide insight into how the evaluation is performed?",
    "3199290": "Thanks for posting! The results are indeed curious. Can you tell us more - or post a link to a public notebook - on the cropping/padding strategy that you’re exploring?",
    "3199550": "rhijudas  \n`cropped_seq = seq[:700],\n remaining_len = len(seq) - 700,\nPredict = cropped_sequence,\nfull_coord = [torch.tensor(predict),torch.zeros(remanining_len)] \n`\nThis is how I do it, I believe this should give better scoring rather than filling with zeros",
    "3199890": "Thanks for the explanation. Have you tried to zero-center the prediction for the crop?",
    "3200079": "yes i did it",
    "3200403": "OK, then this response to 0.0's seems like a puzzling property of the Tm-score evaluation metric. \n\nMy guess is that it's better to not include any 0.0's in the prediction -- try to get reasonable models for the entire targets, and avoid padding. \n\nLet us know what you find!",
    "3200408": "sure, i will try and let u know",
    "3200437": "there's a minor difference between zero-filled tail vs the centroid-filled tail\n\n```python\nFull-length TM-score:             0.1531\nZero-fill tail (>700) TM-score:   0.1519\nCentroid-fill tail (>700) TM-score: 0.1516\n```\n\nnotebook is here: https://www.kaggle.com/code/jaejohn/rna-3d-folding-evaluation?scriptVersionId=239292902"
  },
  "source": "meta"
}