{
  "id": 567262,
  "title": "kaggle evaluation code: per rna or per rna chain?",
  "url": "/competitions/stanford-rna-3d-folding/discussion/567262",
  "author_name": "",
  "post_date": "2025-03-09T10:22:52.890275300Z",
  "votes": 11,
  "comment_count": 6,
  "views": 0,
  "content": "<p>the metric code is given at:<br>\n<a href=\"https://www.kaggle.com/code/metric/ribonanza-tm-score/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/ribonanza-tm-score/notebook</a></p>\n<p>the score code is:</p>\n<pre><code>def (: pd.DataFrame, : pd.DataFrame, : str) -&gt; :\n\n    \n    solution[] = solution[].(lambda : x.()[])\n    submission[] = submission[].(lambda : x.()[])\n</code></pre>\n<p>now we know that, the submission ID column has the format:</p>\n<pre><code>,resname,resid,x_1,y_1,z_1,x_2,y_2,z_2,x_3,y_3,z_3,x_4,y_4,z_4,x_5,y_5,z_5\n,G,,-.,.,.,,,,,,,,,,,,\n,G,,-.,.,.,,,,,,,,,,,,\n,C,,-.,-.,.,,,,,,,,,,,,\n,A,,-.,-.,.,,,,,,,,,,,,\n,G,,-.,-.,.,,,,,,,,,,,,\n\n\n =  pdb_id + chain_id + residue_id\n</code></pre>\n<p>the score code will ignore the chain and combine all chains into one. is this the intended behaviour?<br>\nif so, our model must align the multiple chains per rna? </p>",
  "messages": [
    {
      "id": "3145090",
      "postDate": "03/09/2025 10:22:52",
      "content": "<p>the metric code is given at:<br>\n<a href=\"https://www.kaggle.com/code/metric/ribonanza-tm-score/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/ribonanza-tm-score/notebook</a></p>\n<p>the score code is:</p>\n<pre><code>def (: pd.DataFrame, : pd.DataFrame, : str) -&gt; :\n\n    \n    solution[] = solution[].(lambda : x.()[])\n    submission[] = submission[].(lambda : x.()[])\n</code></pre>\n<p>now we know that, the submission ID column has the format:</p>\n<pre><code>,resname,resid,x_1,y_1,z_1,x_2,y_2,z_2,x_3,y_3,z_3,x_4,y_4,z_4,x_5,y_5,z_5\n,G,,-.,.,.,,,,,,,,,,,,\n,G,,-.,.,.,,,,,,,,,,,,\n,C,,-.,-.,.,,,,,,,,,,,,\n,A,,-.,-.,.,,,,,,,,,,,,\n,G,,-.,-.,.,,,,,,,,,,,,\n\n\n =  pdb_id + chain_id + residue_id\n</code></pre>\n<p>the score code will ignore the chain and combine all chains into one. is this the intended behaviour?<br>\nif so, our model must align the multiple chains per rna? </p>",
      "rawMarkdown": "the metric code is given at:\nhttps://www.kaggle.com/code/metric/ribonanza-tm-score/notebook\n\nthe score code is:\n```\ndef score(solution: pd.DataFrame, submission: pd.DataFrame, row_id_column_name: str) -> float:\n\n    # Extract target_id from ID (target_resid)\n    solution['target_id'] = solution['ID'].apply(lambda x: x.split('_')[0])\n    submission['target_id'] = submission['ID'].apply(lambda x: x.split('_')[0])\n```\n\nnow we know that, the submission ID column has the format:\n```\nID,resname,resid,x_1,y_1,z_1,x_2,y_2,z_2,x_3,y_3,z_3,x_4,y_4,z_4,x_5,y_5,z_5\n6SDW_B_1,G,1,-22.83503,6.876176,17.790787,0,0,0,0,0,0,0,0,0,0,0,0\n6SDW_B_2,G,2,-20.73482,1.7411711,17.335567,0,0,0,0,0,0,0,0,0,0,0,0\n6SDW_B_3,C,3,-20.854864,-1.736068,14.085384,0,0,0,0,0,0,0,0,0,0,0,0\n6SDW_B_4,A,4,-21.762539,-4.106943,9.645739,0,0,0,0,0,0,0,0,0,0,0,0\n6SDW_B_5,G,5,-23.610292,-3.844714,4.6971865,0,0,0,0,0,0,0,0,0,0,0,0\n\n\nID =  pdb_id + chain_id + residue_id\n```\n\nthe score code will ignore the chain and combine all chains into one. is this the intended behaviour?\nif so, our model must align the multiple chains per rna?",
      "votes": null
    },
    {
      "id": "3145302",
      "postDate": "03/09/2025 17:16:12",
      "content": "<p>Not sure if is the answer you need but:<br>\n\"For a few targets, multiple slightly different structures have been captured experimentally; your predictions' scores will be based on the best TM-score compared to each of these reference structures.\"</p>",
      "rawMarkdown": "Not sure if is the answer you need but:\n\"For a few targets, multiple slightly different structures have been captured experimentally; your predictions' scores will be based on the best TM-score compared to each of these reference structures.\"",
      "votes": null
    },
    {
      "id": "3145476",
      "postDate": "03/09/2025 21:05:40",
      "content": "<p>Good spot - I'm a complete newbie to this field, but it seems that RNA is a single-stranded chain, so maybe there is no ambiguity here <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.</p>",
      "rawMarkdown": "Good spot - I'm a complete newbie to this field, but it seems that RNA is a single-stranded chain, so maybe there is no ambiguity here @hengck23.",
      "votes": null
    },
    {
      "id": "3147170",
      "postDate": "03/11/2025 18:24:07",
      "content": "<p>there are monomers and complex RNA targets in casp. I now presume that we are only interested in monomers case (until the host clarifies). In that case the correct code for getting target id in dataframe should have been</p>\n<pre><code>df[] = submission[].apply(lambda x: ,join(x.split()[:-]))\n</code></pre>",
      "rawMarkdown": "there are monomers and complex RNA targets in casp. I now presume that we are only interested in monomers case (until the host clarifies). In that case the correct code for getting target id in dataframe should have been\n\n```\ndf['target_id'] = submission['ID'].apply(lambda x: '_',join(x.split('_')[:-1]))\n\n```",
      "votes": null
    },
    {
      "id": "3147193",
      "postDate": "03/11/2025 18:43:10",
      "content": "<p>Bruhv is good, I want to be as good as you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. 😪</p>",
      "rawMarkdown": "Bruhv is good, I want to be as good as you @hengck23. 😪",
      "votes": null
    },
    {
      "id": "3147215",
      "postDate": "03/11/2025 19:07:58",
      "content": "<p>This is not intended and it's because of splitting multi-chain structures into separate monomers for training. For train, you may use something like this instead</p>\n<pre><code> re\n\n ():\n     = re.(, input_str)\n     :\n         .group(), .group()\n     input_str,   \n\n\n(split_pdb_chain())  \n(split_pdb_chain())    \n(split_pdb_chain())    \n</code></pre>\n<p>All test targets will be monomers.</p>",
      "rawMarkdown": "This is not intended and it's because of splitting multi-chain structures into separate monomers for training. For train, you may use something like this instead\n\n```python\nimport re\n\ndef split_pdb_chain(input_str):\n    match = re.match(r\"(.+?)_(\\d+)$\", input_str)\n    if match:\n        return match.group(1), match.group(2)\n    return input_str, None  # If no match, return the input and None\n\n# Example cases\nprint(split_pdb_chain(\"6SDW_B_1\"))  # Expected: ('6SDW_B', '1')\nprint(split_pdb_chain(\"6SDW_1\"))    # Expected: ('6SDW', '1')\nprint(split_pdb_chain(\"6SDW_B\"))    # Expected: ('6SDW_B', None)\n```\n\n\nAll test targets will be monomers.",
      "votes": null
    },
    {
      "id": "3147226",
      "postDate": "03/11/2025 19:17:59",
      "content": "<p>thanks for the clarification!</p>",
      "rawMarkdown": "thanks for the clarification!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3145302,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "03/09/2025 17:16:12",
      "content": "<p>Not sure if is the answer you need but:<br>\n\"For a few targets, multiple slightly different structures have been captured experimentally; your predictions' scores will be based on the best TM-score compared to each of these reference structures.\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3145476,
      "author_name": "jagofc",
      "author_url": "",
      "post_date": "03/09/2025 21:05:40",
      "content": "<p>Good spot - I'm a complete newbie to this field, but it seems that RNA is a single-stranded chain, so maybe there is no ambiguity here <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3147170,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/11/2025 18:24:07",
      "content": "<p>there are monomers and complex RNA targets in casp. I now presume that we are only interested in monomers case (until the host clarifies). In that case the correct code for getting target id in dataframe should have been</p>\n<pre><code>df[] = submission[].apply(lambda x: ,join(x.split()[:-]))\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3147193,
          "author_name": "theexaltedone",
          "author_url": "",
          "post_date": "03/11/2025 18:43:10",
          "content": "<p>Bruhv is good, I want to be as good as you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. 😪</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3147215,
      "author_name": "shujun717",
      "author_url": "",
      "post_date": "03/11/2025 19:07:58",
      "content": "<p>This is not intended and it's because of splitting multi-chain structures into separate monomers for training. For train, you may use something like this instead</p>\n<pre><code> re\n\n ():\n     = re.(, input_str)\n     :\n         .group(), .group()\n     input_str,   \n\n\n(split_pdb_chain())  \n(split_pdb_chain())    \n(split_pdb_chain())    \n</code></pre>\n<p>All test targets will be monomers.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3147226,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/11/2025 19:17:59",
          "content": "<p>thanks for the clarification!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3145090": "the metric code is given at:\nhttps://www.kaggle.com/code/metric/ribonanza-tm-score/notebook\n\nthe score code is:\n```\ndef score(solution: pd.DataFrame, submission: pd.DataFrame, row_id_column_name: str) -> float:\n\n    # Extract target_id from ID (target_resid)\n    solution['target_id'] = solution['ID'].apply(lambda x: x.split('_')[0])\n    submission['target_id'] = submission['ID'].apply(lambda x: x.split('_')[0])\n```\n\nnow we know that, the submission ID column has the format:\n```\nID,resname,resid,x_1,y_1,z_1,x_2,y_2,z_2,x_3,y_3,z_3,x_4,y_4,z_4,x_5,y_5,z_5\n6SDW_B_1,G,1,-22.83503,6.876176,17.790787,0,0,0,0,0,0,0,0,0,0,0,0\n6SDW_B_2,G,2,-20.73482,1.7411711,17.335567,0,0,0,0,0,0,0,0,0,0,0,0\n6SDW_B_3,C,3,-20.854864,-1.736068,14.085384,0,0,0,0,0,0,0,0,0,0,0,0\n6SDW_B_4,A,4,-21.762539,-4.106943,9.645739,0,0,0,0,0,0,0,0,0,0,0,0\n6SDW_B_5,G,5,-23.610292,-3.844714,4.6971865,0,0,0,0,0,0,0,0,0,0,0,0\n\n\nID =  pdb_id + chain_id + residue_id\n```\n\nthe score code will ignore the chain and combine all chains into one. is this the intended behaviour?\nif so, our model must align the multiple chains per rna?",
    "3145302": "Not sure if is the answer you need but:\n\"For a few targets, multiple slightly different structures have been captured experimentally; your predictions' scores will be based on the best TM-score compared to each of these reference structures.\"",
    "3145476": "Good spot - I'm a complete newbie to this field, but it seems that RNA is a single-stranded chain, so maybe there is no ambiguity here @hengck23.",
    "3147170": "there are monomers and complex RNA targets in casp. I now presume that we are only interested in monomers case (until the host clarifies). In that case the correct code for getting target id in dataframe should have been\n\n```\ndf['target_id'] = submission['ID'].apply(lambda x: '_',join(x.split('_')[:-1]))\n\n```",
    "3147193": "Bruhv is good, I want to be as good as you @hengck23. 😪",
    "3147215": "This is not intended and it's because of splitting multi-chain structures into separate monomers for training. For train, you may use something like this instead\n\n```python\nimport re\n\ndef split_pdb_chain(input_str):\n    match = re.match(r\"(.+?)_(\\d+)$\", input_str)\n    if match:\n        return match.group(1), match.group(2)\n    return input_str, None  # If no match, return the input and None\n\n# Example cases\nprint(split_pdb_chain(\"6SDW_B_1\"))  # Expected: ('6SDW_B', '1')\nprint(split_pdb_chain(\"6SDW_1\"))    # Expected: ('6SDW', '1')\nprint(split_pdb_chain(\"6SDW_B\"))    # Expected: ('6SDW_B', None)\n```\n\n\nAll test targets will be monomers.",
    "3147226": "thanks for the clarification!"
  },
  "source": "meta"
}