{
  "id": 565831,
  "title": "NaN values in train labels set",
  "url": "/competitions/stanford-rna-3d-folding/discussion/565831",
  "author_name": "",
  "post_date": "2025-03-02T09:58:02.776455700Z",
  "votes": 9,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi<br>\nI noticed some missing values in the training labels. The missing values are often the positions of the first/last residues of the sequences.</p>\n<p>Some examples:<br>\n9FO9_A_32,C,32,25.631000518798828,-6.853000164031982,-0.8730000257492065<br>\n9FO9_A_33,C,33,29.976999282836918,-3.688999891281128,1.368000030517578<br>\n8XPP_B_1,G,1,,,<br>\n8XPP_B_2,G,2,9.82800006866455,-17.54800033569336,23.549999237060547<br>\n8XPP_B_3,G,3,11.680999755859377,-18.09499931335449,18.58300018310547<br>\n8XPP_B_4,A,4,16.104999542236328,-19.54599952697754,15.269000053405762<br>\n8XPP_B_5,G,5,20.43600082397461,-21.989999771118164,14.366000175476074</p>\n<p>8Z1F_T_58,U,58,95.94200134277344,107.73699951171876,119.09300231933594<br>\n8Z1F_T_59,U,59,97.62100219726562,111.9990005493164,121.96900177001952<br>\n8Z1F_T_60,U,60,101.26000213623048,115.4540023803711,122.81999969482422<br>\n8Z1F_T_61,U,61,107.6780014038086,117.89700317382812,122.1959991455078<br>\n8Z1F_T_62,U,62,112.51699829101562,117.88099670410156,119.24500274658205<br>\n8Z1F_T_63,A,63,115.29299926757812,116.5719985961914,114.8270034790039<br>\n8Z1F_T_64,C,64,115.85700225830078,114.59500122070312,109.50900268554688<br>\n8Z1F_T_65,C,65,113.81600189208984,113.23600006103516,104.33999633789062<br>\n8Z1F_T_66,A,66,118.27999877929688,113.88300323486328,98.0719985961914<br>\n8Z1F_T_67,G,67,,,<br>\n8Z1F_T_68,C,68,,,<br>\n8Z1F_T_69,U,69,,,<br>\n8Z1F_T_70,C,70,,,<br>\n8Z1F_T_71,C,71,,,<br>\n8Z1F_T_72,G,72,,,<br>\n8Z1F_T_73,A,73,,,<br>\n8Z1F_T_74,G,74,,,<br>\n8Z1F_T_75,G,75,,,<br>\n8Z1F_T_76,U,76,,,<br>\n8Z1F_T_77,G,77,,,<br>\n8Z1F_T_78,A,78,,,<br>\n8Z1F_T_79,U,79,,,<br>\n8Z1F_T_80,U,80,,,<br>\n8Z1F_T_81,U,81,,,<br>\n8Z1F_T_82,U,82,,,<br>\n8Z1F_T_83,C,83,,,<br>\n8Z1F_T_84,A,84,,,<br>\n8Z1F_T_85,U,85,,,<br>\n8Z1F_T_86,A,86,,,</p>\n<p>Missing values can also be positions of in-the-middle residues:<br>\n1FOQ_A_3,A,3,-38.89099884033203,-23.388999938964844,2.510999917984009<br>\n1FOQ_A_4,A,4,-34.917999267578125,-27.291000366210938,2.7679998874664307<br>\n1FOQ_A_5,U,5,,,<br>\n1FOQ_A_6,G,6,-33.224998474121094,-32.5359992980957,3.601999998092652<br>\n1FOQ_A_7,G,7,-34.20399856567383,-38.005001068115234,4.061999797821045</p>\n<h3>Questions</h3>\n<ul>\n<li>Why are some positions missing? (my 2 cents is that some residues may move a lot and thus have no clear static position?)</li>\n<li>Should we drop these NaN, or does it make sense to keep them ? (if they add some information)</li>\n</ul>\n<p>Thank you for your help</p>",
  "messages": [
    {
      "id": "3138165",
      "postDate": "03/02/2025 09:58:02",
      "content": "<p>Hi<br>\nI noticed some missing values in the training labels. The missing values are often the positions of the first/last residues of the sequences.</p>\n<p>Some examples:<br>\n9FO9_A_32,C,32,25.631000518798828,-6.853000164031982,-0.8730000257492065<br>\n9FO9_A_33,C,33,29.976999282836918,-3.688999891281128,1.368000030517578<br>\n8XPP_B_1,G,1,,,<br>\n8XPP_B_2,G,2,9.82800006866455,-17.54800033569336,23.549999237060547<br>\n8XPP_B_3,G,3,11.680999755859377,-18.09499931335449,18.58300018310547<br>\n8XPP_B_4,A,4,16.104999542236328,-19.54599952697754,15.269000053405762<br>\n8XPP_B_5,G,5,20.43600082397461,-21.989999771118164,14.366000175476074</p>\n<p>8Z1F_T_58,U,58,95.94200134277344,107.73699951171876,119.09300231933594<br>\n8Z1F_T_59,U,59,97.62100219726562,111.9990005493164,121.96900177001952<br>\n8Z1F_T_60,U,60,101.26000213623048,115.4540023803711,122.81999969482422<br>\n8Z1F_T_61,U,61,107.6780014038086,117.89700317382812,122.1959991455078<br>\n8Z1F_T_62,U,62,112.51699829101562,117.88099670410156,119.24500274658205<br>\n8Z1F_T_63,A,63,115.29299926757812,116.5719985961914,114.8270034790039<br>\n8Z1F_T_64,C,64,115.85700225830078,114.59500122070312,109.50900268554688<br>\n8Z1F_T_65,C,65,113.81600189208984,113.23600006103516,104.33999633789062<br>\n8Z1F_T_66,A,66,118.27999877929688,113.88300323486328,98.0719985961914<br>\n8Z1F_T_67,G,67,,,<br>\n8Z1F_T_68,C,68,,,<br>\n8Z1F_T_69,U,69,,,<br>\n8Z1F_T_70,C,70,,,<br>\n8Z1F_T_71,C,71,,,<br>\n8Z1F_T_72,G,72,,,<br>\n8Z1F_T_73,A,73,,,<br>\n8Z1F_T_74,G,74,,,<br>\n8Z1F_T_75,G,75,,,<br>\n8Z1F_T_76,U,76,,,<br>\n8Z1F_T_77,G,77,,,<br>\n8Z1F_T_78,A,78,,,<br>\n8Z1F_T_79,U,79,,,<br>\n8Z1F_T_80,U,80,,,<br>\n8Z1F_T_81,U,81,,,<br>\n8Z1F_T_82,U,82,,,<br>\n8Z1F_T_83,C,83,,,<br>\n8Z1F_T_84,A,84,,,<br>\n8Z1F_T_85,U,85,,,<br>\n8Z1F_T_86,A,86,,,</p>\n<p>Missing values can also be positions of in-the-middle residues:<br>\n1FOQ_A_3,A,3,-38.89099884033203,-23.388999938964844,2.510999917984009<br>\n1FOQ_A_4,A,4,-34.917999267578125,-27.291000366210938,2.7679998874664307<br>\n1FOQ_A_5,U,5,,,<br>\n1FOQ_A_6,G,6,-33.224998474121094,-32.5359992980957,3.601999998092652<br>\n1FOQ_A_7,G,7,-34.20399856567383,-38.005001068115234,4.061999797821045</p>\n<h3>Questions</h3>\n<ul>\n<li>Why are some positions missing? (my 2 cents is that some residues may move a lot and thus have no clear static position?)</li>\n<li>Should we drop these NaN, or does it make sense to keep them ? (if they add some information)</li>\n</ul>\n<p>Thank you for your help</p>",
      "rawMarkdown": "Hi\nI noticed some missing values in the training labels. The missing values are often the positions of the first/last residues of the sequences.\n\nSome examples:\n9FO9_A_32,C,32,25.631000518798828,-6.853000164031982,-0.8730000257492065\n9FO9_A_33,C,33,29.976999282836918,-3.688999891281128,1.368000030517578\n8XPP_B_1,G,1,,,\n8XPP_B_2,G,2,9.82800006866455,-17.54800033569336,23.549999237060547\n8XPP_B_3,G,3,11.680999755859377,-18.09499931335449,18.58300018310547\n8XPP_B_4,A,4,16.104999542236328,-19.54599952697754,15.269000053405762\n8XPP_B_5,G,5,20.43600082397461,-21.989999771118164,14.366000175476074\n\n\n8Z1F_T_58,U,58,95.94200134277344,107.73699951171876,119.09300231933594\n8Z1F_T_59,U,59,97.62100219726562,111.9990005493164,121.96900177001952\n8Z1F_T_60,U,60,101.26000213623048,115.4540023803711,122.81999969482422\n8Z1F_T_61,U,61,107.6780014038086,117.89700317382812,122.1959991455078\n8Z1F_T_62,U,62,112.51699829101562,117.88099670410156,119.24500274658205\n8Z1F_T_63,A,63,115.29299926757812,116.5719985961914,114.8270034790039\n8Z1F_T_64,C,64,115.85700225830078,114.59500122070312,109.50900268554688\n8Z1F_T_65,C,65,113.81600189208984,113.23600006103516,104.33999633789062\n8Z1F_T_66,A,66,118.27999877929688,113.88300323486328,98.0719985961914\n8Z1F_T_67,G,67,,,\n8Z1F_T_68,C,68,,,\n8Z1F_T_69,U,69,,,\n8Z1F_T_70,C,70,,,\n8Z1F_T_71,C,71,,,\n8Z1F_T_72,G,72,,,\n8Z1F_T_73,A,73,,,\n8Z1F_T_74,G,74,,,\n8Z1F_T_75,G,75,,,\n8Z1F_T_76,U,76,,,\n8Z1F_T_77,G,77,,,\n8Z1F_T_78,A,78,,,\n8Z1F_T_79,U,79,,,\n8Z1F_T_80,U,80,,,\n8Z1F_T_81,U,81,,,\n8Z1F_T_82,U,82,,,\n8Z1F_T_83,C,83,,,\n8Z1F_T_84,A,84,,,\n8Z1F_T_85,U,85,,,\n8Z1F_T_86,A,86,,,\n\nMissing values can also be positions of in-the-middle residues:\n1FOQ_A_3,A,3,-38.89099884033203,-23.388999938964844,2.510999917984009\n1FOQ_A_4,A,4,-34.917999267578125,-27.291000366210938,2.7679998874664307\n1FOQ_A_5,U,5,,,\n1FOQ_A_6,G,6,-33.224998474121094,-32.5359992980957,3.601999998092652\n1FOQ_A_7,G,7,-34.20399856567383,-38.005001068115234,4.061999797821045\n\n### Questions \n\n- Why are some positions missing? (my 2 cents is that some residues may move a lot and thus have no clear static position?)\n- Should we drop these NaN, or does it make sense to keep them ? (if they add some information)\n\nThank you for your help",
      "votes": null
    },
    {
      "id": "3138269",
      "postDate": "03/02/2025 12:08:33",
      "content": "<p>For now, I’ve decided to \"extract\" the correct sequences if they are at least 6 elements long. I’m not sure if this is the right approach, but I don’t want to lose them. On the other hand, replacing them (NaN) with something else doesn’t seem right to me. What other options are there?</p>",
      "rawMarkdown": "For now, I’ve decided to \"extract\" the correct sequences if they are at least 6 elements long. I’m not sure if this is the right approach, but I don’t want to lose them. On the other hand, replacing them (NaN) with something else doesn’t seem right to me. What other options are there?",
      "votes": null
    },
    {
      "id": "3138322",
      "postDate": "03/02/2025 12:52:49",
      "content": "<p>If it comes from a physical phenomenon, I am pretty sure there are similar issues with proteins</p>",
      "rawMarkdown": "If it comes from a physical phenomenon, I am pretty sure there are similar issues with proteins",
      "votes": null
    },
    {
      "id": "3138523",
      "postDate": "03/02/2025 16:33:04",
      "content": "<p>Hi! nan/null at a position means that experimentally, that part of the RNA could not be resolved. <br>\nIt could be due to flexibility of those segments, which blurs out their density in experimental techniques like crystallography or cryo-EM. <br>\nIt's up to you to explore how to best handle this information -- looking forward to the discussion!</p>",
      "rawMarkdown": "Hi! nan/null at a position means that experimentally, that part of the RNA could not be resolved. \nIt could be due to flexibility of those segments, which blurs out their density in experimental techniques like crystallography or cryo-EM. \nIt's up to you to explore how to best handle this information -- looking forward to the discussion!",
      "votes": null
    },
    {
      "id": "3138603",
      "postDate": "03/02/2025 18:21:50",
      "content": "<p>thank you for the answer!</p>",
      "rawMarkdown": "thank you for the answer!",
      "votes": null
    },
    {
      "id": "3139205",
      "postDate": "03/03/2025 09:08:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>, do the test set labels (public/private) contain nan values?</p>",
      "rawMarkdown": "Hi @rhijudas, do the test set labels (public/private) contain nan values?",
      "votes": null
    },
    {
      "id": "3140255",
      "postDate": "03/04/2025 12:35:35",
      "content": "<p>Thanks for the answer. Another note on the data - at least one of these seems a little wild:</p>\n<p>target id == 6EVJ_V</p>\n<pre><code>  resname resid   x_1 y_1 z_1\n</code></pre>\n<p>81005    6EVJ_V_1    A   1   96.825996   795.460022  -324.712006<br>\n81006    6EVJ_V_2    G   93.570999   796.059998  -320.483002<br>\n81007    6EVJ_V_3    U   3   89.374001   796.434021  -317.442993<br>\n81008    6EVJ_V_4    A   85.195999   793.974976  -317.768005<br>\n81009    6EVJ_V_5    G   5   81.110001            789.883972 -318.579010<br>\n81010    6EVJ_V_6    U   6   75.205002   788.846008  -321.894012<br>\n81011    6EVJ_V_7    A   7   77.397003   795.304993  -323.020996<br>\n81012    6EVJ_V_8    A   8   80.192001   798.466980  -322.545013<br>\n81013    6EVJ_V_9    C   9   84.644997   801.317017  -321.303009<br>\n81014    6EVJ_V_10   A   89.483002   802.109009  -323.695007<br>\n81015    6EVJ_V_11   A   86.855003   810.286011  -323.160004<br>\n81016    6EVJ_V_12   G   82.238998   812.549988  -325.160004<br>\n81017    6EVJ_V_13   A   80.735001   817.119019  -327.610992<br>\n81018    6EVJ_V_14   G   81.975998   821.786011  -330.247986<br>\n81019    6EVJ_V_15   G   84.862000   826.000000  -333.403992<br>\n81020    6EVJ_V_16   G   NaN NaN NaN</p>\n<p>The y and z co-ordinates seem massive relative to other samples (typically in the range of -30 - +30). Is this expected? </p>",
      "rawMarkdown": "Thanks for the answer. Another note on the data - at least one of these seems a little wild:\n\ntarget id == 6EVJ_V\n\n\tID\tresname\tresid\tx_1\ty_1\tz_1\n81005\t6EVJ_V_1\tA\t1\t96.825996\t795.460022\t-324.712006\n81006\t6EVJ_V_2\tG\t93.570999\t796.059998\t-320.483002\n81007\t6EVJ_V_3\tU\t3\t89.374001\t796.434021\t-317.442993\n81008\t6EVJ_V_4\tA\t85.195999\t793.974976\t-317.768005\n81009\t6EVJ_V_5\tG\t5\t81.110001\t         789.883972\t-318.579010\n81010\t6EVJ_V_6\tU\t6\t75.205002\t788.846008\t-321.894012\n81011\t6EVJ_V_7\tA\t7\t77.397003\t795.304993\t-323.020996\n81012\t6EVJ_V_8\tA\t8\t80.192001\t798.466980\t-322.545013\n81013\t6EVJ_V_9\tC\t9\t84.644997\t801.317017\t-321.303009\n81014\t6EVJ_V_10\tA\t89.483002\t802.109009\t-323.695007\n81015\t6EVJ_V_11\tA\t86.855003\t810.286011\t-323.160004\n81016\t6EVJ_V_12\tG\t82.238998\t812.549988\t-325.160004\n81017\t6EVJ_V_13\tA\t80.735001\t817.119019\t-327.610992\n81018\t6EVJ_V_14\tG\t81.975998\t821.786011\t-330.247986\n81019\t6EVJ_V_15\tG\t84.862000\t826.000000\t-333.403992\n81020\t6EVJ_V_16\tG\tNaN\tNaN\tNaN\n\n\nThe y and z co-ordinates seem massive relative to other samples (typically in the range of -30 - +30). Is this expected?",
      "votes": null
    },
    {
      "id": "3140315",
      "postDate": "03/04/2025 13:42:46",
      "content": "<p>Nevermind, I think I found my answer:</p>\n<p>Molecules are rotated and translated in 3D space to minimize the Root Mean Square Deviation (RMSD). RMSD is a good proxy for the TM-score and they are inversely correlated.</p>\n<p>The orientation you submit is irrelevant, as the coordinate space is arbitrary. Coordinates will always be manipulated to create the best alignment with target structures.</p>\n<p>From: <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566056\" target=\"_blank\">Here</a></p>",
      "rawMarkdown": "Nevermind, I think I found my answer:\n\n\nMolecules are rotated and translated in 3D space to minimize the Root Mean Square Deviation (RMSD). RMSD is a good proxy for the TM-score and they are inversely correlated.\n\nThe orientation you submit is irrelevant, as the coordinate space is arbitrary. Coordinates will always be manipulated to create the best alignment with target structures.\n\nFrom: [Here](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566056)",
      "votes": null
    },
    {
      "id": "3140977",
      "postDate": "03/05/2025 05:21:35",
      "content": "<p>I am thinking do delete all types of null values rows. Just work with full data. Is it a valid point? or should I work around filling data ? <br>\nI think it might be unrealistic on bio data to auto fill it by any technique. </p>\n<p>Please share your perspectives. Thanks!</p>",
      "rawMarkdown": "I am thinking do delete all types of null values rows. Just work with full data. Is it a valid point? or should I work around filling data ? \nI think it might be unrealistic on bio data to auto fill it by any technique. \n\nPlease share your perspectives. Thanks!",
      "votes": null
    },
    {
      "id": "3145596",
      "postDate": "03/10/2025 02:53:31",
      "content": "<p>I am interested to see how others approach this. As missing rows made up 4.5% of all the data, for now I decided to simply drop them, but when I eventually submit and get a score, I plan to see if I can push up the score by working with this missing data. I noted that most RNA molecules x, y, z coordinates are fairly similar, implying you could fill the missing values by finding a pattern within each RNA molecule.</p>\n<p>A blanket approach, e.g. applying the mean seems wrong in this context as each RNA is a seperate molecule.</p>",
      "rawMarkdown": "I am interested to see how others approach this. As missing rows made up 4.5% of all the data, for now I decided to simply drop them, but when I eventually submit and get a score, I plan to see if I can push up the score by working with this missing data. I noted that most RNA molecules x, y, z coordinates are fairly similar, implying you could fill the missing values by finding a pattern within each RNA molecule.\n\nA blanket approach, e.g. applying the mean seems wrong in this context as each RNA is a seperate molecule.",
      "votes": null
    },
    {
      "id": "3158002",
      "postDate": "03/24/2025 06:16:07",
      "content": "<p>According to Rhiju, NaN values in our dataset represent structurally flexible regions that couldn't be resolved experimentally, not just missing data. Rather than simply imputing these values, i think we can implement a specialized approach: (1) maintaining separate tracking of flexible positions, (2) excluding these positions from loss calculation during training via masked loss functions, (3) considering only resolved regions for normalization statistics, and (4) potentially predicting flexibility itself as an auxiliary task. This biologically-informed method preserves the important distinction between truly missing data and intrinsically flexible RNA regions that may have functional significance.</p>",
      "rawMarkdown": "According to Rhiju, NaN values in our dataset represent structurally flexible regions that couldn't be resolved experimentally, not just missing data. Rather than simply imputing these values, i think we can implement a specialized approach: (1) maintaining separate tracking of flexible positions, (2) excluding these positions from loss calculation during training via masked loss functions, (3) considering only resolved regions for normalization statistics, and (4) potentially predicting flexibility itself as an auxiliary task. This biologically-informed method preserves the important distinction between truly missing data and intrinsically flexible RNA regions that may have functional significance.",
      "votes": null
    },
    {
      "id": "3179007",
      "postDate": "04/14/2025 22:05:49",
      "content": "<p>I'm thinking create 2 separate training datasets with different learning rates.  Higher lr for #1 and lower lr for #2.</p>\n<ol>\n<li>A clean dataset that drop any nan/null groupby pdb_id.  </li>\n<li>A synthetic dataset with multiple variations where you'd fill nan/null with values predicted by an LLM fine-tuned with the clean dataset from #1 using something like self-supervised learning.</li>\n</ol>",
      "rawMarkdown": "I'm thinking create 2 separate training datasets with different learning rates.  Higher lr for #1 and lower lr for #2.\n1. A clean dataset that drop any nan/null groupby pdb_id.  \n2. A synthetic dataset with multiple variations where you'd fill nan/null with values predicted by an LLM fine-tuned with the clean dataset from #1 using something like self-supervised learning.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3138269,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "03/02/2025 12:08:33",
      "content": "<p>For now, I’ve decided to \"extract\" the correct sequences if they are at least 6 elements long. I’m not sure if this is the right approach, but I don’t want to lose them. On the other hand, replacing them (NaN) with something else doesn’t seem right to me. What other options are there?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3138322,
          "author_name": "louisstefanuto",
          "author_url": "",
          "post_date": "03/02/2025 12:52:49",
          "content": "<p>If it comes from a physical phenomenon, I am pretty sure there are similar issues with proteins</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3138523,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "03/02/2025 16:33:04",
      "content": "<p>Hi! nan/null at a position means that experimentally, that part of the RNA could not be resolved. <br>\nIt could be due to flexibility of those segments, which blurs out their density in experimental techniques like crystallography or cryo-EM. <br>\nIt's up to you to explore how to best handle this information -- looking forward to the discussion!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3138603,
          "author_name": "louisstefanuto",
          "author_url": "",
          "post_date": "03/02/2025 18:21:50",
          "content": "<p>thank you for the answer!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3139205,
          "author_name": "minhtu123",
          "author_url": "",
          "post_date": "03/03/2025 09:08:26",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>, do the test set labels (public/private) contain nan values?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3140255,
          "author_name": "seangormann",
          "author_url": "",
          "post_date": "03/04/2025 12:35:35",
          "content": "<p>Thanks for the answer. Another note on the data - at least one of these seems a little wild:</p>\n<p>target id == 6EVJ_V</p>\n<pre><code>  resname resid   x_1 y_1 z_1\n</code></pre>\n<p>81005    6EVJ_V_1    A   1   96.825996   795.460022  -324.712006<br>\n81006    6EVJ_V_2    G   93.570999   796.059998  -320.483002<br>\n81007    6EVJ_V_3    U   3   89.374001   796.434021  -317.442993<br>\n81008    6EVJ_V_4    A   85.195999   793.974976  -317.768005<br>\n81009    6EVJ_V_5    G   5   81.110001            789.883972 -318.579010<br>\n81010    6EVJ_V_6    U   6   75.205002   788.846008  -321.894012<br>\n81011    6EVJ_V_7    A   7   77.397003   795.304993  -323.020996<br>\n81012    6EVJ_V_8    A   8   80.192001   798.466980  -322.545013<br>\n81013    6EVJ_V_9    C   9   84.644997   801.317017  -321.303009<br>\n81014    6EVJ_V_10   A   89.483002   802.109009  -323.695007<br>\n81015    6EVJ_V_11   A   86.855003   810.286011  -323.160004<br>\n81016    6EVJ_V_12   G   82.238998   812.549988  -325.160004<br>\n81017    6EVJ_V_13   A   80.735001   817.119019  -327.610992<br>\n81018    6EVJ_V_14   G   81.975998   821.786011  -330.247986<br>\n81019    6EVJ_V_15   G   84.862000   826.000000  -333.403992<br>\n81020    6EVJ_V_16   G   NaN NaN NaN</p>\n<p>The y and z co-ordinates seem massive relative to other samples (typically in the range of -30 - +30). Is this expected? </p>",
          "votes": null,
          "replies": [
            {
              "id": 3140315,
              "author_name": "seangormann",
              "author_url": "",
              "post_date": "03/04/2025 13:42:46",
              "content": "<p>Nevermind, I think I found my answer:</p>\n<p>Molecules are rotated and translated in 3D space to minimize the Root Mean Square Deviation (RMSD). RMSD is a good proxy for the TM-score and they are inversely correlated.</p>\n<p>The orientation you submit is irrelevant, as the coordinate space is arbitrary. Coordinates will always be manipulated to create the best alignment with target structures.</p>\n<p>From: <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566056\" target=\"_blank\">Here</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3140977,
      "author_name": "waqas288",
      "author_url": "",
      "post_date": "03/05/2025 05:21:35",
      "content": "<p>I am thinking do delete all types of null values rows. Just work with full data. Is it a valid point? or should I work around filling data ? <br>\nI think it might be unrealistic on bio data to auto fill it by any technique. </p>\n<p>Please share your perspectives. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3145596,
      "author_name": "explorersanchorage",
      "author_url": "",
      "post_date": "03/10/2025 02:53:31",
      "content": "<p>I am interested to see how others approach this. As missing rows made up 4.5% of all the data, for now I decided to simply drop them, but when I eventually submit and get a score, I plan to see if I can push up the score by working with this missing data. I noted that most RNA molecules x, y, z coordinates are fairly similar, implying you could fill the missing values by finding a pattern within each RNA molecule.</p>\n<p>A blanket approach, e.g. applying the mean seems wrong in this context as each RNA is a seperate molecule.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3158002,
      "author_name": "ivantang86",
      "author_url": "",
      "post_date": "03/24/2025 06:16:07",
      "content": "<p>According to Rhiju, NaN values in our dataset represent structurally flexible regions that couldn't be resolved experimentally, not just missing data. Rather than simply imputing these values, i think we can implement a specialized approach: (1) maintaining separate tracking of flexible positions, (2) excluding these positions from loss calculation during training via masked loss functions, (3) considering only resolved regions for normalization statistics, and (4) potentially predicting flexibility itself as an auxiliary task. This biologically-informed method preserves the important distinction between truly missing data and intrinsically flexible RNA regions that may have functional significance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3179007,
      "author_name": "christie",
      "author_url": "",
      "post_date": "04/14/2025 22:05:49",
      "content": "<p>I'm thinking create 2 separate training datasets with different learning rates.  Higher lr for #1 and lower lr for #2.</p>\n<ol>\n<li>A clean dataset that drop any nan/null groupby pdb_id.  </li>\n<li>A synthetic dataset with multiple variations where you'd fill nan/null with values predicted by an LLM fine-tuned with the clean dataset from #1 using something like self-supervised learning.</li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3138165": "Hi\nI noticed some missing values in the training labels. The missing values are often the positions of the first/last residues of the sequences.\n\nSome examples:\n9FO9_A_32,C,32,25.631000518798828,-6.853000164031982,-0.8730000257492065\n9FO9_A_33,C,33,29.976999282836918,-3.688999891281128,1.368000030517578\n8XPP_B_1,G,1,,,\n8XPP_B_2,G,2,9.82800006866455,-17.54800033569336,23.549999237060547\n8XPP_B_3,G,3,11.680999755859377,-18.09499931335449,18.58300018310547\n8XPP_B_4,A,4,16.104999542236328,-19.54599952697754,15.269000053405762\n8XPP_B_5,G,5,20.43600082397461,-21.989999771118164,14.366000175476074\n\n\n8Z1F_T_58,U,58,95.94200134277344,107.73699951171876,119.09300231933594\n8Z1F_T_59,U,59,97.62100219726562,111.9990005493164,121.96900177001952\n8Z1F_T_60,U,60,101.26000213623048,115.4540023803711,122.81999969482422\n8Z1F_T_61,U,61,107.6780014038086,117.89700317382812,122.1959991455078\n8Z1F_T_62,U,62,112.51699829101562,117.88099670410156,119.24500274658205\n8Z1F_T_63,A,63,115.29299926757812,116.5719985961914,114.8270034790039\n8Z1F_T_64,C,64,115.85700225830078,114.59500122070312,109.50900268554688\n8Z1F_T_65,C,65,113.81600189208984,113.23600006103516,104.33999633789062\n8Z1F_T_66,A,66,118.27999877929688,113.88300323486328,98.0719985961914\n8Z1F_T_67,G,67,,,\n8Z1F_T_68,C,68,,,\n8Z1F_T_69,U,69,,,\n8Z1F_T_70,C,70,,,\n8Z1F_T_71,C,71,,,\n8Z1F_T_72,G,72,,,\n8Z1F_T_73,A,73,,,\n8Z1F_T_74,G,74,,,\n8Z1F_T_75,G,75,,,\n8Z1F_T_76,U,76,,,\n8Z1F_T_77,G,77,,,\n8Z1F_T_78,A,78,,,\n8Z1F_T_79,U,79,,,\n8Z1F_T_80,U,80,,,\n8Z1F_T_81,U,81,,,\n8Z1F_T_82,U,82,,,\n8Z1F_T_83,C,83,,,\n8Z1F_T_84,A,84,,,\n8Z1F_T_85,U,85,,,\n8Z1F_T_86,A,86,,,\n\nMissing values can also be positions of in-the-middle residues:\n1FOQ_A_3,A,3,-38.89099884033203,-23.388999938964844,2.510999917984009\n1FOQ_A_4,A,4,-34.917999267578125,-27.291000366210938,2.7679998874664307\n1FOQ_A_5,U,5,,,\n1FOQ_A_6,G,6,-33.224998474121094,-32.5359992980957,3.601999998092652\n1FOQ_A_7,G,7,-34.20399856567383,-38.005001068115234,4.061999797821045\n\n### Questions \n\n- Why are some positions missing? (my 2 cents is that some residues may move a lot and thus have no clear static position?)\n- Should we drop these NaN, or does it make sense to keep them ? (if they add some information)\n\nThank you for your help",
    "3138269": "For now, I’ve decided to \"extract\" the correct sequences if they are at least 6 elements long. I’m not sure if this is the right approach, but I don’t want to lose them. On the other hand, replacing them (NaN) with something else doesn’t seem right to me. What other options are there?",
    "3138322": "If it comes from a physical phenomenon, I am pretty sure there are similar issues with proteins",
    "3138523": "Hi! nan/null at a position means that experimentally, that part of the RNA could not be resolved. \nIt could be due to flexibility of those segments, which blurs out their density in experimental techniques like crystallography or cryo-EM. \nIt's up to you to explore how to best handle this information -- looking forward to the discussion!",
    "3138603": "thank you for the answer!",
    "3139205": "Hi @rhijudas, do the test set labels (public/private) contain nan values?",
    "3140255": "Thanks for the answer. Another note on the data - at least one of these seems a little wild:\n\ntarget id == 6EVJ_V\n\n\tID\tresname\tresid\tx_1\ty_1\tz_1\n81005\t6EVJ_V_1\tA\t1\t96.825996\t795.460022\t-324.712006\n81006\t6EVJ_V_2\tG\t93.570999\t796.059998\t-320.483002\n81007\t6EVJ_V_3\tU\t3\t89.374001\t796.434021\t-317.442993\n81008\t6EVJ_V_4\tA\t85.195999\t793.974976\t-317.768005\n81009\t6EVJ_V_5\tG\t5\t81.110001\t         789.883972\t-318.579010\n81010\t6EVJ_V_6\tU\t6\t75.205002\t788.846008\t-321.894012\n81011\t6EVJ_V_7\tA\t7\t77.397003\t795.304993\t-323.020996\n81012\t6EVJ_V_8\tA\t8\t80.192001\t798.466980\t-322.545013\n81013\t6EVJ_V_9\tC\t9\t84.644997\t801.317017\t-321.303009\n81014\t6EVJ_V_10\tA\t89.483002\t802.109009\t-323.695007\n81015\t6EVJ_V_11\tA\t86.855003\t810.286011\t-323.160004\n81016\t6EVJ_V_12\tG\t82.238998\t812.549988\t-325.160004\n81017\t6EVJ_V_13\tA\t80.735001\t817.119019\t-327.610992\n81018\t6EVJ_V_14\tG\t81.975998\t821.786011\t-330.247986\n81019\t6EVJ_V_15\tG\t84.862000\t826.000000\t-333.403992\n81020\t6EVJ_V_16\tG\tNaN\tNaN\tNaN\n\n\nThe y and z co-ordinates seem massive relative to other samples (typically in the range of -30 - +30). Is this expected?",
    "3140315": "Nevermind, I think I found my answer:\n\n\nMolecules are rotated and translated in 3D space to minimize the Root Mean Square Deviation (RMSD). RMSD is a good proxy for the TM-score and they are inversely correlated.\n\nThe orientation you submit is irrelevant, as the coordinate space is arbitrary. Coordinates will always be manipulated to create the best alignment with target structures.\n\nFrom: [Here](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566056)",
    "3140977": "I am thinking do delete all types of null values rows. Just work with full data. Is it a valid point? or should I work around filling data ? \nI think it might be unrealistic on bio data to auto fill it by any technique. \n\nPlease share your perspectives. Thanks!",
    "3145596": "I am interested to see how others approach this. As missing rows made up 4.5% of all the data, for now I decided to simply drop them, but when I eventually submit and get a score, I plan to see if I can push up the score by working with this missing data. I noted that most RNA molecules x, y, z coordinates are fairly similar, implying you could fill the missing values by finding a pattern within each RNA molecule.\n\nA blanket approach, e.g. applying the mean seems wrong in this context as each RNA is a seperate molecule.",
    "3158002": "According to Rhiju, NaN values in our dataset represent structurally flexible regions that couldn't be resolved experimentally, not just missing data. Rather than simply imputing these values, i think we can implement a specialized approach: (1) maintaining separate tracking of flexible positions, (2) excluding these positions from loss calculation during training via masked loss functions, (3) considering only resolved regions for normalization statistics, and (4) potentially predicting flexibility itself as an auxiliary task. This biologically-informed method preserves the important distinction between truly missing data and intrinsically flexible RNA regions that may have functional significance.",
    "3179007": "I'm thinking create 2 separate training datasets with different learning rates.  Higher lr for #1 and lower lr for #2.\n1. A clean dataset that drop any nan/null groupby pdb_id.  \n2. A synthetic dataset with multiple variations where you'd fill nan/null with values predicted by an LLM fine-tuned with the clean dataset from #1 using something like self-supervised learning."
  },
  "source": "meta"
}