{
  "id": 582295,
  "title": "Brief Explanation of My Temporary 1st-place Solution (Expect Significant Drop)",
  "url": "/competitions/stanford-rna-3d-folding/discussion/582295",
  "author_name": "",
  "post_date": "2025-05-30T09:30:02.144253100Z",
  "votes": 20,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I performed template based modeling with very recent (2025-05-21) pdb dataset.<br>\n(+missing regions were filled by boltz-1 or chai-1.)<br>\nTherefore, a significant drop in ranking is expected.<br>\nI'm looking forward to see the final results with more robust solutions. Good luck to everyone!</p>\n<p>Update 2025/06/17<br>\nI have attached slides with more detailed explanations. Please feel free to leave any comments if you have any questions.</p>",
  "messages": [
    {
      "id": "3213645",
      "postDate": "05/30/2025 09:30:02",
      "content": "<p>I performed template based modeling with very recent (2025-05-21) pdb dataset.<br>\n(+missing regions were filled by boltz-1 or chai-1.)<br>\nTherefore, a significant drop in ranking is expected.<br>\nI'm looking forward to see the final results with more robust solutions. Good luck to everyone!</p>\n<p>Update 2025/06/17<br>\nI have attached slides with more detailed explanations. Please feel free to leave any comments if you have any questions.</p>",
      "rawMarkdown": "I performed template based modeling with very recent (2025-05-21) pdb dataset.\n(+missing regions were filled by boltz-1 or chai-1.)\nTherefore, a significant drop in ranking is expected.\nI'm looking forward to see the final results with more robust solutions. Good luck to everyone!\n\n\nUpdate 2025/06/17\nI have attached slides with more detailed explanations. Please feel free to leave any comments if you have any questions.",
      "votes": null
    },
    {
      "id": "3213655",
      "postDate": "05/30/2025 09:34:44",
      "content": "<p>&amp; The highest public leader board score of my submission is 0.605.<br>\nI didn't think my last submission made such a large improvement.</p>",
      "rawMarkdown": "& The highest public leader board score of my submission is 0.605.\nI didn't think my last submission made such a large improvement.",
      "votes": null
    },
    {
      "id": "3213678",
      "postDate": "05/30/2025 10:17:18",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/odat1248\" target=\"_blank\">@odat1248</a> Thanks for sharing. For template matching, did you use this code?<br>\n<a href=\"https://www.kaggle.com/code/fernandosr85/rna-3d-structure\" target=\"_blank\">https://www.kaggle.com/code/fernandosr85/rna-3d-structure</a></p>\n<p>Or did you have some other script?</p>",
      "rawMarkdown": "Hi @odat1248 Thanks for sharing. For template matching, did you use this code?\nhttps://www.kaggle.com/code/fernandosr85/rna-3d-structure\n\nOr did you have some other script?",
      "votes": null
    },
    {
      "id": "3213681",
      "postDate": "05/30/2025 10:20:59",
      "content": "<p>No. I wrote the code by myself.</p>",
      "rawMarkdown": "No. I wrote the code by myself.",
      "votes": null
    },
    {
      "id": "3213698",
      "postDate": "05/30/2025 10:44:22",
      "content": "<blockquote>\n  <p>I performed template based modeling with very recent (2025-05-21) pdb dataset.</p>\n</blockquote>\n<p>hahahaha, thanks for your honesty, I figured it out a while ago. No offense, just kidding.</p>\n<p>I realized that a significant portion of the current public test dataset is from CASP16, so I used the RNA sequences from CASP16 and over 600 training sets from Kaggle for SFT, based on protenix and drfold. My boss chewed me out for using up too much of the company's server resources, LOL. My actual training time was very limited, and most of the work involved post-processing techniques. The key is robustness, absolutely.</p>\n<p>I checked the test results of CASP16, so I know what the upper limit of performance looks like in this competition.</p>",
      "rawMarkdown": ">I performed template based modeling with very recent (2025-05-21) pdb dataset.\n\nhahahaha, thanks for your honesty, I figured it out a while ago. No offense, just kidding.\n\nI realized that a significant portion of the current public test dataset is from CASP16, so I used the RNA sequences from CASP16 and over 600 training sets from Kaggle for SFT, based on protenix and drfold. My boss chewed me out for using up too much of the company's server resources, LOL. My actual training time was very limited, and most of the work involved post-processing techniques. The key is robustness, absolutely.\n\nI checked the test results of CASP16, so I know what the upper limit of performance looks like in this competition.",
      "votes": null
    },
    {
      "id": "3213700",
      "postDate": "05/30/2025 10:46:03",
      "content": "<p>I was expecting it, but thank you for confirming it clearly.</p>",
      "rawMarkdown": "I was expecting it, but thank you for confirming it clearly.",
      "votes": null
    },
    {
      "id": "3213714",
      "postDate": "05/30/2025 11:04:21",
      "content": "<p>hi, i think the submission went through despite exceeding the deadline? (not too sure though), probably because this competition didn't have a private test set.</p>\n<p>here's a screenshot from <a href=\"https://www.kaggle.com/competitions/drawing-with-llms/overview\" target=\"_blank\">Drawing with LLMs</a>, where I learned about this. and for RNA competition, I miscalculated how many submissions i had left, so i had to stay up till the deadline to make sure all submissions were processed in time. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14464632%2Fb496d2c1cd42e79532e85b06716b5876%2FScreenshot%202025-05-30%20161420.png?generation=1748601872059961&amp;alt=media\" alt=\"\"></p>\n<p>either way, learned this the hard way, i think more people should know about it.</p>",
      "rawMarkdown": "hi, i think the submission went through despite exceeding the deadline? (not too sure though), probably because this competition didn't have a private test set.\n\nhere's a screenshot from [Drawing with LLMs](https://www.kaggle.com/competitions/drawing-with-llms/overview), where I learned about this. and for RNA competition, I miscalculated how many submissions i had left, so i had to stay up till the deadline to make sure all submissions were processed in time. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14464632%2Fb496d2c1cd42e79532e85b06716b5876%2FScreenshot%202025-05-30%20161420.png?generation=1748601872059961&alt=media)\n\neither way, learned this the hard way, i think more people should know about it.",
      "votes": null
    },
    {
      "id": "3213727",
      "postDate": "05/30/2025 11:15:40",
      "content": "<p> Ah, the highest scoring notebook had potential bug thus I had to manually select 2 notebooks.<br>\nAnd the selection can not be changed after deadline.<br>\nBut I submitted the last submission thinking it would have lower score because I used the lessor scoring function.</p>",
      "rawMarkdown": "~~I learned it while ago (CAFA5 challenge).~~ Ah, the highest scoring notebook had potential bug thus I had to manually select 2 notebooks.\nAnd the selection can not be changed after deadline.\nBut I submitted the last submission thinking it would have lower score because I used the lessor scoring function.",
      "votes": null
    },
    {
      "id": "3213740",
      "postDate": "05/30/2025 11:28:58",
      "content": "<p>If I may ask, what kind of post processings did you perform? I tried ranking 20 structures and taking top5 with LGBM ranker (evaluated map@5). This worked way better than random locally but didnt seem to have an effect on the LB despite my local validation had no leaks</p>",
      "rawMarkdown": "If I may ask, what kind of post processings did you perform? I tried ranking 20 structures and taking top5 with LGBM ranker (evaluated map@5). This worked way better than random locally but didnt seem to have an effect on the LB despite my local validation had no leaks",
      "votes": null
    },
    {
      "id": "3213745",
      "postDate": "05/30/2025 11:33:23",
      "content": "<p>Simply put, I used various models to generate a total of 15 sets of coordinates and selected the best 5 coords for the corresponding RNA bases through energy calculations.</p>",
      "rawMarkdown": "Simply put, I used various models to generate a total of 15 sets of coordinates and selected the best 5 coords for the corresponding RNA bases through energy calculations.",
      "votes": null
    },
    {
      "id": "3216161",
      "postDate": "06/03/2025 07:26:49",
      "content": "<p>i think the public test data has some pdb RNA after the allowed cutoff date.<br>\nfrom our experiments, we think this is current state of art:</p>\n<ol>\n<li>fully automatic (e.g. ensemble openoursce models like Rosetta, blotz-1, proteinx … and finetune version of them:<br>\npublic lb: 0.38 to 0.40</li>\n<li>(1) above to zero filling : 0.40 to 0.50  </li>\n<li>(1) + human konwledge : 0.50 to 0.65</li>\n<li>(1) + dataleak : 0.55 to 0.60</li>\n</ol>",
      "rawMarkdown": "i think the public test data has some pdb RNA after the allowed cutoff date.\nfrom our experiments, we think this is current state of art:\n1. fully automatic (e.g. ensemble openoursce models like Rosetta, blotz-1, proteinx ... and finetune version of them:\npublic lb: 0.38 to 0.40\n2. (1) above to zero filling : 0.40 to 0.50  \n3. (1) + human konwledge : 0.50 to 0.65\n4. (1) + dataleak : 0.55 to 0.60",
      "votes": null
    },
    {
      "id": "3216639",
      "postDate": "06/03/2025 21:55:40",
      "content": "<p>Why do you need to zero fill? Models like proteinX can make predictions for full sequences already</p>",
      "rawMarkdown": "Why do you need to zero fill? Models like proteinX can make predictions for full sequences already",
      "votes": null
    },
    {
      "id": "3216896",
      "postDate": "06/04/2025 08:53:27",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> sequence length more than 1000 cuda out of memory</p>",
      "rawMarkdown": "shujun717 sequence length more than 1000 cuda out of memory",
      "votes": null
    },
    {
      "id": "3217358",
      "postDate": "06/04/2025 23:51:55",
      "content": "<p>I got OOM after length ~850 with Protenix (using MSA), idk the length without MSA</p>",
      "rawMarkdown": "I got OOM after length ~850 with Protenix (using MSA), idk the length without MSA",
      "votes": null
    },
    {
      "id": "3220850",
      "postDate": "06/10/2025 04:00:09",
      "content": "<p>As I expected, your score is horrible.😂</p>",
      "rawMarkdown": "As I expected, your score is horrible.😂",
      "votes": null
    },
    {
      "id": "3225633",
      "postDate": "06/16/2025 17:06:12",
      "content": "<p>Thank Oda (@odat1248) for sharing your solution especially your clever way to use embedding from RibonanzaNet for similar sequence search! Have you also tried the embeddings from other models like AF3, DrFold, etc. ?</p>",
      "rawMarkdown": "Thank Oda (@odat1248) for sharing your solution especially your clever way to use embedding from RibonanzaNet for similar sequence search! Have you also tried the embeddings from other models like AF3, DrFold, etc. ?",
      "votes": null
    },
    {
      "id": "3225650",
      "postDate": "06/16/2025 17:21:52",
      "content": "<p>I tried RNAErnie provided by multimolecules <a href=\"https://huggingface.co/multimolecule/rnaernie\" target=\"_blank\">https://huggingface.co/multimolecule/rnaernie</a> . (Official RNAErnie implementation is <a href=\"https://github.com/CatIIIIIIII/RNAErnie\" target=\"_blank\">https://github.com/CatIIIIIIII/RNAErnie</a> .)<br>\nIn my preliminary experiments RNAErnie was slightly superior, however, since RNAErnie requires more disk space due to its larger embedding size, I selected RibonanzaNet.</p>",
      "rawMarkdown": "I tried RNAErnie provided by multimolecules https://huggingface.co/multimolecule/rnaernie . (Official RNAErnie implementation is https://github.com/CatIIIIIIII/RNAErnie .)\nIn my preliminary experiments RNAErnie was slightly superior, however, since RNAErnie requires more disk space due to its larger embedding size, I selected RibonanzaNet.",
      "votes": null
    },
    {
      "id": "3225785",
      "postDate": "06/16/2025 20:44:23",
      "content": "<p>Thanks again for sharing! It’s really great approach for RNA template search! I’ve learned a lot!</p>",
      "rawMarkdown": "Thanks again for sharing! It’s really great approach for RNA template search! I’ve learned a lot!",
      "votes": null
    },
    {
      "id": "3295530",
      "postDate": "09/29/2025 03:39:50",
      "content": "<p>My last submission notebook (0.671 in the public leaderboard <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/608277\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/608277</a> ) scored 0.56966 in the private leaderboard.<br>\nhm, nice.</p>",
      "rawMarkdown": "My last submission notebook (0.671 in the public leaderboard https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/608277 ) scored 0.56966 in the private leaderboard.\nhm, nice.",
      "votes": null
    },
    {
      "id": "3295532",
      "postDate": "09/29/2025 03:43:42",
      "content": "<p>Good work! It shows that previous results is not overfitting and work generally well! You probably created a new direction in rna structure prediction.</p>",
      "rawMarkdown": "Good work! It shows that previous results is not overfitting and work generally well! You probably created a new direction in rna structure prediction.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3213655,
      "author_name": "odat1248",
      "author_url": "",
      "post_date": "05/30/2025 09:34:44",
      "content": "<p>&amp; The highest public leader board score of my submission is 0.605.<br>\nI didn't think my last submission made such a large improvement.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3295530,
          "author_name": "odat1248",
          "author_url": "",
          "post_date": "09/29/2025 03:39:50",
          "content": "<p>My last submission notebook (0.671 in the public leaderboard <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/608277\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/608277</a> ) scored 0.56966 in the private leaderboard.<br>\nhm, nice.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3295532,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "09/29/2025 03:43:42",
              "content": "<p>Good work! It shows that previous results is not overfitting and work generally well! You probably created a new direction in rna structure prediction.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3213678,
      "author_name": "kaiserm",
      "author_url": "",
      "post_date": "05/30/2025 10:17:18",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/odat1248\" target=\"_blank\">@odat1248</a> Thanks for sharing. For template matching, did you use this code?<br>\n<a href=\"https://www.kaggle.com/code/fernandosr85/rna-3d-structure\" target=\"_blank\">https://www.kaggle.com/code/fernandosr85/rna-3d-structure</a></p>\n<p>Or did you have some other script?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3213681,
          "author_name": "odat1248",
          "author_url": "",
          "post_date": "05/30/2025 10:20:59",
          "content": "<p>No. I wrote the code by myself.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3213698,
      "author_name": "sweetyheehee",
      "author_url": "",
      "post_date": "05/30/2025 10:44:22",
      "content": "<blockquote>\n  <p>I performed template based modeling with very recent (2025-05-21) pdb dataset.</p>\n</blockquote>\n<p>hahahaha, thanks for your honesty, I figured it out a while ago. No offense, just kidding.</p>\n<p>I realized that a significant portion of the current public test dataset is from CASP16, so I used the RNA sequences from CASP16 and over 600 training sets from Kaggle for SFT, based on protenix and drfold. My boss chewed me out for using up too much of the company's server resources, LOL. My actual training time was very limited, and most of the work involved post-processing techniques. The key is robustness, absolutely.</p>\n<p>I checked the test results of CASP16, so I know what the upper limit of performance looks like in this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3213740,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "05/30/2025 11:28:58",
          "content": "<p>If I may ask, what kind of post processings did you perform? I tried ranking 20 structures and taking top5 with LGBM ranker (evaluated map@5). This worked way better than random locally but didnt seem to have an effect on the LB despite my local validation had no leaks</p>",
          "votes": null,
          "replies": [
            {
              "id": 3213745,
              "author_name": "sweetyheehee",
              "author_url": "",
              "post_date": "05/30/2025 11:33:23",
              "content": "<p>Simply put, I used various models to generate a total of 15 sets of coordinates and selected the best 5 coords for the corresponding RNA bases through energy calculations.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3213700,
      "author_name": "junhanzangai",
      "author_url": "",
      "post_date": "05/30/2025 10:46:03",
      "content": "<p>I was expecting it, but thank you for confirming it clearly.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3213714,
      "author_name": "jaejohn",
      "author_url": "",
      "post_date": "05/30/2025 11:04:21",
      "content": "<p>hi, i think the submission went through despite exceeding the deadline? (not too sure though), probably because this competition didn't have a private test set.</p>\n<p>here's a screenshot from <a href=\"https://www.kaggle.com/competitions/drawing-with-llms/overview\" target=\"_blank\">Drawing with LLMs</a>, where I learned about this. and for RNA competition, I miscalculated how many submissions i had left, so i had to stay up till the deadline to make sure all submissions were processed in time. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14464632%2Fb496d2c1cd42e79532e85b06716b5876%2FScreenshot%202025-05-30%20161420.png?generation=1748601872059961&amp;alt=media\" alt=\"\"></p>\n<p>either way, learned this the hard way, i think more people should know about it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3213727,
          "author_name": "odat1248",
          "author_url": "",
          "post_date": "05/30/2025 11:15:40",
          "content": "<p> Ah, the highest scoring notebook had potential bug thus I had to manually select 2 notebooks.<br>\nAnd the selection can not be changed after deadline.<br>\nBut I submitted the last submission thinking it would have lower score because I used the lessor scoring function.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3216161,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/03/2025 07:26:49",
      "content": "<p>i think the public test data has some pdb RNA after the allowed cutoff date.<br>\nfrom our experiments, we think this is current state of art:</p>\n<ol>\n<li>fully automatic (e.g. ensemble openoursce models like Rosetta, blotz-1, proteinx … and finetune version of them:<br>\npublic lb: 0.38 to 0.40</li>\n<li>(1) above to zero filling : 0.40 to 0.50  </li>\n<li>(1) + human konwledge : 0.50 to 0.65</li>\n<li>(1) + dataleak : 0.55 to 0.60</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 3216639,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "06/03/2025 21:55:40",
          "content": "<p>Why do you need to zero fill? Models like proteinX can make predictions for full sequences already</p>",
          "votes": null,
          "replies": [
            {
              "id": 3216896,
              "author_name": "arunodhayan",
              "author_url": "",
              "post_date": "06/04/2025 08:53:27",
              "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> sequence length more than 1000 cuda out of memory</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3217358,
                  "author_name": "alejopaullier",
                  "author_url": "",
                  "post_date": "06/04/2025 23:51:55",
                  "content": "<p>I got OOM after length ~850 with Protenix (using MSA), idk the length without MSA</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3220850,
      "author_name": "daoheliu",
      "author_url": "",
      "post_date": "06/10/2025 04:00:09",
      "content": "<p>As I expected, your score is horrible.😂</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3225633,
      "author_name": "nguyenhoa",
      "author_url": "",
      "post_date": "06/16/2025 17:06:12",
      "content": "<p>Thank Oda (@odat1248) for sharing your solution especially your clever way to use embedding from RibonanzaNet for similar sequence search! Have you also tried the embeddings from other models like AF3, DrFold, etc. ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3225650,
          "author_name": "odat1248",
          "author_url": "",
          "post_date": "06/16/2025 17:21:52",
          "content": "<p>I tried RNAErnie provided by multimolecules <a href=\"https://huggingface.co/multimolecule/rnaernie\" target=\"_blank\">https://huggingface.co/multimolecule/rnaernie</a> . (Official RNAErnie implementation is <a href=\"https://github.com/CatIIIIIIII/RNAErnie\" target=\"_blank\">https://github.com/CatIIIIIIII/RNAErnie</a> .)<br>\nIn my preliminary experiments RNAErnie was slightly superior, however, since RNAErnie requires more disk space due to its larger embedding size, I selected RibonanzaNet.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3225785,
              "author_name": "nguyenhoa",
              "author_url": "",
              "post_date": "06/16/2025 20:44:23",
              "content": "<p>Thanks again for sharing! It’s really great approach for RNA template search! I’ve learned a lot!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3213645": "I performed template based modeling with very recent (2025-05-21) pdb dataset.\n(+missing regions were filled by boltz-1 or chai-1.)\nTherefore, a significant drop in ranking is expected.\nI'm looking forward to see the final results with more robust solutions. Good luck to everyone!\n\n\nUpdate 2025/06/17\nI have attached slides with more detailed explanations. Please feel free to leave any comments if you have any questions.",
    "3213655": "& The highest public leader board score of my submission is 0.605.\nI didn't think my last submission made such a large improvement.",
    "3213678": "Hi @odat1248 Thanks for sharing. For template matching, did you use this code?\nhttps://www.kaggle.com/code/fernandosr85/rna-3d-structure\n\nOr did you have some other script?",
    "3213681": "No. I wrote the code by myself.",
    "3213698": ">I performed template based modeling with very recent (2025-05-21) pdb dataset.\n\nhahahaha, thanks for your honesty, I figured it out a while ago. No offense, just kidding.\n\nI realized that a significant portion of the current public test dataset is from CASP16, so I used the RNA sequences from CASP16 and over 600 training sets from Kaggle for SFT, based on protenix and drfold. My boss chewed me out for using up too much of the company's server resources, LOL. My actual training time was very limited, and most of the work involved post-processing techniques. The key is robustness, absolutely.\n\nI checked the test results of CASP16, so I know what the upper limit of performance looks like in this competition.",
    "3213700": "I was expecting it, but thank you for confirming it clearly.",
    "3213714": "hi, i think the submission went through despite exceeding the deadline? (not too sure though), probably because this competition didn't have a private test set.\n\nhere's a screenshot from [Drawing with LLMs](https://www.kaggle.com/competitions/drawing-with-llms/overview), where I learned about this. and for RNA competition, I miscalculated how many submissions i had left, so i had to stay up till the deadline to make sure all submissions were processed in time. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14464632%2Fb496d2c1cd42e79532e85b06716b5876%2FScreenshot%202025-05-30%20161420.png?generation=1748601872059961&alt=media)\n\neither way, learned this the hard way, i think more people should know about it.",
    "3213727": "~~I learned it while ago (CAFA5 challenge).~~ Ah, the highest scoring notebook had potential bug thus I had to manually select 2 notebooks.\nAnd the selection can not be changed after deadline.\nBut I submitted the last submission thinking it would have lower score because I used the lessor scoring function.",
    "3213740": "If I may ask, what kind of post processings did you perform? I tried ranking 20 structures and taking top5 with LGBM ranker (evaluated map@5). This worked way better than random locally but didnt seem to have an effect on the LB despite my local validation had no leaks",
    "3213745": "Simply put, I used various models to generate a total of 15 sets of coordinates and selected the best 5 coords for the corresponding RNA bases through energy calculations.",
    "3216161": "i think the public test data has some pdb RNA after the allowed cutoff date.\nfrom our experiments, we think this is current state of art:\n1. fully automatic (e.g. ensemble openoursce models like Rosetta, blotz-1, proteinx ... and finetune version of them:\npublic lb: 0.38 to 0.40\n2. (1) above to zero filling : 0.40 to 0.50  \n3. (1) + human konwledge : 0.50 to 0.65\n4. (1) + dataleak : 0.55 to 0.60",
    "3216639": "Why do you need to zero fill? Models like proteinX can make predictions for full sequences already",
    "3216896": "shujun717 sequence length more than 1000 cuda out of memory",
    "3217358": "I got OOM after length ~850 with Protenix (using MSA), idk the length without MSA",
    "3220850": "As I expected, your score is horrible.😂",
    "3225633": "Thank Oda (@odat1248) for sharing your solution especially your clever way to use embedding from RibonanzaNet for similar sequence search! Have you also tried the embeddings from other models like AF3, DrFold, etc. ?",
    "3225650": "I tried RNAErnie provided by multimolecules https://huggingface.co/multimolecule/rnaernie . (Official RNAErnie implementation is https://github.com/CatIIIIIIII/RNAErnie .)\nIn my preliminary experiments RNAErnie was slightly superior, however, since RNAErnie requires more disk space due to its larger embedding size, I selected RibonanzaNet.",
    "3225785": "Thanks again for sharing! It’s really great approach for RNA template search! I’ve learned a lot!",
    "3295530": "My last submission notebook (0.671 in the public leaderboard https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/608277 ) scored 0.56966 in the private leaderboard.\nhm, nice.",
    "3295532": "Good work! It shows that previous results is not overfitting and work generally well! You probably created a new direction in rna structure prediction."
  },
  "source": "meta"
}