{
  "id": 584487,
  "title": "Public leaderboard 3rd place solution",
  "url": "/competitions/stanford-rna-3d-folding/discussion/584487",
  "author_name": "",
  "post_date": "2025-06-13T17:33:38.711054500Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thank you hosts for this amazing competition.</p>\n<p>To be a little help, we briefly share the strategies we used below.</p>\n<ol>\n<li>Drfold2 + Protenix (Public 0.589)</li>\n</ol>\n<p>▪ Seq. &lt; 400</p>\n<pre><code>  •   drfold2\n  •   Protenix finetuned RNA   MSA ()\n</code></pre>\n<p>▪ Seq. &gt;= 400</p>\n<pre><code>  •   Protenix finetuned RNA   MSA ()\n  •   Protenix finetuned   MSA ()\n  •   Protenix baseline model\n</code></pre>\n<ol>\n<li>Protenix + Boltz (Public 0.615)<br>\n▪ 2 from Protenix fine-tuned RNA only + MSA (&lt;800)<br>\n▪ 1 from Protenix fine-tuned ALL + MSA (&lt;800)<br>\n▪ 1 from Protenix baseline<br>\n▪ 1 from Boltz baseline</li>\n</ol>\n<p>Below 800 because OOM on 96GB VRAM: GH200.</p>\n<p>RNA only: polymer composition RNA<br>\nALL: v2 data + recent pdb release<br>\nWe included Casp 16 data so that bumps up the public score.<br>\nIf we don't use the casp 16 data, we get about 0.42.</p>\n<p>rMSA pipeline MSA data (used the codes hosts provided):<br>\n<a href=\"https://drive.google.com/drive/folders/15bhXWsR6QuDQo6U4j8Ii8-OE4B4SmC7i?usp=drive_link\" target=\"_blank\">https://drive.google.com/drive/folders/15bhXWsR6QuDQo6U4j8Ii8-OE4B4SmC7i?usp=drive_link</a></p>\n<p>Training Data (for protenix finetuning) and labels:<br>\n<a href=\"https://drive.google.com/drive/folders/1XKYzk2oCcHPt6DB_wLNYL-7w3s1R793n?usp=sharing\" target=\"_blank\">https://drive.google.com/drive/folders/1XKYzk2oCcHPt6DB_wLNYL-7w3s1R793n?usp=sharing</a></p>\n<p>Codes used from :<br>\nProtenix inference<br>\n<a href=\"https://www.kaggle.com/code/geraseva/protenix\" target=\"_blank\">https://www.kaggle.com/code/geraseva/protenix</a><br>\n<a href=\"https://www.kaggle.com/geraseva\" target=\"_blank\">@geraseva</a> </p>\n<p>Protenix finetuning<br>\n<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495</a><br>\n<a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> </p>\n<p>Boltz inference<br>\n<a href=\"https://www.kaggle.com/code/youhanlee/boltz-1-inference-submission\" target=\"_blank\">https://www.kaggle.com/code/youhanlee/boltz-1-inference-submission</a><br>\n<a href=\"https://www.kaggle.com/youhanlee\" target=\"_blank\">@youhanlee</a> </p>",
  "messages": [
    {
      "id": "3223690",
      "postDate": "06/13/2025 17:33:38",
      "content": "<p>Thank you hosts for this amazing competition.</p>\n<p>To be a little help, we briefly share the strategies we used below.</p>\n<ol>\n<li>Drfold2 + Protenix (Public 0.589)</li>\n</ol>\n<p>▪ Seq. &lt; 400</p>\n<pre><code>  •   drfold2\n  •   Protenix finetuned RNA   MSA ()\n</code></pre>\n<p>▪ Seq. &gt;= 400</p>\n<pre><code>  •   Protenix finetuned RNA   MSA ()\n  •   Protenix finetuned   MSA ()\n  •   Protenix baseline model\n</code></pre>\n<ol>\n<li>Protenix + Boltz (Public 0.615)<br>\n▪ 2 from Protenix fine-tuned RNA only + MSA (&lt;800)<br>\n▪ 1 from Protenix fine-tuned ALL + MSA (&lt;800)<br>\n▪ 1 from Protenix baseline<br>\n▪ 1 from Boltz baseline</li>\n</ol>\n<p>Below 800 because OOM on 96GB VRAM: GH200.</p>\n<p>RNA only: polymer composition RNA<br>\nALL: v2 data + recent pdb release<br>\nWe included Casp 16 data so that bumps up the public score.<br>\nIf we don't use the casp 16 data, we get about 0.42.</p>\n<p>rMSA pipeline MSA data (used the codes hosts provided):<br>\n<a href=\"https://drive.google.com/drive/folders/15bhXWsR6QuDQo6U4j8Ii8-OE4B4SmC7i?usp=drive_link\" target=\"_blank\">https://drive.google.com/drive/folders/15bhXWsR6QuDQo6U4j8Ii8-OE4B4SmC7i?usp=drive_link</a></p>\n<p>Training Data (for protenix finetuning) and labels:<br>\n<a href=\"https://drive.google.com/drive/folders/1XKYzk2oCcHPt6DB_wLNYL-7w3s1R793n?usp=sharing\" target=\"_blank\">https://drive.google.com/drive/folders/1XKYzk2oCcHPt6DB_wLNYL-7w3s1R793n?usp=sharing</a></p>\n<p>Codes used from :<br>\nProtenix inference<br>\n<a href=\"https://www.kaggle.com/code/geraseva/protenix\" target=\"_blank\">https://www.kaggle.com/code/geraseva/protenix</a><br>\n<a href=\"https://www.kaggle.com/geraseva\" target=\"_blank\">@geraseva</a> </p>\n<p>Protenix finetuning<br>\n<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495</a><br>\n<a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> </p>\n<p>Boltz inference<br>\n<a href=\"https://www.kaggle.com/code/youhanlee/boltz-1-inference-submission\" target=\"_blank\">https://www.kaggle.com/code/youhanlee/boltz-1-inference-submission</a><br>\n<a href=\"https://www.kaggle.com/youhanlee\" target=\"_blank\">@youhanlee</a> </p>",
      "rawMarkdown": "Thank you hosts for this amazing competition.\n\nTo be a little help, we briefly share the strategies we used below.\n\n1. Drfold2 + Protenix (Public 0.589)\n\n▪ Seq. < 400\n\n      • 3 from drfold2\n      • 2 from Protenix fine-tuned RNA only + MSA (<800)\n\n▪ Seq. >= 400\n\n      • 2 from Protenix fine-tuned RNA only + MSA (<800)\n      • 2 from Protenix fine-tuned All + MSA (<800)\n      • 1 from Protenix baseline model\n\n2. Protenix + Boltz (Public 0.615)\n▪ 2 from Protenix fine-tuned RNA only + MSA (<800)\n▪ 1 from Protenix fine-tuned ALL + MSA (<800)\n▪ 1 from Protenix baseline\n▪ 1 from Boltz baseline\n\nBelow 800 because OOM on 96GB VRAM: GH200.\n\nRNA only: polymer composition RNA\nALL: v2 data + recent pdb release\nWe included Casp 16 data so that bumps up the public score.\nIf we don't use the casp 16 data, we get about 0.42.\n\nrMSA pipeline MSA data (used the codes hosts provided):\nhttps://drive.google.com/drive/folders/15bhXWsR6QuDQo6U4j8Ii8-OE4B4SmC7i?usp=drive_link\n\nTraining Data (for protenix finetuning) and labels:\nhttps://drive.google.com/drive/folders/1XKYzk2oCcHPt6DB_wLNYL-7w3s1R793n?usp=sharing\n\n\nCodes used from :\nProtenix inference\nhttps://www.kaggle.com/code/geraseva/protenix\n@geraseva \n\nProtenix finetuning\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495\n@lihaoweicvch \n\nBoltz inference\nhttps://www.kaggle.com/code/youhanlee/boltz-1-inference-submission\n@youhanlee",
      "votes": null
    },
    {
      "id": "3223918",
      "postDate": "06/14/2025 05:46:25",
      "content": "<p>Very helpful, your solution is amazing! thx</p>",
      "rawMarkdown": "Very helpful, your solution is amazing! thx",
      "votes": null
    },
    {
      "id": "3225269",
      "postDate": "06/16/2025 07:54:16",
      "content": "<p>Thank you for sharing, but could your fine-tuning strategy potentially lead to excessive leakage? It's possible that only <strong>3 from DrFold2</strong> and <strong>1 from the ProteinX baseline model</strong>  actually works, hard to say, live or die?</p>",
      "rawMarkdown": "Thank you for sharing, but could your fine-tuning strategy potentially lead to excessive leakage? It's possible that only **3 from DrFold2** and **1 from the ProteinX baseline model**  actually works, hard to say, live or die?",
      "votes": null
    },
    {
      "id": "3225312",
      "postDate": "06/16/2025 08:40:30",
      "content": "<p>Well, we did try fine-tuning without the leakage data and it still improved.</p>\n<p>Baseline model with no fine-tuning had a public test score of 3.7,<br>\nwhile fine-tuned model with v1 data + MSA had a public test score of 4.12.<br>\nSo this suggests that fine-tuning with MSAs does work, so we used all the available data for training (more data the better..?)</p>\n<p>One thing to notice was if we fine-tune just 1-2 epoch (1 or 2 full iterations of the training dataset) the performance dropped (possibility due to the optimizer as <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> suggested). <br>\nSo we had to train for more than 5 epochs for performance boost.</p>\n<p>Meanwhile we did work on 2 strategies but the drfold2 one didn't go through for submissions because we didn't select that one :(</p>",
      "rawMarkdown": "Well, we did try fine-tuning without the leakage data and it still improved.\n\nBaseline model with no fine-tuning had a public test score of 3.7,\nwhile fine-tuned model with v1 data + MSA had a public test score of 4.12.\nSo this suggests that fine-tuning with MSAs does work, so we used all the available data for training (more data the better..?)\n\nOne thing to notice was if we fine-tune just 1-2 epoch (1 or 2 full iterations of the training dataset) the performance dropped (possibility due to the optimizer as @shujun717 suggested). \nSo we had to train for more than 5 epochs for performance boost.\n\nMeanwhile we did work on 2 strategies but the drfold2 one didn't go through for submissions because we didn't select that one :(",
      "votes": null
    },
    {
      "id": "3226043",
      "postDate": "06/17/2025 07:22:59",
      "content": "<p>Very helpful work</p>",
      "rawMarkdown": "Very helpful work",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3223918,
      "author_name": "",
      "author_url": "",
      "post_date": "06/14/2025 05:46:25",
      "content": "<p>Very helpful, your solution is amazing! thx</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3225269,
      "author_name": "sweetyheehee",
      "author_url": "",
      "post_date": "06/16/2025 07:54:16",
      "content": "<p>Thank you for sharing, but could your fine-tuning strategy potentially lead to excessive leakage? It's possible that only <strong>3 from DrFold2</strong> and <strong>1 from the ProteinX baseline model</strong>  actually works, hard to say, live or die?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3225312,
          "author_name": "yekim102",
          "author_url": "",
          "post_date": "06/16/2025 08:40:30",
          "content": "<p>Well, we did try fine-tuning without the leakage data and it still improved.</p>\n<p>Baseline model with no fine-tuning had a public test score of 3.7,<br>\nwhile fine-tuned model with v1 data + MSA had a public test score of 4.12.<br>\nSo this suggests that fine-tuning with MSAs does work, so we used all the available data for training (more data the better..?)</p>\n<p>One thing to notice was if we fine-tune just 1-2 epoch (1 or 2 full iterations of the training dataset) the performance dropped (possibility due to the optimizer as <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> suggested). <br>\nSo we had to train for more than 5 epochs for performance boost.</p>\n<p>Meanwhile we did work on 2 strategies but the drfold2 one didn't go through for submissions because we didn't select that one :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3226043,
      "author_name": "starlitdreamstar",
      "author_url": "",
      "post_date": "06/17/2025 07:22:59",
      "content": "<p>Very helpful work</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3223690": "Thank you hosts for this amazing competition.\n\nTo be a little help, we briefly share the strategies we used below.\n\n1. Drfold2 + Protenix (Public 0.589)\n\n▪ Seq. < 400\n\n      • 3 from drfold2\n      • 2 from Protenix fine-tuned RNA only + MSA (<800)\n\n▪ Seq. >= 400\n\n      • 2 from Protenix fine-tuned RNA only + MSA (<800)\n      • 2 from Protenix fine-tuned All + MSA (<800)\n      • 1 from Protenix baseline model\n\n2. Protenix + Boltz (Public 0.615)\n▪ 2 from Protenix fine-tuned RNA only + MSA (<800)\n▪ 1 from Protenix fine-tuned ALL + MSA (<800)\n▪ 1 from Protenix baseline\n▪ 1 from Boltz baseline\n\nBelow 800 because OOM on 96GB VRAM: GH200.\n\nRNA only: polymer composition RNA\nALL: v2 data + recent pdb release\nWe included Casp 16 data so that bumps up the public score.\nIf we don't use the casp 16 data, we get about 0.42.\n\nrMSA pipeline MSA data (used the codes hosts provided):\nhttps://drive.google.com/drive/folders/15bhXWsR6QuDQo6U4j8Ii8-OE4B4SmC7i?usp=drive_link\n\nTraining Data (for protenix finetuning) and labels:\nhttps://drive.google.com/drive/folders/1XKYzk2oCcHPt6DB_wLNYL-7w3s1R793n?usp=sharing\n\n\nCodes used from :\nProtenix inference\nhttps://www.kaggle.com/code/geraseva/protenix\n@geraseva \n\nProtenix finetuning\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495\n@lihaoweicvch \n\nBoltz inference\nhttps://www.kaggle.com/code/youhanlee/boltz-1-inference-submission\n@youhanlee",
    "3223918": "Very helpful, your solution is amazing! thx",
    "3225269": "Thank you for sharing, but could your fine-tuning strategy potentially lead to excessive leakage? It's possible that only **3 from DrFold2** and **1 from the ProteinX baseline model**  actually works, hard to say, live or die?",
    "3225312": "Well, we did try fine-tuning without the leakage data and it still improved.\n\nBaseline model with no fine-tuning had a public test score of 3.7,\nwhile fine-tuned model with v1 data + MSA had a public test score of 4.12.\nSo this suggests that fine-tuning with MSAs does work, so we used all the available data for training (more data the better..?)\n\nOne thing to notice was if we fine-tune just 1-2 epoch (1 or 2 full iterations of the training dataset) the performance dropped (possibility due to the optimizer as @shujun717 suggested). \nSo we had to train for more than 5 epochs for performance boost.\n\nMeanwhile we did work on 2 strategies but the drfold2 one didn't go through for submissions because we didn't select that one :(",
    "3226043": "Very helpful work"
  },
  "source": "meta"
}