{
  "id": 573495,
  "title": "ProteinX finetune result and code",
  "url": "/competitions/stanford-rna-3d-folding/discussion/573495",
  "author_name": "lhwcv",
  "post_date": "2025-04-16T00:34:32.665000",
  "votes": 32,
  "comment_count": 29,
  "views": 0,
  "content": "<p>code:   <a href=\"https://github.com/lhwcv/Protenix-RNA-Kaggle\" target=\"_blank\">https://github.com/lhwcv/Protenix-RNA-Kaggle</a></p>\n<h1>exp</h1>\n<p>overfit: casp16 (44 samples)</p>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>steps</th>\n<th>seq_len_crop</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>overfit vfold top1</td>\n<td>1000</td>\n<td>416</td>\n<td>0.454</td>\n</tr>\n<tr>\n<td>overfit vfold top5 dynamic match</td>\n<td>1000</td>\n<td>384 (416 OOM on 5090)</td>\n<td>0.467</td>\n</tr>\n</tbody>\n</table>\n<p>comp data: 799 samples (len&lt;=416) </p>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>steps</th>\n<th>seq_len</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>without msa</td>\n<td>4000</td>\n<td>416</td>\n<td>0.293</td>\n</tr>\n<tr>\n<td>with msa</td>\n<td>4000</td>\n<td>416</td>\n<td>xxx</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 3179945,
      "postDate": "2025-04-16T00:34:32.667Z",
      "content": "<p>code:   <a href=\"https://github.com/lhwcv/Protenix-RNA-Kaggle\" target=\"_blank\">https://github.com/lhwcv/Protenix-RNA-Kaggle</a></p>\n<h1>exp</h1>\n<p>overfit: casp16 (44 samples)</p>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>steps</th>\n<th>seq_len_crop</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>overfit vfold top1</td>\n<td>1000</td>\n<td>416</td>\n<td>0.454</td>\n</tr>\n<tr>\n<td>overfit vfold top5 dynamic match</td>\n<td>1000</td>\n<td>384 (416 OOM on 5090)</td>\n<td>0.467</td>\n</tr>\n</tbody>\n</table>\n<p>comp data: 799 samples (len&lt;=416) </p>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>steps</th>\n<th>seq_len</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>without msa</td>\n<td>4000</td>\n<td>416</td>\n<td>0.293</td>\n</tr>\n<tr>\n<td>with msa</td>\n<td>4000</td>\n<td>416</td>\n<td>xxx</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "code:   https://github.com/lhwcv/Protenix-RNA-Kaggle\n\n# exp\n\noverfit: casp16 (44 samples)\n\n| exp                              | steps | seq_len_crop          | LB    |\n|----------------------------------|-------|-----------------------|-------|\n| overfit vfold top1               | 1000  | 416                   | 0.454 |\n| overfit vfold top5 dynamic match | 1000  | 384 (416 OOM on 5090) | 0.467 |\n\ncomp data: 799 samples (len<=416) \n\n| exp         | steps | seq_len      | LB    |\n|-------------|-------|--------------|-------|\n| without msa | 4000  | 416          | 0.293 |\n| with msa    | 4000  | 416          | xxx   |\n",
      "votes": 32
    },
    {
      "id": 3181517,
      "postDate": "2025-04-18T01:58:08.183Z",
      "content": "<p>let me explain about the logic of our work:<br>\n(1) we want to investigate discriminative capability/capacity of proteinX transformer </p>\n<ul>\n<li>is it universial enough to fit anything?</li>\n<li>if its capacity is not large enough, then it will underfit if we have more data, or more MSA etc …</li>\n<li>our litmus test is to see if it can \"overfit\" some difficult data in shortest number of training steps</li>\n<li>so we overfit ProteinX on vfold (which we think is somehigh uncorrelated to af3)</li>\n</ul>\n<p>(2) is rMSA useful for proteinX?</p>\n<ul>\n<li>rMSA is useful for AF3, but is it useful for proteinX?</li>\n<li>we overfit proteinX with and without MSA as input.</li>\n<li>if MSA is useful, it should converge faster (and lower) in training, even if we have small data</li>\n</ul>\n<p>we are trying to estimate results without full training and also debug our code.<br>\nonce we think it is ok, we will proceed to large scale training/finetuning</p>\n<hr>\n<p>overfiting test code and pipline, etc. but overfitting don't test generalisation.<br>\ngeneralisation is tested by large scale training/finetuning</p>",
      "rawMarkdown": "let me explain about the logic of our work:\n(1) we want to investigate discriminative capability/capacity of proteinX transformer \n- is it universial enough to fit anything?\n- if its capacity is not large enough, then it will underfit if we have more data, or more MSA etc ...\n- our litmus test is to see if it can \"overfit\" some difficult data in shortest number of training steps\n- so we overfit ProteinX on vfold (which we think is somehigh uncorrelated to af3)\n\n(2) is rMSA useful for proteinX?\n- rMSA is useful for AF3, but is it useful for proteinX?\n- we overfit proteinX with and without MSA as input.\n- if MSA is useful, it should converge faster (and lower) in training, even if we have small data\n\nwe are trying to estimate results without full training and also debug our code.\nonce we think it is ok, we will proceed to large scale training/finetuning\n\n---\n\noverfiting test code and pipline, etc. but overfitting don't test generalisation.\ngeneralisation is tested by large scale training/finetuning",
      "votes": 3
    },
    {
      "id": 3207671,
      "postDate": "2025-05-23T05:42:41.723Z",
      "content": "<p>Very useful suggestions! </p>",
      "rawMarkdown": "Very useful suggestions! ",
      "votes": 1
    },
    {
      "id": 3179988,
      "postDate": "2025-04-16T02:01:14.313Z",
      "content": "<p>Similar to yours.</p>\n<p>Be careful of overfit, I once asked my senior labmate for some data that hadn't been submitted to the PDB to verify my approach, but in reality, my own finetuned model performed not well.  </p>\n<p>Different RNA structures may vary too much.</p>",
      "rawMarkdown": "Similar to yours.\n\nBe careful of overfit, I once asked my senior labmate for some data that hadn't been submitted to the PDB to verify my approach, but in reality, my own finetuned model performed not well.  \n\nDifferent RNA structures may vary too much.",
      "votes": 1,
      "replies": [
        {
          "id": 3180265,
          "postDate": "2025-04-16T11:08:54.020Z",
          "content": "<p>My own finetuned model also performed not well without MSA, how about your result with MSA?</p>",
          "rawMarkdown": "My own finetuned model also performed not well without MSA, how about your result with MSA?",
          "replies": [
            {
              "id": 3180324,
              "postDate": "2025-04-16T13:04:53.207Z",
              "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> </p>\n<p>I used the CASP16 data for finetuning, performed 5-fold validation with the kaggle's train dataset, and also incorporated my own data. The results are as follows: </p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>TMscore without MSA</th>\n<th></th>\n<th>TMscore with MSA</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Fold0</td>\n<td>0.301</td>\n<td>0.337</td>\n<td></td>\n</tr>\n<tr>\n<td>Fold1</td>\n<td>0.328</td>\n<td>0.352</td>\n<td></td>\n</tr>\n<tr>\n<td>Fold2</td>\n<td>0.296</td>\n<td>0.318</td>\n<td></td>\n</tr>\n<tr>\n<td>Fold3</td>\n<td>0.287</td>\n<td>0.305</td>\n<td></td>\n</tr>\n<tr>\n<td>Fold4</td>\n<td>0.347</td>\n<td>0.374</td>\n<td></td>\n</tr>\n<tr>\n<td>my private 14 sequences</td>\n<td>0.267</td>\n<td>0.317</td>\n<td></td>\n</tr>\n</tbody>\n</table>",
              "rawMarkdown": "@lihaoweicvch \n\nI used the CASP16 data for finetuning, performed 5-fold validation with the kaggle's train dataset, and also incorporated my own data. The results are as follows: \n\n|       | TMscore without MSA| |TMscore with MSA|\n| ----- | ------- |------- |\n| Fold0 | 0.301   | 0.337   |\n| Fold1 | 0.328   | 0.352   |\n| Fold2 | 0.296   | 0.318   |\n| Fold3 | 0.287   | 0.305   |\n| Fold4 | 0.347   | 0.374   |\n| my private 14 sequences | 0.267   | 0.317   |",
              "votes": 2
            },
            {
              "id": 3181859,
              "postDate": "2025-04-18T12:39:57.430Z",
              "content": "<p><a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a> how did you generate MSA?</p>",
              "rawMarkdown": "@sweetyheehee how did you generate MSA?"
            },
            {
              "id": 3181917,
              "postDate": "2025-04-18T14:23:02.420Z",
              "content": "<p><a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a> I tried some NN like RfamGen and modifying its backbone,  and I think it isn't very accurate, but at least it works.</p>",
              "rawMarkdown": "@arunodhayan I tried some NN like RfamGen and modifying its backbone,  and I think it isn't very accurate, but at least it works.",
              "votes": 2
            },
            {
              "id": 3181990,
              "postDate": "2025-04-18T16:06:12.447Z",
              "content": "<p><a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a>  thankyou for the info</p>",
              "rawMarkdown": "@sweetyheehee  thankyou for the info"
            }
          ]
        }
      ]
    },
    {
      "id": 3181730,
      "postDate": "2025-04-18T08:59:43.663Z",
      "content": "<p>Adding MSA and training on competition data LB score of 0.311</p>",
      "rawMarkdown": "Adding MSA and training on competition data LB score of 0.311",
      "votes": 2,
      "replies": [
        {
          "id": 3182509,
          "postDate": "2025-04-19T12:46:41.077Z",
          "content": "<p>are you using rMSA and based on ProteinX? wo MSA: 0.296 with MSA: 0.311</p>",
          "rawMarkdown": "are you using rMSA and based on ProteinX? wo MSA: 0.296 with MSA: 0.311",
          "votes": 1,
          "replies": [
            {
              "id": 3182573,
              "postDate": "2025-04-19T15:00:57.740Z",
              "content": "<p>The rMSA form kaggle prov provided by the organizers</p>",
              "rawMarkdown": "The rMSA form kaggle prov provided by the organizers"
            }
          ]
        }
      ]
    },
    {
      "id": 3181020,
      "postDate": "2025-04-17T11:23:57.550Z",
      "content": "<p>I trained Protinex with a sequence length of 800, achieved  LB score of 0.296  and on casp 15  local test  0.454</p>",
      "rawMarkdown": "I trained Protinex with a sequence length of 800, achieved  LB score of 0.296  and on casp 15  local test  0.454",
      "votes": 2,
      "replies": [
        {
          "id": 3197815,
          "postDate": "2025-05-08T16:20:09.147Z",
          "content": "<p>Why would it be like that?<br>\nIf I remember correctly, Protenix could achieve an LB score of over 0.3 even without fine-tuning.<br>\nBut in CASP15, fine-tuned models still showed better performance, so that gives me a bit of hope.</p>",
          "rawMarkdown": "Why would it be like that?\nIf I remember correctly, Protenix could achieve an LB score of over 0.3 even without fine-tuning.\nBut in CASP15, fine-tuned models still showed better performance, so that gives me a bit of hope."
        }
      ]
    },
    {
      "id": 3191629,
      "postDate": "2025-05-01T23:47:07.613Z",
      "content": "<p>Protenix is for proteins, which are nothing to do with this task. It will unlikely help too much for RNAs, AFAIK.</p>",
      "rawMarkdown": "Protenix is for proteins, which are nothing to do with this task. It will unlikely help too much for RNAs, AFAIK.",
      "votes": -3,
      "replies": [
        {
          "id": 3191992,
          "postDate": "2025-05-02T10:38:30.820Z",
          "content": "<p>From my perspective, this represents a transfer learning strategy that involves transferring knowledge from protein data to RNA data. However, despite fine-tuning on the RNA dataset, the model achieved relatively low scores. I believe this is due to overfitting, as predicted by Hengck23.</p>",
          "rawMarkdown": "From my perspective, this represents a transfer learning strategy that involves transferring knowledge from protein data to RNA data. However, despite fine-tuning on the RNA dataset, the model achieved relatively low scores. I believe this is due to overfitting, as predicted by Hengck23.",
          "replies": [
            {
              "id": 3194489,
              "postDate": "2025-05-06T00:30:31.987Z",
              "content": "<p>OK, good luck with it.</p>",
              "rawMarkdown": "OK, good luck with it.",
              "votes": -2
            }
          ]
        }
      ]
    },
    {
      "id": 3191061,
      "postDate": "2025-05-01T10:36:23.707Z",
      "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> Thank you for sharing the git repo and results. Can you please share the hardware configurations you guys used for fine tuning?</p>",
      "rawMarkdown": "@lihaoweicvch Thank you for sharing the git repo and results. Can you please share the hardware configurations you guys used for fine tuning?",
      "replies": [
        {
          "id": 3193962,
          "postDate": "2025-05-05T09:04:48.790Z",
          "content": "<p>RTX5090、RTX6000、 A100、H100 for different team members</p>",
          "rawMarkdown": "RTX5090、RTX6000、 A100、H100 for different team members"
        }
      ]
    },
    {
      "id": 3180881,
      "postDate": "2025-04-17T08:32:42.570Z",
      "content": "<p>Thank you so much for sharing the code and results!</p>\n<p>For the experiments above, could you clarify the following questions:</p>\n<ol>\n<li>For the overfit CASP16 targets experiment,  do you use CASP16 samples as the training set to finetune protenix, using the vfold predictions as labels, or use it as the validation set?</li>\n<li>What is the performance when using the Kaggle data to finetune protenix with MSA? </li>\n</ol>",
      "rawMarkdown": "Thank you so much for sharing the code and results!\n\nFor the experiments above, could you clarify the following questions:\n1. For the overfit CASP16 targets experiment,  do you use CASP16 samples as the training set to finetune protenix, using the vfold predictions as labels, or use it as the validation set?\n2. What is the performance when using the Kaggle data to finetune protenix with MSA? ",
      "replies": [
        {
          "id": 3182510,
          "postDate": "2025-04-19T12:49:05.310Z",
          "content": "<ol>\n<li>CASP16 samples as the training set to finetune protenix</li>\n<li>I'm also referring other's result currently 😆</li>\n</ol>",
          "rawMarkdown": "1. CASP16 samples as the training set to finetune protenix\n2. I'm also referring other's result currently 😆"
        }
      ]
    },
    {
      "id": 3180297,
      "postDate": "2025-04-16T11:58:13.860Z",
      "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a>  casp16 (44 samples)?  Did you download via the CASP16 website ? </p>",
      "rawMarkdown": "@lihaoweicvch  casp16 (44 samples)?  Did you download via the CASP16 website ? \n",
      "replies": [
        {
          "id": 3182506,
          "postDate": "2025-04-19T12:41:52.807Z",
          "content": "<p>yes  I downloaded via the CASP16 website</p>",
          "rawMarkdown": "yes  I downloaded via the CASP16 website",
          "replies": [
            {
              "id": 3182648,
              "postDate": "2025-04-19T17:03:22.803Z",
              "content": "<p>May I ask how can you get the label for casp16 samples? From vfold prediction?</p>",
              "rawMarkdown": "May I ask how can you get the label for casp16 samples? From vfold prediction?"
            },
            {
              "id": 3204408,
              "postDate": "2025-05-18T08:53:53.233Z",
              "content": "<p>Hi, thank you for sharing! If possible, could you share the CASP16 dataset HP?</p>",
              "rawMarkdown": "Hi, thank you for sharing! If possible, could you share the CASP16 dataset HP?"
            }
          ]
        }
      ]
    },
    {
      "id": 3180257,
      "postDate": "2025-04-16T11:00:10.053Z",
      "content": "<p>It seems that just inferencing with the pretrained Protenix model result(LB 0.344) is better than fine-tuning Protenix model result(as in your case 0.293). Is this what you pretend to tell us?</p>",
      "rawMarkdown": "It seems that just inferencing with the pretrained Protenix model result(LB 0.344) is better than fine-tuning Protenix model result(as in your case 0.293). Is this what you pretend to tell us?",
      "replies": [
        {
          "id": 3180270,
          "postDate": "2025-04-16T11:11:16.617Z",
          "content": "<p>Based on the results, it seems to be like that—using only the competition data without MSA to finetune, performs worse than the original version. I'm not sure if it's the same on everyone's end.</p>",
          "rawMarkdown": "Based on the results, it seems to be like that—using only the competition data without MSA to finetune, performs worse than the original version. I'm not sure if it's the same on everyone's end.",
          "votes": 3,
          "replies": [
            {
              "id": 3183122,
              "postDate": "2025-04-20T11:57:11.430Z",
              "content": "<p>The top2 player stated he used the sample from casp16 in other discussion and got his result.</p>",
              "rawMarkdown": "The top2 player stated he used the sample from casp16 in other discussion and got his result.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3180085,
      "postDate": "2025-04-16T05:59:11.503Z",
      "content": "<p>Hi, did you only use the cas16 to fine tune the proteinx model and get 0.454 lb and get 0.293 by using the training samples provided in the game. What is the result if you use both?</p>",
      "rawMarkdown": "Hi, did you only use the cas16 to fine tune the proteinx model and get 0.454 lb and get 0.293 by using the training samples provided in the game. What is the result if you use both?"
    },
    {
      "id": 3182615,
      "postDate": "2025-04-19T16:14:06.333Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3181517,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-18T01:58:08.183000",
      "content": "<p>let me explain about the logic of our work:<br>\n(1) we want to investigate discriminative capability/capacity of proteinX transformer </p>\n<ul>\n<li>is it universial enough to fit anything?</li>\n<li>if its capacity is not large enough, then it will underfit if we have more data, or more MSA etc …</li>\n<li>our litmus test is to see if it can \"overfit\" some difficult data in shortest number of training steps</li>\n<li>so we overfit ProteinX on vfold (which we think is somehigh uncorrelated to af3)</li>\n</ul>\n<p>(2) is rMSA useful for proteinX?</p>\n<ul>\n<li>rMSA is useful for AF3, but is it useful for proteinX?</li>\n<li>we overfit proteinX with and without MSA as input.</li>\n<li>if MSA is useful, it should converge faster (and lower) in training, even if we have small data</li>\n</ul>\n<p>we are trying to estimate results without full training and also debug our code.<br>\nonce we think it is ok, we will proceed to large scale training/finetuning</p>\n<hr>\n<p>overfiting test code and pipline, etc. but overfitting don't test generalisation.<br>\ngeneralisation is tested by large scale training/finetuning</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3207671,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-23T05:42:41.723000",
      "content": "<p>Very useful suggestions! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3179988,
      "author_name": "Timmy Juicehouse",
      "author_url": "",
      "post_date": "2025-04-16T02:01:14.313000",
      "content": "<p>Similar to yours.</p>\n<p>Be careful of overfit, I once asked my senior labmate for some data that hadn't been submitted to the PDB to verify my approach, but in reality, my own finetuned model performed not well.  </p>\n<p>Different RNA structures may vary too much.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3180265,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2025-04-16T11:08:54.020000",
          "content": "<p>My own finetuned model also performed not well without MSA, how about your result with MSA?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3180324,
              "author_name": "Timmy Juicehouse",
              "author_url": "",
              "post_date": "2025-04-16T13:04:53.207000",
              "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> </p>\n<p>I used the CASP16 data for finetuning, performed 5-fold validation with the kaggle's train dataset, and also incorporated my own data. The results are as follows: </p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>TMscore without MSA</th>\n<th></th>\n<th>TMscore with MSA</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Fold0</td>\n<td>0.301</td>\n<td>0.337</td>\n<td></td>\n</tr>\n<tr>\n<td>Fold1</td>\n<td>0.328</td>\n<td>0.352</td>\n<td></td>\n</tr>\n<tr>\n<td>Fold2</td>\n<td>0.296</td>\n<td>0.318</td>\n<td></td>\n</tr>\n<tr>\n<td>Fold3</td>\n<td>0.287</td>\n<td>0.305</td>\n<td></td>\n</tr>\n<tr>\n<td>Fold4</td>\n<td>0.347</td>\n<td>0.374</td>\n<td></td>\n</tr>\n<tr>\n<td>my private 14 sequences</td>\n<td>0.267</td>\n<td>0.317</td>\n<td></td>\n</tr>\n</tbody>\n</table>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3181859,
              "author_name": "Arunodhayan",
              "author_url": "",
              "post_date": "2025-04-18T12:39:57.430000",
              "content": "<p><a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a> how did you generate MSA?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3181917,
              "author_name": "Timmy Juicehouse",
              "author_url": "",
              "post_date": "2025-04-18T14:23:02.420000",
              "content": "<p><a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a> I tried some NN like RfamGen and modifying its backbone,  and I think it isn't very accurate, but at least it works.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3181990,
              "author_name": "Arunodhayan",
              "author_url": "",
              "post_date": "2025-04-18T16:06:12.447000",
              "content": "<p><a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a>  thankyou for the info</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3181730,
      "author_name": "Arunodhayan",
      "author_url": "",
      "post_date": "2025-04-18T08:59:43.663000",
      "content": "<p>Adding MSA and training on competition data LB score of 0.311</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3182509,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2025-04-19T12:46:41.077000",
          "content": "<p>are you using rMSA and based on ProteinX? wo MSA: 0.296 with MSA: 0.311</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3182573,
              "author_name": "Arunodhayan",
              "author_url": "",
              "post_date": "2025-04-19T15:00:57.740000",
              "content": "<p>The rMSA form kaggle prov provided by the organizers</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3181020,
      "author_name": "Arunodhayan",
      "author_url": "",
      "post_date": "2025-04-17T11:23:57.550000",
      "content": "<p>I trained Protinex with a sequence length of 800, achieved  LB score of 0.296  and on casp 15  local test  0.454</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3197815,
          "author_name": "doheon114",
          "author_url": "",
          "post_date": "2025-05-08T16:20:09.147000",
          "content": "<p>Why would it be like that?<br>\nIf I remember correctly, Protenix could achieve an LB score of over 0.3 even without fine-tuning.<br>\nBut in CASP15, fine-tuned models still showed better performance, so that gives me a bit of hope.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3191629,
      "author_name": "Ilya Kupchenko",
      "author_url": "",
      "post_date": "2025-05-01T23:47:07.613000",
      "content": "<p>Protenix is for proteins, which are nothing to do with this task. It will unlikely help too much for RNAs, AFAIK.</p>",
      "votes": -3,
      "replies": [
        {
          "id": 3191992,
          "author_name": "Nigmat Rahim",
          "author_url": "",
          "post_date": "2025-05-02T10:38:30.820000",
          "content": "<p>From my perspective, this represents a transfer learning strategy that involves transferring knowledge from protein data to RNA data. However, despite fine-tuning on the RNA dataset, the model achieved relatively low scores. I believe this is due to overfitting, as predicted by Hengck23.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3194489,
              "author_name": "Ilya Kupchenko",
              "author_url": "",
              "post_date": "2025-05-06T00:30:31.987000",
              "content": "<p>OK, good luck with it.</p>",
              "votes": -2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3191061,
      "author_name": "Jaydev",
      "author_url": "",
      "post_date": "2025-05-01T10:36:23.707000",
      "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> Thank you for sharing the git repo and results. Can you please share the hardware configurations you guys used for fine tuning?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3193962,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2025-05-05T09:04:48.790000",
          "content": "<p>RTX5090、RTX6000、 A100、H100 for different team members</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3180881,
      "author_name": "Shuxian Zou",
      "author_url": "",
      "post_date": "2025-04-17T08:32:42.570000",
      "content": "<p>Thank you so much for sharing the code and results!</p>\n<p>For the experiments above, could you clarify the following questions:</p>\n<ol>\n<li>For the overfit CASP16 targets experiment,  do you use CASP16 samples as the training set to finetune protenix, using the vfold predictions as labels, or use it as the validation set?</li>\n<li>What is the performance when using the Kaggle data to finetune protenix with MSA? </li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 3182510,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2025-04-19T12:49:05.310000",
          "content": "<ol>\n<li>CASP16 samples as the training set to finetune protenix</li>\n<li>I'm also referring other's result currently 😆</li>\n</ol>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3180297,
      "author_name": "Arunodhayan",
      "author_url": "",
      "post_date": "2025-04-16T11:58:13.860000",
      "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a>  casp16 (44 samples)?  Did you download via the CASP16 website ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3182506,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2025-04-19T12:41:52.807000",
          "content": "<p>yes  I downloaded via the CASP16 website</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3182648,
              "author_name": "minhtu.mt.mt",
              "author_url": "",
              "post_date": "2025-04-19T17:03:22.803000",
              "content": "<p>May I ask how can you get the label for casp16 samples? From vfold prediction?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3204408,
              "author_name": "T41K1",
              "author_url": "",
              "post_date": "2025-05-18T08:53:53.233000",
              "content": "<p>Hi, thank you for sharing! If possible, could you share the CASP16 dataset HP?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3180257,
      "author_name": "doheon114",
      "author_url": "",
      "post_date": "2025-04-16T11:00:10.053000",
      "content": "<p>It seems that just inferencing with the pretrained Protenix model result(LB 0.344) is better than fine-tuning Protenix model result(as in your case 0.293). Is this what you pretend to tell us?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3180270,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2025-04-16T11:11:16.617000",
          "content": "<p>Based on the results, it seems to be like that—using only the competition data without MSA to finetune, performs worse than the original version. I'm not sure if it's the same on everyone's end.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3183122,
              "author_name": "Wenxuan Ye",
              "author_url": "",
              "post_date": "2025-04-20T11:57:11.430000",
              "content": "<p>The top2 player stated he used the sample from casp16 in other discussion and got his result.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3180085,
      "author_name": "Wenxuan Ye",
      "author_url": "",
      "post_date": "2025-04-16T05:59:11.503000",
      "content": "<p>Hi, did you only use the cas16 to fine tune the proteinx model and get 0.454 lb and get 0.293 by using the training samples provided in the game. What is the result if you use both?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3182615,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-19T16:14:06.333000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3179945": "code:   https://github.com/lhwcv/Protenix-RNA-Kaggle\n\n# exp\n\noverfit: casp16 (44 samples)\n\n| exp                              | steps | seq_len_crop          | LB    |\n|----------------------------------|-------|-----------------------|-------|\n| overfit vfold top1               | 1000  | 416                   | 0.454 |\n| overfit vfold top5 dynamic match | 1000  | 384 (416 OOM on 5090) | 0.467 |\n\ncomp data: 799 samples (len<=416) \n\n| exp         | steps | seq_len      | LB    |\n|-------------|-------|--------------|-------|\n| without msa | 4000  | 416          | 0.293 |\n| with msa    | 4000  | 416          | xxx   |\n",
    "3181517": "let me explain about the logic of our work:\n(1) we want to investigate discriminative capability/capacity of proteinX transformer \n- is it universial enough to fit anything?\n- if its capacity is not large enough, then it will underfit if we have more data, or more MSA etc ...\n- our litmus test is to see if it can \"overfit\" some difficult data in shortest number of training steps\n- so we overfit ProteinX on vfold (which we think is somehigh uncorrelated to af3)\n\n(2) is rMSA useful for proteinX?\n- rMSA is useful for AF3, but is it useful for proteinX?\n- we overfit proteinX with and without MSA as input.\n- if MSA is useful, it should converge faster (and lower) in training, even if we have small data\n\nwe are trying to estimate results without full training and also debug our code.\nonce we think it is ok, we will proceed to large scale training/finetuning\n\n---\n\noverfiting test code and pipline, etc. but overfitting don't test generalisation.\ngeneralisation is tested by large scale training/finetuning",
    "3207671": "Very useful suggestions! ",
    "3179988": "Similar to yours.\n\nBe careful of overfit, I once asked my senior labmate for some data that hadn't been submitted to the PDB to verify my approach, but in reality, my own finetuned model performed not well.  \n\nDifferent RNA structures may vary too much.",
    "3181730": "Adding MSA and training on competition data LB score of 0.311",
    "3181020": "I trained Protinex with a sequence length of 800, achieved  LB score of 0.296  and on casp 15  local test  0.454",
    "3191629": "Protenix is for proteins, which are nothing to do with this task. It will unlikely help too much for RNAs, AFAIK.",
    "3191061": "@lihaoweicvch Thank you for sharing the git repo and results. Can you please share the hardware configurations you guys used for fine tuning?",
    "3180881": "Thank you so much for sharing the code and results!\n\nFor the experiments above, could you clarify the following questions:\n1. For the overfit CASP16 targets experiment,  do you use CASP16 samples as the training set to finetune protenix, using the vfold predictions as labels, or use it as the validation set?\n2. What is the performance when using the Kaggle data to finetune protenix with MSA? ",
    "3180297": "@lihaoweicvch  casp16 (44 samples)?  Did you download via the CASP16 website ? \n",
    "3180257": "It seems that just inferencing with the pretrained Protenix model result(LB 0.344) is better than fine-tuning Protenix model result(as in your case 0.293). Is this what you pretend to tell us?",
    "3180085": "Hi, did you only use the cas16 to fine tune the proteinx model and get 0.454 lb and get 0.293 by using the training samples provided in the game. What is the result if you use both?",
    "3182615": ""
  }
}