{
  "id": 566875,
  "title": "I am little confused with 5 predictions",
  "url": "/competitions/stanford-rna-3d-folding/discussion/566875",
  "author_name": "",
  "post_date": "2025-03-07T07:55:11.024001900Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I am little confused with making 5 predictions for each residue.</p>\n<p>I understand that in real world, experiment results show different 3D structure for the same RNA sequence.</p>\n<p>But in the provided dataset, we can find that there are only one coordinates(x_1, y_1, z_1) for each residue.</p>\n<p>It is stated in the PDB database, there are different structure data for the same sequence in the train dataset. This will help the model to predict different structure for the test dataset. </p>\n<p>Then, before applying PDF database datas to training, how are other kagglers making 5 prediction for each residues? </p>\n<p>And, How is TM-score applied for the 5 different predictions for one RNA? Is the TM-score calculated for each prediction and the maximum TM-score chosen between the predictions for the final TM-score of one RNA? </p>\n<p>Thank you</p>",
  "messages": [
    {
      "id": "3143403",
      "postDate": "03/07/2025 07:55:11",
      "content": "<p>Hello,</p>\n<p>I am little confused with making 5 predictions for each residue.</p>\n<p>I understand that in real world, experiment results show different 3D structure for the same RNA sequence.</p>\n<p>But in the provided dataset, we can find that there are only one coordinates(x_1, y_1, z_1) for each residue.</p>\n<p>It is stated in the PDB database, there are different structure data for the same sequence in the train dataset. This will help the model to predict different structure for the test dataset. </p>\n<p>Then, before applying PDF database datas to training, how are other kagglers making 5 prediction for each residues? </p>\n<p>And, How is TM-score applied for the 5 different predictions for one RNA? Is the TM-score calculated for each prediction and the maximum TM-score chosen between the predictions for the final TM-score of one RNA? </p>\n<p>Thank you</p>",
      "rawMarkdown": "Hello,\n\nI am little confused with making 5 predictions for each residue.\n\nI understand that in real world, experiment results show different 3D structure for the same RNA sequence.\n \nBut in the provided dataset, we can find that there are only one coordinates(x_1, y_1, z_1) for each residue.\n\nIt is stated in the PDB database, there are different structure data for the same sequence in the train dataset. This will help the model to predict different structure for the test dataset. \n\nThen, before applying PDF database datas to training, how are other kagglers making 5 prediction for each residues? \n\nAnd, How is TM-score applied for the 5 different predictions for one RNA? Is the TM-score calculated for each prediction and the maximum TM-score chosen between the predictions for the final TM-score of one RNA? \n\nThank you",
      "votes": null
    },
    {
      "id": "3143409",
      "postDate": "03/07/2025 08:01:06",
      "content": "<blockquote>\n  <p>Then, before applying PDF database datas to training, how are other kagglers making 5 prediction for each residues?</p>\n</blockquote>\n<p>You can use a generative model (VAE/GAN/Diffusion) to generate 5 structures from a single sequence but with changing seeds (as you would for image generation)</p>\n<blockquote>\n  <p>And, How is TM-score applied for the 5 different predictions for one RNA? Is the TM-score calculated for each prediction and the maximum TM-score chosen between the predictions for the final TM-score of one RNA?</p>\n</blockquote>\n<p>Yes. For each sequence, take the best-of-5 prediction (w.r.t to one of the RNA &lt;40 conformations for ex. in the provided test set). Then average over the sequences</p>\n<p>\"For each target RNA sequence, you will submit 5 predictions and your final score will be the average of best-of-5 TM-scores of all targets. For a few targets, multiple slightly different structures have been captured experimentally; your predictions' scores will be based on the best TM-score compared to each of these reference structures.\"</p>",
      "rawMarkdown": ">Then, before applying PDF database datas to training, how are other kagglers making 5 prediction for each residues?\n\nYou can use a generative model (VAE/GAN/Diffusion) to generate 5 structures from a single sequence but with changing seeds (as you would for image generation)\n\n> And, How is TM-score applied for the 5 different predictions for one RNA? Is the TM-score calculated for each prediction and the maximum TM-score chosen between the predictions for the final TM-score of one RNA?\n\nYes. For each sequence, take the best-of-5 prediction (w.r.t to one of the RNA <40 conformations for ex. in the provided test set). Then average over the sequences\n\n\"For each target RNA sequence, you will submit 5 predictions and your final score will be the average of best-of-5 TM-scores of all targets. For a few targets, multiple slightly different structures have been captured experimentally; your predictions' scores will be based on the best TM-score compared to each of these reference structures.\"",
      "votes": null
    },
    {
      "id": "3143432",
      "postDate": "03/07/2025 08:40:48",
      "content": "<p>Thank you very much. Never done Generative models before. I will definitely try changing seeds. </p>",
      "rawMarkdown": "Thank you very much. Never done Generative models before. I will definitely try changing seeds.",
      "votes": null
    },
    {
      "id": "3144071",
      "postDate": "03/07/2025 22:40:20",
      "content": "<p>This has confused me also, the comments here have helped - so from my understanding we are using the same model but changing seeds and then regenerating the predictions. Then the data will take the average value across all five different models when evaluating the performance if I am correct? I still  don't really understand why.</p>",
      "rawMarkdown": "This has confused me also, the comments here have helped - so from my understanding we are using the same model but changing seeds and then regenerating the predictions. Then the data will take the average value across all five different models when evaluating the performance if I am correct? I still  don't really understand why.",
      "votes": null
    },
    {
      "id": "3144278",
      "postDate": "03/08/2025 06:21:07",
      "content": "<p>To be clear, I think the data which has the best(max) TM-score will be chosen for each RNA</p>",
      "rawMarkdown": "To be clear, I think the data which has the best(max) TM-score will be chosen for each RNA",
      "votes": null
    },
    {
      "id": "3144595",
      "postDate": "03/08/2025 15:16:45",
      "content": "<p>Oh ok, that helps thanks.</p>",
      "rawMarkdown": "Oh ok, that helps thanks.",
      "votes": null
    },
    {
      "id": "3145029",
      "postDate": "03/09/2025 08:35:52",
      "content": "<p><a href=\"https://www.kaggle.com/peterhopkinson\" target=\"_blank\">@peterhopkinson</a> it can also be 5 predictions from 5 models. Mono model + changing seeds is just one option. I guess that ensemble would lead to better perf (as often)</p>",
      "rawMarkdown": "peterhopkinson it can also be 5 predictions from 5 models. Mono model + changing seeds is just one option. I guess that ensemble would lead to better perf (as often)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3143409,
      "author_name": "louisstefanuto",
      "author_url": "",
      "post_date": "03/07/2025 08:01:06",
      "content": "<blockquote>\n  <p>Then, before applying PDF database datas to training, how are other kagglers making 5 prediction for each residues?</p>\n</blockquote>\n<p>You can use a generative model (VAE/GAN/Diffusion) to generate 5 structures from a single sequence but with changing seeds (as you would for image generation)</p>\n<blockquote>\n  <p>And, How is TM-score applied for the 5 different predictions for one RNA? Is the TM-score calculated for each prediction and the maximum TM-score chosen between the predictions for the final TM-score of one RNA?</p>\n</blockquote>\n<p>Yes. For each sequence, take the best-of-5 prediction (w.r.t to one of the RNA &lt;40 conformations for ex. in the provided test set). Then average over the sequences</p>\n<p>\"For each target RNA sequence, you will submit 5 predictions and your final score will be the average of best-of-5 TM-scores of all targets. For a few targets, multiple slightly different structures have been captured experimentally; your predictions' scores will be based on the best TM-score compared to each of these reference structures.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 3143432,
          "author_name": "kendontcare11",
          "author_url": "",
          "post_date": "03/07/2025 08:40:48",
          "content": "<p>Thank you very much. Never done Generative models before. I will definitely try changing seeds. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3144071,
      "author_name": "peterhopkinson",
      "author_url": "",
      "post_date": "03/07/2025 22:40:20",
      "content": "<p>This has confused me also, the comments here have helped - so from my understanding we are using the same model but changing seeds and then regenerating the predictions. Then the data will take the average value across all five different models when evaluating the performance if I am correct? I still  don't really understand why.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3144278,
          "author_name": "kendontcare11",
          "author_url": "",
          "post_date": "03/08/2025 06:21:07",
          "content": "<p>To be clear, I think the data which has the best(max) TM-score will be chosen for each RNA</p>",
          "votes": null,
          "replies": [
            {
              "id": 3144595,
              "author_name": "peterhopkinson",
              "author_url": "",
              "post_date": "03/08/2025 15:16:45",
              "content": "<p>Oh ok, that helps thanks.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3145029,
          "author_name": "louisstefanuto",
          "author_url": "",
          "post_date": "03/09/2025 08:35:52",
          "content": "<p><a href=\"https://www.kaggle.com/peterhopkinson\" target=\"_blank\">@peterhopkinson</a> it can also be 5 predictions from 5 models. Mono model + changing seeds is just one option. I guess that ensemble would lead to better perf (as often)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3143403": "Hello,\n\nI am little confused with making 5 predictions for each residue.\n\nI understand that in real world, experiment results show different 3D structure for the same RNA sequence.\n \nBut in the provided dataset, we can find that there are only one coordinates(x_1, y_1, z_1) for each residue.\n\nIt is stated in the PDB database, there are different structure data for the same sequence in the train dataset. This will help the model to predict different structure for the test dataset. \n\nThen, before applying PDF database datas to training, how are other kagglers making 5 prediction for each residues? \n\nAnd, How is TM-score applied for the 5 different predictions for one RNA? Is the TM-score calculated for each prediction and the maximum TM-score chosen between the predictions for the final TM-score of one RNA? \n\nThank you",
    "3143409": ">Then, before applying PDF database datas to training, how are other kagglers making 5 prediction for each residues?\n\nYou can use a generative model (VAE/GAN/Diffusion) to generate 5 structures from a single sequence but with changing seeds (as you would for image generation)\n\n> And, How is TM-score applied for the 5 different predictions for one RNA? Is the TM-score calculated for each prediction and the maximum TM-score chosen between the predictions for the final TM-score of one RNA?\n\nYes. For each sequence, take the best-of-5 prediction (w.r.t to one of the RNA <40 conformations for ex. in the provided test set). Then average over the sequences\n\n\"For each target RNA sequence, you will submit 5 predictions and your final score will be the average of best-of-5 TM-scores of all targets. For a few targets, multiple slightly different structures have been captured experimentally; your predictions' scores will be based on the best TM-score compared to each of these reference structures.\"",
    "3143432": "Thank you very much. Never done Generative models before. I will definitely try changing seeds.",
    "3144071": "This has confused me also, the comments here have helped - so from my understanding we are using the same model but changing seeds and then regenerating the predictions. Then the data will take the average value across all five different models when evaluating the performance if I am correct? I still  don't really understand why.",
    "3144278": "To be clear, I think the data which has the best(max) TM-score will be chosen for each RNA",
    "3144595": "Oh ok, that helps thanks.",
    "3145029": "peterhopkinson it can also be 5 predictions from 5 models. Mono model + changing seeds is just one option. I guess that ensemble would lead to better perf (as often)"
  },
  "source": "meta"
}