{
  "id": 569085,
  "title": "Competition data processing pipeline",
  "url": "/competitions/stanford-rna-3d-folding/discussion/569085",
  "author_name": "",
  "post_date": "2025-03-19T22:05:04.611767700Z",
  "votes": 38,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I have uploaded the competition data processing pipeline that downloads pdbs from the Protein Data Bank and processes them into competition format training data: <a href=\"https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing\" target=\"_blank\">https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing</a>, so you may try to experiment with different data processing procedures</p>\n<p>let me know if you have questions!</p>",
  "messages": [
    {
      "id": "3154391",
      "postDate": "03/19/2025 22:05:04",
      "content": "<p>I have uploaded the competition data processing pipeline that downloads pdbs from the Protein Data Bank and processes them into competition format training data: <a href=\"https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing\" target=\"_blank\">https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing</a>, so you may try to experiment with different data processing procedures</p>\n<p>let me know if you have questions!</p>",
      "rawMarkdown": "I have uploaded the competition data processing pipeline that downloads pdbs from the Protein Data Bank and processes them into competition format training data: https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing, so you may try to experiment with different data processing procedures\n\nlet me know if you have questions!",
      "votes": null
    },
    {
      "id": "3154475",
      "postDate": "03/20/2025 01:49:40",
      "content": "<p>thansk a lot! that is very helpful</p>",
      "rawMarkdown": "thansk a lot! that is very helpful",
      "votes": null
    },
    {
      "id": "3154650",
      "postDate": "03/20/2025 07:49:13",
      "content": "<p>It looks like the code in your <code>get_xyz_data.py</code> is a bit messy with the multiprocessing :D <br>\nAlso, the <code>get_pdb_publication_date</code> method is not found in the repo. I think we need the <code>utils</code> module (the <code>get_xyz_data.py</code> import utils module)</p>",
      "rawMarkdown": "It looks like the code in your `get_xyz_data.py` is a bit messy with the multiprocessing :D \nAlso, the `get_pdb_publication_date` method is not found in the repo. I think we need the `utils` module (the `get_xyz_data.py` import utils module)",
      "votes": null
    },
    {
      "id": "3154657",
      "postDate": "03/20/2025 08:12:34",
      "content": "<p>Here is the way to workaround:</p>\n<pre><code> ():\n    pdb_info = MMCIF2Dict()\n     pdb_info[][]\n</code></pre>",
      "rawMarkdown": "Here is the way to workaround:\n```\ndef get_pdb_publication_date(pdb_id):\n    pdb_info = MMCIF2Dict(f\"./rna_structures/{pdb_id}.cif\")\n    return pdb_info['_pdbx_audit_revision_history.revision_date'][0]\n```",
      "votes": null
    },
    {
      "id": "3156529",
      "postDate": "03/22/2025 08:41:35",
      "content": "<p><code>run_cdhit_clustering</code> is also unknown without a published <code>utils</code> file</p>",
      "rawMarkdown": "`run_cdhit_clustering` is also unknown without a published `utils` file",
      "votes": null
    },
    {
      "id": "3158772",
      "postDate": "03/24/2025 21:32:52",
      "content": "<p>Just added utils.py! Sorry somehow it was not committed earlier</p>",
      "rawMarkdown": "Just added utils.py! Sorry somehow it was not committed earlier",
      "votes": null
    },
    {
      "id": "3158935",
      "postDate": "03/25/2025 04:34:25",
      "content": "<p>Thanks a lot .</p>",
      "rawMarkdown": "Thanks a lot .",
      "votes": null
    },
    {
      "id": "3167533",
      "postDate": "04/01/2025 16:07:00",
      "content": "<p>Thank you so much for sharing this pipeline. It helps a lot!<br>\nI have a question about the filtering pipeline. In step 3 to get_xyz_data, you filtered the dataset by \"structuredness\".  What does \"structuredness\" mean? And how do you get the \"extracted_structures.csv\" file? </p>\n<pre><code>stats=pd.read_csv()\nstats=stats.set_index()\nfiltered_cif_files=[f  f  cif_files  f  stats.index  stats.loc[f,]&gt;]\n</code></pre>",
      "rawMarkdown": "Thank you so much for sharing this pipeline. It helps a lot!\nI have a question about the filtering pipeline. In step 3 to get_xyz_data, you filtered the dataset by \"structuredness\".  What does \"structuredness\" mean? And how do you get the \"extracted_structures.csv\" file? \n\n```python\nstats=pd.read_csv(\"extracted_structures.csv\")\nstats=stats.set_index('cif')\nfiltered_cif_files=[f for f in cif_files if f in stats.index and stats.loc[f,'structuredness']>0.2]\n```",
      "votes": null
    },
    {
      "id": "3170760",
      "postDate": "04/05/2025 01:24:45",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> can you reply this question? I think we need this information to generate more training data</p>",
      "rawMarkdown": "shujun717 can you reply this question? I think we need this information to generate more training data",
      "votes": null
    },
    {
      "id": "3170770",
      "postDate": "04/05/2025 01:49:54",
      "content": "<p>It's the fraction of paired residues in an RNA structure</p>",
      "rawMarkdown": "It's the fraction of paired residues in an RNA structure",
      "votes": null
    },
    {
      "id": "3177372",
      "postDate": "04/12/2025 16:27:39",
      "content": "<p>Super helpful!</p>",
      "rawMarkdown": "Super helpful!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3154475,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/20/2025 01:49:40",
      "content": "<p>thansk a lot! that is very helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3154650,
      "author_name": "minhtu123",
      "author_url": "",
      "post_date": "03/20/2025 07:49:13",
      "content": "<p>It looks like the code in your <code>get_xyz_data.py</code> is a bit messy with the multiprocessing :D <br>\nAlso, the <code>get_pdb_publication_date</code> method is not found in the repo. I think we need the <code>utils</code> module (the <code>get_xyz_data.py</code> import utils module)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3154657,
          "author_name": "minhtu123",
          "author_url": "",
          "post_date": "03/20/2025 08:12:34",
          "content": "<p>Here is the way to workaround:</p>\n<pre><code> ():\n    pdb_info = MMCIF2Dict()\n     pdb_info[][]\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3156529,
          "author_name": "zacchaeus",
          "author_url": "",
          "post_date": "03/22/2025 08:41:35",
          "content": "<p><code>run_cdhit_clustering</code> is also unknown without a published <code>utils</code> file</p>",
          "votes": null,
          "replies": [
            {
              "id": 3158772,
              "author_name": "shujun717",
              "author_url": "",
              "post_date": "03/24/2025 21:32:52",
              "content": "<p>Just added utils.py! Sorry somehow it was not committed earlier</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3158935,
      "author_name": "",
      "author_url": "",
      "post_date": "03/25/2025 04:34:25",
      "content": "<p>Thanks a lot .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3167533,
      "author_name": "zoushuxian",
      "author_url": "",
      "post_date": "04/01/2025 16:07:00",
      "content": "<p>Thank you so much for sharing this pipeline. It helps a lot!<br>\nI have a question about the filtering pipeline. In step 3 to get_xyz_data, you filtered the dataset by \"structuredness\".  What does \"structuredness\" mean? And how do you get the \"extracted_structures.csv\" file? </p>\n<pre><code>stats=pd.read_csv()\nstats=stats.set_index()\nfiltered_cif_files=[f  f  cif_files  f  stats.index  stats.loc[f,]&gt;]\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3170760,
          "author_name": "minhtu123",
          "author_url": "",
          "post_date": "04/05/2025 01:24:45",
          "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> can you reply this question? I think we need this information to generate more training data</p>",
          "votes": null,
          "replies": [
            {
              "id": 3170770,
              "author_name": "shujun717",
              "author_url": "",
              "post_date": "04/05/2025 01:49:54",
              "content": "<p>It's the fraction of paired residues in an RNA structure</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3177372,
      "author_name": "jagofc",
      "author_url": "",
      "post_date": "04/12/2025 16:27:39",
      "content": "<p>Super helpful!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3154391": "I have uploaded the competition data processing pipeline that downloads pdbs from the Protein Data Bank and processes them into competition format training data: https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing, so you may try to experiment with different data processing procedures\n\nlet me know if you have questions!",
    "3154475": "thansk a lot! that is very helpful",
    "3154650": "It looks like the code in your `get_xyz_data.py` is a bit messy with the multiprocessing :D \nAlso, the `get_pdb_publication_date` method is not found in the repo. I think we need the `utils` module (the `get_xyz_data.py` import utils module)",
    "3154657": "Here is the way to workaround:\n```\ndef get_pdb_publication_date(pdb_id):\n    pdb_info = MMCIF2Dict(f\"./rna_structures/{pdb_id}.cif\")\n    return pdb_info['_pdbx_audit_revision_history.revision_date'][0]\n```",
    "3156529": "`run_cdhit_clustering` is also unknown without a published `utils` file",
    "3158772": "Just added utils.py! Sorry somehow it was not committed earlier",
    "3158935": "Thanks a lot .",
    "3167533": "Thank you so much for sharing this pipeline. It helps a lot!\nI have a question about the filtering pipeline. In step 3 to get_xyz_data, you filtered the dataset by \"structuredness\".  What does \"structuredness\" mean? And how do you get the \"extracted_structures.csv\" file? \n\n```python\nstats=pd.read_csv(\"extracted_structures.csv\")\nstats=stats.set_index('cif')\nfiltered_cif_files=[f for f in cif_files if f in stats.index and stats.loc[f,'structuredness']>0.2]\n```",
    "3170760": "shujun717 can you reply this question? I think we need this information to generate more training data",
    "3170770": "It's the fraction of paired residues in an RNA structure",
    "3177372": "Super helpful!"
  },
  "source": "meta"
}