{
  "id": 567959,
  "title": "Convert UW synthetic dataset of 440k synthetic RNAs",
  "url": "/competitions/stanford-rna-3d-folding/discussion/567959",
  "author_name": "tomoo inubushi",
  "post_date": "2025-03-13T02:33:38.101000",
  "votes": 20,
  "comment_count": 10,
  "views": 0,
  "content": "<p>The host shared <a href=\"https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-final\" target=\"_blank\">synthetic RNA structure dataset</a> including 447,402 RNA sequences.</p>\n<p>I converted this dataset into the same format of this competition. You can download it from <a href=\"https://www.kaggle.com/datasets/tomooinubushi/converted-uw-synthetic-rna-final\" target=\"_blank\">here</a>.<br>\nThe conversion script is <a href=\"https://www.kaggle.com/code/tomooinubushi/convert-uw-synthetic-dataset\" target=\"_blank\">here</a>.</p>\n<p>I am not sure it works for this competition. I appreciate if anyone share us the result.</p>\n<p>p.s.<br>\nDear <a href=\"https://www.kaggle.com/andrewfavor\" target=\"_blank\">@andrewfavor</a> </p>\n<ul>\n<li> Already answered in <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567539\" target=\"_blank\">here</a>. They are identical.</li>\n<li> Already answered in <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567539\" target=\"_blank\">here</a>. We should just use final version.</li>\n<li>What are the licenses of these datasets? Is my dataset OK?</li>\n</ul>\n<p><strong>[UPDATE 20240314]</strong> I rewrited my conversion script and reuploaded the dataset. Please use the latest version from <a href=\"https://www.kaggle.com/datasets/tomooinubushi/converted-uw-synthetic-rna-final\" target=\"_blank\">here</a><br>\nI also converted <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">Stanford Ribonanza RNA Folding</a> (<a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/451853\" target=\"_blank\">data description</a>) including 124k RNAs. The conversion script is <a href=\"https://www.kaggle.com/code/tomooinubushi/convert-stanford-ribonanza-rna-folding-dataset\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": 3148328,
      "postDate": "2025-03-13T02:33:38.100Z",
      "content": "<p>The host shared <a href=\"https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-final\" target=\"_blank\">synthetic RNA structure dataset</a> including 447,402 RNA sequences.</p>\n<p>I converted this dataset into the same format of this competition. You can download it from <a href=\"https://www.kaggle.com/datasets/tomooinubushi/converted-uw-synthetic-rna-final\" target=\"_blank\">here</a>.<br>\nThe conversion script is <a href=\"https://www.kaggle.com/code/tomooinubushi/convert-uw-synthetic-dataset\" target=\"_blank\">here</a>.</p>\n<p>I am not sure it works for this competition. I appreciate if anyone share us the result.</p>\n<p>p.s.<br>\nDear <a href=\"https://www.kaggle.com/andrewfavor\" target=\"_blank\">@andrewfavor</a> </p>\n<ul>\n<li> Already answered in <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567539\" target=\"_blank\">here</a>. They are identical.</li>\n<li> Already answered in <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567539\" target=\"_blank\">here</a>. We should just use final version.</li>\n<li>What are the licenses of these datasets? Is my dataset OK?</li>\n</ul>\n<p><strong>[UPDATE 20240314]</strong> I rewrited my conversion script and reuploaded the dataset. Please use the latest version from <a href=\"https://www.kaggle.com/datasets/tomooinubushi/converted-uw-synthetic-rna-final\" target=\"_blank\">here</a><br>\nI also converted <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">Stanford Ribonanza RNA Folding</a> (<a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/451853\" target=\"_blank\">data description</a>) including 124k RNAs. The conversion script is <a href=\"https://www.kaggle.com/code/tomooinubushi/convert-stanford-ribonanza-rna-folding-dataset\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "The host shared [synthetic RNA structure dataset](https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-final) including 447,402 RNA sequences.\n\nI converted this dataset into the same format of this competition. You can download it from [here](https://www.kaggle.com/datasets/tomooinubushi/converted-uw-synthetic-rna-final).\nThe conversion script is [here](https://www.kaggle.com/code/tomooinubushi/convert-uw-synthetic-dataset).\n\nI am not sure it works for this competition. I appreciate if anyone share us the result.\n\np.s.\nDear @andrewfavor \n- ~~What is the difference between [uw-synthetic-rna-final](https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-final) and [uw-synthetic-rna-structures](https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-structures) ?~~ Already answered in [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567539). They are identical.\n- ~~Should we concatenate these datasets? or just use final?~~ Already answered in [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567539). We should just use final version.\n- What are the licenses of these datasets? Is my dataset OK?\n\n\n**[UPDATE 20240314]** I rewrited my conversion script and reuploaded the dataset. Please use the latest version from [here](https://www.kaggle.com/datasets/tomooinubushi/converted-uw-synthetic-rna-final)\nI also converted [Stanford Ribonanza RNA Folding](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data) ([data description](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/451853)) including 124k RNAs. The conversion script is [here](https://www.kaggle.com/code/tomooinubushi/convert-stanford-ribonanza-rna-folding-dataset).",
      "votes": 20
    },
    {
      "id": 3148517,
      "postDate": "2025-03-13T08:13:36.307Z",
      "content": "<p>i did not check your code in details but:</p>\n<ol>\n<li>we only need C1 atom (we predict backbone structure only)</li>\n<li>resname and resid should be extracted from pdb file (and not running number)</li>\n</ol>\n<p>while it is safe to assume that there is only one strcuture and the chain is continous for synthetic data, it may be better to use assert check in your code for these 2 points.</p>\n<p>finally, you can verify your conversion code is correct or not by:</p>\n<ul>\n<li>choose a pdb id from kaggle train or validation csv</li>\n<li>download pdb file from protein bank</li>\n<li>compare your conversion results with that from the given kaggle label csv</li>\n</ul>",
      "rawMarkdown": "i did not check your code in details but:\n1. we only need C1 atom (we predict backbone structure only)\n2. resname and resid should be extracted from pdb file (and not running number)\n\nwhile it is safe to assume that there is only one strcuture and the chain is continous for synthetic data, it may be better to use assert check in your code for these 2 points.\n\nfinally, you can verify your conversion code is correct or not by:\n- choose a pdb id from kaggle train or validation csv\n- download pdb file from protein bank\n- compare your conversion results with that from the given kaggle label csv",
      "votes": 4,
      "replies": [
        {
          "id": 3148520,
          "postDate": "2025-03-13T08:20:55.750Z",
          "content": "<pre><code>def parse_pdb_to_df(pdb_file, target_id):\n    parser = PDBParser()\n\n    structure = parser.get_structure('', pdb_file)\n\n    df=[]\n    for model in structure:\n        for chain in model:\n            print(chain)\n            chain_data = []\n            for residue in chain:\n                 print(residue)\n                if residue.get_resname() in ['A', 'U', 'G', 'C']:\n\n                     Check if the residue has a C1' atom\n                    if 'C1\\'' in residue:\n                        atom = residue['C1\\'']\n                        xyz = atom.get_coord()\n                        resname = residue.get_resname()\n                        resid = residue.get_id()[1]\n\n                            resname resid   x_1 y_1 z_1\n                         detect discontinous: resid = previous resid +1\n                        chain_data.append(dict(\n                            ID=target_id+'_'+str(resid),  change to target_id+'_'+str(chainid)+'_'+str(resid)\n                            resname=resname,\n                            resid=resid,\n                            x_1=xyz[0],\n                            y_1=xyz[1],\n                            z_1=xyz[2],\n                        ))\n                        \n\n            if len(chain_data)!=0:\n                chain_df = pd.DataFrame(chain_data)\n                df.append(chain_df)\n                \n    return df\n</code></pre>",
          "rawMarkdown": "```\n\n\ndef parse_pdb_to_df(pdb_file, target_id):\n    parser = PDBParser()\n\n    structure = parser.get_structure('', pdb_file)\n\n    df=[]\n    for model in structure:\n        for chain in model:\n            print(chain)\n            chain_data = []\n            for residue in chain:\n                # print(residue)\n                if residue.get_resname() in ['A', 'U', 'G', 'C']:\n\n                    # Check if the residue has a C1' atom\n                    if 'C1\\'' in residue:\n                        atom = residue['C1\\'']\n                        xyz = atom.get_coord()\n                        resname = residue.get_resname()\n                        resid = residue.get_id()[1]\n\n                        #ID\tresname\tresid\tx_1\ty_1\tz_1\n                        #todo detect discontinous: resid = previous resid +1\n                        chain_data.append(dict(\n                            ID=target_id+'_'+str(resid), #todo: change to target_id+'_'+str(chainid)+'_'+str(resid)\n                            resname=resname,\n                            resid=resid,\n                            x_1=xyz[0],\n                            y_1=xyz[1],\n                            z_1=xyz[2],\n                        ))\n                        ##print(f\"Residue {resname} {resid}, Atom: {atom.get_name()}, xyz: {xyz}\")\n\n            if len(chain_data)!=0:\n                chain_df = pd.DataFrame(chain_data)\n                df.append(chain_df)\n                ##print(chain_df)\n    return df\n\n\n```",
          "votes": 4,
          "replies": [
            {
              "id": 3148580,
              "postDate": "2025-03-13T10:05:33.610Z",
              "content": "<p>Thank you for your comment. <br>\nI am <a href=\"https://www.kaggle.com/code/tomooinubushi/convert-uw-synthetic-dataset\" target=\"_blank\">now rewriting my conversion code</a>.</p>\n<p>I downloaded 1scl.pdb from <a href=\"https://www.rcsb.org/structure/1scl\" target=\"_blank\">here</a> and uploaded to <a href=\"https://www.kaggle.com/datasets/tomooinubushi/1scl-pdb\" target=\"_blank\">here</a>, and compared the data.</p>\n<p>I found that the sequences of these data are the same, but x, y, and z coordinates are different.</p>\n<p>According to 1scl.pdb, x, y, and z of first c1' are 11.334, -27.307,  -2.027.<br>\nOn the other hand, according to 1SCL_A data in train_labels.csv, x, y, and z are 13.760,  -25.974001,  0.102.<br>\nDo anybody give me a suggestion? Maybe I downloaded the wrong pdb file.</p>\n<p>1SCL_A data in train_labels.csv </p>\n<table>\n<thead>\n<tr>\n<th>ID</th>\n<th>resname</th>\n<th>resid</th>\n<th>x_1</th>\n<th>y_1</th>\n<th>z_1</th>\n<th>pdb_id</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1SCL_A_1</td>\n<td>G</td>\n<td>1</td>\n<td>13.760</td>\n<td>-25.974001</td>\n<td>0.102</td>\n<td>1SCL_A</td>\n</tr>\n<tr>\n<td>1SCL_A_2</td>\n<td>G</td>\n<td>2</td>\n<td>9.310</td>\n<td>-29.638000</td>\n<td>2.669</td>\n<td>1SCL_A</td>\n</tr>\n<tr>\n<td>1SCL_A_3</td>\n<td>G</td>\n<td>3</td>\n<td>5.529</td>\n<td>-27.813000</td>\n<td>5.878</td>\n<td>1SCL_A</td>\n</tr>\n<tr>\n<td>1SCL_A_4</td>\n<td>U</td>\n<td>4</td>\n<td>2.678</td>\n<td>-24.900999</td>\n<td>9.793</td>\n<td>1SCL_A</td>\n</tr>\n<tr>\n<td>1SCL_A_5</td>\n<td>G</td>\n<td>5</td>\n<td>1.827</td>\n<td>-20.136000</td>\n<td>11.793</td>\n<td>1SCL_A</td>\n</tr>\n</tbody>\n</table>\n<p><a href=\"https://www.kaggle.com/code/tomooinubushi/convert-uw-synthetic-dataset\" target=\"_blank\">The result of my current script</a> from 1scl.pdb</p>\n<table>\n<thead>\n<tr>\n<th>resname</th>\n<th>resid</th>\n<th>x_1</th>\n<th>y_1</th>\n<th>z_1</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>G</td>\n<td>1</td>\n<td>11.334</td>\n<td>-27.306999</td>\n<td>-2.027</td>\n</tr>\n<tr>\n<td>G</td>\n<td>2</td>\n<td>6.168</td>\n<td>-29.639000</td>\n<td>0.850</td>\n</tr>\n<tr>\n<td>G</td>\n<td>3</td>\n<td>3.555</td>\n<td>-27.528999</td>\n<td>5.061</td>\n</tr>\n<tr>\n<td>U</td>\n<td>4</td>\n<td>1.614</td>\n<td>-24.337999</td>\n<td>9.446</td>\n</tr>\n<tr>\n<td>G</td>\n<td>5</td>\n<td>2.020</td>\n<td>-19.667000</td>\n<td>11.685</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2F8a5080341d56ff7b3ce2555db7edd7a2%2F__results___6_0.png?generation=1741860220469441&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Thank you for your comment. \nI am [now rewriting my conversion code](https://www.kaggle.com/code/tomooinubushi/convert-uw-synthetic-dataset).\n\nI downloaded 1scl.pdb from [here](https://www.rcsb.org/structure/1scl) and uploaded to [here](https://www.kaggle.com/datasets/tomooinubushi/1scl-pdb), and compared the data.\n\nI found that the sequences of these data are the same, but x, y, and z coordinates are different.\n\nAccording to 1scl.pdb, x, y, and z of first c1' are 11.334, -27.307,  -2.027.\nOn the other hand, according to 1SCL_A data in train_labels.csv, x, y, and z are 13.760,  -25.974001,  0.102.\nDo anybody give me a suggestion? Maybe I downloaded the wrong pdb file.\n\n\n1SCL_A data in train_labels.csv \n\n|ID | resname | resid | x_1 | y_1  | z_1 | pdb_id|\n| --- | --- |  --- | --- | --- |  --- |\n|  1SCL_A_1|  G \t|  1 \t|  13.760|  -25.974001 |  0.102|  1SCL_A|  \n|  1SCL_A_2 |  G \t|  2 \t|  9.310|  -29.638000|  2.669|  1SCL_A|  \n|  1SCL_A_3 |  G \t|  3 \t|  5.529|  -27.813000 |  5.878|  1SCL_A|  \n|  1SCL_A_4 |  U \t|  4 \t|  2.678|  -24.900999 |  9.793|  1SCL_A|  \n|  1SCL_A_5 |  G \t|  5 \t|  1.827 |  -20.136000 |  11.793|  1SCL_A|  \n\n\n [The result of my current script](https://www.kaggle.com/code/tomooinubushi/convert-uw-synthetic-dataset) from 1scl.pdb\n\n| resname | resid | x_1 | y_1  | z_1 | \n| --- |  --- | --- | --- |  --- |\n|  G \t|  1 \t|  11.334 |-27.306999|-2.027|\n|  G \t|  2 \t|  6.168|-29.639000|0.850|\n|  G \t|  3 \t|  3.555 |-27.528999|5.061|\n|  U \t|  4 \t|  1.614 |-24.337999|9.446|\n|  G \t|  5 \t|  2.020 |-19.667000|11.685|\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2F8a5080341d56ff7b3ce2555db7edd7a2%2F__results___6_0.png?generation=1741860220469441&alt=media)"
            },
            {
              "id": 3148583,
              "postDate": "2025-03-13T10:09:45.403Z",
              "content": "<p>not all will target be the same. if there are some target the same, then your code is correct.</p>\n<p>e.g. in my conversion, i have<br>\nkaggle traget id1  = same (all residue has extact same xyz from downloaded pdb)<br>\nkaggle traget id2  = not same (athough there is only one coformation in downloaded pdb)</p>\n<p>(i am not sure why they are different but it could be due to different version updated? of pdb file, or maybe conformation, etc)</p>",
              "rawMarkdown": "not all will target be the same. if there are some target the same, then your code is correct.\n\ne.g. in my conversion, i have\nkaggle traget id1  = same (all residue has extact same xyz from downloaded pdb)\nkaggle traget id2  = not same (athough there is only one coformation in downloaded pdb)\n\n\n(i am not sure why they are different but it could be due to different version updated? of pdb file, or maybe conformation, etc)",
              "votes": 1
            },
            {
              "id": 3148632,
              "postDate": "2025-03-13T11:11:17.037Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 3148656,
              "postDate": "2025-03-13T11:37:08.440Z",
              "content": "<p>Thank you for your comment. </p>\n<p>So, train_labels.csv with only one conformation would be noisy labels.</p>",
              "rawMarkdown": "Thank you for your comment. \n\nSo, train_labels.csv with only one conformation would be noisy labels."
            },
            {
              "id": 3149593,
              "postDate": "2025-03-14T12:23:03.693Z",
              "content": "<p>I  downloaded 1rnk.pdb from <a href=\"https://www.rcsb.org/structure/1rnk\" target=\"_blank\">here</a>, and confirmed successful conversion.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2Fa58369a3497da65a7fa6d7cc31d6e149%2F__results___7_0.png?generation=1741954946178236&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "I  downloaded 1rnk.pdb from [here](https://www.rcsb.org/structure/1rnk), and confirmed successful conversion.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2Fa58369a3497da65a7fa6d7cc31d6e149%2F__results___7_0.png?generation=1741954946178236&alt=media)"
            }
          ]
        }
      ]
    },
    {
      "id": 3148387,
      "postDate": "2025-03-13T03:58:18.663Z",
      "content": "<p>Some of your questions are already answered <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567539\" target=\"_blank\"><strong>here</strong></a>.</p>",
      "rawMarkdown": "Some of your questions are already answered [**here**](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567539).",
      "votes": 2,
      "replies": [
        {
          "id": 3148390,
          "postDate": "2025-03-13T04:06:45.270Z",
          "content": "<p>Thank you for your comment. I didn't check it.</p>",
          "rawMarkdown": "Thank you for your comment. I didn't check it."
        }
      ]
    },
    {
      "id": 3160535,
      "postDate": "2025-03-26T22:14:10.547Z",
      "content": "<p>Thanks, this is helpful to me as well.</p>",
      "rawMarkdown": "Thanks, this is helpful to me as well.",
      "votes": 12
    }
  ],
  "comments": [
    {
      "id": 3148517,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-13T08:13:36.307000",
      "content": "<p>i did not check your code in details but:</p>\n<ol>\n<li>we only need C1 atom (we predict backbone structure only)</li>\n<li>resname and resid should be extracted from pdb file (and not running number)</li>\n</ol>\n<p>while it is safe to assume that there is only one strcuture and the chain is continous for synthetic data, it may be better to use assert check in your code for these 2 points.</p>\n<p>finally, you can verify your conversion code is correct or not by:</p>\n<ul>\n<li>choose a pdb id from kaggle train or validation csv</li>\n<li>download pdb file from protein bank</li>\n<li>compare your conversion results with that from the given kaggle label csv</li>\n</ul>",
      "votes": 4,
      "replies": [
        {
          "id": 3148520,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-03-13T08:20:55.750000",
          "content": "<pre><code>def parse_pdb_to_df(pdb_file, target_id):\n    parser = PDBParser()\n\n    structure = parser.get_structure('', pdb_file)\n\n    df=[]\n    for model in structure:\n        for chain in model:\n            print(chain)\n            chain_data = []\n            for residue in chain:\n                 print(residue)\n                if residue.get_resname() in ['A', 'U', 'G', 'C']:\n\n                     Check if the residue has a C1' atom\n                    if 'C1\\'' in residue:\n                        atom = residue['C1\\'']\n                        xyz = atom.get_coord()\n                        resname = residue.get_resname()\n                        resid = residue.get_id()[1]\n\n                            resname resid   x_1 y_1 z_1\n                         detect discontinous: resid = previous resid +1\n                        chain_data.append(dict(\n                            ID=target_id+'_'+str(resid),  change to target_id+'_'+str(chainid)+'_'+str(resid)\n                            resname=resname,\n                            resid=resid,\n                            x_1=xyz[0],\n                            y_1=xyz[1],\n                            z_1=xyz[2],\n                        ))\n                        \n\n            if len(chain_data)!=0:\n                chain_df = pd.DataFrame(chain_data)\n                df.append(chain_df)\n                \n    return df\n</code></pre>",
          "votes": 4,
          "replies": [
            {
              "id": 3148580,
              "author_name": "tomoo inubushi",
              "author_url": "",
              "post_date": "2025-03-13T10:05:33.610000",
              "content": "<p>Thank you for your comment. <br>\nI am <a href=\"https://www.kaggle.com/code/tomooinubushi/convert-uw-synthetic-dataset\" target=\"_blank\">now rewriting my conversion code</a>.</p>\n<p>I downloaded 1scl.pdb from <a href=\"https://www.rcsb.org/structure/1scl\" target=\"_blank\">here</a> and uploaded to <a href=\"https://www.kaggle.com/datasets/tomooinubushi/1scl-pdb\" target=\"_blank\">here</a>, and compared the data.</p>\n<p>I found that the sequences of these data are the same, but x, y, and z coordinates are different.</p>\n<p>According to 1scl.pdb, x, y, and z of first c1' are 11.334, -27.307,  -2.027.<br>\nOn the other hand, according to 1SCL_A data in train_labels.csv, x, y, and z are 13.760,  -25.974001,  0.102.<br>\nDo anybody give me a suggestion? Maybe I downloaded the wrong pdb file.</p>\n<p>1SCL_A data in train_labels.csv </p>\n<table>\n<thead>\n<tr>\n<th>ID</th>\n<th>resname</th>\n<th>resid</th>\n<th>x_1</th>\n<th>y_1</th>\n<th>z_1</th>\n<th>pdb_id</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1SCL_A_1</td>\n<td>G</td>\n<td>1</td>\n<td>13.760</td>\n<td>-25.974001</td>\n<td>0.102</td>\n<td>1SCL_A</td>\n</tr>\n<tr>\n<td>1SCL_A_2</td>\n<td>G</td>\n<td>2</td>\n<td>9.310</td>\n<td>-29.638000</td>\n<td>2.669</td>\n<td>1SCL_A</td>\n</tr>\n<tr>\n<td>1SCL_A_3</td>\n<td>G</td>\n<td>3</td>\n<td>5.529</td>\n<td>-27.813000</td>\n<td>5.878</td>\n<td>1SCL_A</td>\n</tr>\n<tr>\n<td>1SCL_A_4</td>\n<td>U</td>\n<td>4</td>\n<td>2.678</td>\n<td>-24.900999</td>\n<td>9.793</td>\n<td>1SCL_A</td>\n</tr>\n<tr>\n<td>1SCL_A_5</td>\n<td>G</td>\n<td>5</td>\n<td>1.827</td>\n<td>-20.136000</td>\n<td>11.793</td>\n<td>1SCL_A</td>\n</tr>\n</tbody>\n</table>\n<p><a href=\"https://www.kaggle.com/code/tomooinubushi/convert-uw-synthetic-dataset\" target=\"_blank\">The result of my current script</a> from 1scl.pdb</p>\n<table>\n<thead>\n<tr>\n<th>resname</th>\n<th>resid</th>\n<th>x_1</th>\n<th>y_1</th>\n<th>z_1</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>G</td>\n<td>1</td>\n<td>11.334</td>\n<td>-27.306999</td>\n<td>-2.027</td>\n</tr>\n<tr>\n<td>G</td>\n<td>2</td>\n<td>6.168</td>\n<td>-29.639000</td>\n<td>0.850</td>\n</tr>\n<tr>\n<td>G</td>\n<td>3</td>\n<td>3.555</td>\n<td>-27.528999</td>\n<td>5.061</td>\n</tr>\n<tr>\n<td>U</td>\n<td>4</td>\n<td>1.614</td>\n<td>-24.337999</td>\n<td>9.446</td>\n</tr>\n<tr>\n<td>G</td>\n<td>5</td>\n<td>2.020</td>\n<td>-19.667000</td>\n<td>11.685</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2F8a5080341d56ff7b3ce2555db7edd7a2%2F__results___6_0.png?generation=1741860220469441&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3148583,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-13T10:09:45.403000",
              "content": "<p>not all will target be the same. if there are some target the same, then your code is correct.</p>\n<p>e.g. in my conversion, i have<br>\nkaggle traget id1  = same (all residue has extact same xyz from downloaded pdb)<br>\nkaggle traget id2  = not same (athough there is only one coformation in downloaded pdb)</p>\n<p>(i am not sure why they are different but it could be due to different version updated? of pdb file, or maybe conformation, etc)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3148632,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-13T11:11:17.037000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3148656,
              "author_name": "tomoo inubushi",
              "author_url": "",
              "post_date": "2025-03-13T11:37:08.440000",
              "content": "<p>Thank you for your comment. </p>\n<p>So, train_labels.csv with only one conformation would be noisy labels.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3149593,
              "author_name": "tomoo inubushi",
              "author_url": "",
              "post_date": "2025-03-14T12:23:03.693000",
              "content": "<p>I  downloaded 1rnk.pdb from <a href=\"https://www.rcsb.org/structure/1rnk\" target=\"_blank\">here</a>, and confirmed successful conversion.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2Fa58369a3497da65a7fa6d7cc31d6e149%2F__results___7_0.png?generation=1741954946178236&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3148387,
      "author_name": "Tilii",
      "author_url": "",
      "post_date": "2025-03-13T03:58:18.663000",
      "content": "<p>Some of your questions are already answered <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567539\" target=\"_blank\"><strong>here</strong></a>.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3148390,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2025-03-13T04:06:45.270000",
          "content": "<p>Thank you for your comment. I didn't check it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3160535,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-26T22:14:10.547000",
      "content": "<p>Thanks, this is helpful to me as well.</p>",
      "votes": 12,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3148328": "The host shared [synthetic RNA structure dataset](https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-final) including 447,402 RNA sequences.\n\nI converted this dataset into the same format of this competition. You can download it from [here](https://www.kaggle.com/datasets/tomooinubushi/converted-uw-synthetic-rna-final).\nThe conversion script is [here](https://www.kaggle.com/code/tomooinubushi/convert-uw-synthetic-dataset).\n\nI am not sure it works for this competition. I appreciate if anyone share us the result.\n\np.s.\nDear @andrewfavor \n- ~~What is the difference between [uw-synthetic-rna-final](https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-final) and [uw-synthetic-rna-structures](https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-structures) ?~~ Already answered in [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567539). They are identical.\n- ~~Should we concatenate these datasets? or just use final?~~ Already answered in [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567539). We should just use final version.\n- What are the licenses of these datasets? Is my dataset OK?\n\n\n**[UPDATE 20240314]** I rewrited my conversion script and reuploaded the dataset. Please use the latest version from [here](https://www.kaggle.com/datasets/tomooinubushi/converted-uw-synthetic-rna-final)\nI also converted [Stanford Ribonanza RNA Folding](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data) ([data description](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/451853)) including 124k RNAs. The conversion script is [here](https://www.kaggle.com/code/tomooinubushi/convert-stanford-ribonanza-rna-folding-dataset).",
    "3148517": "i did not check your code in details but:\n1. we only need C1 atom (we predict backbone structure only)\n2. resname and resid should be extracted from pdb file (and not running number)\n\nwhile it is safe to assume that there is only one strcuture and the chain is continous for synthetic data, it may be better to use assert check in your code for these 2 points.\n\nfinally, you can verify your conversion code is correct or not by:\n- choose a pdb id from kaggle train or validation csv\n- download pdb file from protein bank\n- compare your conversion results with that from the given kaggle label csv",
    "3148387": "Some of your questions are already answered [**here**](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567539).",
    "3160535": "Thanks, this is helpful to me as well."
  }
}