{
  "id": 497826,
  "title": "Protein crystal structures",
  "url": "/competitions/leash-BELKA/discussion/497826",
  "author_name": "KirkDCO",
  "post_date": "2024-04-25T21:46:58.631000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<p>For anyone wanting protein structural information, I've just created a dataset with 270 crystal structures for the proteins associated with this competition.  <a href=\"https://www.kaggle.com/datasets/kirkdco/leashbio-belka-proteins\" target=\"_blank\">LeashBio_BELKA_Proteins</a></p>\n<ul>\n<li>sEH - 91 stuctures</li>\n<li>BRD4 - 47 structures</li>\n<li>HSA - 132 structures</li>\n</ul>\n<p>Each protein has its own directory with Original and Aligned subdirectories containing the original PDB.gz files and versions aligned to the structure described by the competition hosts, resp.  I've also included 3 <a href=\"https://pymol.org/\" target=\"_blank\">PyMOL</a> files (*.pse) that contain all the structures for each protein aligned.</p>\n<p>This is a work in progress and only the first version.  The various structures contain monomers, dimers, trimers, tetramers, etc. as well as other proteins along with the ones of interest.  Many files also have bound small molecules, which could be useful to enrich the training set.  I will work on updates to this dataset as the competition progresses and post \"Release Notes\" in this thread.  </p>\n<p><strong>24.04.25 Initial Release</strong></p>\n<ul>\n<li>Actual version 2 of the original as one set of proteins was in the wrong directory.</li>\n</ul>\n<p><strong>24.04.27</strong></p>\n<ul>\n<li>I found that in the dataset here on Kaggle, the .pdb.gz files were extracted by Kaggle, which is not what I wanted.  So, I extracted them all myself and replaced the .pdb.gz files with the .pdb files.  Bigger, but less confusing.</li>\n<li>Based on <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/498333\" target=\"_blank\">my response to this post</a>, I extracted all the ligands from the sEH structures and added them to the <code>sEH/Ligands</code> directory.</li>\n<li>Updated various details on dataset homepage.</li>\n</ul>",
  "messages": [
    {
      "id": 2775905,
      "postDate": "2024-04-25T21:46:58.630Z",
      "content": "<p>For anyone wanting protein structural information, I've just created a dataset with 270 crystal structures for the proteins associated with this competition.  <a href=\"https://www.kaggle.com/datasets/kirkdco/leashbio-belka-proteins\" target=\"_blank\">LeashBio_BELKA_Proteins</a></p>\n<ul>\n<li>sEH - 91 stuctures</li>\n<li>BRD4 - 47 structures</li>\n<li>HSA - 132 structures</li>\n</ul>\n<p>Each protein has its own directory with Original and Aligned subdirectories containing the original PDB.gz files and versions aligned to the structure described by the competition hosts, resp.  I've also included 3 <a href=\"https://pymol.org/\" target=\"_blank\">PyMOL</a> files (*.pse) that contain all the structures for each protein aligned.</p>\n<p>This is a work in progress and only the first version.  The various structures contain monomers, dimers, trimers, tetramers, etc. as well as other proteins along with the ones of interest.  Many files also have bound small molecules, which could be useful to enrich the training set.  I will work on updates to this dataset as the competition progresses and post \"Release Notes\" in this thread.  </p>\n<p><strong>24.04.25 Initial Release</strong></p>\n<ul>\n<li>Actual version 2 of the original as one set of proteins was in the wrong directory.</li>\n</ul>\n<p><strong>24.04.27</strong></p>\n<ul>\n<li>I found that in the dataset here on Kaggle, the .pdb.gz files were extracted by Kaggle, which is not what I wanted.  So, I extracted them all myself and replaced the .pdb.gz files with the .pdb files.  Bigger, but less confusing.</li>\n<li>Based on <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/498333\" target=\"_blank\">my response to this post</a>, I extracted all the ligands from the sEH structures and added them to the <code>sEH/Ligands</code> directory.</li>\n<li>Updated various details on dataset homepage.</li>\n</ul>",
      "rawMarkdown": "For anyone wanting protein structural information, I've just created a dataset with 270 crystal structures for the proteins associated with this competition.  [LeashBio_BELKA_Proteins](https://www.kaggle.com/datasets/kirkdco/leashbio-belka-proteins)\n\n* sEH - 91 stuctures\n* BRD4 - 47 structures\n* HSA - 132 structures\n\nEach protein has its own directory with Original and Aligned subdirectories containing the original PDB.gz files and versions aligned to the structure described by the competition hosts, resp.  I've also included 3 [PyMOL](https://pymol.org/) files (*.pse) that contain all the structures for each protein aligned.\n\nThis is a work in progress and only the first version.  The various structures contain monomers, dimers, trimers, tetramers, etc. as well as other proteins along with the ones of interest.  Many files also have bound small molecules, which could be useful to enrich the training set.  I will work on updates to this dataset as the competition progresses and post \"Release Notes\" in this thread.  \n\n**24.04.25 Initial Release**\n* Actual version 2 of the original as one set of proteins was in the wrong directory.\n\n**24.04.27**\n* I found that in the dataset here on Kaggle, the .pdb.gz files were extracted by Kaggle, which is not what I wanted.  So, I extracted them all myself and replaced the .pdb.gz files with the .pdb files.  Bigger, but less confusing.\n* Based on [my response to this post](https://www.kaggle.com/competitions/leash-BELKA/discussion/498333), I extracted all the ligands from the sEH structures and added them to the `sEH/Ligands` directory.\n* Updated various details on dataset homepage.",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2775905": "For anyone wanting protein structural information, I've just created a dataset with 270 crystal structures for the proteins associated with this competition.  [LeashBio_BELKA_Proteins](https://www.kaggle.com/datasets/kirkdco/leashbio-belka-proteins)\n\n* sEH - 91 stuctures\n* BRD4 - 47 structures\n* HSA - 132 structures\n\nEach protein has its own directory with Original and Aligned subdirectories containing the original PDB.gz files and versions aligned to the structure described by the competition hosts, resp.  I've also included 3 [PyMOL](https://pymol.org/) files (*.pse) that contain all the structures for each protein aligned.\n\nThis is a work in progress and only the first version.  The various structures contain monomers, dimers, trimers, tetramers, etc. as well as other proteins along with the ones of interest.  Many files also have bound small molecules, which could be useful to enrich the training set.  I will work on updates to this dataset as the competition progresses and post \"Release Notes\" in this thread.  \n\n**24.04.25 Initial Release**\n* Actual version 2 of the original as one set of proteins was in the wrong directory.\n\n**24.04.27**\n* I found that in the dataset here on Kaggle, the .pdb.gz files were extracted by Kaggle, which is not what I wanted.  So, I extracted them all myself and replaced the .pdb.gz files with the .pdb files.  Bigger, but less confusing.\n* Based on [my response to this post](https://www.kaggle.com/competitions/leash-BELKA/discussion/498333), I extracted all the ligands from the sEH structures and added them to the `sEH/Ligands` directory.\n* Updated various details on dataset homepage."
  }
}