{
  "id": 492458,
  "title": "is Protein have SMILES string format? [solved - answer no]",
  "url": "/competitions/leash-BELKA/discussion/492458",
  "author_name": "",
  "post_date": "2024-04-09T16:34:05.515203800Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<blockquote>\n  <p>is this right ? for converting 3 Proteins to SMILES format using following site <br>\n  <a href=\"https://www.novoprolabs.com/tools/convert-peptide-to-smiles-string\" target=\"_blank\">PepSMI: Convert Peptide to SMILES string</a></p>\n  <table>\n  <thead>\n  <tr>\n  <th>Protein</th>\n  <th>SMILES</th>\n  </tr>\n  </thead>\n  <tbody>\n  <tr>\n  <td>BRD4</td>\n  <td>N<a href=\"[H]\" target=\"_blank\">C@@</a>(CCCNC(=N)N)C(=O)N<a href=\"[H]\" target=\"_blank\">C@@</a>(CC(=O)O)C(=O)O</td>\n  </tr>\n  <tr>\n  <td>HSA</td>\n  <td>N<a href=\"[H]\" target=\"_blank\">C@@</a>(CC1=CN=C-N1)C(=O)N<a href=\"[H]\" target=\"_blank\">C@@</a>(CO)C(=O)N<a href=\"[H]\" target=\"_blank\">C@@</a>(C)C(=O)O</td>\n  </tr>\n  <tr>\n  <td>sEH</td>\n  <td>N<a href=\"[H]\" target=\"_blank\">C@</a>(CO)C(=O)N<a href=\"[H]\" target=\"_blank\">C@@</a>(CCC(=O)O)C(=O)N<a href=\"[H]\" target=\"_blank\">C@@</a>(CC1=CN=C-N1)C(=O)O</td>\n  </tr>\n  </tbody>\n  </table>\n</blockquote>\n<h1>Above process is wrong</h1>",
  "messages": [
    {
      "id": "2743869",
      "postDate": "04/09/2024 16:34:05",
      "content": "<blockquote>\n  <p>is this right ? for converting 3 Proteins to SMILES format using following site <br>\n  <a href=\"https://www.novoprolabs.com/tools/convert-peptide-to-smiles-string\" target=\"_blank\">PepSMI: Convert Peptide to SMILES string</a></p>\n  <table>\n  <thead>\n  <tr>\n  <th>Protein</th>\n  <th>SMILES</th>\n  </tr>\n  </thead>\n  <tbody>\n  <tr>\n  <td>BRD4</td>\n  <td>N<a href=\"[H]\" target=\"_blank\">C@@</a>(CCCNC(=N)N)C(=O)N<a href=\"[H]\" target=\"_blank\">C@@</a>(CC(=O)O)C(=O)O</td>\n  </tr>\n  <tr>\n  <td>HSA</td>\n  <td>N<a href=\"[H]\" target=\"_blank\">C@@</a>(CC1=CN=C-N1)C(=O)N<a href=\"[H]\" target=\"_blank\">C@@</a>(CO)C(=O)N<a href=\"[H]\" target=\"_blank\">C@@</a>(C)C(=O)O</td>\n  </tr>\n  <tr>\n  <td>sEH</td>\n  <td>N<a href=\"[H]\" target=\"_blank\">C@</a>(CO)C(=O)N<a href=\"[H]\" target=\"_blank\">C@@</a>(CCC(=O)O)C(=O)N<a href=\"[H]\" target=\"_blank\">C@@</a>(CC1=CN=C-N1)C(=O)O</td>\n  </tr>\n  </tbody>\n  </table>\n</blockquote>\n<h1>Above process is wrong</h1>",
      "rawMarkdown": "> is this right ? for converting 3 Proteins to SMILES format using following site \n[PepSMI: Convert Peptide to SMILES string](https://www.novoprolabs.com/tools/convert-peptide-to-smiles-string)\n| Protein | SMILES |\n| --- | --- |\n|  BRD4 | N[C@@]([H])(CCCNC(=N)N)C(=O)N[C@@]([H])(CC(=O)O)C(=O)O  |\n|  HSA |  N[C@@]([H])(CC1=CN=C-N1)C(=O)N[C@@]([H])(CO)C(=O)N[C@@]([H])(C)C(=O)O |\n|  sEH |  N[C@]([H])(CO)C(=O)N[C@@]([H])(CCC(=O)O)C(=O)N[C@@]([H])(CC1=CN=C-N1)C(=O)O |~~\n\n# Above process is wrong",
      "votes": null
    },
    {
      "id": "2743894",
      "postDate": "04/09/2024 16:44:04",
      "content": "<p>no, this is treating the letters in the name of the protein as if they were 1-letter codes for amino acids.</p>\n<p>You would first need to get the amino acid sequence for each of these proteins. You can find a link to the uniprot for each of the proteins in the data tab</p>\n<p><a href=\"https://www.uniprot.org/uniprotkb/P34913/entry#sequences\" target=\"_blank\">https://www.uniprot.org/uniprotkb/P34913/entry#sequences</a></p>",
      "rawMarkdown": "no, this is treating the letters in the name of the protein as if they were 1-letter codes for amino acids.\n\nYou would first need to get the amino acid sequence for each of these proteins. You can find a link to the uniprot for each of the proteins in the data tab\n\nhttps://www.uniprot.org/uniprotkb/P34913/entry#sequences",
      "votes": null
    },
    {
      "id": "2744033",
      "postDate": "04/09/2024 18:05:13",
      "content": "<p>use fasta<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe95fa70ce95d2a019a2e0f1e8b0fa764%2FSelection_014.png?generation=1712685906170035&amp;alt=media\"></p>",
      "rawMarkdown": "use fasta\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe95fa70ce95d2a019a2e0f1e8b0fa764%2FSelection_014.png?generation=1712685906170035&alt=media)",
      "votes": null
    },
    {
      "id": "2744383",
      "postDate": "04/09/2024 21:38:19",
      "content": "<p>It's usual to represent proteins differently from the small/medium-sized ligand molecules that may bind to them in pharmaceutal chemistry.</p>\n<p>SMILES is great for representing the structures of small to medium-sized molecules, say Mol WT &lt; 500 or a few tens of atoms. However, it is a really clunky and inefficient way to represent a protein structure, where for almost all purposes we want either the amino acid sequence or the 3D structure. A typical protein has hundreds of residues and hence thousands of atoms. Margaret Dayhoff's one-letter code is an order of magnitude more efficient than SMILES for representing the chemical structure of a protein.</p>",
      "rawMarkdown": "It's usual to represent proteins differently from the small/medium-sized ligand molecules that may bind to them in pharmaceutal chemistry.\n\nSMILES is great for representing the structures of small to medium-sized molecules, say Mol WT < 500 or a few tens of atoms. However, it is a really clunky and inefficient way to represent a protein structure, where for almost all purposes we want either the amino acid sequence or the 3D structure. A typical protein has hundreds of residues and hence thousands of atoms. Margaret Dayhoff's one-letter code is an order of magnitude more efficient than SMILES for representing the chemical structure of a protein.",
      "votes": null
    },
    {
      "id": "2758720",
      "postDate": "04/18/2024 10:05:23",
      "content": "<p>Proteins can be represented in SMILES, but it woulud be super long.</p>",
      "rawMarkdown": "Proteins can be represented in SMILES, but it woulud be super long.",
      "votes": null
    },
    {
      "id": "2758744",
      "postDate": "04/18/2024 10:18:28",
      "content": "<p><a href=\"https://www.kaggle.com/gyulamaloveczky4\" target=\"_blank\">@gyulamaloveczky4</a> will you share the SMILES format for following 3 proteins<br>\n<strong>BRD4</strong>, <strong>HSA</strong>, <strong>sEH</strong></p>",
      "rawMarkdown": "gyulamaloveczky4 will you share the SMILES format for following 3 proteins\n**BRD4**, **HSA**, **sEH**",
      "votes": null
    },
    {
      "id": "2758778",
      "postDate": "04/18/2024 10:41:26",
      "content": "<p>The converter you were using can not deal with such long peptides (maximum 60 peptides). The SMILES encoding would be longer than 2000 characters in the case of these proteins. Might not be the best way to represent them tbh.</p>",
      "rawMarkdown": "The converter you were using can not deal with such long peptides (maximum 60 peptides). The SMILES encoding would be longer than 2000 characters in the case of these proteins. Might not be the best way to represent them tbh.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2743894,
      "author_name": "andrewdblevins",
      "author_url": "",
      "post_date": "04/09/2024 16:44:04",
      "content": "<p>no, this is treating the letters in the name of the protein as if they were 1-letter codes for amino acids.</p>\n<p>You would first need to get the amino acid sequence for each of these proteins. You can find a link to the uniprot for each of the proteins in the data tab</p>\n<p><a href=\"https://www.uniprot.org/uniprotkb/P34913/entry#sequences\" target=\"_blank\">https://www.uniprot.org/uniprotkb/P34913/entry#sequences</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744033,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/09/2024 18:05:13",
      "content": "<p>use fasta<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe95fa70ce95d2a019a2e0f1e8b0fa764%2FSelection_014.png?generation=1712685906170035&amp;alt=media\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744383,
      "author_name": "jbomitchell",
      "author_url": "",
      "post_date": "04/09/2024 21:38:19",
      "content": "<p>It's usual to represent proteins differently from the small/medium-sized ligand molecules that may bind to them in pharmaceutal chemistry.</p>\n<p>SMILES is great for representing the structures of small to medium-sized molecules, say Mol WT &lt; 500 or a few tens of atoms. However, it is a really clunky and inefficient way to represent a protein structure, where for almost all purposes we want either the amino acid sequence or the 3D structure. A typical protein has hundreds of residues and hence thousands of atoms. Margaret Dayhoff's one-letter code is an order of magnitude more efficient than SMILES for representing the chemical structure of a protein.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2758720,
      "author_name": "gyulamaloveczky4",
      "author_url": "",
      "post_date": "04/18/2024 10:05:23",
      "content": "<p>Proteins can be represented in SMILES, but it woulud be super long.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2758744,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "04/18/2024 10:18:28",
          "content": "<p><a href=\"https://www.kaggle.com/gyulamaloveczky4\" target=\"_blank\">@gyulamaloveczky4</a> will you share the SMILES format for following 3 proteins<br>\n<strong>BRD4</strong>, <strong>HSA</strong>, <strong>sEH</strong></p>",
          "votes": null,
          "replies": [
            {
              "id": 2758778,
              "author_name": "gyulamaloveczky4",
              "author_url": "",
              "post_date": "04/18/2024 10:41:26",
              "content": "<p>The converter you were using can not deal with such long peptides (maximum 60 peptides). The SMILES encoding would be longer than 2000 characters in the case of these proteins. Might not be the best way to represent them tbh.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2743869": "> is this right ? for converting 3 Proteins to SMILES format using following site \n[PepSMI: Convert Peptide to SMILES string](https://www.novoprolabs.com/tools/convert-peptide-to-smiles-string)\n| Protein | SMILES |\n| --- | --- |\n|  BRD4 | N[C@@]([H])(CCCNC(=N)N)C(=O)N[C@@]([H])(CC(=O)O)C(=O)O  |\n|  HSA |  N[C@@]([H])(CC1=CN=C-N1)C(=O)N[C@@]([H])(CO)C(=O)N[C@@]([H])(C)C(=O)O |\n|  sEH |  N[C@]([H])(CO)C(=O)N[C@@]([H])(CCC(=O)O)C(=O)N[C@@]([H])(CC1=CN=C-N1)C(=O)O |~~\n\n# Above process is wrong",
    "2743894": "no, this is treating the letters in the name of the protein as if they were 1-letter codes for amino acids.\n\nYou would first need to get the amino acid sequence for each of these proteins. You can find a link to the uniprot for each of the proteins in the data tab\n\nhttps://www.uniprot.org/uniprotkb/P34913/entry#sequences",
    "2744033": "use fasta\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe95fa70ce95d2a019a2e0f1e8b0fa764%2FSelection_014.png?generation=1712685906170035&alt=media)",
    "2744383": "It's usual to represent proteins differently from the small/medium-sized ligand molecules that may bind to them in pharmaceutal chemistry.\n\nSMILES is great for representing the structures of small to medium-sized molecules, say Mol WT < 500 or a few tens of atoms. However, it is a really clunky and inefficient way to represent a protein structure, where for almost all purposes we want either the amino acid sequence or the 3D structure. A typical protein has hundreds of residues and hence thousands of atoms. Margaret Dayhoff's one-letter code is an order of magnitude more efficient than SMILES for representing the chemical structure of a protein.",
    "2758720": "Proteins can be represented in SMILES, but it woulud be super long.",
    "2758744": "gyulamaloveczky4 will you share the SMILES format for following 3 proteins\n**BRD4**, **HSA**, **sEH**",
    "2758778": "The converter you were using can not deal with such long peptides (maximum 60 peptides). The SMILES encoding would be longer than 2000 characters in the case of these proteins. Might not be the best way to represent them tbh."
  },
  "source": "meta"
}