{
  "id": 513137,
  "title": "Building blocks don't add up? ",
  "url": "/competitions/leash-BELKA/discussion/513137",
  "author_name": "",
  "post_date": "2024-06-18T18:07:27.695118800Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<pre><code>!pip install dask rdkit matplotlib\n dask.dataframe  dd\n rdkit  Chem\n rdkit.Chem  Draw\n matplotlib.pyplot  plt\n PIL  Image\n\ntrain_data = dd.read_parquet(, engine=, columns=[,, , ])\n\ncolumns = [, , , ]\n\nimages = []\nlabels = []\n\n col  columns:\n    el = train_data[col].head(, compute=).iloc[]\n\n    molecule = Chem.MolFromSmiles(el)\n\n    img = Draw.MolToImage(molecule)\n\n    img.save()\n\n    images.append(Image.())\n    labels.append(col)\n\nfig, axs = plt.subplots(, (images), figsize=(, ))\n\n ax, img, label  (axs, images, labels):\n    ax.imshow(img)\n    ax.set_title(label)\n    ax.axis()\n\nplt.tight_layout()\nplt.savefig()\nplt.show()\n\n()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10578610%2Fb640c14b9bfb8dbb9ceee1d46f713ece%2FCaptura%20de%20ecra%202024-06-18%20as%2018.43.40.png?generation=1718733887741893&amp;alt=media\"></p>\n<p>Hey! As you can see above BB3 is the part below on the molecule, BB2 clearly the part on the left, but check BB1 it is a bit different from the molecule… they don't add up to the molecule SMILE. Is something wrong here? From this molecule the section that binds would be the section with triple connection?</p>\n<p>Thank you and sorry for the lack in chem knowledge.</p>",
  "messages": [
    {
      "id": "2878081",
      "postDate": "06/18/2024 18:07:27",
      "content": "<pre><code>!pip install dask rdkit matplotlib\n dask.dataframe  dd\n rdkit  Chem\n rdkit.Chem  Draw\n matplotlib.pyplot  plt\n PIL  Image\n\ntrain_data = dd.read_parquet(, engine=, columns=[,, , ])\n\ncolumns = [, , , ]\n\nimages = []\nlabels = []\n\n col  columns:\n    el = train_data[col].head(, compute=).iloc[]\n\n    molecule = Chem.MolFromSmiles(el)\n\n    img = Draw.MolToImage(molecule)\n\n    img.save()\n\n    images.append(Image.())\n    labels.append(col)\n\nfig, axs = plt.subplots(, (images), figsize=(, ))\n\n ax, img, label  (axs, images, labels):\n    ax.imshow(img)\n    ax.set_title(label)\n    ax.axis()\n\nplt.tight_layout()\nplt.savefig()\nplt.show()\n\n()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10578610%2Fb640c14b9bfb8dbb9ceee1d46f713ece%2FCaptura%20de%20ecra%202024-06-18%20as%2018.43.40.png?generation=1718733887741893&amp;alt=media\"></p>\n<p>Hey! As you can see above BB3 is the part below on the molecule, BB2 clearly the part on the left, but check BB1 it is a bit different from the molecule… they don't add up to the molecule SMILE. Is something wrong here? From this molecule the section that binds would be the section with triple connection?</p>\n<p>Thank you and sorry for the lack in chem knowledge.</p>",
      "rawMarkdown": "```python\n\n!pip install dask rdkit matplotlib\nimport dask.dataframe as dd\nfrom rdkit import Chem\nfrom rdkit.Chem import Draw\nimport matplotlib.pyplot as plt\nfrom PIL import Image\n\ntrain_data = dd.read_parquet('dataset/train.parquet', engine=\"pyarrow\", columns=[\"molecule_smiles\",\"buildingblock1_smiles\", \"buildingblock2_smiles\", \"buildingblock3_smiles\"])\n\ncolumns = [\"molecule_smiles\", \"buildingblock1_smiles\", \"buildingblock2_smiles\", \"buildingblock3_smiles\"]\n\nimages = []\nlabels = []\n\nfor col in columns:\n    el = train_data[col].head(1, compute=True).iloc[0]\n    \n    molecule = Chem.MolFromSmiles(el)\n\n    img = Draw.MolToImage(molecule)\n    \n    img.save(f'molecule_visualization_{col}.png')\n\n    images.append(Image.open(f'molecule_visualization_{col}.png'))\n    labels.append(col)\n\nfig, axs = plt.subplots(1, len(images), figsize=(15, 5))\n\nfor ax, img, label in zip(axs, images, labels):\n    ax.imshow(img)\n    ax.set_title(label)\n    ax.axis('off')\n\nplt.tight_layout()\nplt.savefig('combined_molecule_visualization.png')\nplt.show()\n\nprint(\"The combined molecule visualization has been saved as 'combined_molecule_visualization.png'.\")\n\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10578610%2Fb640c14b9bfb8dbb9ceee1d46f713ece%2FCaptura%20de%20ecra%202024-06-18%20as%2018.43.40.png?generation=1718733887741893&alt=media)\n\nHey! As you can see above BB3 is the part below on the molecule, BB2 clearly the part on the left, but check BB1 it is a bit different from the molecule... they don't add up to the molecule SMILE. Is something wrong here? From this molecule the section that binds would be the section with triple connection?\n\nThank you and sorry for the lack in chem knowledge.",
      "votes": null
    },
    {
      "id": "2878177",
      "postDate": "06/18/2024 19:09:48",
      "content": "<p>BB1 have what is known as a \"protecting group\" which prevents unwanted chemical reactions from occurring.  This particular group is call <a href=\"https://en.m.wikipedia.org/wiki/Fluorenylmethyloxycarbonyl_protecting_group\" target=\"_blank\">fmoc</a>.  Protecting group are removed prior to our as part of the reaction with the core pyrazine.</p>\n<p>All BB1 structures in the training set have this fmoc group, and I don't think other BBs have any protecting groups.  I didn't know about the test set. </p>\n<p>RDKit has functionality to remove those groups.</p>",
      "rawMarkdown": "BB1 have what is known as a \"protecting group\" which prevents unwanted chemical reactions from occurring.  This particular group is call [fmoc](https://en.m.wikipedia.org/wiki/Fluorenylmethyloxycarbonyl_protecting_group).  Protecting group are removed prior to our as part of the reaction with the core pyrazine.\n\nAll BB1 structures in the training set have this fmoc group, and I don't think other BBs have any protecting groups.  I didn't know about the test set. \n\nRDKit has functionality to remove those groups.",
      "votes": null
    },
    {
      "id": "2878806",
      "postDate": "06/19/2024 06:58:38",
      "content": "<blockquote>\n  <p>BB1 have what is known as a \"protecting group\" which prevents unwanted chemical reactions from occurring.  This particular group is call <a href=\"https://en.m.wikipedia.org/wiki/Fluorenylmethyloxycarbonyl_protecting_group\" target=\"_blank\">fmoc</a>.  Protecting group are removed prior to our as part of the reaction with the core pyrazine.</p>\n  <p>All BB1 structures in the training set have this fmoc group, and I don't think other BBs have any protecting groups.  I didn't know about the test set. </p>\n  <p>RDKit has functionality to remove those groups.</p>\n</blockquote>\n<p>Interesting! So the building block SMILES are represented before the reaction that creates the final SMILE. Notice that OH on BB1 is replaced by Dy an NH. The FMOC is also lost, so the molecule changes a lot to the final SMILE. There aren't any BB1 that appear in the BB2 or BB3 position this can explain why. </p>\n<p>I was trying to understand what parts of the molecule belong to which BB for better representation. </p>\n<ol>\n<li>Do you think that with triazine there is also a FMOC?</li>\n<li>Why does this only happen with BB1? </li>\n</ol>",
      "rawMarkdown": "> BB1 have what is known as a \"protecting group\" which prevents unwanted chemical reactions from occurring.  This particular group is call [fmoc](https://en.m.wikipedia.org/wiki/Fluorenylmethyloxycarbonyl_protecting_group).  Protecting group are removed prior to our as part of the reaction with the core pyrazine.\n> \n> All BB1 structures in the training set have this fmoc group, and I don't think other BBs have any protecting groups.  I didn't know about the test set. \n> \n> RDKit has functionality to remove those groups.\n\nInteresting! So the building block SMILES are represented before the reaction that creates the final SMILE. Notice that OH on BB1 is replaced by Dy an NH. The FMOC is also lost, so the molecule changes a lot to the final SMILE. There aren't any BB1 that appear in the BB2 or BB3 position this can explain why. \n\nI was trying to understand what parts of the molecule belong to which BB for better representation. \n\n1. Do you think that with triazine there is also a FMOC?\n2. Why does this only happen with BB1?",
      "votes": null
    },
    {
      "id": "2879586",
      "postDate": "06/19/2024 17:25:22",
      "content": "<p>The [Dy] is just a placeholder to show where the DNA attachment occurs.  In other discussion threads people have suggested replacing it with a methyl (just a single carbon in SMILES notation).  </p>\n<ol>\n<li>I'm not sure what the reactions are the connect to the triazine.  There may be protecting groups on it that are removed one-by-one in order to direct where the different BBs go, but I'm not sure about that.</li>\n<li>I'm also not sure why only BB1 has a protecting group added.  It may have something to do with the synthesis pathway - maybe it is the only one that needs a protecting group given how the final compounds are constructed.  Regardless, each BB group appears to be very consistent.  BB1 has the same protecting group on each BB, while BB2 and BB3 appear to have the same amine terminal group.  With that consistency, those groups will not affect any modeling activity since they are constant.  HUGE thanks to the organizers for making a clean dataset for us!!</li>\n</ol>",
      "rawMarkdown": "The [Dy] is just a placeholder to show where the DNA attachment occurs.  In other discussion threads people have suggested replacing it with a methyl (just a single carbon in SMILES notation).  \n\n1.  I'm not sure what the reactions are the connect to the triazine.  There may be protecting groups on it that are removed one-by-one in order to direct where the different BBs go, but I'm not sure about that.\n2.  I'm also not sure why only BB1 has a protecting group added.  It may have something to do with the synthesis pathway - maybe it is the only one that needs a protecting group given how the final compounds are constructed.  Regardless, each BB group appears to be very consistent.  BB1 has the same protecting group on each BB, while BB2 and BB3 appear to have the same amine terminal group.  With that consistency, those groups will not affect any modeling activity since they are constant.  HUGE thanks to the organizers for making a clean dataset for us!!",
      "votes": null
    },
    {
      "id": "2889751",
      "postDate": "06/25/2024 16:46:29",
      "content": "<p>thank you for your sharing</p>",
      "rawMarkdown": "thank you for your sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2878177,
      "author_name": "kirkdco",
      "author_url": "",
      "post_date": "06/18/2024 19:09:48",
      "content": "<p>BB1 have what is known as a \"protecting group\" which prevents unwanted chemical reactions from occurring.  This particular group is call <a href=\"https://en.m.wikipedia.org/wiki/Fluorenylmethyloxycarbonyl_protecting_group\" target=\"_blank\">fmoc</a>.  Protecting group are removed prior to our as part of the reaction with the core pyrazine.</p>\n<p>All BB1 structures in the training set have this fmoc group, and I don't think other BBs have any protecting groups.  I didn't know about the test set. </p>\n<p>RDKit has functionality to remove those groups.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2878806,
          "author_name": "manuelsokolov",
          "author_url": "",
          "post_date": "06/19/2024 06:58:38",
          "content": "<blockquote>\n  <p>BB1 have what is known as a \"protecting group\" which prevents unwanted chemical reactions from occurring.  This particular group is call <a href=\"https://en.m.wikipedia.org/wiki/Fluorenylmethyloxycarbonyl_protecting_group\" target=\"_blank\">fmoc</a>.  Protecting group are removed prior to our as part of the reaction with the core pyrazine.</p>\n  <p>All BB1 structures in the training set have this fmoc group, and I don't think other BBs have any protecting groups.  I didn't know about the test set. </p>\n  <p>RDKit has functionality to remove those groups.</p>\n</blockquote>\n<p>Interesting! So the building block SMILES are represented before the reaction that creates the final SMILE. Notice that OH on BB1 is replaced by Dy an NH. The FMOC is also lost, so the molecule changes a lot to the final SMILE. There aren't any BB1 that appear in the BB2 or BB3 position this can explain why. </p>\n<p>I was trying to understand what parts of the molecule belong to which BB for better representation. </p>\n<ol>\n<li>Do you think that with triazine there is also a FMOC?</li>\n<li>Why does this only happen with BB1? </li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 2879586,
              "author_name": "kirkdco",
              "author_url": "",
              "post_date": "06/19/2024 17:25:22",
              "content": "<p>The [Dy] is just a placeholder to show where the DNA attachment occurs.  In other discussion threads people have suggested replacing it with a methyl (just a single carbon in SMILES notation).  </p>\n<ol>\n<li>I'm not sure what the reactions are the connect to the triazine.  There may be protecting groups on it that are removed one-by-one in order to direct where the different BBs go, but I'm not sure about that.</li>\n<li>I'm also not sure why only BB1 has a protecting group added.  It may have something to do with the synthesis pathway - maybe it is the only one that needs a protecting group given how the final compounds are constructed.  Regardless, each BB group appears to be very consistent.  BB1 has the same protecting group on each BB, while BB2 and BB3 appear to have the same amine terminal group.  With that consistency, those groups will not affect any modeling activity since they are constant.  HUGE thanks to the organizers for making a clean dataset for us!!</li>\n</ol>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2889751,
      "author_name": "traditionistjj",
      "author_url": "",
      "post_date": "06/25/2024 16:46:29",
      "content": "<p>thank you for your sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2878081": "```python\n\n!pip install dask rdkit matplotlib\nimport dask.dataframe as dd\nfrom rdkit import Chem\nfrom rdkit.Chem import Draw\nimport matplotlib.pyplot as plt\nfrom PIL import Image\n\ntrain_data = dd.read_parquet('dataset/train.parquet', engine=\"pyarrow\", columns=[\"molecule_smiles\",\"buildingblock1_smiles\", \"buildingblock2_smiles\", \"buildingblock3_smiles\"])\n\ncolumns = [\"molecule_smiles\", \"buildingblock1_smiles\", \"buildingblock2_smiles\", \"buildingblock3_smiles\"]\n\nimages = []\nlabels = []\n\nfor col in columns:\n    el = train_data[col].head(1, compute=True).iloc[0]\n    \n    molecule = Chem.MolFromSmiles(el)\n\n    img = Draw.MolToImage(molecule)\n    \n    img.save(f'molecule_visualization_{col}.png')\n\n    images.append(Image.open(f'molecule_visualization_{col}.png'))\n    labels.append(col)\n\nfig, axs = plt.subplots(1, len(images), figsize=(15, 5))\n\nfor ax, img, label in zip(axs, images, labels):\n    ax.imshow(img)\n    ax.set_title(label)\n    ax.axis('off')\n\nplt.tight_layout()\nplt.savefig('combined_molecule_visualization.png')\nplt.show()\n\nprint(\"The combined molecule visualization has been saved as 'combined_molecule_visualization.png'.\")\n\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10578610%2Fb640c14b9bfb8dbb9ceee1d46f713ece%2FCaptura%20de%20ecra%202024-06-18%20as%2018.43.40.png?generation=1718733887741893&alt=media)\n\nHey! As you can see above BB3 is the part below on the molecule, BB2 clearly the part on the left, but check BB1 it is a bit different from the molecule... they don't add up to the molecule SMILE. Is something wrong here? From this molecule the section that binds would be the section with triple connection?\n\nThank you and sorry for the lack in chem knowledge.",
    "2878177": "BB1 have what is known as a \"protecting group\" which prevents unwanted chemical reactions from occurring.  This particular group is call [fmoc](https://en.m.wikipedia.org/wiki/Fluorenylmethyloxycarbonyl_protecting_group).  Protecting group are removed prior to our as part of the reaction with the core pyrazine.\n\nAll BB1 structures in the training set have this fmoc group, and I don't think other BBs have any protecting groups.  I didn't know about the test set. \n\nRDKit has functionality to remove those groups.",
    "2878806": "> BB1 have what is known as a \"protecting group\" which prevents unwanted chemical reactions from occurring.  This particular group is call [fmoc](https://en.m.wikipedia.org/wiki/Fluorenylmethyloxycarbonyl_protecting_group).  Protecting group are removed prior to our as part of the reaction with the core pyrazine.\n> \n> All BB1 structures in the training set have this fmoc group, and I don't think other BBs have any protecting groups.  I didn't know about the test set. \n> \n> RDKit has functionality to remove those groups.\n\nInteresting! So the building block SMILES are represented before the reaction that creates the final SMILE. Notice that OH on BB1 is replaced by Dy an NH. The FMOC is also lost, so the molecule changes a lot to the final SMILE. There aren't any BB1 that appear in the BB2 or BB3 position this can explain why. \n\nI was trying to understand what parts of the molecule belong to which BB for better representation. \n\n1. Do you think that with triazine there is also a FMOC?\n2. Why does this only happen with BB1?",
    "2879586": "The [Dy] is just a placeholder to show where the DNA attachment occurs.  In other discussion threads people have suggested replacing it with a methyl (just a single carbon in SMILES notation).  \n\n1.  I'm not sure what the reactions are the connect to the triazine.  There may be protecting groups on it that are removed one-by-one in order to direct where the different BBs go, but I'm not sure about that.\n2.  I'm also not sure why only BB1 has a protecting group added.  It may have something to do with the synthesis pathway - maybe it is the only one that needs a protecting group given how the final compounds are constructed.  Regardless, each BB group appears to be very consistent.  BB1 has the same protecting group on each BB, while BB2 and BB3 appear to have the same amine terminal group.  With that consistency, those groups will not affect any modeling activity since they are constant.  HUGE thanks to the organizers for making a clean dataset for us!!",
    "2889751": "thank you for your sharing"
  },
  "source": "meta"
}