{
  "id": 493294,
  "title": "Scaffold Hopping",
  "url": "/competitions/leash-BELKA/discussion/493294",
  "author_name": "",
  "post_date": "2024-04-12T20:01:36.938457500Z",
  "votes": 48,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi All,</p>\n<p>I wanted to introduce the topic of scaffold hopping for folks that aren't aware of what it is. Chemists working on developing drugs often have a \"core\" of their molecule that remains relatively constant and they tend to change the things attached to that core. However, sometimes this core has some fundamental flaws and the team will wish to find a new core. The search for alternative cores is called \"scaffold hopping\" and it's common that models trained with one \"chemical series\" (i.e. molecules sharing a core) will not translate well to new cores.</p>\n<p>I've published a notebook that clearly shows one of the main goals of this competition is to build a model based on chemical representations that allow for SAR translation between cores the model has never seen, as our train set is 100% based on triazine cores and the test set is about 50:50 triazines and others. It's going to be really interesting to see how folks approach this and I hope this helps clarify the task at hand :).</p>\n<p>Feel free to have a look for yourself: <a href=\"https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration\" target=\"_blank\">https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration</a></p>",
  "messages": [
    {
      "id": "2749067",
      "postDate": "04/12/2024 20:01:36",
      "content": "<p>Hi All,</p>\n<p>I wanted to introduce the topic of scaffold hopping for folks that aren't aware of what it is. Chemists working on developing drugs often have a \"core\" of their molecule that remains relatively constant and they tend to change the things attached to that core. However, sometimes this core has some fundamental flaws and the team will wish to find a new core. The search for alternative cores is called \"scaffold hopping\" and it's common that models trained with one \"chemical series\" (i.e. molecules sharing a core) will not translate well to new cores.</p>\n<p>I've published a notebook that clearly shows one of the main goals of this competition is to build a model based on chemical representations that allow for SAR translation between cores the model has never seen, as our train set is 100% based on triazine cores and the test set is about 50:50 triazines and others. It's going to be really interesting to see how folks approach this and I hope this helps clarify the task at hand :).</p>\n<p>Feel free to have a look for yourself: <a href=\"https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration\" target=\"_blank\">https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration</a></p>",
      "rawMarkdown": "Hi All,\n\nI wanted to introduce the topic of scaffold hopping for folks that aren't aware of what it is. Chemists working on developing drugs often have a \"core\" of their molecule that remains relatively constant and they tend to change the things attached to that core. However, sometimes this core has some fundamental flaws and the team will wish to find a new core. The search for alternative cores is called \"scaffold hopping\" and it's common that models trained with one \"chemical series\" (i.e. molecules sharing a core) will not translate well to new cores.\n\nI've published a notebook that clearly shows one of the main goals of this competition is to build a model based on chemical representations that allow for SAR translation between cores the model has never seen, as our train set is 100% based on triazine cores and the test set is about 50:50 triazines and others. It's going to be really interesting to see how folks approach this and I hope this helps clarify the task at hand :).\n\nFeel free to have a look for yourself: https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration",
      "votes": null
    },
    {
      "id": "2751030",
      "postDate": "04/14/2024 01:46:33",
      "content": "<p>Some added data on the non-triazines:</p>\n<p>Even though non-triazines are 55% of the test set molecules, they only make up 30% of the test set rows, because all of the triazine molecules are in the test set against all 3 proteins, and the vast majority of the non-triazines you are only asked to predict against 1 protein.</p>\n<p>There's 36 by 377 by 446 unique building blocks in the non-triazine subset of test. I found that none of those 36, 377, 446 are found in train, and no overlap between the 377 with the 446 either. There's about 1/12th of the max possible permutations of ABC. Most individual molecules are only predicted against one protein, so then about 1/36th of the max possible permutations of all possible ABC:P for a single protein target.</p>\n<pre><code>(, test[test.triazine == ].shape[])\n\nnew_core = test[test.triazine == ]\n\n(, new_core.groupby().size().describe())\n\n(, new_core.groupby().size().describe())\n(, new_core.groupby().size().describe())\n\n()\n(new_core[[, , ]].())\n</code></pre>\n<pre><code>Non-Triazine molecules (test): \n\nbb1\n count       \nmean     \nstd        \n      \n%      \n%      \n%      \n      \ndtype: float64\n\nbb2\n count     \nmean     \nstd        \n      \n%      \n%      \n%      \n      \ndtype: float64\n\nbb3\n count     \nmean     \nstd        \n      \n%      \n%      \n%      \n      \ndtype: float64\n\nis_BRD4    \nis_HSA     \nis_sEH     \n</code></pre>",
      "rawMarkdown": "Some added data on the non-triazines:\n\nEven though non-triazines are 55% of the test set molecules, they only make up 30% of the test set rows, because all of the triazine molecules are in the test set against all 3 proteins, and the vast majority of the non-triazines you are only asked to predict against 1 protein.\n\nThere's 36 by 377 by 446 unique building blocks in the non-triazine subset of test. I found that none of those 36, 377, 446 are found in train, and no overlap between the 377 with the 446 either. There's about 1/12th of the max possible permutations of ABC. Most individual molecules are only predicted against one protein, so then about 1/36th of the max possible permutations of all possible ABC:P for a single protein target.\n\n```python\nprint(\"Non-Triazine molecules (test):\", test[test.triazine == False].shape[0])\n\nnew_core = test[test.triazine == False]\n\nprint(\"\\nbb1\\n\", new_core.groupby('buildingblock1_smiles').size().describe())\n# print(\"\\nbb1\\n\", new_core.groupby('buildingblock1_smiles')[['is_BRD4', 'is_HSA', 'is_sEH']].sum().sum(axis=1).describe())\nprint(\"\\nbb2\\n\", new_core.groupby('buildingblock2_smiles').size().describe())\nprint(\"\\nbb3\\n\", new_core.groupby('buildingblock3_smiles').size().describe())\n\nprint()\nprint(new_core[['is_BRD4', 'is_HSA', 'is_sEH']].sum())\n\n```\n\n```python\nNon-Triazine molecules (test): 486390\n\nbb1\n count       36.000000\nmean     13510.833333\nstd        109.154543\nmin      13313.000000\n25%      13429.500000\n50%      13500.000000\n75%      13614.500000\nmax      13708.000000\ndtype: float64\n\nbb2\n count     377.000000\nmean     1290.159151\nstd        36.462472\nmin      1183.000000\n25%      1267.000000\n50%      1292.000000\n75%      1314.000000\nmax      1386.000000\ndtype: float64\n\nbb3\n count     446.000000\nmean     1090.560538\nstd        30.478059\nmin      1006.000000\n25%      1070.250000\n50%      1091.000000\n75%      1110.750000\nmax      1185.000000\ndtype: float64\n\nis_BRD4    167227\nis_HSA     166263\nis_sEH     166510\n```",
      "votes": null
    },
    {
      "id": "2751057",
      "postDate": "04/14/2024 02:22:06",
      "content": "<p>Nice observations <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> </p>\n<p>Good point on the # of rows, I was a little focused on the molecules XD.</p>\n<p>It's going to be interesting to see the creative strategies folks use to tackle this problem. </p>",
      "rawMarkdown": "Nice observations @roberthatch \n\nGood point on the # of rows, I was a little focused on the molecules XD.\n\nIt's going to be interesting to see the creative strategies folks use to tackle this problem.",
      "votes": null
    },
    {
      "id": "2756366",
      "postDate": "04/17/2024 02:26:08",
      "content": "<p>So for much of the train dataset, I can pretty easily use rdkit to find building blocks in the molecule from molecule_smiles column, with maybe a couple small corner cases.</p>\n<p>However, for the test set, specifically the non-triazines, I'm wondering how much change is expected between building blocks and the molecule? Is there a logical explanation for the deltas, and do these particular deltas tell a biochemist anything important?</p>\n<p>Here's a couple consecutive case studies: (if you search test df on BB2 string then it's the first two results).</p>\n<pre><code>BB1: CC(C)(C)OC(=O)N1CC(c2ccccc2Br)=C[C@H]1C(=O)O\nBB2: C/C(=C\\B1OC(C)(C)C(C)(C)O1)c1ccnn1C\nBB3: CC(C)(C)c1ccnc2c(C(=O)O)cnn12\nMol: CC(=Cc1ccccc1C1=C[C@@H](C(=O)N[Dy])N(C(=O)c2cnn3c(C(C)(C)C)ccnc23)C1)c1ccnn1C\n</code></pre>\n<p>xxx</p>\n<pre><code>BB1: CC(C)(C)OC(=O)N1CC(c2ccccc2Br)=C[C@H]1C(=O)O\nBB2: C/C(=C\\B1OC(C)(C)C(C)(C)O1)c1ccnn1C\nBB3: CC(C)=CC(=O)O\nMol: CC(C)=CC(=O)N1CC(c2ccccc2C=C(C)c2ccnn2C)=C[C@H]1C(=O)N[Dy]\n</code></pre>\n<p>So if I match up the lower one, and eliminate really obvious visual matches:<br>\nBB1: \"CC(C)(C)OC(=O)\" is just gone? Did \"Br\" become \"C\"?<br>\nBB2: All this \"C/C(=C\\B1OC(C)(C)C(C)(C)O1)\" just became \"=C(C)\"?</p>\n<p>I took some time to understand the SMILES format, and it was starting to all make sense… but ~1/3rd of all the atoms in the BBs vanishing in the molecule_smiles is a bit much for me.</p>",
      "rawMarkdown": "So for much of the train dataset, I can pretty easily use rdkit to find building blocks in the molecule from molecule_smiles column, with maybe a couple small corner cases.\n\nHowever, for the test set, specifically the non-triazines, I'm wondering how much change is expected between building blocks and the molecule? Is there a logical explanation for the deltas, and do these particular deltas tell a biochemist anything important?\n\nHere's a couple consecutive case studies: (if you search test df on BB2 string then it's the first two results).\n```python\nBB1: CC(C)(C)OC(=O)N1CC(c2ccccc2Br)=C[C@H]1C(=O)O\nBB2: C/C(=C\\B1OC(C)(C)C(C)(C)O1)c1ccnn1C\nBB3: CC(C)(C)c1ccnc2c(C(=O)O)cnn12\nMol: CC(=Cc1ccccc1C1=C[C@@H](C(=O)N[Dy])N(C(=O)c2cnn3c(C(C)(C)C)ccnc23)C1)c1ccnn1C\n\n```\nxxx\n\n```python\nBB1: CC(C)(C)OC(=O)N1CC(c2ccccc2Br)=C[C@H]1C(=O)O\nBB2: C/C(=C\\B1OC(C)(C)C(C)(C)O1)c1ccnn1C\nBB3: CC(C)=CC(=O)O\nMol: CC(C)=CC(=O)N1CC(c2ccccc2C=C(C)c2ccnn2C)=C[C@H]1C(=O)N[Dy]\n\n```\nSo if I match up the lower one, and eliminate really obvious visual matches:\nBB1: \"CC(C)(C)OC(=O)\" is just gone? Did \"Br\" become \"C\"?\nBB2: All this \"C/C(=C\\B1OC(C)(C)C(C)(C)O1)\" just became \"=C(C)\"?\n\nI took some time to understand the SMILES format, and it was starting to all make sense... but ~1/3rd of all the atoms in the BBs vanishing in the molecule_smiles is a bit much for me.",
      "votes": null
    },
    {
      "id": "2756422",
      "postDate": "04/17/2024 03:14:36",
      "content": "<p>Great question! I don't have access to my PC at the moment, just my phone, so I'll come back and double check this answer later to make sure I answered what you were actually asking.</p>\n<p>You are correct in identifying that not all of the components of the building blocks are present in the final molecule. This is because some of those bits are necessary to form the new bond you want, but get discarded as a byproduct in the process (often acting as the thermodynamic driving force of the net reaction).</p>\n<p>The reactions in the train set are quite simple. An amine reacts with a chloro triazine and to balance the equation all you lose is HCl. However, many reactions are much more complicated than this. If I am reading your SMILES right, I think you identified a Suzuki reaction (2010 Nobel prize) where the boron nucleophile is a pinacolboronate.</p>\n<p>Ideally, nothing about the building block itself would matter for binding as the byproducts are discarded along the synthesis. Only the structure of the final product would matter for a purified compound. Where the building block information might potentially have value in a DEL context is when the reaction didn't actually form the product you suspect it did. </p>\n<p>Just to be clear: you can make the exact same molecule many different ways from many different building blocks generating different amounts of byproducts along the way. However, in an ideal case where you have 100% pure material, it wouldn't matter how you made it. The final molecule made by each route would behave identically.</p>",
      "rawMarkdown": "Great question! I don't have access to my PC at the moment, just my phone, so I'll come back and double check this answer later to make sure I answered what you were actually asking.\n\nYou are correct in identifying that not all of the components of the building blocks are present in the final molecule. This is because some of those bits are necessary to form the new bond you want, but get discarded as a byproduct in the process (often acting as the thermodynamic driving force of the net reaction).\n\nThe reactions in the train set are quite simple. An amine reacts with a chloro triazine and to balance the equation all you lose is HCl. However, many reactions are much more complicated than this. If I am reading your SMILES right, I think you identified a Suzuki reaction (2010 Nobel prize) where the boron nucleophile is a pinacolboronate.\n\nIdeally, nothing about the building block itself would matter for binding as the byproducts are discarded along the synthesis. Only the structure of the final product would matter for a purified compound. Where the building block information might potentially have value in a DEL context is when the reaction didn't actually form the product you suspect it did. \n\nJust to be clear: you can make the exact same molecule many different ways from many different building blocks generating different amounts of byproducts along the way. However, in an ideal case where you have 100% pure material, it wouldn't matter how you made it. The final molecule made by each route would behave identically.",
      "votes": null
    },
    {
      "id": "2757587",
      "postDate": "04/17/2024 15:46:33",
      "content": "<p>Yeah, it's really just for (train) efficiency (~1000 building blocks vs ~100000000 full mol smiles). But if going down that path and trading off some accuracy for efficiency, then it becomes important to be able to apply the same methods to test dataset (even though efficiency on test set is not so crucial).</p>\n<p>Maybe I can still go down this route if the change (reaction) for a given building block is consistent everywhere it's used. If so I can investigate the few hundred building blocks in the non-triazine set, learn and pre-apply the change to the building block and still use any \"building block\"-based features I end up having in my final model(s).</p>\n<p>Thanks as always for the insight! It's good to at least partially understand what's going on with the data. :)</p>",
      "rawMarkdown": "Yeah, it's really just for (train) efficiency (~1000 building blocks vs ~100000000 full mol smiles). But if going down that path and trading off some accuracy for efficiency, then it becomes important to be able to apply the same methods to test dataset (even though efficiency on test set is not so crucial).\n\nMaybe I can still go down this route if the change (reaction) for a given building block is consistent everywhere it's used. If so I can investigate the few hundred building blocks in the non-triazine set, learn and pre-apply the change to the building block and still use any \"building block\"-based features I end up having in my final model(s).\n\nThanks as always for the insight! It's good to at least partially understand what's going on with the data. :)",
      "votes": null
    },
    {
      "id": "2757932",
      "postDate": "04/17/2024 20:02:13",
      "content": "<p>I'm not quite following. Are you training on building blocks?</p>",
      "rawMarkdown": "I'm not quite following. Are you training on building blocks?",
      "votes": null
    },
    {
      "id": "2757960",
      "postDate": "04/17/2024 20:29:30",
      "content": "<p>I'm considering doing some (of the more time consuming) feature generation on building blocks, then lookup and pull in those building block features for each row of train. </p>\n<p>Clearly it would be better to do on full molecule, but given time and compute constraints using BB features might be better than nothing. </p>",
      "rawMarkdown": "I'm considering doing some (of the more time consuming) feature generation on building blocks, then lookup and pull in those building block features for each row of train. \n\nClearly it would be better to do on full molecule, but given time and compute constraints using BB features might be better than nothing.",
      "votes": null
    },
    {
      "id": "2758067",
      "postDate": "04/17/2024 23:16:09",
      "content": "<p>Ah, I see. Good luck :)</p>",
      "rawMarkdown": "Ah, I see. Good luck :)",
      "votes": null
    },
    {
      "id": "2758832",
      "postDate": "04/18/2024 11:17:43",
      "content": "<p>Good point !!</p>",
      "rawMarkdown": "Good point !!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2751030,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "04/14/2024 01:46:33",
      "content": "<p>Some added data on the non-triazines:</p>\n<p>Even though non-triazines are 55% of the test set molecules, they only make up 30% of the test set rows, because all of the triazine molecules are in the test set against all 3 proteins, and the vast majority of the non-triazines you are only asked to predict against 1 protein.</p>\n<p>There's 36 by 377 by 446 unique building blocks in the non-triazine subset of test. I found that none of those 36, 377, 446 are found in train, and no overlap between the 377 with the 446 either. There's about 1/12th of the max possible permutations of ABC. Most individual molecules are only predicted against one protein, so then about 1/36th of the max possible permutations of all possible ABC:P for a single protein target.</p>\n<pre><code>(, test[test.triazine == ].shape[])\n\nnew_core = test[test.triazine == ]\n\n(, new_core.groupby().size().describe())\n\n(, new_core.groupby().size().describe())\n(, new_core.groupby().size().describe())\n\n()\n(new_core[[, , ]].())\n</code></pre>\n<pre><code>Non-Triazine molecules (test): \n\nbb1\n count       \nmean     \nstd        \n      \n%      \n%      \n%      \n      \ndtype: float64\n\nbb2\n count     \nmean     \nstd        \n      \n%      \n%      \n%      \n      \ndtype: float64\n\nbb3\n count     \nmean     \nstd        \n      \n%      \n%      \n%      \n      \ndtype: float64\n\nis_BRD4    \nis_HSA     \nis_sEH     \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2751057,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "04/14/2024 02:22:06",
          "content": "<p>Nice observations <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> </p>\n<p>Good point on the # of rows, I was a little focused on the molecules XD.</p>\n<p>It's going to be interesting to see the creative strategies folks use to tackle this problem. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2756366,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "04/17/2024 02:26:08",
      "content": "<p>So for much of the train dataset, I can pretty easily use rdkit to find building blocks in the molecule from molecule_smiles column, with maybe a couple small corner cases.</p>\n<p>However, for the test set, specifically the non-triazines, I'm wondering how much change is expected between building blocks and the molecule? Is there a logical explanation for the deltas, and do these particular deltas tell a biochemist anything important?</p>\n<p>Here's a couple consecutive case studies: (if you search test df on BB2 string then it's the first two results).</p>\n<pre><code>BB1: CC(C)(C)OC(=O)N1CC(c2ccccc2Br)=C[C@H]1C(=O)O\nBB2: C/C(=C\\B1OC(C)(C)C(C)(C)O1)c1ccnn1C\nBB3: CC(C)(C)c1ccnc2c(C(=O)O)cnn12\nMol: CC(=Cc1ccccc1C1=C[C@@H](C(=O)N[Dy])N(C(=O)c2cnn3c(C(C)(C)C)ccnc23)C1)c1ccnn1C\n</code></pre>\n<p>xxx</p>\n<pre><code>BB1: CC(C)(C)OC(=O)N1CC(c2ccccc2Br)=C[C@H]1C(=O)O\nBB2: C/C(=C\\B1OC(C)(C)C(C)(C)O1)c1ccnn1C\nBB3: CC(C)=CC(=O)O\nMol: CC(C)=CC(=O)N1CC(c2ccccc2C=C(C)c2ccnn2C)=C[C@H]1C(=O)N[Dy]\n</code></pre>\n<p>So if I match up the lower one, and eliminate really obvious visual matches:<br>\nBB1: \"CC(C)(C)OC(=O)\" is just gone? Did \"Br\" become \"C\"?<br>\nBB2: All this \"C/C(=C\\B1OC(C)(C)C(C)(C)O1)\" just became \"=C(C)\"?</p>\n<p>I took some time to understand the SMILES format, and it was starting to all make sense… but ~1/3rd of all the atoms in the BBs vanishing in the molecule_smiles is a bit much for me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2756422,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "04/17/2024 03:14:36",
          "content": "<p>Great question! I don't have access to my PC at the moment, just my phone, so I'll come back and double check this answer later to make sure I answered what you were actually asking.</p>\n<p>You are correct in identifying that not all of the components of the building blocks are present in the final molecule. This is because some of those bits are necessary to form the new bond you want, but get discarded as a byproduct in the process (often acting as the thermodynamic driving force of the net reaction).</p>\n<p>The reactions in the train set are quite simple. An amine reacts with a chloro triazine and to balance the equation all you lose is HCl. However, many reactions are much more complicated than this. If I am reading your SMILES right, I think you identified a Suzuki reaction (2010 Nobel prize) where the boron nucleophile is a pinacolboronate.</p>\n<p>Ideally, nothing about the building block itself would matter for binding as the byproducts are discarded along the synthesis. Only the structure of the final product would matter for a purified compound. Where the building block information might potentially have value in a DEL context is when the reaction didn't actually form the product you suspect it did. </p>\n<p>Just to be clear: you can make the exact same molecule many different ways from many different building blocks generating different amounts of byproducts along the way. However, in an ideal case where you have 100% pure material, it wouldn't matter how you made it. The final molecule made by each route would behave identically.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2757587,
              "author_name": "roberthatch",
              "author_url": "",
              "post_date": "04/17/2024 15:46:33",
              "content": "<p>Yeah, it's really just for (train) efficiency (~1000 building blocks vs ~100000000 full mol smiles). But if going down that path and trading off some accuracy for efficiency, then it becomes important to be able to apply the same methods to test dataset (even though efficiency on test set is not so crucial).</p>\n<p>Maybe I can still go down this route if the change (reaction) for a given building block is consistent everywhere it's used. If so I can investigate the few hundred building blocks in the non-triazine set, learn and pre-apply the change to the building block and still use any \"building block\"-based features I end up having in my final model(s).</p>\n<p>Thanks as always for the insight! It's good to at least partially understand what's going on with the data. :)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2757932,
                  "author_name": "chemdatafarmer",
                  "author_url": "",
                  "post_date": "04/17/2024 20:02:13",
                  "content": "<p>I'm not quite following. Are you training on building blocks?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2757960,
                      "author_name": "roberthatch",
                      "author_url": "",
                      "post_date": "04/17/2024 20:29:30",
                      "content": "<p>I'm considering doing some (of the more time consuming) feature generation on building blocks, then lookup and pull in those building block features for each row of train. </p>\n<p>Clearly it would be better to do on full molecule, but given time and compute constraints using BB features might be better than nothing. </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2758067,
                          "author_name": "chemdatafarmer",
                          "author_url": "",
                          "post_date": "04/17/2024 23:16:09",
                          "content": "<p>Ah, I see. Good luck :)</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2758832,
      "author_name": "sonialikhan",
      "author_url": "",
      "post_date": "04/18/2024 11:17:43",
      "content": "<p>Good point !!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2749067": "Hi All,\n\nI wanted to introduce the topic of scaffold hopping for folks that aren't aware of what it is. Chemists working on developing drugs often have a \"core\" of their molecule that remains relatively constant and they tend to change the things attached to that core. However, sometimes this core has some fundamental flaws and the team will wish to find a new core. The search for alternative cores is called \"scaffold hopping\" and it's common that models trained with one \"chemical series\" (i.e. molecules sharing a core) will not translate well to new cores.\n\nI've published a notebook that clearly shows one of the main goals of this competition is to build a model based on chemical representations that allow for SAR translation between cores the model has never seen, as our train set is 100% based on triazine cores and the test set is about 50:50 triazines and others. It's going to be really interesting to see how folks approach this and I hope this helps clarify the task at hand :).\n\nFeel free to have a look for yourself: https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration",
    "2751030": "Some added data on the non-triazines:\n\nEven though non-triazines are 55% of the test set molecules, they only make up 30% of the test set rows, because all of the triazine molecules are in the test set against all 3 proteins, and the vast majority of the non-triazines you are only asked to predict against 1 protein.\n\nThere's 36 by 377 by 446 unique building blocks in the non-triazine subset of test. I found that none of those 36, 377, 446 are found in train, and no overlap between the 377 with the 446 either. There's about 1/12th of the max possible permutations of ABC. Most individual molecules are only predicted against one protein, so then about 1/36th of the max possible permutations of all possible ABC:P for a single protein target.\n\n```python\nprint(\"Non-Triazine molecules (test):\", test[test.triazine == False].shape[0])\n\nnew_core = test[test.triazine == False]\n\nprint(\"\\nbb1\\n\", new_core.groupby('buildingblock1_smiles').size().describe())\n# print(\"\\nbb1\\n\", new_core.groupby('buildingblock1_smiles')[['is_BRD4', 'is_HSA', 'is_sEH']].sum().sum(axis=1).describe())\nprint(\"\\nbb2\\n\", new_core.groupby('buildingblock2_smiles').size().describe())\nprint(\"\\nbb3\\n\", new_core.groupby('buildingblock3_smiles').size().describe())\n\nprint()\nprint(new_core[['is_BRD4', 'is_HSA', 'is_sEH']].sum())\n\n```\n\n```python\nNon-Triazine molecules (test): 486390\n\nbb1\n count       36.000000\nmean     13510.833333\nstd        109.154543\nmin      13313.000000\n25%      13429.500000\n50%      13500.000000\n75%      13614.500000\nmax      13708.000000\ndtype: float64\n\nbb2\n count     377.000000\nmean     1290.159151\nstd        36.462472\nmin      1183.000000\n25%      1267.000000\n50%      1292.000000\n75%      1314.000000\nmax      1386.000000\ndtype: float64\n\nbb3\n count     446.000000\nmean     1090.560538\nstd        30.478059\nmin      1006.000000\n25%      1070.250000\n50%      1091.000000\n75%      1110.750000\nmax      1185.000000\ndtype: float64\n\nis_BRD4    167227\nis_HSA     166263\nis_sEH     166510\n```",
    "2751057": "Nice observations @roberthatch \n\nGood point on the # of rows, I was a little focused on the molecules XD.\n\nIt's going to be interesting to see the creative strategies folks use to tackle this problem.",
    "2756366": "So for much of the train dataset, I can pretty easily use rdkit to find building blocks in the molecule from molecule_smiles column, with maybe a couple small corner cases.\n\nHowever, for the test set, specifically the non-triazines, I'm wondering how much change is expected between building blocks and the molecule? Is there a logical explanation for the deltas, and do these particular deltas tell a biochemist anything important?\n\nHere's a couple consecutive case studies: (if you search test df on BB2 string then it's the first two results).\n```python\nBB1: CC(C)(C)OC(=O)N1CC(c2ccccc2Br)=C[C@H]1C(=O)O\nBB2: C/C(=C\\B1OC(C)(C)C(C)(C)O1)c1ccnn1C\nBB3: CC(C)(C)c1ccnc2c(C(=O)O)cnn12\nMol: CC(=Cc1ccccc1C1=C[C@@H](C(=O)N[Dy])N(C(=O)c2cnn3c(C(C)(C)C)ccnc23)C1)c1ccnn1C\n\n```\nxxx\n\n```python\nBB1: CC(C)(C)OC(=O)N1CC(c2ccccc2Br)=C[C@H]1C(=O)O\nBB2: C/C(=C\\B1OC(C)(C)C(C)(C)O1)c1ccnn1C\nBB3: CC(C)=CC(=O)O\nMol: CC(C)=CC(=O)N1CC(c2ccccc2C=C(C)c2ccnn2C)=C[C@H]1C(=O)N[Dy]\n\n```\nSo if I match up the lower one, and eliminate really obvious visual matches:\nBB1: \"CC(C)(C)OC(=O)\" is just gone? Did \"Br\" become \"C\"?\nBB2: All this \"C/C(=C\\B1OC(C)(C)C(C)(C)O1)\" just became \"=C(C)\"?\n\nI took some time to understand the SMILES format, and it was starting to all make sense... but ~1/3rd of all the atoms in the BBs vanishing in the molecule_smiles is a bit much for me.",
    "2756422": "Great question! I don't have access to my PC at the moment, just my phone, so I'll come back and double check this answer later to make sure I answered what you were actually asking.\n\nYou are correct in identifying that not all of the components of the building blocks are present in the final molecule. This is because some of those bits are necessary to form the new bond you want, but get discarded as a byproduct in the process (often acting as the thermodynamic driving force of the net reaction).\n\nThe reactions in the train set are quite simple. An amine reacts with a chloro triazine and to balance the equation all you lose is HCl. However, many reactions are much more complicated than this. If I am reading your SMILES right, I think you identified a Suzuki reaction (2010 Nobel prize) where the boron nucleophile is a pinacolboronate.\n\nIdeally, nothing about the building block itself would matter for binding as the byproducts are discarded along the synthesis. Only the structure of the final product would matter for a purified compound. Where the building block information might potentially have value in a DEL context is when the reaction didn't actually form the product you suspect it did. \n\nJust to be clear: you can make the exact same molecule many different ways from many different building blocks generating different amounts of byproducts along the way. However, in an ideal case where you have 100% pure material, it wouldn't matter how you made it. The final molecule made by each route would behave identically.",
    "2757587": "Yeah, it's really just for (train) efficiency (~1000 building blocks vs ~100000000 full mol smiles). But if going down that path and trading off some accuracy for efficiency, then it becomes important to be able to apply the same methods to test dataset (even though efficiency on test set is not so crucial).\n\nMaybe I can still go down this route if the change (reaction) for a given building block is consistent everywhere it's used. If so I can investigate the few hundred building blocks in the non-triazine set, learn and pre-apply the change to the building block and still use any \"building block\"-based features I end up having in my final model(s).\n\nThanks as always for the insight! It's good to at least partially understand what's going on with the data. :)",
    "2757932": "I'm not quite following. Are you training on building blocks?",
    "2757960": "I'm considering doing some (of the more time consuming) feature generation on building blocks, then lookup and pull in those building block features for each row of train. \n\nClearly it would be better to do on full molecule, but given time and compute constraints using BB features might be better than nothing.",
    "2758067": "Ah, I see. Good luck :)",
    "2758832": "Good point !!"
  },
  "source": "meta"
}