{
  "id": 494885,
  "title": "Features for training",
  "url": "/competitions/leash-BELKA/discussion/494885",
  "author_name": "Nader E Afshar",
  "post_date": "2024-04-18T19:37:55.122000",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>The overview states :</p>\n<blockquote>\n  <p>develop machine learning (ML) models to predict the binding affinity of small molecules to specific protein targets</p>\n</blockquote>\n<p>is the intended use of this model to be fed a molecule (in SMILEs format) and see if that molecule would have a \"binding affinity\" to one of the three proteins?  If so do we need the building blocks for training?</p>",
  "messages": [
    {
      "id": 2759623,
      "postDate": "2024-04-18T19:37:55.123Z",
      "content": "<p>The overview states :</p>\n<blockquote>\n  <p>develop machine learning (ML) models to predict the binding affinity of small molecules to specific protein targets</p>\n</blockquote>\n<p>is the intended use of this model to be fed a molecule (in SMILEs format) and see if that molecule would have a \"binding affinity\" to one of the three proteins?  If so do we need the building blocks for training?</p>",
      "rawMarkdown": "The overview states :\n\n>develop machine learning (ML) models to predict the binding affinity of small molecules to specific protein targets\n\nis the intended use of this model to be fed a molecule (in SMILEs format) and see if that molecule would have a \"binding affinity\" to one of the three proteins?  If so do we need the building blocks for training?",
      "votes": 5
    },
    {
      "id": 2761384,
      "postDate": "2024-04-19T19:55:32.210Z",
      "content": "<p>The short answer to the second question is no, the building blocks are not necessary. The molecule_smiles tells you the final result after putting together the blocks. </p>",
      "rawMarkdown": "The short answer to the second question is no, the building blocks are not necessary. The molecule_smiles tells you the final result after putting together the blocks. ",
      "votes": 1
    },
    {
      "id": 2773534,
      "postDate": "2024-04-24T19:15:17.597Z",
      "content": "<p>you can use similarity between two building blocks as feature.</p>\n<p>i often see this type of cube diagram for chemical space exploration.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9f7415dc7042686549f34cd730ebea56%2FSelection_037.png?generation=1713986236181224&amp;alt=media\"></p>\n<p>by i think it should be plotted on the block similarity space. e.g.<br>\n1) you have 100 reference block in train<br>\n2) you have 50 test block<br>\n3) then you have 50x100 similarity feature<br>\n4) perform tsne on 50x100 features and plot the pos and neg binding cases</p>",
      "rawMarkdown": "you can use similarity between two building blocks as feature.\n\ni often see this type of cube diagram for chemical space exploration.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9f7415dc7042686549f34cd730ebea56%2FSelection_037.png?generation=1713986236181224&alt=media)\n\nby i think it should be plotted on the block similarity space. e.g.\n1) you have 100 reference block in train\n2) you have 50 test block\n3) then you have 50x100 similarity feature\n4) perform tsne on 50x100 features and plot the pos and neg binding cases",
      "replies": [
        {
          "id": 2773620,
          "postDate": "2024-04-24T19:59:49.997Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - This is a great idea, but it seems the molecule already has this information embedded in it, which is why I questions the need for individual Building Blocks. After all it is the molecule as a whole that binds to the protein. </p>",
          "rawMarkdown": "@hengck23 - This is a great idea, but it seems the molecule already has this information embedded in it, which is why I questions the need for individual Building Blocks. After all it is the molecule as a whole that binds to the protein. ",
          "replies": [
            {
              "id": 2773719,
              "postDate": "2024-04-24T21:15:34.523Z",
              "content": "<p>refer to this post:<br>\n<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491908\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/491908</a></p>\n<p>just by KNN of building builds, you can get good precision-recall curve.</p>\n<p>the issue is that test and train blocks are difficult.<br>\nanother method (like vae, diffusion that encode the test/train blocks/molecules to overlapping latent spaces should work)</p>",
              "rawMarkdown": "refer to this post:\nhttps://www.kaggle.com/competitions/leash-BELKA/discussion/491908\n\njust by KNN of building builds, you can get good precision-recall curve.\n\nthe issue is that test and train blocks are difficult.\nanother method (like vae, diffusion that encode the test/train blocks/molecules to overlapping latent spaces should work)"
            }
          ]
        }
      ]
    },
    {
      "id": 2761542,
      "postDate": "2024-04-19T22:20:52.057Z",
      "content": "<p>i have a question for the chemist. does the knowledge of how the building blocks react and form the final compound molecules helps in prediction of binding?</p>\n<p>e.g. i build a model as:<br>\nmodel (bb1,bb2,bb3) = molecule smiles + binding</p>\n<p>if it helps, then i can do self supervised learning on test data.<br>\nmodel (bb1,bb2,bb3) = molecule smiles</p>",
      "rawMarkdown": "i have a question for the chemist. does the knowledge of how the building blocks react and form the final compound molecules helps in prediction of binding?\n\ne.g. i build a model as:\nmodel (bb1,bb2,bb3) = molecule smiles + binding\n\nif it helps, then i can do self supervised learning on test data.\nmodel (bb1,bb2,bb3) = molecule smiles\n\n",
      "replies": [
        {
          "id": 2761576,
          "postDate": "2024-04-19T23:49:56.627Z",
          "content": "<p><strong>Knowledge of how the building blocks react shouldn't matter for the binding prediction</strong>… with a few major caveats.</p>\n<p>when we screen via DELs, we just get to read the DNA tag of the molecules that stuck. However, the chemical synthesis isn't perfect. for the most part we pretend each building block reacted, and we ended up with the full molecule, in reality each of these reactions have a yield around 70%-100%, so you end up with a truncated tree of possible products for the reaction, and if you have 3 reactions with 70% yield, you can end up with less than 50% of the dna-tags attached to your full molecule by the end. See Figure 2 <a href=\"https://pubs.acs.org/doi/epdf/10.1021/acscombsci.6b00001\" target=\"_blank\">here</a></p>\n<p>Unfortunately, we do not have yield information about this DEL, so it's hard to estimate how much of the shared effect from building blocks are due to full molecules binding in similar ways or to shared truncated products driving binding. If you had additional chemistry knowledge, or a good yield/reaction prediction model, you could start modelling each molecule_smiles as a bag of truncate_smiles, which might or might not open up new patterns in the data.</p>",
          "rawMarkdown": "**Knowledge of how the building blocks react shouldn't matter for the binding prediction**... with a few major caveats.\n\nwhen we screen via DELs, we just get to read the DNA tag of the molecules that stuck. However, the chemical synthesis isn't perfect. for the most part we pretend each building block reacted, and we ended up with the full molecule, in reality each of these reactions have a yield around 70%-100%, so you end up with a truncated tree of possible products for the reaction, and if you have 3 reactions with 70% yield, you can end up with less than 50% of the dna-tags attached to your full molecule by the end. See Figure 2 [here](https://pubs.acs.org/doi/epdf/10.1021/acscombsci.6b00001)\n\nUnfortunately, we do not have yield information about this DEL, so it's hard to estimate how much of the shared effect from building blocks are due to full molecules binding in similar ways or to shared truncated products driving binding. If you had additional chemistry knowledge, or a good yield/reaction prediction model, you could start modelling each molecule_smiles as a bag of truncate_smiles, which might or might not open up new patterns in the data.",
          "votes": 5
        }
      ]
    },
    {
      "id": 2761133,
      "postDate": "2024-04-19T17:09:30.947Z",
      "content": "<p>The main task with this dataset is to predict the binding affinity of each molecule with all three proteins. The** 'molecule_smiles'** column contains our final molecule. You need to test this molecule with the protein and predict its binding affinity. To accomplish this, you can utilize information from the <strong>'buildingblock1_smiles', 'buildingblock2_smiles', and 'buildingblock3_smiles</strong>' columns by incorporating them as features. Additionally, leveraging domain knowledge can enhance the accuracy of your predictions.</p>\n<p>Note : The true difficulty of this competition lies in accurately predicting outcomes for building blocks that were not included in the training phase.</p>",
      "rawMarkdown": "The main task with this dataset is to predict the binding affinity of each molecule with all three proteins. The** 'molecule_smiles'** column contains our final molecule. You need to test this molecule with the protein and predict its binding affinity. To accomplish this, you can utilize information from the **'buildingblock1_smiles', 'buildingblock2_smiles', and 'buildingblock3_smiles**' columns by incorporating them as features. Additionally, leveraging domain knowledge can enhance the accuracy of your predictions.\n\nNote : The true difficulty of this competition lies in accurately predicting outcomes for building blocks that were not included in the training phase."
    },
    {
      "id": 2766725,
      "postDate": "2024-04-21T21:08:48.727Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2761384,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2024-04-19T19:55:32.210000",
      "content": "<p>The short answer to the second question is no, the building blocks are not necessary. The molecule_smiles tells you the final result after putting together the blocks. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2773534,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-24T19:15:17.597000",
      "content": "<p>you can use similarity between two building blocks as feature.</p>\n<p>i often see this type of cube diagram for chemical space exploration.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9f7415dc7042686549f34cd730ebea56%2FSelection_037.png?generation=1713986236181224&amp;alt=media\"></p>\n<p>by i think it should be plotted on the block similarity space. e.g.<br>\n1) you have 100 reference block in train<br>\n2) you have 50 test block<br>\n3) then you have 50x100 similarity feature<br>\n4) perform tsne on 50x100 features and plot the pos and neg binding cases</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2773620,
          "author_name": "Nader E Afshar",
          "author_url": "",
          "post_date": "2024-04-24T19:59:49.997000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - This is a great idea, but it seems the molecule already has this information embedded in it, which is why I questions the need for individual Building Blocks. After all it is the molecule as a whole that binds to the protein. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2773719,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-24T21:15:34.523000",
              "content": "<p>refer to this post:<br>\n<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491908\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/491908</a></p>\n<p>just by KNN of building builds, you can get good precision-recall curve.</p>\n<p>the issue is that test and train blocks are difficult.<br>\nanother method (like vae, diffusion that encode the test/train blocks/molecules to overlapping latent spaces should work)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2761542,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-19T22:20:52.057000",
      "content": "<p>i have a question for the chemist. does the knowledge of how the building blocks react and form the final compound molecules helps in prediction of binding?</p>\n<p>e.g. i build a model as:<br>\nmodel (bb1,bb2,bb3) = molecule smiles + binding</p>\n<p>if it helps, then i can do self supervised learning on test data.<br>\nmodel (bb1,bb2,bb3) = molecule smiles</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2761576,
          "author_name": "Andrew D. Blevins",
          "author_url": "",
          "post_date": "2024-04-19T23:49:56.627000",
          "content": "<p><strong>Knowledge of how the building blocks react shouldn't matter for the binding prediction</strong>… with a few major caveats.</p>\n<p>when we screen via DELs, we just get to read the DNA tag of the molecules that stuck. However, the chemical synthesis isn't perfect. for the most part we pretend each building block reacted, and we ended up with the full molecule, in reality each of these reactions have a yield around 70%-100%, so you end up with a truncated tree of possible products for the reaction, and if you have 3 reactions with 70% yield, you can end up with less than 50% of the dna-tags attached to your full molecule by the end. See Figure 2 <a href=\"https://pubs.acs.org/doi/epdf/10.1021/acscombsci.6b00001\" target=\"_blank\">here</a></p>\n<p>Unfortunately, we do not have yield information about this DEL, so it's hard to estimate how much of the shared effect from building blocks are due to full molecules binding in similar ways or to shared truncated products driving binding. If you had additional chemistry knowledge, or a good yield/reaction prediction model, you could start modelling each molecule_smiles as a bag of truncate_smiles, which might or might not open up new patterns in the data.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 2761133,
      "author_name": "Md. Shakil Hossen",
      "author_url": "",
      "post_date": "2024-04-19T17:09:30.947000",
      "content": "<p>The main task with this dataset is to predict the binding affinity of each molecule with all three proteins. The** 'molecule_smiles'** column contains our final molecule. You need to test this molecule with the protein and predict its binding affinity. To accomplish this, you can utilize information from the <strong>'buildingblock1_smiles', 'buildingblock2_smiles', and 'buildingblock3_smiles</strong>' columns by incorporating them as features. Additionally, leveraging domain knowledge can enhance the accuracy of your predictions.</p>\n<p>Note : The true difficulty of this competition lies in accurately predicting outcomes for building blocks that were not included in the training phase.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2766725,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-21T21:08:48.727000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2759623": "The overview states :\n\n>develop machine learning (ML) models to predict the binding affinity of small molecules to specific protein targets\n\nis the intended use of this model to be fed a molecule (in SMILEs format) and see if that molecule would have a \"binding affinity\" to one of the three proteins?  If so do we need the building blocks for training?",
    "2761384": "The short answer to the second question is no, the building blocks are not necessary. The molecule_smiles tells you the final result after putting together the blocks. ",
    "2773534": "you can use similarity between two building blocks as feature.\n\ni often see this type of cube diagram for chemical space exploration.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9f7415dc7042686549f34cd730ebea56%2FSelection_037.png?generation=1713986236181224&alt=media)\n\nby i think it should be plotted on the block similarity space. e.g.\n1) you have 100 reference block in train\n2) you have 50 test block\n3) then you have 50x100 similarity feature\n4) perform tsne on 50x100 features and plot the pos and neg binding cases",
    "2761542": "i have a question for the chemist. does the knowledge of how the building blocks react and form the final compound molecules helps in prediction of binding?\n\ne.g. i build a model as:\nmodel (bb1,bb2,bb3) = molecule smiles + binding\n\nif it helps, then i can do self supervised learning on test data.\nmodel (bb1,bb2,bb3) = molecule smiles\n\n",
    "2761133": "The main task with this dataset is to predict the binding affinity of each molecule with all three proteins. The** 'molecule_smiles'** column contains our final molecule. You need to test this molecule with the protein and predict its binding affinity. To accomplish this, you can utilize information from the **'buildingblock1_smiles', 'buildingblock2_smiles', and 'buildingblock3_smiles**' columns by incorporating them as features. Additionally, leveraging domain knowledge can enhance the accuracy of your predictions.\n\nNote : The true difficulty of this competition lies in accurately predicting outcomes for building blocks that were not included in the training phase.",
    "2766725": ""
  }
}