{
  "id": 501053,
  "title": "train a model to predict 3d locations of atoms",
  "url": "/competitions/leash-BELKA/discussion/501053",
  "author_name": "",
  "post_date": "2024-05-07T22:52:43.459911800Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>It is too slow to use rdkit ETKDG and it is not accurate.</p>\n<p>since our training data comes from the same building blocks, i think maybe <br>\n1) we train a model on subset of molecule data + all building blocks to predict 3d coordinates<br>\n2) use the model to label the rest of the data </p>\n<p>I am interested in molecules of sharing blocks at the moment.</p>\n<p>i have not done any work like this before. any experienced kagglers can comment on this?<br>\nwe need pretty accurate results for step2.</p>\n<p>--</p>\n<p>Note that in  AlphaFold, it searches known 3d location from database using multiple sequence alignment (MSA) to get an initial estimate. I wonder does it apply to our case, since we have the building blocks</p>",
  "messages": [
    {
      "id": "2799757",
      "postDate": "05/07/2024 22:52:43",
      "content": "<p>It is too slow to use rdkit ETKDG and it is not accurate.</p>\n<p>since our training data comes from the same building blocks, i think maybe <br>\n1) we train a model on subset of molecule data + all building blocks to predict 3d coordinates<br>\n2) use the model to label the rest of the data </p>\n<p>I am interested in molecules of sharing blocks at the moment.</p>\n<p>i have not done any work like this before. any experienced kagglers can comment on this?<br>\nwe need pretty accurate results for step2.</p>\n<p>--</p>\n<p>Note that in  AlphaFold, it searches known 3d location from database using multiple sequence alignment (MSA) to get an initial estimate. I wonder does it apply to our case, since we have the building blocks</p>",
      "rawMarkdown": "It is too slow to use rdkit ETKDG and it is not accurate.\n\nsince our training data comes from the same building blocks, i think maybe \n1) we train a model on subset of molecule data + all building blocks to predict 3d coordinates\n2) use the model to label the rest of the data \n\nI am interested in molecules of sharing blocks at the moment.\n\ni have not done any work like this before. any experienced kagglers can comment on this?\nwe need pretty accurate results for step2.\n\n--\n\nNote that in  AlphaFold, it searches known 3d location from database using multiple sequence alignment (MSA) to get an initial estimate. I wonder does it apply to our case, since we have the building blocks",
      "votes": null
    },
    {
      "id": "2800878",
      "postDate": "05/08/2024 11:46:43",
      "content": "<p>Somewhat related question here - we're given SMILES strings for the 3 building blocks but we're not given the SMILES for the \"core\". It's not an issue for the train set, since all molecules here have triazine core, but I wonder if all molecules from the test set have triazine core too? If not - we likely need to extract it somehow from the molecule to have all the necessary components for your experiment.</p>\n<p>Perhaps, this is more a question to domain experts <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> and competition hosts <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> </p>",
      "rawMarkdown": "Somewhat related question here - we're given SMILES strings for the 3 building blocks but we're not given the SMILES for the \"core\". It's not an issue for the train set, since all molecules here have triazine core, but I wonder if all molecules from the test set have triazine core too? If not - we likely need to extract it somehow from the molecule to have all the necessary components for your experiment.\n\nPerhaps, this is more a question to domain experts @chemdatafarmer @roberthatch and competition hosts @addisonhoward",
      "votes": null
    },
    {
      "id": "2801241",
      "postDate": "05/08/2024 14:54:09",
      "content": "<p>The test set definitely contains non-triazine cores.  <a href=\"https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration\" target=\"_blank\">This notebook</a> illustrates this.  I think it is mentioned somewhere in the competition details as well, but I can't seem to find that right now.</p>",
      "rawMarkdown": "The test set definitely contains non-triazine cores.  [This notebook](https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration) illustrates this.  I think it is mentioned somewhere in the competition details as well, but I can't seem to find that right now.",
      "votes": null
    },
    {
      "id": "2801417",
      "postDate": "05/08/2024 16:03:36",
      "content": "<p>Note that a single molecule can have multiple set of valid 3D coordinates (i.e. conformational isomers, tautomerism, etc). To select the most relevant conformation, docking softwares use the protein structure to rank each possible conformers in terms of favorable interactions, torsions and clashes (but then some protein targets can have multiple binding pockets, i.e. orthosteric sites and allosteric sites).  </p>\n<p>Indeed, one possible strategy is to dock a sample of molecules (in training and test set) and then build a surrogate regression model that can predict the docking scores from the smiles string which can then be used as features for the classification task.</p>",
      "rawMarkdown": "Note that a single molecule can have multiple set of valid 3D coordinates (i.e. conformational isomers, tautomerism, etc). To select the most relevant conformation, docking softwares use the protein structure to rank each possible conformers in terms of favorable interactions, torsions and clashes (but then some protein targets can have multiple binding pockets, i.e. orthosteric sites and allosteric sites).  \n\nIndeed, one possible strategy is to dock a sample of molecules (in training and test set) and then build a surrogate regression model that can predict the docking scores from the smiles string which can then be used as features for the classification task.",
      "votes": null
    },
    {
      "id": "2801530",
      "postDate": "05/08/2024 16:45:24",
      "content": "<p>It may be completely stupid what I am going to say, however I think that the problem mostly lies with protein representation. I am using a RoBERTa model trained on the UniProt dataset and I have noticed a (marginal) improvement when removing from the training set one building block after having it going up in dimensionality (at the expense of my RAM)</p>",
      "rawMarkdown": "It may be completely stupid what I am going to say, however I think that the problem mostly lies with protein representation. I am using a RoBERTa model trained on the UniProt dataset and I have noticed a (marginal) improvement when removing from the training set one building block after having it going up in dimensionality (at the expense of my RAM)",
      "votes": null
    },
    {
      "id": "2801714",
      "postDate": "05/08/2024 17:48:50",
      "content": "<p>Than we miss the part of the quiz here… I know how to check if the assembled molecule has a certain component - Triazine, Benzene, Furan, whatever, but have no idea how to get a single \"core\" out of it. Being able to do so would be helpful… </p>",
      "rawMarkdown": "Than we miss the part of the quiz here... I know how to check if the assembled molecule has a certain component - Triazine, Benzene, Furan, whatever, but have no idea how to get a single \"core\" out of it. Being able to do so would be helpful...",
      "votes": null
    },
    {
      "id": "2801773",
      "postDate": "05/08/2024 18:15:13",
      "content": "<p>\"but I wonder if all molecules from the test set have triazine core too?\" As <a href=\"https://www.kaggle.com/kirkdco\" target=\"_blank\">@kirkdco</a> mentioned, the test set contains non-triazine cores. </p>\n<p>To systematically get the core, you need to have the SMIRKS strings of the reactions used (which is not provided). You can however visually subtract the building blocks from the final compound to approximate the core.  </p>",
      "rawMarkdown": "\"but I wonder if all molecules from the test set have triazine core too?\" As @kirkdco mentioned, the test set contains non-triazine cores. \n\nTo systematically get the core, you need to have the SMIRKS strings of the reactions used (which is not provided). You can however visually subtract the building blocks from the final compound to approximate the core.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2800878,
      "author_name": "victorshlepov",
      "author_url": "",
      "post_date": "05/08/2024 11:46:43",
      "content": "<p>Somewhat related question here - we're given SMILES strings for the 3 building blocks but we're not given the SMILES for the \"core\". It's not an issue for the train set, since all molecules here have triazine core, but I wonder if all molecules from the test set have triazine core too? If not - we likely need to extract it somehow from the molecule to have all the necessary components for your experiment.</p>\n<p>Perhaps, this is more a question to domain experts <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> and competition hosts <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2801241,
          "author_name": "kirkdco",
          "author_url": "",
          "post_date": "05/08/2024 14:54:09",
          "content": "<p>The test set definitely contains non-triazine cores.  <a href=\"https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration\" target=\"_blank\">This notebook</a> illustrates this.  I think it is mentioned somewhere in the competition details as well, but I can't seem to find that right now.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2801714,
              "author_name": "victorshlepov",
              "author_url": "",
              "post_date": "05/08/2024 17:48:50",
              "content": "<p>Than we miss the part of the quiz here… I know how to check if the assembled molecule has a certain component - Triazine, Benzene, Furan, whatever, but have no idea how to get a single \"core\" out of it. Being able to do so would be helpful… </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2801773,
          "author_name": "ruelcedeno",
          "author_url": "",
          "post_date": "05/08/2024 18:15:13",
          "content": "<p>\"but I wonder if all molecules from the test set have triazine core too?\" As <a href=\"https://www.kaggle.com/kirkdco\" target=\"_blank\">@kirkdco</a> mentioned, the test set contains non-triazine cores. </p>\n<p>To systematically get the core, you need to have the SMIRKS strings of the reactions used (which is not provided). You can however visually subtract the building blocks from the final compound to approximate the core.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2801417,
      "author_name": "ruelcedeno",
      "author_url": "",
      "post_date": "05/08/2024 16:03:36",
      "content": "<p>Note that a single molecule can have multiple set of valid 3D coordinates (i.e. conformational isomers, tautomerism, etc). To select the most relevant conformation, docking softwares use the protein structure to rank each possible conformers in terms of favorable interactions, torsions and clashes (but then some protein targets can have multiple binding pockets, i.e. orthosteric sites and allosteric sites).  </p>\n<p>Indeed, one possible strategy is to dock a sample of molecules (in training and test set) and then build a surrogate regression model that can predict the docking scores from the smiles string which can then be used as features for the classification task.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2801530,
      "author_name": "giorgiomicaletto",
      "author_url": "",
      "post_date": "05/08/2024 16:45:24",
      "content": "<p>It may be completely stupid what I am going to say, however I think that the problem mostly lies with protein representation. I am using a RoBERTa model trained on the UniProt dataset and I have noticed a (marginal) improvement when removing from the training set one building block after having it going up in dimensionality (at the expense of my RAM)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2799757": "It is too slow to use rdkit ETKDG and it is not accurate.\n\nsince our training data comes from the same building blocks, i think maybe \n1) we train a model on subset of molecule data + all building blocks to predict 3d coordinates\n2) use the model to label the rest of the data \n\nI am interested in molecules of sharing blocks at the moment.\n\ni have not done any work like this before. any experienced kagglers can comment on this?\nwe need pretty accurate results for step2.\n\n--\n\nNote that in  AlphaFold, it searches known 3d location from database using multiple sequence alignment (MSA) to get an initial estimate. I wonder does it apply to our case, since we have the building blocks",
    "2800878": "Somewhat related question here - we're given SMILES strings for the 3 building blocks but we're not given the SMILES for the \"core\". It's not an issue for the train set, since all molecules here have triazine core, but I wonder if all molecules from the test set have triazine core too? If not - we likely need to extract it somehow from the molecule to have all the necessary components for your experiment.\n\nPerhaps, this is more a question to domain experts @chemdatafarmer @roberthatch and competition hosts @addisonhoward",
    "2801241": "The test set definitely contains non-triazine cores.  [This notebook](https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration) illustrates this.  I think it is mentioned somewhere in the competition details as well, but I can't seem to find that right now.",
    "2801417": "Note that a single molecule can have multiple set of valid 3D coordinates (i.e. conformational isomers, tautomerism, etc). To select the most relevant conformation, docking softwares use the protein structure to rank each possible conformers in terms of favorable interactions, torsions and clashes (but then some protein targets can have multiple binding pockets, i.e. orthosteric sites and allosteric sites).  \n\nIndeed, one possible strategy is to dock a sample of molecules (in training and test set) and then build a surrogate regression model that can predict the docking scores from the smiles string which can then be used as features for the classification task.",
    "2801530": "It may be completely stupid what I am going to say, however I think that the problem mostly lies with protein representation. I am using a RoBERTa model trained on the UniProt dataset and I have noticed a (marginal) improvement when removing from the training set one building block after having it going up in dimensionality (at the expense of my RAM)",
    "2801714": "Than we miss the part of the quiz here... I know how to check if the assembled molecule has a certain component - Triazine, Benzene, Furan, whatever, but have no idea how to get a single \"core\" out of it. Being able to do so would be helpful...",
    "2801773": "\"but I wonder if all molecules from the test set have triazine core too?\" As @kirkdco mentioned, the test set contains non-triazine cores. \n\nTo systematically get the core, you need to have the SMIRKS strings of the reactions used (which is not provided). You can however visually subtract the building blocks from the final compound to approximate the core."
  },
  "source": "meta"
}