{
  "id": 491908,
  "title": "background reading on DEL + binding prediction",
  "url": "/competitions/leash-BELKA/discussion/491908",
  "author_name": "",
  "post_date": "2024-04-07T20:11:45.534717300Z",
  "votes": 48,
  "comment_count": 12,
  "views": 0,
  "content": "<p>paper: Building Block-Based Binding Predictions for DNA-Encoded Libraries<br>\ncode, etc : <a href=\"https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57\" target=\"_blank\">https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57</a></p>\n<p>\"Application to Holdout Data. In this final section, we demonstrate how we would apply this method in a practical setting, where we would work to guide the design of a new DEL using information available from a prior screen. Here, we model this design process by testing the performance of our model using a holdout set (using building blocks not seen previously) to mimic a new set of building blocks to test.\"</p>\n<p>related: Compositional Deep Probabilistic Models of DNA-Encoded Libraries<br>\n<a href=\"https://arxiv.org/abs/2310.13769\" target=\"_blank\">https://arxiv.org/abs/2310.13769</a></p>\n<hr>\n<p>Note: the real challenge of this competition is really about prediction on building blocks not used in training.<br>\nas a first step, you should perform separate validation on both blocks in and not in training to measure the \"performance difference\" of the two.</p>",
  "messages": [
    {
      "id": "2740543",
      "postDate": "04/07/2024 20:11:45",
      "content": "<p>paper: Building Block-Based Binding Predictions for DNA-Encoded Libraries<br>\ncode, etc : <a href=\"https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57\" target=\"_blank\">https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57</a></p>\n<p>\"Application to Holdout Data. In this final section, we demonstrate how we would apply this method in a practical setting, where we would work to guide the design of a new DEL using information available from a prior screen. Here, we model this design process by testing the performance of our model using a holdout set (using building blocks not seen previously) to mimic a new set of building blocks to test.\"</p>\n<p>related: Compositional Deep Probabilistic Models of DNA-Encoded Libraries<br>\n<a href=\"https://arxiv.org/abs/2310.13769\" target=\"_blank\">https://arxiv.org/abs/2310.13769</a></p>\n<hr>\n<p>Note: the real challenge of this competition is really about prediction on building blocks not used in training.<br>\nas a first step, you should perform separate validation on both blocks in and not in training to measure the \"performance difference\" of the two.</p>",
      "rawMarkdown": "paper: Building Block-Based Binding Predictions for DNA-Encoded Libraries\ncode, etc : https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57\n\n\"Application to Holdout Data. In this final section, we demonstrate how we would apply this method in a practical setting, where we would work to guide the design of a new DEL using information available from a prior screen. Here, we model this design process by testing the performance of our model using a holdout set (using building blocks not seen previously) to mimic a new set of building blocks to test.\"\n\nrelated: Compositional Deep Probabilistic Models of DNA-Encoded Libraries\nhttps://arxiv.org/abs/2310.13769\n\n----\n\nNote: the real challenge of this competition is really about prediction on building blocks not used in training.\nas a first step, you should perform separate validation on both blocks in and not in training to measure the \"performance difference\" of the two.",
      "votes": null
    },
    {
      "id": "2740545",
      "postDate": "04/07/2024 20:14:56",
      "content": "<p>paper: Machine Learning on DNA-Encoded Libraries: A New Paradigm for Hit Finding<br>\n<a href=\"https://accio.github.io/AMIDD/assets/2020/13-14/McCloskey-2020-ML-DELT.pdf\" target=\"_blank\">https://accio.github.io/AMIDD/assets/2020/13-14/McCloskey-2020-ML-DELT.pdf</a><br>\n<a href=\"https://accio.github.io/AMIDD/\" target=\"_blank\">https://accio.github.io/AMIDD/</a><br>\ncode: <a href=\"https://github.com/google-research/google-research/tree/master/gigamol\" target=\"_blank\">https://github.com/google-research/google-research/tree/master/gigamol</a></p>",
      "rawMarkdown": "paper: Machine Learning on DNA-Encoded Libraries: A New Paradigm for Hit Finding\nhttps://accio.github.io/AMIDD/assets/2020/13-14/McCloskey-2020-ML-DELT.pdf\nhttps://accio.github.io/AMIDD/\ncode: https://github.com/google-research/google-research/tree/master/gigamol",
      "votes": null
    },
    {
      "id": "2740965",
      "postDate": "04/08/2024 03:52:29",
      "content": "<p>nice course for understanding this competition, thanks for sharing <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://accio.github.io/AMIDD/\" target=\"_blank\">Applied Mathematics and Informatics in Drug Discovery - https://accio.github.io/AMIDD/</a></p>",
      "rawMarkdown": "nice course for understanding this competition, thanks for sharing @hengck23 [Applied Mathematics and Informatics in Drug Discovery - https://accio.github.io/AMIDD/](https://accio.github.io/AMIDD/)",
      "votes": null
    },
    {
      "id": "2741310",
      "postDate": "04/08/2024 08:49:59",
      "content": "<p>external data from the paper  for sEH binding:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5f2c6a58a2a803f5ced6169ce431639a%2FSelection_012.png?generation=1712566146875200&amp;alt=media\"></p>\n<p>but there is an issue: how to insert the DNA linker [Py] that is used kaggle dataset?<br>\nCan any expert advise?</p>",
      "rawMarkdown": "external data from the paper  for sEH binding:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5f2c6a58a2a803f5ced6169ce431639a%2FSelection_012.png?generation=1712566146875200&alt=media)\n\nbut there is an issue: how to insert the DNA linker [Py] that is used kaggle dataset?\nCan any expert advise?",
      "votes": null
    },
    {
      "id": "2741574",
      "postDate": "04/08/2024 13:14:30",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> !</p>\n<p>It looks like the authors of the original paper used a methyl group to identify the DNA attachment point. This makes some sense from a modeling perspective, but it's a little tricky to trace back as there could be methyl groups elsewhere on the molecule that are real methyl groups. Ideally, they would have labeled them isotopically or something to signify the attachment point.</p>\n<p>Regardless, I may be able to help here. Give me some time and I'll see if I can post a dataset where the final molecule has [Dy] instead of an ambiguous methyl.</p>",
      "rawMarkdown": "Hi @hengck23 !\n\nIt looks like the authors of the original paper used a methyl group to identify the DNA attachment point. This makes some sense from a modeling perspective, but it's a little tricky to trace back as there could be methyl groups elsewhere on the molecule that are real methyl groups. Ideally, they would have labeled them isotopically or something to signify the attachment point.\n\nRegardless, I may be able to help here. Give me some time and I'll see if I can post a dataset where the final molecule has [Dy] instead of an ambiguous methyl.",
      "votes": null
    },
    {
      "id": "2741593",
      "postDate": "04/08/2024 13:32:08",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!",
      "votes": null
    },
    {
      "id": "2741855",
      "postDate": "04/08/2024 16:15:48",
      "content": "<p>thanks. the dataset can be downloaded at: <a href=\"https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57\" target=\"_blank\">https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57</a></p>\n<p>code for preparing the data: <a href=\"https://github.com/MobleyLab/DEL_analysis\" target=\"_blank\">https://github.com/MobleyLab/DEL_analysis</a></p>\n<p>both smiles and isomeric smiles of the building blocks are provided.</p>",
      "rawMarkdown": "thanks. the dataset can be downloaded at: https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57\n\ncode for preparing the data: https://github.com/MobleyLab/DEL_analysis\n\nboth smiles and isomeric smiles of the building blocks are provided.",
      "votes": null
    },
    {
      "id": "2742204",
      "postDate": "04/08/2024 19:48:21",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Please see the following notebook and the associated dataset: <a href=\"https://www.kaggle.com/code/chemdatafarmer/additional-seh-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/chemdatafarmer/additional-seh-data/notebook</a></p>\n<p>The new structures are under the \"new_structure\" column in the .csv</p>\n<p>Not sure if it'll be useful since these scaffolds aren't triazenes, but who knows! Good luck and let me know if you find some value here.</p>",
      "rawMarkdown": "hengck23 Please see the following notebook and the associated dataset: https://www.kaggle.com/code/chemdatafarmer/additional-seh-data/notebook\n\nThe new structures are under the \"new_structure\" column in the .csv\n\nNot sure if it'll be useful since these scaffolds aren't triazenes, but who knows! Good luck and let me know if you find some value here.",
      "votes": null
    },
    {
      "id": "2743415",
      "postDate": "04/09/2024 12:52:38",
      "content": "<p><a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> <br>\nthank you very much. i will take a look! :)</p>",
      "rawMarkdown": "chemdatafarmer \nthank you very much. i will take a look! :)",
      "votes": null
    },
    {
      "id": "2746040",
      "postDate": "04/11/2024 02:32:11",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> it turns out, I think this extra data may be useful after all.</p>\n<p>From the competition organizers: \"molecule_smiles - The structure of the fully assembled molecule, in SMILES. This includes the three building blocks and the triazine core. Note we use a [Dy] as the stand-in for the DNA linker.\"</p>\n<p>However, I've found that some of the molecules in the test set are not enumerated from a triazine. I need to look more closely at the train to see if that's indeed the case.</p>\n<p>Example ID: 296896118 in test.csv is not enumerated from a triazine…If the train set solely comprises triazines and the test set does not, then our models need to be powerful enough to perform scaffold hopping (a difficult task in medicinal chemistry).</p>",
      "rawMarkdown": "hengck23 it turns out, I think this extra data may be useful after all.\n\nFrom the competition organizers: \"molecule_smiles - The structure of the fully assembled molecule, in SMILES. This includes the three building blocks and the triazine core. Note we use a [Dy] as the stand-in for the DNA linker.\"\n\nHowever, I've found that some of the molecules in the test set are not enumerated from a triazine. I need to look more closely at the train to see if that's indeed the case.\n\nExample ID: 296896118 in test.csv is not enumerated from a triazine...If the train set solely comprises triazines and the test set does not, then our models need to be powerful enough to perform scaffold hopping (a difficult task in medicinal chemistry).",
      "votes": null
    },
    {
      "id": "2760520",
      "postDate": "04/19/2024 10:26:33",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for sharing. Have you already experienced some insights from the papers ? If yes, do they improve your CV score (not LB) ?</p>",
      "rawMarkdown": "Thank you @hengck23 for sharing. Have you already experienced some insights from the papers ? If yes, do they improve your CV score (not LB) ?",
      "votes": null
    },
    {
      "id": "2766880",
      "postDate": "04/22/2024 01:44:30",
      "content": "<p>here are more resources</p>\n<ol>\n<li><p><a href=\"https://chemistry-europe.onlinelibrary.wiley.com/doi/10.1002/cbic.202200776\" target=\"_blank\">https://chemistry-europe.onlinelibrary.wiley.com/doi/10.1002/cbic.202200776</a><br>\nunderstand the tasks Drug-target interaction prediction(DTI), Docking, Binding site detection, De novo design and what are the relevant methods dataset</p></li>\n<li><p><a href=\"https://paperswithcode.com/sota/protein-ligand-affinity-prediction-on-pdbbind\" target=\"_blank\">https://paperswithcode.com/sota/protein-ligand-affinity-prediction-on-pdbbind</a><br>\nbenchmark<br>\nsee also <a href=\"https://paperswithcode.com/dataset/pdbbind-1\" target=\"_blank\">https://paperswithcode.com/dataset/pdbbind-1</a></p></li>\n<li><p>interesting paper: High Performance of Gradient Boosting in Binding Affinity Prediction <br>\n<a href=\"https://paperswithcode.com/paper/high-performance-of-gradient-boosting-in\" target=\"_blank\">https://paperswithcode.com/paper/high-performance-of-gradient-boosting-in</a><br>\n<a href=\"https://github.com/miladrayka/GB_Score\" target=\"_blank\">https://github.com/miladrayka/GB_Score</a> (?)<br>\n<a href=\"https://github.com/agave233/SIGN\" target=\"_blank\">https://github.com/agave233/SIGN</a> (feature extraction)</p></li>\n</ol>",
      "rawMarkdown": "here are more resources\n1. https://chemistry-europe.onlinelibrary.wiley.com/doi/10.1002/cbic.202200776\nunderstand the tasks Drug-target interaction prediction(DTI), Docking, Binding site detection, De novo design and what are the relevant methods dataset\n\n2. https://paperswithcode.com/sota/protein-ligand-affinity-prediction-on-pdbbind\nbenchmark\nsee also https://paperswithcode.com/dataset/pdbbind-1\n\n3. interesting paper: High Performance of Gradient Boosting in Binding Affinity Prediction \nhttps://paperswithcode.com/paper/high-performance-of-gradient-boosting-in\nhttps://github.com/miladrayka/GB_Score (?)\nhttps://github.com/agave233/SIGN (feature extraction)",
      "votes": null
    },
    {
      "id": "2776051",
      "postDate": "04/26/2024 00:46:58",
      "content": "<p>Towards DNA-Encoded Library Generation with GFlowNets<br>\nYoshua Bengio<br>\n<a href=\"https://arxiv.org/html/2404.10094v1\" target=\"_blank\">https://arxiv.org/html/2404.10094v1</a></p>",
      "rawMarkdown": "Towards DNA-Encoded Library Generation with GFlowNets\nYoshua Bengio\nhttps://arxiv.org/html/2404.10094v1",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2740545,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/07/2024 20:14:56",
      "content": "<p>paper: Machine Learning on DNA-Encoded Libraries: A New Paradigm for Hit Finding<br>\n<a href=\"https://accio.github.io/AMIDD/assets/2020/13-14/McCloskey-2020-ML-DELT.pdf\" target=\"_blank\">https://accio.github.io/AMIDD/assets/2020/13-14/McCloskey-2020-ML-DELT.pdf</a><br>\n<a href=\"https://accio.github.io/AMIDD/\" target=\"_blank\">https://accio.github.io/AMIDD/</a><br>\ncode: <a href=\"https://github.com/google-research/google-research/tree/master/gigamol\" target=\"_blank\">https://github.com/google-research/google-research/tree/master/gigamol</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2740965,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "04/08/2024 03:52:29",
          "content": "<p>nice course for understanding this competition, thanks for sharing <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://accio.github.io/AMIDD/\" target=\"_blank\">Applied Mathematics and Informatics in Drug Discovery - https://accio.github.io/AMIDD/</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2741310,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/08/2024 08:49:59",
      "content": "<p>external data from the paper  for sEH binding:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5f2c6a58a2a803f5ced6169ce431639a%2FSelection_012.png?generation=1712566146875200&amp;alt=media\"></p>\n<p>but there is an issue: how to insert the DNA linker [Py] that is used kaggle dataset?<br>\nCan any expert advise?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2741574,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "04/08/2024 13:14:30",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> !</p>\n<p>It looks like the authors of the original paper used a methyl group to identify the DNA attachment point. This makes some sense from a modeling perspective, but it's a little tricky to trace back as there could be methyl groups elsewhere on the molecule that are real methyl groups. Ideally, they would have labeled them isotopically or something to signify the attachment point.</p>\n<p>Regardless, I may be able to help here. Give me some time and I'll see if I can post a dataset where the final molecule has [Dy] instead of an ambiguous methyl.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2741855,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "04/08/2024 16:15:48",
              "content": "<p>thanks. the dataset can be downloaded at: <a href=\"https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57\" target=\"_blank\">https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57</a></p>\n<p>code for preparing the data: <a href=\"https://github.com/MobleyLab/DEL_analysis\" target=\"_blank\">https://github.com/MobleyLab/DEL_analysis</a></p>\n<p>both smiles and isomeric smiles of the building blocks are provided.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2742204,
                  "author_name": "chemdatafarmer",
                  "author_url": "",
                  "post_date": "04/08/2024 19:48:21",
                  "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Please see the following notebook and the associated dataset: <a href=\"https://www.kaggle.com/code/chemdatafarmer/additional-seh-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/chemdatafarmer/additional-seh-data/notebook</a></p>\n<p>The new structures are under the \"new_structure\" column in the .csv</p>\n<p>Not sure if it'll be useful since these scaffolds aren't triazenes, but who knows! Good luck and let me know if you find some value here.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2743415,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "04/09/2024 12:52:38",
                      "content": "<p><a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> <br>\nthank you very much. i will take a look! :)</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2746040,
                          "author_name": "chemdatafarmer",
                          "author_url": "",
                          "post_date": "04/11/2024 02:32:11",
                          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> it turns out, I think this extra data may be useful after all.</p>\n<p>From the competition organizers: \"molecule_smiles - The structure of the fully assembled molecule, in SMILES. This includes the three building blocks and the triazine core. Note we use a [Dy] as the stand-in for the DNA linker.\"</p>\n<p>However, I've found that some of the molecules in the test set are not enumerated from a triazine. I need to look more closely at the train to see if that's indeed the case.</p>\n<p>Example ID: 296896118 in test.csv is not enumerated from a triazine…If the train set solely comprises triazines and the test set does not, then our models need to be powerful enough to perform scaffold hopping (a difficult task in medicinal chemistry).</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2741593,
      "author_name": "davezzq",
      "author_url": "",
      "post_date": "04/08/2024 13:32:08",
      "content": "<p>Thank you very much!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2760520,
      "author_name": "ulrich07",
      "author_url": "",
      "post_date": "04/19/2024 10:26:33",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for sharing. Have you already experienced some insights from the papers ? If yes, do they improve your CV score (not LB) ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2766880,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/22/2024 01:44:30",
      "content": "<p>here are more resources</p>\n<ol>\n<li><p><a href=\"https://chemistry-europe.onlinelibrary.wiley.com/doi/10.1002/cbic.202200776\" target=\"_blank\">https://chemistry-europe.onlinelibrary.wiley.com/doi/10.1002/cbic.202200776</a><br>\nunderstand the tasks Drug-target interaction prediction(DTI), Docking, Binding site detection, De novo design and what are the relevant methods dataset</p></li>\n<li><p><a href=\"https://paperswithcode.com/sota/protein-ligand-affinity-prediction-on-pdbbind\" target=\"_blank\">https://paperswithcode.com/sota/protein-ligand-affinity-prediction-on-pdbbind</a><br>\nbenchmark<br>\nsee also <a href=\"https://paperswithcode.com/dataset/pdbbind-1\" target=\"_blank\">https://paperswithcode.com/dataset/pdbbind-1</a></p></li>\n<li><p>interesting paper: High Performance of Gradient Boosting in Binding Affinity Prediction <br>\n<a href=\"https://paperswithcode.com/paper/high-performance-of-gradient-boosting-in\" target=\"_blank\">https://paperswithcode.com/paper/high-performance-of-gradient-boosting-in</a><br>\n<a href=\"https://github.com/miladrayka/GB_Score\" target=\"_blank\">https://github.com/miladrayka/GB_Score</a> (?)<br>\n<a href=\"https://github.com/agave233/SIGN\" target=\"_blank\">https://github.com/agave233/SIGN</a> (feature extraction)</p></li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2776051,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/26/2024 00:46:58",
      "content": "<p>Towards DNA-Encoded Library Generation with GFlowNets<br>\nYoshua Bengio<br>\n<a href=\"https://arxiv.org/html/2404.10094v1\" target=\"_blank\">https://arxiv.org/html/2404.10094v1</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2740543": "paper: Building Block-Based Binding Predictions for DNA-Encoded Libraries\ncode, etc : https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57\n\n\"Application to Holdout Data. In this final section, we demonstrate how we would apply this method in a practical setting, where we would work to guide the design of a new DEL using information available from a prior screen. Here, we model this design process by testing the performance of our model using a holdout set (using building blocks not seen previously) to mimic a new set of building blocks to test.\"\n\nrelated: Compositional Deep Probabilistic Models of DNA-Encoded Libraries\nhttps://arxiv.org/abs/2310.13769\n\n----\n\nNote: the real challenge of this competition is really about prediction on building blocks not used in training.\nas a first step, you should perform separate validation on both blocks in and not in training to measure the \"performance difference\" of the two.",
    "2740545": "paper: Machine Learning on DNA-Encoded Libraries: A New Paradigm for Hit Finding\nhttps://accio.github.io/AMIDD/assets/2020/13-14/McCloskey-2020-ML-DELT.pdf\nhttps://accio.github.io/AMIDD/\ncode: https://github.com/google-research/google-research/tree/master/gigamol",
    "2740965": "nice course for understanding this competition, thanks for sharing @hengck23 [Applied Mathematics and Informatics in Drug Discovery - https://accio.github.io/AMIDD/](https://accio.github.io/AMIDD/)",
    "2741310": "external data from the paper  for sEH binding:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5f2c6a58a2a803f5ced6169ce431639a%2FSelection_012.png?generation=1712566146875200&alt=media)\n\nbut there is an issue: how to insert the DNA linker [Py] that is used kaggle dataset?\nCan any expert advise?",
    "2741574": "Hi @hengck23 !\n\nIt looks like the authors of the original paper used a methyl group to identify the DNA attachment point. This makes some sense from a modeling perspective, but it's a little tricky to trace back as there could be methyl groups elsewhere on the molecule that are real methyl groups. Ideally, they would have labeled them isotopically or something to signify the attachment point.\n\nRegardless, I may be able to help here. Give me some time and I'll see if I can post a dataset where the final molecule has [Dy] instead of an ambiguous methyl.",
    "2741593": "Thank you very much!",
    "2741855": "thanks. the dataset can be downloaded at: https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57\n\ncode for preparing the data: https://github.com/MobleyLab/DEL_analysis\n\nboth smiles and isomeric smiles of the building blocks are provided.",
    "2742204": "hengck23 Please see the following notebook and the associated dataset: https://www.kaggle.com/code/chemdatafarmer/additional-seh-data/notebook\n\nThe new structures are under the \"new_structure\" column in the .csv\n\nNot sure if it'll be useful since these scaffolds aren't triazenes, but who knows! Good luck and let me know if you find some value here.",
    "2743415": "chemdatafarmer \nthank you very much. i will take a look! :)",
    "2746040": "hengck23 it turns out, I think this extra data may be useful after all.\n\nFrom the competition organizers: \"molecule_smiles - The structure of the fully assembled molecule, in SMILES. This includes the three building blocks and the triazine core. Note we use a [Dy] as the stand-in for the DNA linker.\"\n\nHowever, I've found that some of the molecules in the test set are not enumerated from a triazine. I need to look more closely at the train to see if that's indeed the case.\n\nExample ID: 296896118 in test.csv is not enumerated from a triazine...If the train set solely comprises triazines and the test set does not, then our models need to be powerful enough to perform scaffold hopping (a difficult task in medicinal chemistry).",
    "2760520": "Thank you @hengck23 for sharing. Have you already experienced some insights from the papers ? If yes, do they improve your CV score (not LB) ?",
    "2766880": "here are more resources\n1. https://chemistry-europe.onlinelibrary.wiley.com/doi/10.1002/cbic.202200776\nunderstand the tasks Drug-target interaction prediction(DTI), Docking, Binding site detection, De novo design and what are the relevant methods dataset\n\n2. https://paperswithcode.com/sota/protein-ligand-affinity-prediction-on-pdbbind\nbenchmark\nsee also https://paperswithcode.com/dataset/pdbbind-1\n\n3. interesting paper: High Performance of Gradient Boosting in Binding Affinity Prediction \nhttps://paperswithcode.com/paper/high-performance-of-gradient-boosting-in\nhttps://github.com/miladrayka/GB_Score (?)\nhttps://github.com/agave233/SIGN (feature extraction)",
    "2776051": "Towards DNA-Encoded Library Generation with GFlowNets\nYoshua Bengio\nhttps://arxiv.org/html/2404.10094v1"
  },
  "source": "meta"
}