{
  "id": 519861,
  "title": "Results and representations",
  "url": "/competitions/leash-BELKA/discussion/519861",
  "author_name": "",
  "post_date": "2024-07-13T09:32:49.995414100Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to everyone that has posted a description of their modeling methods.  Reading through them, while there are a variety of modeling methods - embeddings with ChemBerta, MolFormer, etc., bespoke embeddings, XGBoost, 1D-CNN , GNN - there are some very consistent themes with respect to representations.  Specifically, everything I've read so far use the supplied SMILES as input or generated fingerprints using atom identities (ECFP, RDKit, etc.)</p>\n<p>One conclusion from the competition was that generalization to compounds with structures outside the original training remains unsolved.  This really is the challenge in drug discovery - to identify binders from a unique chemical space which is constant changing.</p>\n<p>I'm wondering if this lack of generalization is a symptom of the representations used.  SMILES and ECFP/RDKit fingerprints are very useful for structural similarity, but the question of generalization outside the training set structures is inconsistent with structural similarity.  </p>\n<p>As I mentioned in my write up, I approached the problem using 2D-pharmacophores in hopes of finding better generalization.  My result was disappointingly in the middle of the pack, however, the evaluation metric doesn't allow for evaluation of the non-shared building block set or the non-triazine set independent of the shared BB - triazine core set.  (I'm holding out hope that my models show some generalization.)</p>\n<p>I'm very curious about the degree of true generalization outside the training set seen with the SMILES or atom identity based fingerprint approaches versus a more abstract approach like 2D-pharmacophores.  I plan to follow up investigating my models and extending to other model types.  I will post the results here once I have them.</p>\n<p>There are a lot of very talented researchers here, and I would very much like to hear the thoughts of others on this idea.  </p>",
  "messages": [
    {
      "id": "2919870",
      "postDate": "07/13/2024 09:32:49",
      "content": "<p>Thanks to everyone that has posted a description of their modeling methods.  Reading through them, while there are a variety of modeling methods - embeddings with ChemBerta, MolFormer, etc., bespoke embeddings, XGBoost, 1D-CNN , GNN - there are some very consistent themes with respect to representations.  Specifically, everything I've read so far use the supplied SMILES as input or generated fingerprints using atom identities (ECFP, RDKit, etc.)</p>\n<p>One conclusion from the competition was that generalization to compounds with structures outside the original training remains unsolved.  This really is the challenge in drug discovery - to identify binders from a unique chemical space which is constant changing.</p>\n<p>I'm wondering if this lack of generalization is a symptom of the representations used.  SMILES and ECFP/RDKit fingerprints are very useful for structural similarity, but the question of generalization outside the training set structures is inconsistent with structural similarity.  </p>\n<p>As I mentioned in my write up, I approached the problem using 2D-pharmacophores in hopes of finding better generalization.  My result was disappointingly in the middle of the pack, however, the evaluation metric doesn't allow for evaluation of the non-shared building block set or the non-triazine set independent of the shared BB - triazine core set.  (I'm holding out hope that my models show some generalization.)</p>\n<p>I'm very curious about the degree of true generalization outside the training set seen with the SMILES or atom identity based fingerprint approaches versus a more abstract approach like 2D-pharmacophores.  I plan to follow up investigating my models and extending to other model types.  I will post the results here once I have them.</p>\n<p>There are a lot of very talented researchers here, and I would very much like to hear the thoughts of others on this idea.  </p>",
      "rawMarkdown": "Thanks to everyone that has posted a description of their modeling methods.  Reading through them, while there are a variety of modeling methods - embeddings with ChemBerta, MolFormer, etc., bespoke embeddings, XGBoost, 1D-CNN , GNN - there are some very consistent themes with respect to representations.  Specifically, everything I've read so far use the supplied SMILES as input or generated fingerprints using atom identities (ECFP, RDKit, etc.)\n\nOne conclusion from the competition was that generalization to compounds with structures outside the original training remains unsolved.  This really is the challenge in drug discovery - to identify binders from a unique chemical space which is constant changing.\n\nI'm wondering if this lack of generalization is a symptom of the representations used.  SMILES and ECFP/RDKit fingerprints are very useful for structural similarity, but the question of generalization outside the training set structures is inconsistent with structural similarity.  \n\nAs I mentioned in my write up, I approached the problem using 2D-pharmacophores in hopes of finding better generalization.  My result was disappointingly in the middle of the pack, however, the evaluation metric doesn't allow for evaluation of the non-shared building block set or the non-triazine set independent of the shared BB - triazine core set.  (I'm holding out hope that my models show some generalization.)\n\nI'm very curious about the degree of true generalization outside the training set seen with the SMILES or atom identity based fingerprint approaches versus a more abstract approach like 2D-pharmacophores.  I plan to follow up investigating my models and extending to other model types.  I will post the results here once I have them.\n\nThere are a lot of very talented researchers here, and I would very much like to hear the thoughts of others on this idea.",
      "votes": null
    },
    {
      "id": "2922358",
      "postDate": "07/15/2024 06:19:00",
      "content": "<p>We don't seem to need the representations to have extrapolation capabilities. Perhaps we just need to do a round of pre-training on all possible molecules and have each molecule be reasonably well placed in the representation space. This might work better for molecules consisting of just a few building blocks……？</p>\n<p>Frankly, I have the same encounter as you in that models that perform well in some common benchmarks (e.g., TDC, MolecularNet, LIT-PCBA, etc.) may not be able to perform well in this competition. I would also like to know what is the reason for this result.</p>",
      "rawMarkdown": "We don't seem to need the representations to have extrapolation capabilities. Perhaps we just need to do a round of pre-training on all possible molecules and have each molecule be reasonably well placed in the representation space. This might work better for molecules consisting of just a few building blocks......？\n\nFrankly, I have the same encounter as you in that models that perform well in some common benchmarks (e.g., TDC, MolecularNet, LIT-PCBA, etc.) may not be able to perform well in this competition. I would also like to know what is the reason for this result.",
      "votes": null
    },
    {
      "id": "2924895",
      "postDate": "07/16/2024 19:34:22",
      "content": "<p>Interesting.  </p>\n<p>I can see your point about embeddings made with very large datasets that cover a lot of chemical space, but I would still expect a reduced representation like pharmacophores to generalize better.  Maybe that's just my bias.</p>\n<p>My plan over the next week or so is to modify my CV approach to develop test sets that are very non-similar to the training set and compare different representations.  (Someone in the discussion forum mentioned Tanimoto &lt; 0.40 as a good cutoff for that.)  When the actual key for the public/private LB sets are released, I'll dig into those a bit as well.  </p>\n<p>I doubt that I can solve the generalization problem, but hopefully I can gain a better understanding of the limitations.</p>",
      "rawMarkdown": "Interesting.  \n\nI can see your point about embeddings made with very large datasets that cover a lot of chemical space, but I would still expect a reduced representation like pharmacophores to generalize better.  Maybe that's just my bias.\n\nMy plan over the next week or so is to modify my CV approach to develop test sets that are very non-similar to the training set and compare different representations.  (Someone in the discussion forum mentioned Tanimoto < 0.40 as a good cutoff for that.)  When the actual key for the public/private LB sets are released, I'll dig into those a bit as well.  \n\nI doubt that I can solve the generalization problem, but hopefully I can gain a better understanding of the limitations.",
      "votes": null
    },
    {
      "id": "2995360",
      "postDate": "09/22/2024 08:26:38",
      "content": "<p>I'm sorry I just saw this comment.</p>\n<p>After two months have passed, have you made some progress on this issue? </p>\n<p>During the month, one of our recent models trained based on intermolecular conformational space similarity achieved good results in some real-world applications (first place in a drug-screening competition through intermolecular similarity).</p>\n<p>Our strategy is to provide a continuous similarity annotation for comparative learning by comparing the shape similarity between all conformations in the conformational space of drug small molecules, and this pre-training strategy helps our <a href=\"https://github.com/Wang-Lin-boop/GeminiMol\" target=\"_blank\">GeminiMol </a>to achieve good generalization performance. </p>\n<p>Although this model did not perform well in this competition, I think it may have some value for researchers interested in using AI for drug discovery. </p>",
      "rawMarkdown": "I'm sorry I just saw this comment.\n\nAfter two months have passed, have you made some progress on this issue? \n\nDuring the month, one of our recent models trained based on intermolecular conformational space similarity achieved good results in some real-world applications (first place in a drug-screening competition through intermolecular similarity).\n\nOur strategy is to provide a continuous similarity annotation for comparative learning by comparing the shape similarity between all conformations in the conformational space of drug small molecules, and this pre-training strategy helps our [GeminiMol ](https://github.com/Wang-Lin-boop/GeminiMol)to achieve good generalization performance. \n\nAlthough this model did not perform well in this competition, I think it may have some value for researchers interested in using AI for drug discovery.",
      "votes": null
    },
    {
      "id": "2995603",
      "postDate": "09/22/2024 13:12:54",
      "content": "<p>Unfortunately, I haven't had time to work on this yet.  Ironically, I've been busy with a project working on leads from a DEL screen. </p>\n<p>Once things settle down a bit, I plan to revisit this. </p>\n<p>Thanks for the link to your paper.  Looks very interesting!</p>",
      "rawMarkdown": "Unfortunately, I haven't had time to work on this yet.  Ironically, I've been busy with a project working on leads from a DEL screen. \n\nOnce things settle down a bit, I plan to revisit this. \n\nThanks for the link to your paper.  Looks very interesting!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2922358,
      "author_name": "wanglinboop",
      "author_url": "",
      "post_date": "07/15/2024 06:19:00",
      "content": "<p>We don't seem to need the representations to have extrapolation capabilities. Perhaps we just need to do a round of pre-training on all possible molecules and have each molecule be reasonably well placed in the representation space. This might work better for molecules consisting of just a few building blocks……？</p>\n<p>Frankly, I have the same encounter as you in that models that perform well in some common benchmarks (e.g., TDC, MolecularNet, LIT-PCBA, etc.) may not be able to perform well in this competition. I would also like to know what is the reason for this result.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2924895,
          "author_name": "kirkdco",
          "author_url": "",
          "post_date": "07/16/2024 19:34:22",
          "content": "<p>Interesting.  </p>\n<p>I can see your point about embeddings made with very large datasets that cover a lot of chemical space, but I would still expect a reduced representation like pharmacophores to generalize better.  Maybe that's just my bias.</p>\n<p>My plan over the next week or so is to modify my CV approach to develop test sets that are very non-similar to the training set and compare different representations.  (Someone in the discussion forum mentioned Tanimoto &lt; 0.40 as a good cutoff for that.)  When the actual key for the public/private LB sets are released, I'll dig into those a bit as well.  </p>\n<p>I doubt that I can solve the generalization problem, but hopefully I can gain a better understanding of the limitations.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2995360,
              "author_name": "wanglinboop",
              "author_url": "",
              "post_date": "09/22/2024 08:26:38",
              "content": "<p>I'm sorry I just saw this comment.</p>\n<p>After two months have passed, have you made some progress on this issue? </p>\n<p>During the month, one of our recent models trained based on intermolecular conformational space similarity achieved good results in some real-world applications (first place in a drug-screening competition through intermolecular similarity).</p>\n<p>Our strategy is to provide a continuous similarity annotation for comparative learning by comparing the shape similarity between all conformations in the conformational space of drug small molecules, and this pre-training strategy helps our <a href=\"https://github.com/Wang-Lin-boop/GeminiMol\" target=\"_blank\">GeminiMol </a>to achieve good generalization performance. </p>\n<p>Although this model did not perform well in this competition, I think it may have some value for researchers interested in using AI for drug discovery. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2995603,
                  "author_name": "kirkdco",
                  "author_url": "",
                  "post_date": "09/22/2024 13:12:54",
                  "content": "<p>Unfortunately, I haven't had time to work on this yet.  Ironically, I've been busy with a project working on leads from a DEL screen. </p>\n<p>Once things settle down a bit, I plan to revisit this. </p>\n<p>Thanks for the link to your paper.  Looks very interesting!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2919870": "Thanks to everyone that has posted a description of their modeling methods.  Reading through them, while there are a variety of modeling methods - embeddings with ChemBerta, MolFormer, etc., bespoke embeddings, XGBoost, 1D-CNN , GNN - there are some very consistent themes with respect to representations.  Specifically, everything I've read so far use the supplied SMILES as input or generated fingerprints using atom identities (ECFP, RDKit, etc.)\n\nOne conclusion from the competition was that generalization to compounds with structures outside the original training remains unsolved.  This really is the challenge in drug discovery - to identify binders from a unique chemical space which is constant changing.\n\nI'm wondering if this lack of generalization is a symptom of the representations used.  SMILES and ECFP/RDKit fingerprints are very useful for structural similarity, but the question of generalization outside the training set structures is inconsistent with structural similarity.  \n\nAs I mentioned in my write up, I approached the problem using 2D-pharmacophores in hopes of finding better generalization.  My result was disappointingly in the middle of the pack, however, the evaluation metric doesn't allow for evaluation of the non-shared building block set or the non-triazine set independent of the shared BB - triazine core set.  (I'm holding out hope that my models show some generalization.)\n\nI'm very curious about the degree of true generalization outside the training set seen with the SMILES or atom identity based fingerprint approaches versus a more abstract approach like 2D-pharmacophores.  I plan to follow up investigating my models and extending to other model types.  I will post the results here once I have them.\n\nThere are a lot of very talented researchers here, and I would very much like to hear the thoughts of others on this idea.",
    "2922358": "We don't seem to need the representations to have extrapolation capabilities. Perhaps we just need to do a round of pre-training on all possible molecules and have each molecule be reasonably well placed in the representation space. This might work better for molecules consisting of just a few building blocks......？\n\nFrankly, I have the same encounter as you in that models that perform well in some common benchmarks (e.g., TDC, MolecularNet, LIT-PCBA, etc.) may not be able to perform well in this competition. I would also like to know what is the reason for this result.",
    "2924895": "Interesting.  \n\nI can see your point about embeddings made with very large datasets that cover a lot of chemical space, but I would still expect a reduced representation like pharmacophores to generalize better.  Maybe that's just my bias.\n\nMy plan over the next week or so is to modify my CV approach to develop test sets that are very non-similar to the training set and compare different representations.  (Someone in the discussion forum mentioned Tanimoto < 0.40 as a good cutoff for that.)  When the actual key for the public/private LB sets are released, I'll dig into those a bit as well.  \n\nI doubt that I can solve the generalization problem, but hopefully I can gain a better understanding of the limitations.",
    "2995360": "I'm sorry I just saw this comment.\n\nAfter two months have passed, have you made some progress on this issue? \n\nDuring the month, one of our recent models trained based on intermolecular conformational space similarity achieved good results in some real-world applications (first place in a drug-screening competition through intermolecular similarity).\n\nOur strategy is to provide a continuous similarity annotation for comparative learning by comparing the shape similarity between all conformations in the conformational space of drug small molecules, and this pre-training strategy helps our [GeminiMol ](https://github.com/Wang-Lin-boop/GeminiMol)to achieve good generalization performance. \n\nAlthough this model did not perform well in this competition, I think it may have some value for researchers interested in using AI for drug discovery.",
    "2995603": "Unfortunately, I haven't had time to work on this yet.  Ironically, I've been busy with a project working on leads from a DEL screen. \n\nOnce things settle down a bit, I plan to revisit this. \n\nThanks for the link to your paper.  Looks very interesting!"
  },
  "source": "meta"
}