{
  "id": 508613,
  "title": "how to deal with non-triazine core?",
  "url": "/competitions/leash-BELKA/discussion/508613",
  "author_name": "",
  "post_date": "2024-05-30T08:35:04.360618400Z",
  "votes": 10,
  "comment_count": 23,
  "views": 0,
  "content": "<p>last time i made a bug, train with \"linker [Dy] replaced by C\" and infer with \"[Dy] only\". The model basically collapses and local CV prediction precision drops significantly from 0.65 to 0.22 (which surprises me …  just one atom can make such large difference).</p>\n<p>now if we train with triazine core and test with non-triazine, i don't think the model can perform at all. Anyone has suggestion on this? </p>\n<p>At first i I though of random masking (or mixup) as augmentation, but this is difficult because scaffold itself modify the binding affinity. </p>",
  "messages": [
    {
      "id": "2844800",
      "postDate": "05/30/2024 08:35:04",
      "content": "<p>last time i made a bug, train with \"linker [Dy] replaced by C\" and infer with \"[Dy] only\". The model basically collapses and local CV prediction precision drops significantly from 0.65 to 0.22 (which surprises me …  just one atom can make such large difference).</p>\n<p>now if we train with triazine core and test with non-triazine, i don't think the model can perform at all. Anyone has suggestion on this? </p>\n<p>At first i I though of random masking (or mixup) as augmentation, but this is difficult because scaffold itself modify the binding affinity. </p>",
      "rawMarkdown": "last time i made a bug, train with \"linker [Dy] replaced by C\" and infer with \"[Dy] only\". The model basically collapses and local CV prediction precision drops significantly from 0.65 to 0.22 (which surprises me ...  just one atom can make such large difference).\n\nnow if we train with triazine core and test with non-triazine, i don't think the model can perform at all. Anyone has suggestion on this? \n\nAt first i I though of random masking (or mixup) as augmentation, but this is difficult because scaffold itself modify the binding affinity.",
      "votes": null
    },
    {
      "id": "2844875",
      "postDate": "05/30/2024 09:07:48",
      "content": "<p>Sounds like your model never see [Dy] while training.Maybe its initialization is not at proper scale?</p>",
      "rawMarkdown": "Sounds like your model never see [Dy] while training.Maybe its initialization is not at proper scale?",
      "votes": null
    },
    {
      "id": "2844985",
      "postDate": "05/30/2024 10:22:40",
      "content": "<p>\"Sounds like your model never see [Dy] while training.\"<br>\nThai is correct.</p>\n<p>But [Dy] is not so important. it is the triazine verus other scaffold that i am worried.</p>",
      "rawMarkdown": "\"Sounds like your model never see [Dy] while training.\"\nThai is correct.\n\nBut [Dy] is not so important. it is the triazine verus other scaffold that i am worried.",
      "votes": null
    },
    {
      "id": "2845113",
      "postDate": "05/30/2024 11:56:51",
      "content": "<p>this group should not play crucial role in the whole molecule, otherwise it will not be used as the linker. </p>",
      "rawMarkdown": "this group should not play crucial role in the whole molecule, otherwise it will not be used as the linker.",
      "votes": null
    },
    {
      "id": "2846166",
      "postDate": "05/31/2024 01:42:31",
      "content": "<p>What representation are you using? RDKit assigns <code>[Dy]</code> an atomic number of 66, which it interprets as Dysprosium. This will skew many chemical property features. Fingerprints will also have a number of bits changed.</p>",
      "rawMarkdown": "What representation are you using? RDKit assigns `[Dy]` an atomic number of 66, which it interprets as Dysprosium. This will skew many chemical property features. Fingerprints will also have a number of bits changed.",
      "votes": null
    },
    {
      "id": "2846167",
      "postDate": "05/31/2024 01:43:31",
      "content": "<p>The linker and DNA tag can have a significant effect on the binding assay result. The linker/DNA tag can participate in binding or block binding from occurring, leading to both false positives and false negatives. It's a trade-off for the throughput of DEL screening</p>",
      "rawMarkdown": "The linker and DNA tag can have a significant effect on the binding assay result. The linker/DNA tag can participate in binding or block binding from occurring, leading to both false positives and false negatives. It's a trade-off for the throughput of DEL screening",
      "votes": null
    },
    {
      "id": "2846441",
      "postDate": "05/31/2024 05:50:55",
      "content": "<p>you misunderstand it, the author means the triazine core,not the DNA linker</p>",
      "rawMarkdown": "you misunderstand it, the author means the triazine core,not the DNA linker",
      "votes": null
    },
    {
      "id": "2846568",
      "postDate": "05/31/2024 06:37:50",
      "content": "<blockquote>\n  <p>This will skew many chemical property features</p>\n</blockquote>\n<p>yes, but skew compared to <em>what</em>? We don't have intact molecules w/o linker and targets for such molecules. Replacement of [Dy] by methyl doesn't help, because linker is neither methyl nor Dysprosium.</p>",
      "rawMarkdown": ">This will skew many chemical property features\n\nyes, but skew compared to *what*? We don't have intact molecules w/o linker and targets for such molecules. Replacement of [Dy] by methyl doesn't help, because linker is neither methyl nor Dysprosium.",
      "votes": null
    },
    {
      "id": "2846867",
      "postDate": "05/31/2024 09:53:12",
      "content": "<p>I'd pre-train the model (whatever it is NLP-like, GNN, etc) on a larger set of non-only-triazine inputs (ZINC with 1.8 billion of drug-like molecules). I also plan to test if \"bioactivity\" and \"biogenic\" flags from ZINC could be helpful - I believe they should correlate with binding affinity to a large extent, which gives about 250K of positives with a diverse set of \"cores\". That's the plan, reality might go other way, of course…</p>",
      "rawMarkdown": "I'd pre-train the model (whatever it is NLP-like, GNN, etc) on a larger set of non-only-triazine inputs (ZINC with 1.8 billion of drug-like molecules). I also plan to test if \"bioactivity\" and \"biogenic\" flags from ZINC could be helpful - I believe they should correlate with binding affinity to a large extent, which gives about 250K of positives with a diverse set of \"cores\". That's the plan, reality might go other way, of course...",
      "votes": null
    },
    {
      "id": "2846916",
      "postDate": "05/31/2024 10:48:40",
      "content": "<p>Good to know, back to the start. It's seems I missed to many chemistry classes back in high school.  :)</p>",
      "rawMarkdown": "Good to know, back to the start. It's seems I missed to many chemistry classes back in high school.  :)",
      "votes": null
    },
    {
      "id": "2847613",
      "postDate": "05/31/2024 15:52:06",
      "content": "<p>In this case, skewed compared to the training data. If you train a model using molecular property descriptors on data with <code>[Dy] -&gt; C</code>, then run inference on data with <code>[Dy]</code>, your inference data will have very different properties due to the <code>[Dy]</code> being read as Dysprosium.</p>\n<p>To your broader question, we would need to know how long the barcodes are and what linker was used to fuse the barcode to the molecules. In general DEL barcodes tend to be around 10-20 bp, but can go up to 100 bp. Even at 10 bp, the barcode is ~6100 g/mol which is substantially heavier than the small molecule itself. Linkers are usually 2-5 ethylene glycol units</p>",
      "rawMarkdown": "In this case, skewed compared to the training data. If you train a model using molecular property descriptors on data with `[Dy] -> C`, then run inference on data with `[Dy]`, your inference data will have very different properties due to the `[Dy]` being read as Dysprosium.\n\nTo your broader question, we would need to know how long the barcodes are and what linker was used to fuse the barcode to the molecules. In general DEL barcodes tend to be around 10-20 bp, but can go up to 100 bp. Even at 10 bp, the barcode is ~6100 g/mol which is substantially heavier than the small molecule itself. Linkers are usually 2-5 ethylene glycol units",
      "votes": null
    },
    {
      "id": "2848155",
      "postDate": "05/31/2024 20:11:14",
      "content": "<p>If I may ask - what is the right way of removing linker (to get the right outputs from rdkit) - just drop the \"[Dy]\" fragment from the string or replace it with \"C\" or something more complex?</p>\n<p>PS - problem solved with this notebook from <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> <a href=\"https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations\" target=\"_blank\">https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations</a></p>",
      "rawMarkdown": "If I may ask - what is the right way of removing linker (to get the right outputs from rdkit) - just drop the \"[Dy]\" fragment from the string or replace it with \"C\" or something more complex?\n\nPS - problem solved with this notebook from @chemdatafarmer https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations",
      "votes": null
    },
    {
      "id": "2848170",
      "postDate": "05/31/2024 20:45:23",
      "content": "<p>For a simple replacement like <code>C</code> or <code>[H]</code>, this is sufficient:</p>\n<pre><code>def replace:\n    smile = smile.replace('', repl)\n    smile = Chem.\n    return smile\n</code></pre>\n<p>For replacing with a more complex structure, you want something like this:</p>\n<pre><code> ():\n\n    mol = Chem.MolFromSmiles(input_smile)\n    repl_mol = Chem.MolFromSmiles(replacement_smile)\n\n    pattern = Chem.MolFromSmarts()\n     mol.HasSubstructMatch(pattern) \n\n    updated_mol = Chem.ReplaceSubstructs(mol, \n                                         pattern, \n                                         repl_mol, \n                                         replaceAll=, \n                                         replacementConnectionPoint=connection_point)\n\n     (updated_mol) ==  \n\n    output = updated_mol[]\n     return_smile:\n        output = Chem.MolToSmiles(output)\n\n     output\n</code></pre>\n<p>In the code above, the <code>connection_point</code> argument determines what atom in the replacement structure is used as the attachment point </p>",
      "rawMarkdown": "For a simple replacement like `C` or `[H]`, this is sufficient:\n\n```\ndef replace_dy_simple(smile, repl):\n    smile = smile.replace('[Dy]', repl)\n    smile = Chem.CanonSmiles(smile)\n    return smile\n```\n\nFor replacing with a more complex structure, you want something like this:\n\n```\ndef replace_dy(\n                input_smile, # input smile with `[Dy]`\n                replacement_smile, # smile to replace `[Dy]`\n                connection_point=0, # determines which atom in `replacement_smile` is the attachment point\n                return_smile=True # if True, returns string, else returns RDKit Mol object\n            ):\n    \n    mol = Chem.MolFromSmiles(input_smile)\n    repl_mol = Chem.MolFromSmiles(replacement_smile)\n    \n    pattern = Chem.MolFromSmarts('[Dy]')\n    assert mol.HasSubstructMatch(pattern) # check pattern is in molecule\n    \n    updated_mol = Chem.ReplaceSubstructs(mol, \n                                         pattern, \n                                         repl_mol, \n                                         replaceAll=True, \n                                         replacementConnectionPoint=connection_point)\n    \n    assert len(updated_mol) == 1 # this should only result in one output\n    \n    output = updated_mol[0]\n    if return_smile:\n        output = Chem.MolToSmiles(output)\n        \n    return output\n```\n\nIn the code above, the `connection_point` argument determines what atom in the replacement structure is used as the attachment point",
      "votes": null
    },
    {
      "id": "2848365",
      "postDate": "06/01/2024 02:15:35",
      "content": "<p>this post is actually related to augmentation.</p>\n<p>now assume</p>\n<pre><code>   abcd\npos    pppabcd :    \nneg    nnnabcd :    \n</code></pre>\n<p>`<br>\nthen in testing</p>\n<pre><code>non triazine core:\n\npos test samples = pppxxxx :  score = \nneg test samples = nnnxxxx :  score = \n\nxxx was  seen  training,  it appears  testing the model don know what to \n</code></pre>\n<p>it is like training an autonomous vehicle to avoid collision of animals on road. so it is trained with dog, cat, cow, etc ….</p>\n<p>then one day on road, it sees a a man dressed up in dinosaurs costume for a party. will it collide?</p>",
      "rawMarkdown": "this post is actually related to augmentation.\n\nnow assume\n```\ntriazine smiles = abcd\npos train samples = pppabcd : prediction score = 0.90\nneg train samples = nnnabcd : prediction score = 0.90\n\n````\nthen in testing\n```\nnon triazine core:\n\npos test samples = pppxxxx : prediction score = 0.50\nneg test samples = nnnxxxx : prediction score = 0.50\n\nxxx was not seen in training, when it appears in testing the model don't know what to do\n```\n\nit is like training an autonomous vehicle to avoid collision of animals on road. so it is trained with dog, cat, cow, etc ....\n\nthen one day on road, it sees a a man dressed up in dinosaurs costume for a party. will it collide?",
      "votes": null
    },
    {
      "id": "2848887",
      "postDate": "06/01/2024 09:06:02",
      "content": "<p>Thanks for that. FYI - I came across a strange model  behaviour, it feels like it could be somehow connected to SMILES nuances. Would appreciate your thoughts on that. To make a long story [not so] short:</p>\n<p>1) I've trained an encoder on SMILES from the ZINC dataset 1.8 billion molecules from the drug-like space. <br>\nThe data is shuffled within AND across tranches (no bias/shift between epochs during the training). Nothing special about encoder - it predicts 15% of masked tokens - BERT classics. It reached some 92% accuracy and 98% mAP, and the training is still on - the model has seen less then 5% of the total ZINC dataset so far.</p>\n<p>2) The interesting part - when evaluated  the BELKA dataset, the model has significantly lower accuracy - 75-77%. This is somewhat counterintuitive since ZINC is a much more diverse dataset. I first thought the problem is [Dy] token, which was not present in ZINC (that's the reason for my earlier question). But things has not changed much after I replaced [Dy] with a different structure using your algorithm. </p>\n<p>3) I've also evaluated the model on a different small-molecule-drug candidates dataset from <a href=\"https://doi.org/10.26434/chemrxiv-2023-pq197\" target=\"_blank\">https://doi.org/10.26434/chemrxiv-2023-pq197</a> brought by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> some time ago - it works exactly as one would expect - with some 92% accuracy and 98% mAP, which makes me think there's something related to BELKA SMILES strings here… some different versions, modifications, else…? Could you spot something special here except for DNA linker?</p>\n<p>I think this also somewhat relates to the earlier question by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> too - if you want the SMILES-based model to generalise on non-triazine data, you need to do some pre-training on some larger and more diverse datasets, but they do not connect well now… </p>",
      "rawMarkdown": "Thanks for that. FYI - I came across a strange model  behaviour, it feels like it could be somehow connected to SMILES nuances. Would appreciate your thoughts on that. To make a long story [not so] short:\n\n1) I've trained an encoder on SMILES from the ZINC dataset 1.8 billion molecules from the drug-like space. \nThe data is shuffled within AND across tranches (no bias/shift between epochs during the training). Nothing special about encoder - it predicts 15% of masked tokens - BERT classics. It reached some 92% accuracy and 98% mAP, and the training is still on - the model has seen less then 5% of the total ZINC dataset so far.\n\n2) The interesting part - when evaluated  the BELKA dataset, the model has significantly lower accuracy - 75-77%. This is somewhat counterintuitive since ZINC is a much more diverse dataset. I first thought the problem is [Dy] token, which was not present in ZINC (that's the reason for my earlier question). But things has not changed much after I replaced [Dy] with a different structure using your algorithm. \n\n3) I've also evaluated the model on a different small-molecule-drug candidates dataset from https://doi.org/10.26434/chemrxiv-2023-pq197 brought by @hengck23 and @chemdatafarmer some time ago - it works exactly as one would expect - with some 92% accuracy and 98% mAP, which makes me think there's something related to BELKA SMILES strings here... some different versions, modifications, else...? Could you spot something special here except for DNA linker?\n\nI think this also somewhat relates to the earlier question by @hengck23 too - if you want the SMILES-based model to generalise on non-triazine data, you need to do some pre-training on some larger and more diverse datasets, but they do not connect well now...",
      "votes": null
    },
    {
      "id": "2849049",
      "postDate": "06/01/2024 11:20:36",
      "content": "<p>this site may give you workflow of some COMMERICAL sucessful methods:</p>\n<p>CRITICAL ASSESSMENT OF COMPUTATIONAL HIT-FINDING EXPERIMENTS<br>\n<a href=\"https://cache-challenge.org/\" target=\"_blank\">https://cache-challenge.org/</a></p>\n<p>e.g Challenge #1 : PREDICT HITS FOR THE WDR DOMAIN OF LRRK2<br>\n<a href=\"https://cache-challenge.org/challenges/predict-hits-for-the-wdr-domain-of-lrrk2/computational-methods\" target=\"_blank\">https://cache-challenge.org/challenges/predict-hits-for-the-wdr-domain-of-lrrk2/computational-methods</a><br>\n<a href=\"https://cache-challenge.org/results-cache-challenge-1\" target=\"_blank\">https://cache-challenge.org/results-cache-challenge-1</a></p>\n<p>e.g. Challenge #3 : Finding ligands targeting the macrodomain of SARS-CoV-2 Nsp3<br>\n<a href=\"https://cache-challenge.org/challenges/finding-ligands-targeting-the-macrodomain-of-sars-cov-2-nsp3/computational-methods\" target=\"_blank\">https://cache-challenge.org/challenges/finding-ligands-targeting-the-macrodomain-of-sars-cov-2-nsp3/computational-methods</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F79b9aadca971eb0f823e9258840e112a%2FSelection_172.png?generation=1717241682655990&amp;alt=media\"></p>",
      "rawMarkdown": "this site may give you workflow of some COMMERICAL sucessful methods:\n\nCRITICAL ASSESSMENT OF COMPUTATIONAL HIT-FINDING EXPERIMENTS\nhttps://cache-challenge.org/\n\ne.g Challenge #1 : PREDICT HITS FOR THE WDR DOMAIN OF LRRK2\nhttps://cache-challenge.org/challenges/predict-hits-for-the-wdr-domain-of-lrrk2/computational-methods\nhttps://cache-challenge.org/results-cache-challenge-1\n\ne.g. Challenge #3 : Finding ligands targeting the macrodomain of SARS-CoV-2 Nsp3\nhttps://cache-challenge.org/challenges/finding-ligands-targeting-the-macrodomain-of-sars-cov-2-nsp3/computational-methods\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F79b9aadca971eb0f823e9258840e112a%2FSelection_172.png?generation=1717241682655990&alt=media)",
      "votes": null
    },
    {
      "id": "2849086",
      "postDate": "06/01/2024 11:53:10",
      "content": "<p>one solution is to use tools like this as augmentation<br>\n<a href=\"https://icml-compbio.github.io/2023/papers/WCBICML2023_paper69.pdf\" target=\"_blank\">https://icml-compbio.github.io/2023/papers/WCBICML2023_paper69.pdf</a><br>\nDiffHopp: A Graph Diffusion Model for Novel Drug Design via Scaffold Hopping</p>",
      "rawMarkdown": "one solution is to use tools like this as augmentation\nhttps://icml-compbio.github.io/2023/papers/WCBICML2023_paper69.pdf\nDiffHopp: A Graph Diffusion Model for Novel Drug Design via Scaffold Hopping",
      "votes": null
    },
    {
      "id": "2849672",
      "postDate": "06/01/2024 17:41:46",
      "content": "<p>I've seen similar results with an embedding model trained on ZINC having high loss on BELKA compounds, even with the <code>[Dy]</code> removed. My guess is this is mainly driven by molecular weight differences. The majority of ZINC is &lt;500 g/mol, while the majority of BELKA is &gt;500g/mol</p>",
      "rawMarkdown": "I've seen similar results with an embedding model trained on ZINC having high loss on BELKA compounds, even with the `[Dy]` removed. My guess is this is mainly driven by molecular weight differences. The majority of ZINC is <500 g/mol, while the majority of BELKA is >500g/mol",
      "votes": null
    },
    {
      "id": "2849997",
      "postDate": "06/01/2024 21:07:12",
      "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> </p>\n<p>thanks for your experimental results.<br>\nhere are my comments:</p>\n<ol>\n<li><p>external data: even if i train from scratch (e.g. simple ecfp+DNN, conv1d) performance of external data is much better than kaggle dataset. I conclude external data is a much easier dataset.</p></li>\n<li><p>pretraining with ZINC. my suggestion:</p></li>\n</ol>\n<ul>\n<li>you should use a pretrained model (e.g.ZINC) to project kaggle samples to the pretrained embedding. hopefully this pretrained embedding is better than generic fingerprint Tanimoto, rdkit desriptors, … e.g. you can test by knn retrival or visualise via tsne or measure pairwise distance.</li>\n</ul>\n<ol>\n<li>if pretraining doesn't work, then SSL on test (or validation) set is the next bet. </li>\n</ol>",
      "rawMarkdown": "victorshlepov \n\nthanks for your experimental results.\nhere are my comments:\n1. external data: even if i train from scratch (e.g. simple ecfp+DNN, conv1d) performance of external data is much better than kaggle dataset. I conclude external data is a much easier dataset.\n\n2. pretraining with ZINC. my suggestion:\n- you should use a pretrained model (e.g.ZINC) to project kaggle samples to the pretrained embedding. hopefully this pretrained embedding is better than generic fingerprint Tanimoto, rdkit desriptors, ... e.g. you can test by knn retrival or visualise via tsne or measure pairwise distance.\n\n\n3. if pretraining doesn't work, then SSL on test (or validation) set is the next bet.",
      "votes": null
    },
    {
      "id": "2851370",
      "postDate": "06/02/2024 18:01:47",
      "content": "<p>On ZINK pre-training:</p>\n<p>Yep, that's exactly the plan. I hope that representations learned from a more diverse, but still applicable (\"drug-like\")  Zinc set would be beneficial for predictions of non-triazine molecules.</p>\n<p>I fine-tuned my Zinc-based MLM on a mixed Zinc/Belka data (Belka samples hare are from the test set - 800K molecules - it has no bias towards triazine cores, with a share of Belka samples about 1%, say 5 samples out of 512 per batch) - it solved the issue. I guess the problem was with just a few tokens which are extremely rare in Zinc - \"7\", \"8\" and \"9\" (I guess this somehow correlates with molecular weights mentioned by <a href=\"https://www.kaggle.com/towardsentropy\" target=\"_blank\">@towardsentropy</a>). The model now now yields results above the train metrics for MLM task for the Belka train set (100+ million unseen molecules, but seen blocks) - this is something I was expecting from the start.</p>\n<p>On external data and complexity:</p>\n<p>Given facts above - I'm not so sure Belka is more complex - maybe it's just a few tokens that cause the prob. At the end of the day - Belka is a very repetitive combination of building blocks and a single core.</p>",
      "rawMarkdown": "On ZINK pre-training:\n\nYep, that's exactly the plan. I hope that representations learned from a more diverse, but still applicable (\"drug-like\")  Zinc set would be beneficial for predictions of non-triazine molecules.\n \nI fine-tuned my Zinc-based MLM on a mixed Zinc/Belka data (Belka samples hare are from the test set - 800K molecules - it has no bias towards triazine cores, with a share of Belka samples about 1%, say 5 samples out of 512 per batch) - it solved the issue. I guess the problem was with just a few tokens which are extremely rare in Zinc - \"7\", \"8\" and \"9\" (I guess this somehow correlates with molecular weights mentioned by @towardsentropy). The model now now yields results above the train metrics for MLM task for the Belka train set (100+ million unseen molecules, but seen blocks) - this is something I was expecting from the start.\n\nOn external data and complexity:\n\nGiven facts above - I'm not so sure Belka is more complex - maybe it's just a few tokens that cause the prob. At the end of the day - Belka is a very repetitive combination of building blocks and a single core.",
      "votes": null
    },
    {
      "id": "2851618",
      "postDate": "06/02/2024 20:44:08",
      "content": "<p>Numbers in SMILES strings denote opening/closing of ring systems. <code>1</code> denotes the first ring system, <code>2</code> for the second, and so on. This means <code>7</code>, <code>8</code>, and <code>9</code> will only appear in molecules that have that many ring systems. Number of rings correlates broadly with size, so you would only expect to see these tokens on large, heavy molecules.</p>",
      "rawMarkdown": "Numbers in SMILES strings denote opening/closing of ring systems. `1` denotes the first ring system, `2` for the second, and so on. This means `7`, `8`, and `9` will only appear in molecules that have that many ring systems. Number of rings correlates broadly with size, so you would only expect to see these tokens on large, heavy molecules.",
      "votes": null
    },
    {
      "id": "2851633",
      "postDate": "06/02/2024 21:03:08",
      "content": "<p>Seems to prove you earlier point, right? 7,8, and 9 appear just a few times per 10 million molecules in a \"drug-like\" ZINC dataset - way to infrequent to learn anything useful about these tokens. I should have processed the whole ZINC, i guess, not just \"drug-like\" to pick these tails…</p>",
      "rawMarkdown": "Seems to prove you earlier point, right? 7,8, and 9 appear just a few times per 10 million molecules in a \"drug-like\" ZINC dataset - way to infrequent to learn anything useful about these tokens. I should have processed the whole ZINC, i guess, not just \"drug-like\" to pick these tails...",
      "votes": null
    },
    {
      "id": "2852826",
      "postDate": "06/03/2024 13:26:44",
      "content": "<p>this is just an idea … i still have to figure out how to use it</p>\n<ol>\n<li>instead of one input to model, we can use two model=(core-bb1-bb2-bb3,core-aa1-aa2-aa3)</li>\n<li>here, we are measuring relative score due to difference of the two inputs: how \"bb1-bb2-bb3\" is better than \"aa1-aa2-aa3\"</li>\n<li>hence even if the scaffold core is different, it doesn't matter (! or ?)</li>\n</ol>",
      "rawMarkdown": "this is just an idea ... i still have to figure out how to use it\n1. instead of one input to model, we can use two model=(core-bb1-bb2-bb3,core-aa1-aa2-aa3)\n2. here, we are measuring relative score due to difference of the two inputs: how \"bb1-bb2-bb3\" is better than \"aa1-aa2-aa3\"\n3. hence even if the scaffold core is different, it doesn't matter (! or ?)",
      "votes": null
    },
    {
      "id": "2870741",
      "postDate": "06/13/2024 18:41:30",
      "content": "<p>You can treat triazine+bb1 as bb1. Then for non-triazine core, bb1 already includes core. (Per analysis and per organizer write-up when they changed the scoring)</p>\n<p>So all bb1s in test are novel and different from train, but you are already dealing with same issue even in triazine core with non-shared BB. </p>\n<p>But you still have to see if there's a good way to predict the impact of novel BB substitutions. </p>",
      "rawMarkdown": "You can treat triazine+bb1 as bb1. Then for non-triazine core, bb1 already includes core. (Per analysis and per organizer write-up when they changed the scoring)\n\nSo all bb1s in test are novel and different from train, but you are already dealing with same issue even in triazine core with non-shared BB. \n\nBut you still have to see if there's a good way to predict the impact of novel BB substitutions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2844875,
      "author_name": "w5833946",
      "author_url": "",
      "post_date": "05/30/2024 09:07:48",
      "content": "<p>Sounds like your model never see [Dy] while training.Maybe its initialization is not at proper scale?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2844985,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/30/2024 10:22:40",
          "content": "<p>\"Sounds like your model never see [Dy] while training.\"<br>\nThai is correct.</p>\n<p>But [Dy] is not so important. it is the triazine verus other scaffold that i am worried.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2845113,
      "author_name": "henryzhaowong",
      "author_url": "",
      "post_date": "05/30/2024 11:56:51",
      "content": "<p>this group should not play crucial role in the whole molecule, otherwise it will not be used as the linker. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2846167,
          "author_name": "towardsentropy",
          "author_url": "",
          "post_date": "05/31/2024 01:43:31",
          "content": "<p>The linker and DNA tag can have a significant effect on the binding assay result. The linker/DNA tag can participate in binding or block binding from occurring, leading to both false positives and false negatives. It's a trade-off for the throughput of DEL screening</p>",
          "votes": null,
          "replies": [
            {
              "id": 2846441,
              "author_name": "henryzhaowong",
              "author_url": "",
              "post_date": "05/31/2024 05:50:55",
              "content": "<p>you misunderstand it, the author means the triazine core,not the DNA linker</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2846166,
      "author_name": "towardsentropy",
      "author_url": "",
      "post_date": "05/31/2024 01:42:31",
      "content": "<p>What representation are you using? RDKit assigns <code>[Dy]</code> an atomic number of 66, which it interprets as Dysprosium. This will skew many chemical property features. Fingerprints will also have a number of bits changed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2846568,
          "author_name": "ogurtsov",
          "author_url": "",
          "post_date": "05/31/2024 06:37:50",
          "content": "<blockquote>\n  <p>This will skew many chemical property features</p>\n</blockquote>\n<p>yes, but skew compared to <em>what</em>? We don't have intact molecules w/o linker and targets for such molecules. Replacement of [Dy] by methyl doesn't help, because linker is neither methyl nor Dysprosium.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2847613,
              "author_name": "towardsentropy",
              "author_url": "",
              "post_date": "05/31/2024 15:52:06",
              "content": "<p>In this case, skewed compared to the training data. If you train a model using molecular property descriptors on data with <code>[Dy] -&gt; C</code>, then run inference on data with <code>[Dy]</code>, your inference data will have very different properties due to the <code>[Dy]</code> being read as Dysprosium.</p>\n<p>To your broader question, we would need to know how long the barcodes are and what linker was used to fuse the barcode to the molecules. In general DEL barcodes tend to be around 10-20 bp, but can go up to 100 bp. Even at 10 bp, the barcode is ~6100 g/mol which is substantially heavier than the small molecule itself. Linkers are usually 2-5 ethylene glycol units</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2848155,
                  "author_name": "victorshlepov",
                  "author_url": "",
                  "post_date": "05/31/2024 20:11:14",
                  "content": "<p>If I may ask - what is the right way of removing linker (to get the right outputs from rdkit) - just drop the \"[Dy]\" fragment from the string or replace it with \"C\" or something more complex?</p>\n<p>PS - problem solved with this notebook from <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> <a href=\"https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations\" target=\"_blank\">https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations</a></p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2848170,
                      "author_name": "towardsentropy",
                      "author_url": "",
                      "post_date": "05/31/2024 20:45:23",
                      "content": "<p>For a simple replacement like <code>C</code> or <code>[H]</code>, this is sufficient:</p>\n<pre><code>def replace:\n    smile = smile.replace('', repl)\n    smile = Chem.\n    return smile\n</code></pre>\n<p>For replacing with a more complex structure, you want something like this:</p>\n<pre><code> ():\n\n    mol = Chem.MolFromSmiles(input_smile)\n    repl_mol = Chem.MolFromSmiles(replacement_smile)\n\n    pattern = Chem.MolFromSmarts()\n     mol.HasSubstructMatch(pattern) \n\n    updated_mol = Chem.ReplaceSubstructs(mol, \n                                         pattern, \n                                         repl_mol, \n                                         replaceAll=, \n                                         replacementConnectionPoint=connection_point)\n\n     (updated_mol) ==  \n\n    output = updated_mol[]\n     return_smile:\n        output = Chem.MolToSmiles(output)\n\n     output\n</code></pre>\n<p>In the code above, the <code>connection_point</code> argument determines what atom in the replacement structure is used as the attachment point </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2848887,
                          "author_name": "victorshlepov",
                          "author_url": "",
                          "post_date": "06/01/2024 09:06:02",
                          "content": "<p>Thanks for that. FYI - I came across a strange model  behaviour, it feels like it could be somehow connected to SMILES nuances. Would appreciate your thoughts on that. To make a long story [not so] short:</p>\n<p>1) I've trained an encoder on SMILES from the ZINC dataset 1.8 billion molecules from the drug-like space. <br>\nThe data is shuffled within AND across tranches (no bias/shift between epochs during the training). Nothing special about encoder - it predicts 15% of masked tokens - BERT classics. It reached some 92% accuracy and 98% mAP, and the training is still on - the model has seen less then 5% of the total ZINC dataset so far.</p>\n<p>2) The interesting part - when evaluated  the BELKA dataset, the model has significantly lower accuracy - 75-77%. This is somewhat counterintuitive since ZINC is a much more diverse dataset. I first thought the problem is [Dy] token, which was not present in ZINC (that's the reason for my earlier question). But things has not changed much after I replaced [Dy] with a different structure using your algorithm. </p>\n<p>3) I've also evaluated the model on a different small-molecule-drug candidates dataset from <a href=\"https://doi.org/10.26434/chemrxiv-2023-pq197\" target=\"_blank\">https://doi.org/10.26434/chemrxiv-2023-pq197</a> brought by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> some time ago - it works exactly as one would expect - with some 92% accuracy and 98% mAP, which makes me think there's something related to BELKA SMILES strings here… some different versions, modifications, else…? Could you spot something special here except for DNA linker?</p>\n<p>I think this also somewhat relates to the earlier question by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> too - if you want the SMILES-based model to generalise on non-triazine data, you need to do some pre-training on some larger and more diverse datasets, but they do not connect well now… </p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2849672,
                              "author_name": "towardsentropy",
                              "author_url": "",
                              "post_date": "06/01/2024 17:41:46",
                              "content": "<p>I've seen similar results with an embedding model trained on ZINC having high loss on BELKA compounds, even with the <code>[Dy]</code> removed. My guess is this is mainly driven by molecular weight differences. The majority of ZINC is &lt;500 g/mol, while the majority of BELKA is &gt;500g/mol</p>",
                              "votes": null,
                              "replies": []
                            },
                            {
                              "id": 2849997,
                              "author_name": "hengck23",
                              "author_url": "",
                              "post_date": "06/01/2024 21:07:12",
                              "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> </p>\n<p>thanks for your experimental results.<br>\nhere are my comments:</p>\n<ol>\n<li><p>external data: even if i train from scratch (e.g. simple ecfp+DNN, conv1d) performance of external data is much better than kaggle dataset. I conclude external data is a much easier dataset.</p></li>\n<li><p>pretraining with ZINC. my suggestion:</p></li>\n</ol>\n<ul>\n<li>you should use a pretrained model (e.g.ZINC) to project kaggle samples to the pretrained embedding. hopefully this pretrained embedding is better than generic fingerprint Tanimoto, rdkit desriptors, … e.g. you can test by knn retrival or visualise via tsne or measure pairwise distance.</li>\n</ul>\n<ol>\n<li>if pretraining doesn't work, then SSL on test (or validation) set is the next bet. </li>\n</ol>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2851370,
                                  "author_name": "victorshlepov",
                                  "author_url": "",
                                  "post_date": "06/02/2024 18:01:47",
                                  "content": "<p>On ZINK pre-training:</p>\n<p>Yep, that's exactly the plan. I hope that representations learned from a more diverse, but still applicable (\"drug-like\")  Zinc set would be beneficial for predictions of non-triazine molecules.</p>\n<p>I fine-tuned my Zinc-based MLM on a mixed Zinc/Belka data (Belka samples hare are from the test set - 800K molecules - it has no bias towards triazine cores, with a share of Belka samples about 1%, say 5 samples out of 512 per batch) - it solved the issue. I guess the problem was with just a few tokens which are extremely rare in Zinc - \"7\", \"8\" and \"9\" (I guess this somehow correlates with molecular weights mentioned by <a href=\"https://www.kaggle.com/towardsentropy\" target=\"_blank\">@towardsentropy</a>). The model now now yields results above the train metrics for MLM task for the Belka train set (100+ million unseen molecules, but seen blocks) - this is something I was expecting from the start.</p>\n<p>On external data and complexity:</p>\n<p>Given facts above - I'm not so sure Belka is more complex - maybe it's just a few tokens that cause the prob. At the end of the day - Belka is a very repetitive combination of building blocks and a single core.</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2851618,
                                      "author_name": "towardsentropy",
                                      "author_url": "",
                                      "post_date": "06/02/2024 20:44:08",
                                      "content": "<p>Numbers in SMILES strings denote opening/closing of ring systems. <code>1</code> denotes the first ring system, <code>2</code> for the second, and so on. This means <code>7</code>, <code>8</code>, and <code>9</code> will only appear in molecules that have that many ring systems. Number of rings correlates broadly with size, so you would only expect to see these tokens on large, heavy molecules.</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 2851633,
                                          "author_name": "victorshlepov",
                                          "author_url": "",
                                          "post_date": "06/02/2024 21:03:08",
                                          "content": "<p>Seems to prove you earlier point, right? 7,8, and 9 appear just a few times per 10 million molecules in a \"drug-like\" ZINC dataset - way to infrequent to learn anything useful about these tokens. I should have processed the whole ZINC, i guess, not just \"drug-like\" to pick these tails…</p>",
                                          "votes": null,
                                          "replies": []
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2846916,
          "author_name": "victorshlepov",
          "author_url": "",
          "post_date": "05/31/2024 10:48:40",
          "content": "<p>Good to know, back to the start. It's seems I missed to many chemistry classes back in high school.  :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2846867,
      "author_name": "victorshlepov",
      "author_url": "",
      "post_date": "05/31/2024 09:53:12",
      "content": "<p>I'd pre-train the model (whatever it is NLP-like, GNN, etc) on a larger set of non-only-triazine inputs (ZINC with 1.8 billion of drug-like molecules). I also plan to test if \"bioactivity\" and \"biogenic\" flags from ZINC could be helpful - I believe they should correlate with binding affinity to a large extent, which gives about 250K of positives with a diverse set of \"cores\". That's the plan, reality might go other way, of course…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2848365,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/01/2024 02:15:35",
      "content": "<p>this post is actually related to augmentation.</p>\n<p>now assume</p>\n<pre><code>   abcd\npos    pppabcd :    \nneg    nnnabcd :    \n</code></pre>\n<p>`<br>\nthen in testing</p>\n<pre><code>non triazine core:\n\npos test samples = pppxxxx :  score = \nneg test samples = nnnxxxx :  score = \n\nxxx was  seen  training,  it appears  testing the model don know what to \n</code></pre>\n<p>it is like training an autonomous vehicle to avoid collision of animals on road. so it is trained with dog, cat, cow, etc ….</p>\n<p>then one day on road, it sees a a man dressed up in dinosaurs costume for a party. will it collide?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2849049,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/01/2024 11:20:36",
      "content": "<p>this site may give you workflow of some COMMERICAL sucessful methods:</p>\n<p>CRITICAL ASSESSMENT OF COMPUTATIONAL HIT-FINDING EXPERIMENTS<br>\n<a href=\"https://cache-challenge.org/\" target=\"_blank\">https://cache-challenge.org/</a></p>\n<p>e.g Challenge #1 : PREDICT HITS FOR THE WDR DOMAIN OF LRRK2<br>\n<a href=\"https://cache-challenge.org/challenges/predict-hits-for-the-wdr-domain-of-lrrk2/computational-methods\" target=\"_blank\">https://cache-challenge.org/challenges/predict-hits-for-the-wdr-domain-of-lrrk2/computational-methods</a><br>\n<a href=\"https://cache-challenge.org/results-cache-challenge-1\" target=\"_blank\">https://cache-challenge.org/results-cache-challenge-1</a></p>\n<p>e.g. Challenge #3 : Finding ligands targeting the macrodomain of SARS-CoV-2 Nsp3<br>\n<a href=\"https://cache-challenge.org/challenges/finding-ligands-targeting-the-macrodomain-of-sars-cov-2-nsp3/computational-methods\" target=\"_blank\">https://cache-challenge.org/challenges/finding-ligands-targeting-the-macrodomain-of-sars-cov-2-nsp3/computational-methods</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F79b9aadca971eb0f823e9258840e112a%2FSelection_172.png?generation=1717241682655990&amp;alt=media\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2849086,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/01/2024 11:53:10",
      "content": "<p>one solution is to use tools like this as augmentation<br>\n<a href=\"https://icml-compbio.github.io/2023/papers/WCBICML2023_paper69.pdf\" target=\"_blank\">https://icml-compbio.github.io/2023/papers/WCBICML2023_paper69.pdf</a><br>\nDiffHopp: A Graph Diffusion Model for Novel Drug Design via Scaffold Hopping</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2852826,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/03/2024 13:26:44",
      "content": "<p>this is just an idea … i still have to figure out how to use it</p>\n<ol>\n<li>instead of one input to model, we can use two model=(core-bb1-bb2-bb3,core-aa1-aa2-aa3)</li>\n<li>here, we are measuring relative score due to difference of the two inputs: how \"bb1-bb2-bb3\" is better than \"aa1-aa2-aa3\"</li>\n<li>hence even if the scaffold core is different, it doesn't matter (! or ?)</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2870741,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "06/13/2024 18:41:30",
          "content": "<p>You can treat triazine+bb1 as bb1. Then for non-triazine core, bb1 already includes core. (Per analysis and per organizer write-up when they changed the scoring)</p>\n<p>So all bb1s in test are novel and different from train, but you are already dealing with same issue even in triazine core with non-shared BB. </p>\n<p>But you still have to see if there's a good way to predict the impact of novel BB substitutions. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2844800": "last time i made a bug, train with \"linker [Dy] replaced by C\" and infer with \"[Dy] only\". The model basically collapses and local CV prediction precision drops significantly from 0.65 to 0.22 (which surprises me ...  just one atom can make such large difference).\n\nnow if we train with triazine core and test with non-triazine, i don't think the model can perform at all. Anyone has suggestion on this? \n\nAt first i I though of random masking (or mixup) as augmentation, but this is difficult because scaffold itself modify the binding affinity.",
    "2844875": "Sounds like your model never see [Dy] while training.Maybe its initialization is not at proper scale?",
    "2844985": "\"Sounds like your model never see [Dy] while training.\"\nThai is correct.\n\nBut [Dy] is not so important. it is the triazine verus other scaffold that i am worried.",
    "2845113": "this group should not play crucial role in the whole molecule, otherwise it will not be used as the linker.",
    "2846166": "What representation are you using? RDKit assigns `[Dy]` an atomic number of 66, which it interprets as Dysprosium. This will skew many chemical property features. Fingerprints will also have a number of bits changed.",
    "2846167": "The linker and DNA tag can have a significant effect on the binding assay result. The linker/DNA tag can participate in binding or block binding from occurring, leading to both false positives and false negatives. It's a trade-off for the throughput of DEL screening",
    "2846441": "you misunderstand it, the author means the triazine core,not the DNA linker",
    "2846568": ">This will skew many chemical property features\n\nyes, but skew compared to *what*? We don't have intact molecules w/o linker and targets for such molecules. Replacement of [Dy] by methyl doesn't help, because linker is neither methyl nor Dysprosium.",
    "2846867": "I'd pre-train the model (whatever it is NLP-like, GNN, etc) on a larger set of non-only-triazine inputs (ZINC with 1.8 billion of drug-like molecules). I also plan to test if \"bioactivity\" and \"biogenic\" flags from ZINC could be helpful - I believe they should correlate with binding affinity to a large extent, which gives about 250K of positives with a diverse set of \"cores\". That's the plan, reality might go other way, of course...",
    "2846916": "Good to know, back to the start. It's seems I missed to many chemistry classes back in high school.  :)",
    "2847613": "In this case, skewed compared to the training data. If you train a model using molecular property descriptors on data with `[Dy] -> C`, then run inference on data with `[Dy]`, your inference data will have very different properties due to the `[Dy]` being read as Dysprosium.\n\nTo your broader question, we would need to know how long the barcodes are and what linker was used to fuse the barcode to the molecules. In general DEL barcodes tend to be around 10-20 bp, but can go up to 100 bp. Even at 10 bp, the barcode is ~6100 g/mol which is substantially heavier than the small molecule itself. Linkers are usually 2-5 ethylene glycol units",
    "2848155": "If I may ask - what is the right way of removing linker (to get the right outputs from rdkit) - just drop the \"[Dy]\" fragment from the string or replace it with \"C\" or something more complex?\n\nPS - problem solved with this notebook from @chemdatafarmer https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations",
    "2848170": "For a simple replacement like `C` or `[H]`, this is sufficient:\n\n```\ndef replace_dy_simple(smile, repl):\n    smile = smile.replace('[Dy]', repl)\n    smile = Chem.CanonSmiles(smile)\n    return smile\n```\n\nFor replacing with a more complex structure, you want something like this:\n\n```\ndef replace_dy(\n                input_smile, # input smile with `[Dy]`\n                replacement_smile, # smile to replace `[Dy]`\n                connection_point=0, # determines which atom in `replacement_smile` is the attachment point\n                return_smile=True # if True, returns string, else returns RDKit Mol object\n            ):\n    \n    mol = Chem.MolFromSmiles(input_smile)\n    repl_mol = Chem.MolFromSmiles(replacement_smile)\n    \n    pattern = Chem.MolFromSmarts('[Dy]')\n    assert mol.HasSubstructMatch(pattern) # check pattern is in molecule\n    \n    updated_mol = Chem.ReplaceSubstructs(mol, \n                                         pattern, \n                                         repl_mol, \n                                         replaceAll=True, \n                                         replacementConnectionPoint=connection_point)\n    \n    assert len(updated_mol) == 1 # this should only result in one output\n    \n    output = updated_mol[0]\n    if return_smile:\n        output = Chem.MolToSmiles(output)\n        \n    return output\n```\n\nIn the code above, the `connection_point` argument determines what atom in the replacement structure is used as the attachment point",
    "2848365": "this post is actually related to augmentation.\n\nnow assume\n```\ntriazine smiles = abcd\npos train samples = pppabcd : prediction score = 0.90\nneg train samples = nnnabcd : prediction score = 0.90\n\n````\nthen in testing\n```\nnon triazine core:\n\npos test samples = pppxxxx : prediction score = 0.50\nneg test samples = nnnxxxx : prediction score = 0.50\n\nxxx was not seen in training, when it appears in testing the model don't know what to do\n```\n\nit is like training an autonomous vehicle to avoid collision of animals on road. so it is trained with dog, cat, cow, etc ....\n\nthen one day on road, it sees a a man dressed up in dinosaurs costume for a party. will it collide?",
    "2848887": "Thanks for that. FYI - I came across a strange model  behaviour, it feels like it could be somehow connected to SMILES nuances. Would appreciate your thoughts on that. To make a long story [not so] short:\n\n1) I've trained an encoder on SMILES from the ZINC dataset 1.8 billion molecules from the drug-like space. \nThe data is shuffled within AND across tranches (no bias/shift between epochs during the training). Nothing special about encoder - it predicts 15% of masked tokens - BERT classics. It reached some 92% accuracy and 98% mAP, and the training is still on - the model has seen less then 5% of the total ZINC dataset so far.\n\n2) The interesting part - when evaluated  the BELKA dataset, the model has significantly lower accuracy - 75-77%. This is somewhat counterintuitive since ZINC is a much more diverse dataset. I first thought the problem is [Dy] token, which was not present in ZINC (that's the reason for my earlier question). But things has not changed much after I replaced [Dy] with a different structure using your algorithm. \n\n3) I've also evaluated the model on a different small-molecule-drug candidates dataset from https://doi.org/10.26434/chemrxiv-2023-pq197 brought by @hengck23 and @chemdatafarmer some time ago - it works exactly as one would expect - with some 92% accuracy and 98% mAP, which makes me think there's something related to BELKA SMILES strings here... some different versions, modifications, else...? Could you spot something special here except for DNA linker?\n\nI think this also somewhat relates to the earlier question by @hengck23 too - if you want the SMILES-based model to generalise on non-triazine data, you need to do some pre-training on some larger and more diverse datasets, but they do not connect well now...",
    "2849049": "this site may give you workflow of some COMMERICAL sucessful methods:\n\nCRITICAL ASSESSMENT OF COMPUTATIONAL HIT-FINDING EXPERIMENTS\nhttps://cache-challenge.org/\n\ne.g Challenge #1 : PREDICT HITS FOR THE WDR DOMAIN OF LRRK2\nhttps://cache-challenge.org/challenges/predict-hits-for-the-wdr-domain-of-lrrk2/computational-methods\nhttps://cache-challenge.org/results-cache-challenge-1\n\ne.g. Challenge #3 : Finding ligands targeting the macrodomain of SARS-CoV-2 Nsp3\nhttps://cache-challenge.org/challenges/finding-ligands-targeting-the-macrodomain-of-sars-cov-2-nsp3/computational-methods\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F79b9aadca971eb0f823e9258840e112a%2FSelection_172.png?generation=1717241682655990&alt=media)",
    "2849086": "one solution is to use tools like this as augmentation\nhttps://icml-compbio.github.io/2023/papers/WCBICML2023_paper69.pdf\nDiffHopp: A Graph Diffusion Model for Novel Drug Design via Scaffold Hopping",
    "2849672": "I've seen similar results with an embedding model trained on ZINC having high loss on BELKA compounds, even with the `[Dy]` removed. My guess is this is mainly driven by molecular weight differences. The majority of ZINC is <500 g/mol, while the majority of BELKA is >500g/mol",
    "2849997": "victorshlepov \n\nthanks for your experimental results.\nhere are my comments:\n1. external data: even if i train from scratch (e.g. simple ecfp+DNN, conv1d) performance of external data is much better than kaggle dataset. I conclude external data is a much easier dataset.\n\n2. pretraining with ZINC. my suggestion:\n- you should use a pretrained model (e.g.ZINC) to project kaggle samples to the pretrained embedding. hopefully this pretrained embedding is better than generic fingerprint Tanimoto, rdkit desriptors, ... e.g. you can test by knn retrival or visualise via tsne or measure pairwise distance.\n\n\n3. if pretraining doesn't work, then SSL on test (or validation) set is the next bet.",
    "2851370": "On ZINK pre-training:\n\nYep, that's exactly the plan. I hope that representations learned from a more diverse, but still applicable (\"drug-like\")  Zinc set would be beneficial for predictions of non-triazine molecules.\n \nI fine-tuned my Zinc-based MLM on a mixed Zinc/Belka data (Belka samples hare are from the test set - 800K molecules - it has no bias towards triazine cores, with a share of Belka samples about 1%, say 5 samples out of 512 per batch) - it solved the issue. I guess the problem was with just a few tokens which are extremely rare in Zinc - \"7\", \"8\" and \"9\" (I guess this somehow correlates with molecular weights mentioned by @towardsentropy). The model now now yields results above the train metrics for MLM task for the Belka train set (100+ million unseen molecules, but seen blocks) - this is something I was expecting from the start.\n\nOn external data and complexity:\n\nGiven facts above - I'm not so sure Belka is more complex - maybe it's just a few tokens that cause the prob. At the end of the day - Belka is a very repetitive combination of building blocks and a single core.",
    "2851618": "Numbers in SMILES strings denote opening/closing of ring systems. `1` denotes the first ring system, `2` for the second, and so on. This means `7`, `8`, and `9` will only appear in molecules that have that many ring systems. Number of rings correlates broadly with size, so you would only expect to see these tokens on large, heavy molecules.",
    "2851633": "Seems to prove you earlier point, right? 7,8, and 9 appear just a few times per 10 million molecules in a \"drug-like\" ZINC dataset - way to infrequent to learn anything useful about these tokens. I should have processed the whole ZINC, i guess, not just \"drug-like\" to pick these tails...",
    "2852826": "this is just an idea ... i still have to figure out how to use it\n1. instead of one input to model, we can use two model=(core-bb1-bb2-bb3,core-aa1-aa2-aa3)\n2. here, we are measuring relative score due to difference of the two inputs: how \"bb1-bb2-bb3\" is better than \"aa1-aa2-aa3\"\n3. hence even if the scaffold core is different, it doesn't matter (! or ?)",
    "2870741": "You can treat triazine+bb1 as bb1. Then for non-triazine core, bb1 already includes core. (Per analysis and per organizer write-up when they changed the scoring)\n\nSo all bb1s in test are novel and different from train, but you are already dealing with same issue even in triazine core with non-shared BB. \n\nBut you still have to see if there's a good way to predict the impact of novel BB substitutions."
  },
  "source": "meta"
}