{
  "id": 499890,
  "title": "did anyone try apply block-aware models to validation dataset?",
  "url": "/competitions/leash-BELKA/discussion/499890",
  "author_name": "",
  "post_date": "2024-05-03T12:22:36.466558100Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>my guess could be correct. my experimental results and paper reading shows:</p>\n<ol>\n<li><p>molecule space is not smooth. small change in SMILES can give very different molecular property.</p></li>\n<li><p>learned model  don't gives good results if  test building block is not in training</p></li>\n<li><p>wide neural nets work better than deep ones. i think the are just memorizing molecule motifs/fragments in training.<br>\nhence 98 million samples is NOT a lot, since they cannot represent unseen block. <br>\n(there is probably no difference to the nonshare block test samples  if you are training 10m, 30m, 50m or 100m)</p></li>\n<li><p>hence i carry out the below experiment:</p></li>\n</ol>\n<p>1) divide 271 blocking blocks 1 into 5 group (each approximately 54)<br>\n2) split train smiles into 5 folds, each fold consists of smiles synthesized from each of the group above (i.e. each fold only has 54 blocking blocks 1)<br>\n3) during validation, we only apply model blockwise, i.e. use n-th model if and only if test smile is in n-th group</p>\n<p>here are the results for xgboost + ecfp:</p>\n<pre><code> models groupwise/blockwise:\n         \n(micro,, , )\n\n models by averaging\n        \n\n models by taking \n        \n\n\nreference,  previous experiments, i   split  micro ap  is also about  \n</code></pre>\n<p>implications:</p>\n<ul>\n<li>if your model is complex, e.g. deep GNN or docking, etc … you can just make model for some blocks.</li>\n<li>parallel training</li>\n</ul>",
  "messages": [
    {
      "id": "2790996",
      "postDate": "05/03/2024 12:22:36",
      "content": "<p>my guess could be correct. my experimental results and paper reading shows:</p>\n<ol>\n<li><p>molecule space is not smooth. small change in SMILES can give very different molecular property.</p></li>\n<li><p>learned model  don't gives good results if  test building block is not in training</p></li>\n<li><p>wide neural nets work better than deep ones. i think the are just memorizing molecule motifs/fragments in training.<br>\nhence 98 million samples is NOT a lot, since they cannot represent unseen block. <br>\n(there is probably no difference to the nonshare block test samples  if you are training 10m, 30m, 50m or 100m)</p></li>\n<li><p>hence i carry out the below experiment:</p></li>\n</ol>\n<p>1) divide 271 blocking blocks 1 into 5 group (each approximately 54)<br>\n2) split train smiles into 5 folds, each fold consists of smiles synthesized from each of the group above (i.e. each fold only has 54 blocking blocks 1)<br>\n3) during validation, we only apply model blockwise, i.e. use n-th model if and only if test smile is in n-th group</p>\n<p>here are the results for xgboost + ecfp:</p>\n<pre><code> models groupwise/blockwise:\n         \n(micro,, , )\n\n models by averaging\n        \n\n models by taking \n        \n\n\nreference,  previous experiments, i   split  micro ap  is also about  \n</code></pre>\n<p>implications:</p>\n<ul>\n<li>if your model is complex, e.g. deep GNN or docking, etc … you can just make model for some blocks.</li>\n<li>parallel training</li>\n</ul>",
      "rawMarkdown": "my guess could be correct. my experimental results and paper reading shows:\n\n1. molecule space is not smooth. small change in SMILES can give very different molecular property.\n\n2. learned model  don't gives good results if  test building block is not in training\n\n3. wide neural nets work better than deep ones. i think the are just memorizing molecule motifs/fragments in training.\nhence 98 million samples is NOT a lot, since they cannot represent unseen block. \n(there is probably no difference to the nonshare block test samples  if you are training 10m, 30m, 50m or 100m)\n\n4. hence i carry out the below experiment:\n\n1) divide 271 blocking blocks 1 into 5 group (each approximately 54)\n2) split train smiles into 5 folds, each fold consists of smiles synthesized from each of the group above (i.e. each fold only has 54 blocking blocks 1)\n3) during validation, we only apply model blockwise, i.e. use n-th model if and only if test smile is in n-th group\n\nhere are the results for xgboost + ecfp:\n\n```\napply models groupwise/blockwise:\nSCORE 0.6733886313435921\t0.5756944603438678\t0.3547743388934794\t0.8722992679673244\n(micro,'binds_BRD4', 'binds_HSA', 'binds_sEH')\n\napply models by averaging\nSCORE 0.5468913119251357\t0.4629902134390647\t0.31040167677578356\t0.7413135192018068\n\napply models by taking max\nSCORE 0.5396942034158857\t0.45337941398935977\t0.2756007936103707\t0.7098006405401006\n\n\nreference, in previous experiments, i do random split and micro ap score is also about 0.67 \n\n``` \n\nimplications:\n- if your model is complex, e.g. deep GNN or docking, etc ... you can just make model for some blocks.\n- parallel training",
      "votes": null
    },
    {
      "id": "2791888",
      "postDate": "05/03/2024 22:28:56",
      "content": "<p>This is a very common problem with chemical modeling. I believe it stems from the fact that molecules are not continuous entities. They are discrete entities that have properties we measure as continuous.</p>\n<p>You may enjoy reading about Free-Wilson SAR and matches molecular pairs. These types of modeling explicitly lean into the fact that discrete changes between molecules often have a constant result (slightly oversimplified, but that's the basic idea).</p>\n<p>After that, you may enjoy reading about free energy relationships and Hansch analysis where we try to use the properties that small groups of atoms (often called functional groups) tend to have in order to predict something else.</p>",
      "rawMarkdown": "This is a very common problem with chemical modeling. I believe it stems from the fact that molecules are not continuous entities. They are discrete entities that have properties we measure as continuous.\n\nYou may enjoy reading about Free-Wilson SAR and matches molecular pairs. These types of modeling explicitly lean into the fact that discrete changes between molecules often have a constant result (slightly oversimplified, but that's the basic idea).\n\nAfter that, you may enjoy reading about free energy relationships and Hansch analysis where we try to use the properties that small groups of atoms (often called functional groups) tend to have in order to predict something else.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2791888,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "05/03/2024 22:28:56",
      "content": "<p>This is a very common problem with chemical modeling. I believe it stems from the fact that molecules are not continuous entities. They are discrete entities that have properties we measure as continuous.</p>\n<p>You may enjoy reading about Free-Wilson SAR and matches molecular pairs. These types of modeling explicitly lean into the fact that discrete changes between molecules often have a constant result (slightly oversimplified, but that's the basic idea).</p>\n<p>After that, you may enjoy reading about free energy relationships and Hansch analysis where we try to use the properties that small groups of atoms (often called functional groups) tend to have in order to predict something else.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2790996": "my guess could be correct. my experimental results and paper reading shows:\n\n1. molecule space is not smooth. small change in SMILES can give very different molecular property.\n\n2. learned model  don't gives good results if  test building block is not in training\n\n3. wide neural nets work better than deep ones. i think the are just memorizing molecule motifs/fragments in training.\nhence 98 million samples is NOT a lot, since they cannot represent unseen block. \n(there is probably no difference to the nonshare block test samples  if you are training 10m, 30m, 50m or 100m)\n\n4. hence i carry out the below experiment:\n\n1) divide 271 blocking blocks 1 into 5 group (each approximately 54)\n2) split train smiles into 5 folds, each fold consists of smiles synthesized from each of the group above (i.e. each fold only has 54 blocking blocks 1)\n3) during validation, we only apply model blockwise, i.e. use n-th model if and only if test smile is in n-th group\n\nhere are the results for xgboost + ecfp:\n\n```\napply models groupwise/blockwise:\nSCORE 0.6733886313435921\t0.5756944603438678\t0.3547743388934794\t0.8722992679673244\n(micro,'binds_BRD4', 'binds_HSA', 'binds_sEH')\n\napply models by averaging\nSCORE 0.5468913119251357\t0.4629902134390647\t0.31040167677578356\t0.7413135192018068\n\napply models by taking max\nSCORE 0.5396942034158857\t0.45337941398935977\t0.2756007936103707\t0.7098006405401006\n\n\nreference, in previous experiments, i do random split and micro ap score is also about 0.67 \n\n``` \n\nimplications:\n- if your model is complex, e.g. deep GNN or docking, etc ... you can just make model for some blocks.\n- parallel training",
    "2791888": "This is a very common problem with chemical modeling. I believe it stems from the fact that molecules are not continuous entities. They are discrete entities that have properties we measure as continuous.\n\nYou may enjoy reading about Free-Wilson SAR and matches molecular pairs. These types of modeling explicitly lean into the fact that discrete changes between molecules often have a constant result (slightly oversimplified, but that's the basic idea).\n\nAfter that, you may enjoy reading about free energy relationships and Hansch analysis where we try to use the properties that small groups of atoms (often called functional groups) tend to have in order to predict something else."
  },
  "source": "meta"
}