{
  "id": 491415,
  "title": "CV Strategy ",
  "url": "/competitions/leash-BELKA/discussion/491415",
  "author_name": "SeshuRaju 🧘‍♂️",
  "post_date": "2024-04-05T19:38:26.507000",
  "votes": 16,
  "comment_count": 16,
  "views": 0,
  "content": "<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. Due to the overlapping nature of DEL chemistry, the test-train splits necessarily shrink the amount of data available during the competition (e.g., <strong>for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets</strong> ). These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements)</p>\n</blockquote>\n<pre><code> sklearn.model_selection  StratifiedGroupKFold\ntrain[] = train[]++train[]++train[]++train[]\ntrain[] = train[]++train[].apply( x: (x))\n\nsgkf = StratifiedGroupKFold(n_splits=, random_state=)\n\n train_index, valid_index  sgkf.split(train, train.target, train.buildingblocks):\n   --- \n</code></pre>\n<hr>\n<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. <strong>To test generalizability, the test set contains building blocks that are not in the training set.</strong> These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).</p>\n</blockquote>\n<ul>\n<li><strong>Need to split molecule between folds so it will match with private dataset</strong></li>\n</ul>\n<hr>\n<p><strong>is this right CV strategy as per the requirements?</strong></p>\n<p></p>",
  "messages": [
    {
      "id": 2737517,
      "postDate": "2024-04-05T19:38:26.507Z",
      "content": "<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. Due to the overlapping nature of DEL chemistry, the test-train splits necessarily shrink the amount of data available during the competition (e.g., <strong>for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets</strong> ). These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements)</p>\n</blockquote>\n<pre><code> sklearn.model_selection  StratifiedGroupKFold\ntrain[] = train[]++train[]++train[]++train[]\ntrain[] = train[]++train[].apply( x: (x))\n\nsgkf = StratifiedGroupKFold(n_splits=, random_state=)\n\n train_index, valid_index  sgkf.split(train, train.target, train.buildingblocks):\n   --- \n</code></pre>\n<hr>\n<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. <strong>To test generalizability, the test set contains building blocks that are not in the training set.</strong> These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).</p>\n</blockquote>\n<ul>\n<li><strong>Need to split molecule between folds so it will match with private dataset</strong></li>\n</ul>\n<hr>\n<p><strong>is this right CV strategy as per the requirements?</strong></p>\n<p></p>",
      "rawMarkdown": "> All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. Due to the overlapping nature of DEL chemistry, the test-train splits necessarily shrink the amount of data available during the competition (e.g., **for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets** ). These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements)\n\n```python\nfrom sklearn.model_selection import StratifiedGroupKFold\ntrain['buildingblocks'] = train['buildingblock1_smiles']+\"-\"+train['buildingblock2_smiles']+\"-\"+train['buildingblock3_smiles']+\"-\"+train[\"protein_name\"]\ntrain['target'] = train[\"protein_name\"]+\"-\"+train[\"binds\"].apply(lambda x: str(x))\n\nsgkf = StratifiedGroupKFold(n_splits=5, random_state=42)\n\nfor train_index, valid_index in sgkf.split(train, train.target, train.buildingblocks):\n   --- \n```\n\n---\n\n> All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. **To test generalizability, the test set contains building blocks that are not in the training set.** These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).\n\n- **Need to split molecule between folds so it will match with private dataset**\n\n---\n\n**is this right CV strategy as per the requirements?**\n\n~~i miss understood initially it is a code competition, after it understood not~~",
      "votes": 16
    },
    {
      "id": 2745246,
      "postDate": "2024-04-10T13:51:11.107Z",
      "content": "<p>My split:</p>\n<pre><code>df = df(\n    (\n        ~((pl() &lt; ) &amp;\n        (pl() &lt; ) &amp;\n        (pl() &lt; ))\n    )()(pl.Int8)\n)\n</code></pre>\n<p>I'm using this <a href=\"https://www.kaggle.com/datasets/shlomoron/belka-shrunken-train-set\" target=\"_blank\">shrinked dataset</a></p>\n<p>Upd:<br>\nsince blocks are not unique per-column, this split is not correct, in order to get right split you should first re-enumerate all unique blocks in all three columns and then do this split again</p>",
      "rawMarkdown": "My split:\n```\ndf = df.with_columns(\n    (\n        ~((pl.col('buildingblock1_smiles') < 250) &\n        (pl.col('buildingblock2_smiles') < 600) &\n        (pl.col('buildingblock3_smiles') < 800))\n    ).alias('fold').cast(pl.Int8)\n)\n```\nI'm using this [shrinked dataset](https://www.kaggle.com/datasets/shlomoron/belka-shrunken-train-set)\n\nUpd:\nsince blocks are not unique per-column, this split is not correct, in order to get right split you should first re-enumerate all unique blocks in all three columns and then do this split again",
      "votes": 7
    },
    {
      "id": 2737565,
      "postDate": "2024-04-05T19:59:19.760Z",
      "content": "<p>I think splitting on all 3 building blocks concatenated together is (basically) the same as splitting on the molecule.</p>\n<p>Here are a few resources to help think about ways to split your data:<br>\n<a href=\"https://deepchem.readthedocs.io/en/latest/api_reference/splitters.html\" target=\"_blank\">https://deepchem.readthedocs.io/en/latest/api_reference/splitters.html</a><br>\n<a href=\"https://tdcommons.ai/functions/data_split/\" target=\"_blank\">https://tdcommons.ai/functions/data_split/</a><br>\n<a href=\"https://lifesci.dgl.ai/api/utils.splitters.html\" target=\"_blank\">https://lifesci.dgl.ai/api/utils.splitters.html</a></p>",
      "rawMarkdown": "I think splitting on all 3 building blocks concatenated together is (basically) the same as splitting on the molecule.\n\nHere are a few resources to help think about ways to split your data:\nhttps://deepchem.readthedocs.io/en/latest/api_reference/splitters.html\nhttps://tdcommons.ai/functions/data_split/\nhttps://lifesci.dgl.ai/api/utils.splitters.html",
      "votes": 5,
      "replies": [
        {
          "id": 2737587,
          "postDate": "2024-04-05T20:07:54.017Z",
          "content": "<p><a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> as i understand molecule is together of 3 building blocks, could you share any resource how it combined since it is not just string concatenations and is the order of 3 building blocks sequence also random?.</p>\n<ul>\n<li>Thanks for quick response and clarity.</li>\n</ul>",
          "rawMarkdown": "@andrewdblevins as i understand molecule is together of 3 building blocks, could you share any resource how it combined since it is not just string concatenations and is the order of 3 building blocks sequence also random?.\n- Thanks for quick response and clarity.",
          "votes": 2,
          "replies": [
            {
              "id": 2737653,
              "postDate": "2024-04-05T21:24:37.883Z",
              "content": "<p>it's not a simple operation. You have to simulate all the chemical reactions. </p>\n<p>Later I will post the nitty gritty details on all the chemical reactions that go into constructing the DELs. But that knowledge shouldn't be necessary for this competition.</p>",
              "rawMarkdown": "it's not a simple operation. You have to simulate all the chemical reactions. \n\nLater I will post the nitty gritty details on all the chemical reactions that go into constructing the DELs. But that knowledge shouldn't be necessary for this competition.",
              "votes": 3
            },
            {
              "id": 2737662,
              "postDate": "2024-04-05T21:28:17.440Z",
              "content": "<p>Ok <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> thank you.</p>",
              "rawMarkdown": "Ok @andrewdblevins thank you.",
              "votes": 1
            },
            {
              "id": 2739776,
              "postDate": "2024-04-07T09:19:47.783Z",
              "content": "<p><a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> as molecule have chemical combined of 3 building blocks<br>\n<strong>- binds=0/1 is based on molecule and protein name - is this correct ?</strong><br>\n<strong>- are building blocks helpful? - any reference mater to understand building blocks and molecule relationship?</strong></p>\n<table>\n<thead>\n<tr>\n<th>Molecules/building blocks</th>\n<th>Train  (total samples)</th>\n<th>Test (total samples)</th>\n<th>Common with Test</th>\n<th>Common with Train</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Molecules</td>\n<td>98,415,610</td>\n<td>878,022</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>3 building blocks concat</td>\n<td>98,415,610</td>\n<td>878,022</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>3 building blocks combinations</td>\n<td>98,415,610</td>\n<td>878,022</td>\n<td>98,415,068 =&gt; <strong>99%</strong></td>\n<td>508,983 =&gt; <strong>58%</strong></td>\n</tr>\n</tbody>\n</table>",
              "rawMarkdown": "@andrewdblevins as molecule have chemical combined of 3 building blocks\n**- binds=0/1 is based on molecule and protein name - is this correct ?**\n**- are building blocks helpful? - any reference mater to understand building blocks and molecule relationship?**\n\n| Molecules/building blocks | Train  (total samples) | Test (total samples) | Common with Test | Common with Train |\n| --- | --- | --- | --- |\n| Molecules |  98,415,610 |  878,022 | 0 | 0 |\n| 3 building blocks concat | 98,415,610 | 878,022 | 0 | 0 |\n| 3 building blocks combinations |  98,415,610 | 878,022 |  98,415,068 => **99%**  | 508,983 => **58%** |\n",
              "votes": 3
            },
            {
              "id": 2740123,
              "postDate": "2024-04-07T15:38:03.073Z",
              "content": "<p>Building Blocks were included because much of the previous work predicting DEL data depends on them.<br>\nSee this paper by <a href=\"https://pubs.acs.org/doi/10.1021/acs.jmedchem.0c00452\" target=\"_blank\">google/xchem</a> and this one out of <a href=\"https://arxiv.org/pdf/2205.08020.pdf\" target=\"_blank\">anagenex</a></p>",
              "rawMarkdown": "Building Blocks were included because much of the previous work predicting DEL data depends on them.\nSee this paper by [google/xchem](https://pubs.acs.org/doi/10.1021/acs.jmedchem.0c00452) and this one out of [anagenex](https://arxiv.org/pdf/2205.08020.pdf)",
              "votes": 1
            },
            {
              "id": 2740134,
              "postDate": "2024-04-07T15:48:55.263Z",
              "content": "<p>Thanks for the clarity, I gone through this paper <a href=\"https://arxiv.org/pdf/2205.08020.pdf\" target=\"_blank\">anagenex</a> and shared <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491401\" target=\"_blank\">Discussion - Papers: DNA-encoded chemical library (DEL)</a>, will go through other paper too.</p>",
              "rawMarkdown": "Thanks for the clarity, I gone through this paper [anagenex](https://arxiv.org/pdf/2205.08020.pdf) and shared [Discussion - Papers: DNA-encoded chemical library (DEL)](https://www.kaggle.com/competitions/leash-BELKA/discussion/491401), will go through other paper too."
            }
          ]
        },
        {
          "id": 2740984,
          "postDate": "2024-04-08T04:31:00.837Z",
          "content": "<p><a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">Andrew</a></p>\n<p>Your comment on 3 building blocks suggests that no molecule was created with different initial 3 blocks.  Am I reading into your comment correctly, or can molecules end up being the  same but with a different creation path in train or test?</p>",
          "rawMarkdown": "[Andrew](https://www.kaggle.com/andrewdblevins)\n\nYour comment on 3 building blocks suggests that no molecule was created with different initial 3 blocks.  Am I reading into your comment correctly, or can molecules end up being the  same but with a different creation path in train or test?\n",
          "votes": 1,
          "replies": [
            {
              "id": 2741040,
              "postDate": "2024-04-08T05:07:47.217Z",
              "content": "<p>I believe there is none of that in this set of molecules, but in general, it's possible to create the same molecule with different sets of 3 building blocks.</p>",
              "rawMarkdown": "I believe there is none of that in this set of molecules, but in general, it's possible to create the same molecule with different sets of 3 building blocks.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2739986,
      "postDate": "2024-04-07T12:48:24.707Z",
      "content": "<p><a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> </p>\n<p>I have a question about how the public and private leaderboard data was split. Is the split  random? Or is the public leaderboard correspond to a validation set, and private leaderboard to a test set which have purposely different distributions of building blocks?</p>\n<p>Thanks for being active and giving references and links they have been useful in getting started with this competition. </p>",
      "rawMarkdown": "@andrewdblevins \n\nI have a question about how the public and private leaderboard data was split. Is the split  random? Or is the public leaderboard correspond to a validation set, and private leaderboard to a test set which have purposely different distributions of building blocks?\n\nThanks for being active and giving references and links they have been useful in getting started with this competition. ",
      "votes": 3,
      "replies": [
        {
          "id": 2740129,
          "postDate": "2024-04-07T15:44:42.823Z",
          "content": "<p>For now, I am not going to say anything about the composition of the private test set other than both test sets were both chosen to test generalizability to new molecules.</p>\n<p>If these models are successful, you should be able to start evaluating material from chemistry catalogs like wuxi or enamine, order 100 top hits to be tested, and expect a good percentage of those compounds to come back as active. Since chemical space is so big, it is very unlikely that the catalogs will have any molecules in common with your training set, so generalization is crucial.</p>",
          "rawMarkdown": "For now, I am not going to say anything about the composition of the private test set other than both test sets were both chosen to test generalizability to new molecules.\n\nIf these models are successful, you should be able to start evaluating material from chemistry catalogs like wuxi or enamine, order 100 top hits to be tested, and expect a good percentage of those compounds to come back as active. Since chemical space is so big, it is very unlikely that the catalogs will have any molecules in common with your training set, so generalization is crucial.",
          "votes": 8,
          "replies": [
            {
              "id": 2740206,
              "postDate": "2024-04-07T17:06:31.387Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2740195,
          "postDate": "2024-04-07T16:58:28.147Z",
          "content": "<p>Thanks for info. I think I am getting a handle on validation set up. Definitely a little complicated. </p>",
          "rawMarkdown": "Thanks for info. I think I am getting a handle on validation set up. Definitely a little complicated. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2745178,
      "postDate": "2024-04-10T12:50:52.640Z",
      "content": "<p>So far I am using scaffold split of individual BBs, this gives you fold 0, 1, 2, etc for BB1, fold 0, 1, 2, etc for BB2 and same for BB3. <br>\nIn that case you can make a val fold where all BBs don't overlap with train BBs, and structurally different </p>",
      "rawMarkdown": "So far I am using scaffold split of individual BBs, this gives you fold 0, 1, 2, etc for BB1, fold 0, 1, 2, etc for BB2 and same for BB3. \nIn that case you can make a val fold where all BBs don't overlap with train BBs, and structurally different ",
      "votes": 1,
      "replies": [
        {
          "id": 2745200,
          "postDate": "2024-04-10T13:09:19.683Z",
          "content": "<p><a href=\"https://www.kaggle.com/kvigly55\" target=\"_blank\">@kvigly55</a> as scaffold on BBs will make each fold will have the binds/no binds distributions and # of samples</p>",
          "rawMarkdown": "@kvigly55 as scaffold on BBs will make each fold will have the binds/no binds distributions and # of samples"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2745246,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-04-10T13:51:11.107000",
      "content": "<p>My split:</p>\n<pre><code>df = df(\n    (\n        ~((pl() &lt; ) &amp;\n        (pl() &lt; ) &amp;\n        (pl() &lt; ))\n    )()(pl.Int8)\n)\n</code></pre>\n<p>I'm using this <a href=\"https://www.kaggle.com/datasets/shlomoron/belka-shrunken-train-set\" target=\"_blank\">shrinked dataset</a></p>\n<p>Upd:<br>\nsince blocks are not unique per-column, this split is not correct, in order to get right split you should first re-enumerate all unique blocks in all three columns and then do this split again</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2737565,
      "author_name": "Andrew D. Blevins",
      "author_url": "",
      "post_date": "2024-04-05T19:59:19.760000",
      "content": "<p>I think splitting on all 3 building blocks concatenated together is (basically) the same as splitting on the molecule.</p>\n<p>Here are a few resources to help think about ways to split your data:<br>\n<a href=\"https://deepchem.readthedocs.io/en/latest/api_reference/splitters.html\" target=\"_blank\">https://deepchem.readthedocs.io/en/latest/api_reference/splitters.html</a><br>\n<a href=\"https://tdcommons.ai/functions/data_split/\" target=\"_blank\">https://tdcommons.ai/functions/data_split/</a><br>\n<a href=\"https://lifesci.dgl.ai/api/utils.splitters.html\" target=\"_blank\">https://lifesci.dgl.ai/api/utils.splitters.html</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 2737587,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-04-05T20:07:54.017000",
          "content": "<p><a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> as i understand molecule is together of 3 building blocks, could you share any resource how it combined since it is not just string concatenations and is the order of 3 building blocks sequence also random?.</p>\n<ul>\n<li>Thanks for quick response and clarity.</li>\n</ul>",
          "votes": 2,
          "replies": [
            {
              "id": 2737653,
              "author_name": "Andrew D. Blevins",
              "author_url": "",
              "post_date": "2024-04-05T21:24:37.883000",
              "content": "<p>it's not a simple operation. You have to simulate all the chemical reactions. </p>\n<p>Later I will post the nitty gritty details on all the chemical reactions that go into constructing the DELs. But that knowledge shouldn't be necessary for this competition.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2737662,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-04-05T21:28:17.440000",
              "content": "<p>Ok <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> thank you.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2739776,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-04-07T09:19:47.783000",
              "content": "<p><a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> as molecule have chemical combined of 3 building blocks<br>\n<strong>- binds=0/1 is based on molecule and protein name - is this correct ?</strong><br>\n<strong>- are building blocks helpful? - any reference mater to understand building blocks and molecule relationship?</strong></p>\n<table>\n<thead>\n<tr>\n<th>Molecules/building blocks</th>\n<th>Train  (total samples)</th>\n<th>Test (total samples)</th>\n<th>Common with Test</th>\n<th>Common with Train</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Molecules</td>\n<td>98,415,610</td>\n<td>878,022</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>3 building blocks concat</td>\n<td>98,415,610</td>\n<td>878,022</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>3 building blocks combinations</td>\n<td>98,415,610</td>\n<td>878,022</td>\n<td>98,415,068 =&gt; <strong>99%</strong></td>\n<td>508,983 =&gt; <strong>58%</strong></td>\n</tr>\n</tbody>\n</table>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2740123,
              "author_name": "Andrew D. Blevins",
              "author_url": "",
              "post_date": "2024-04-07T15:38:03.073000",
              "content": "<p>Building Blocks were included because much of the previous work predicting DEL data depends on them.<br>\nSee this paper by <a href=\"https://pubs.acs.org/doi/10.1021/acs.jmedchem.0c00452\" target=\"_blank\">google/xchem</a> and this one out of <a href=\"https://arxiv.org/pdf/2205.08020.pdf\" target=\"_blank\">anagenex</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2740134,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-04-07T15:48:55.263000",
              "content": "<p>Thanks for the clarity, I gone through this paper <a href=\"https://arxiv.org/pdf/2205.08020.pdf\" target=\"_blank\">anagenex</a> and shared <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491401\" target=\"_blank\">Discussion - Papers: DNA-encoded chemical library (DEL)</a>, will go through other paper too.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2740984,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2024-04-08T04:31:00.837000",
          "content": "<p><a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">Andrew</a></p>\n<p>Your comment on 3 building blocks suggests that no molecule was created with different initial 3 blocks.  Am I reading into your comment correctly, or can molecules end up being the  same but with a different creation path in train or test?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2741040,
              "author_name": "Andrew D. Blevins",
              "author_url": "",
              "post_date": "2024-04-08T05:07:47.217000",
              "content": "<p>I believe there is none of that in this set of molecules, but in general, it's possible to create the same molecule with different sets of 3 building blocks.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2739986,
      "author_name": "Devin Anzelmo",
      "author_url": "",
      "post_date": "2024-04-07T12:48:24.707000",
      "content": "<p><a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> </p>\n<p>I have a question about how the public and private leaderboard data was split. Is the split  random? Or is the public leaderboard correspond to a validation set, and private leaderboard to a test set which have purposely different distributions of building blocks?</p>\n<p>Thanks for being active and giving references and links they have been useful in getting started with this competition. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2740129,
          "author_name": "Andrew D. Blevins",
          "author_url": "",
          "post_date": "2024-04-07T15:44:42.823000",
          "content": "<p>For now, I am not going to say anything about the composition of the private test set other than both test sets were both chosen to test generalizability to new molecules.</p>\n<p>If these models are successful, you should be able to start evaluating material from chemistry catalogs like wuxi or enamine, order 100 top hits to be tested, and expect a good percentage of those compounds to come back as active. Since chemical space is so big, it is very unlikely that the catalogs will have any molecules in common with your training set, so generalization is crucial.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2740206,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-07T17:06:31.387000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2740195,
          "author_name": "Devin Anzelmo",
          "author_url": "",
          "post_date": "2024-04-07T16:58:28.147000",
          "content": "<p>Thanks for info. I think I am getting a handle on validation set up. Definitely a little complicated. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2745178,
      "author_name": "kvigly",
      "author_url": "",
      "post_date": "2024-04-10T12:50:52.640000",
      "content": "<p>So far I am using scaffold split of individual BBs, this gives you fold 0, 1, 2, etc for BB1, fold 0, 1, 2, etc for BB2 and same for BB3. <br>\nIn that case you can make a val fold where all BBs don't overlap with train BBs, and structurally different </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2745200,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-04-10T13:09:19.683000",
          "content": "<p><a href=\"https://www.kaggle.com/kvigly55\" target=\"_blank\">@kvigly55</a> as scaffold on BBs will make each fold will have the binds/no binds distributions and # of samples</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2737517": "> All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. Due to the overlapping nature of DEL chemistry, the test-train splits necessarily shrink the amount of data available during the competition (e.g., **for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets** ). These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements)\n\n```python\nfrom sklearn.model_selection import StratifiedGroupKFold\ntrain['buildingblocks'] = train['buildingblock1_smiles']+\"-\"+train['buildingblock2_smiles']+\"-\"+train['buildingblock3_smiles']+\"-\"+train[\"protein_name\"]\ntrain['target'] = train[\"protein_name\"]+\"-\"+train[\"binds\"].apply(lambda x: str(x))\n\nsgkf = StratifiedGroupKFold(n_splits=5, random_state=42)\n\nfor train_index, valid_index in sgkf.split(train, train.target, train.buildingblocks):\n   --- \n```\n\n---\n\n> All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. **To test generalizability, the test set contains building blocks that are not in the training set.** These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).\n\n- **Need to split molecule between folds so it will match with private dataset**\n\n---\n\n**is this right CV strategy as per the requirements?**\n\n~~i miss understood initially it is a code competition, after it understood not~~",
    "2745246": "My split:\n```\ndf = df.with_columns(\n    (\n        ~((pl.col('buildingblock1_smiles') < 250) &\n        (pl.col('buildingblock2_smiles') < 600) &\n        (pl.col('buildingblock3_smiles') < 800))\n    ).alias('fold').cast(pl.Int8)\n)\n```\nI'm using this [shrinked dataset](https://www.kaggle.com/datasets/shlomoron/belka-shrunken-train-set)\n\nUpd:\nsince blocks are not unique per-column, this split is not correct, in order to get right split you should first re-enumerate all unique blocks in all three columns and then do this split again",
    "2737565": "I think splitting on all 3 building blocks concatenated together is (basically) the same as splitting on the molecule.\n\nHere are a few resources to help think about ways to split your data:\nhttps://deepchem.readthedocs.io/en/latest/api_reference/splitters.html\nhttps://tdcommons.ai/functions/data_split/\nhttps://lifesci.dgl.ai/api/utils.splitters.html",
    "2739986": "@andrewdblevins \n\nI have a question about how the public and private leaderboard data was split. Is the split  random? Or is the public leaderboard correspond to a validation set, and private leaderboard to a test set which have purposely different distributions of building blocks?\n\nThanks for being active and giving references and links they have been useful in getting started with this competition. ",
    "2745178": "So far I am using scaffold split of individual BBs, this gives you fold 0, 1, 2, etc for BB1, fold 0, 1, 2, etc for BB2 and same for BB3. \nIn that case you can make a val fold where all BBs don't overlap with train BBs, and structurally different "
  }
}