{
  "id": 491424,
  "title": "Private dataset questions [Solved]",
  "url": "/competitions/leash-BELKA/discussion/491424",
  "author_name": "",
  "post_date": "2024-04-05T20:02:57.804256400Z",
  "votes": 14,
  "comment_count": 2,
  "views": 0,
  "content": "<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein</p>\n</blockquote>\n<table>\n<thead>\n<tr>\n<th>protein_name</th>\n<th>train counts</th>\n<th>test count</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>HSA</td>\n<td>98,415,610</td>\n<td>557,895</td>\n</tr>\n<tr>\n<td>BRD4</td>\n<td>98,415,610</td>\n<td>558,859</td>\n</tr>\n<tr>\n<td>sEH</td>\n<td>98,415,610</td>\n<td>558,142</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>As we have 98M training examples per protein</li>\n</ul>\n<p>So, validation examples per protein is 200k so 600k samples ( total test samples = 878,022 )</p>\n<table>\n<thead>\n<tr>\n<th>Public/Private</th>\n<th>Samples</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Public 35%</td>\n<td>307,308</td>\n</tr>\n<tr>\n<td>Private 65%</td>\n<td>570,714</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<table>\n<thead>\n<tr>\n<th>Building Blocks</th>\n<th>Train</th>\n<th>Test</th>\n<th>Common</th>\n<th>Test New</th>\n<th>Train Extra</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>buildingblock1_smiles</td>\n<td>271</td>\n<td>340</td>\n<td>271</td>\n<td>69</td>\n<td>0</td>\n</tr>\n<tr>\n<td>buildingblock2_smiles</td>\n<td>693</td>\n<td>1,139</td>\n<td>693</td>\n<td>446</td>\n<td>0</td>\n</tr>\n<tr>\n<td>buildingblock3_smiles</td>\n<td>871</td>\n<td>1,388</td>\n<td>870</td>\n<td>517</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<table>\n<thead>\n<tr>\n<th>Molecules/building blocks</th>\n<th>Train  (total samples)</th>\n<th>Test (total samples)</th>\n<th>Common with Train</th>\n<th>Common with Test</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Molecules</td>\n<td>98,415,610</td>\n<td>878,022</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>3 building blocks concat</td>\n<td>98,415,610</td>\n<td>878,022</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>3 building blocks combinations</td>\n<td>98,415,610</td>\n<td>878,022</td>\n<td>98,415,068 =&gt; <strong>99%</strong></td>\n<td>508,983 =&gt; <strong>58%</strong></td>\n</tr>\n<tr>\n<td><strong>- no common molecule smiels between train and test datasets.</strong></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>- 58% of the building blocks of test set are in train's building blocks</strong></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. <strong>To test generalizability, the test set contains building blocks that are not in the training set.</strong> These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).</p>\n</blockquote>\n<ul>\n<li><strong>all building blocks are new ( not part of train ) =&gt; all new molecules</strong></li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103\" target=\"_blank\">Answers - Discussion - deleted topic</a></p>\n<blockquote>\n  <p>This is a good catch and something we weren't super clear on. <strong>The test set is made from the combination of several different splitting strategies.</strong> That statement was only meant to describe one of the strategies: the <strong>bb-split</strong>. for this split we hold out certain building blocks. But the overall test set also includes molecules from a <strong>scaffold split</strong> and a <strong>random split</strong>, hence the overlap. by <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> </p>\n  <ul>\n  <li><strong>bb-split</strong> =&gt; new molecules</li>\n  <li><strong>scaffold split</strong> =&gt; based on scaffold of molecules diversity ( nice article about scaffold split -&gt; <a href=\"https://www.oloren.ai/blog/scaff-split\" target=\"_blank\">https://www.oloren.ai/blog/scaff-split</a> )</li>\n  <li><strong>random split</strong> =&gt; -- </li>\n  </ul>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491415#2737565\" target=\"_blank\">Answers - Discussion - CV Strategy</a></p>\n<blockquote>\n  <p>I think splitting on all 3 building blocks concatenated together is (basically) the same as splitting on the molecule.<br>\n  Here are a few resources to help think about ways to split your data:<br>\n  <a href=\"https://deepchem.readthedocs.io/en/latest/api_reference/splitters.html\" target=\"_blank\">https://deepchem.readthedocs.io/en/latest/api_reference/splitters.html</a><br>\n  <a href=\"https://tdcommons.ai/functions/data_split/\" target=\"_blank\">https://tdcommons.ai/functions/data_split/</a><br>\n  <a href=\"https://lifesci.dgl.ai/api/utils.splitters.html\" target=\"_blank\">https://lifesci.dgl.ai/api/utils.splitters.html</a> </p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: <strong>roughly 0.5% of examples are classified as binders</strong>; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).</p>\n</blockquote>\n<ul>\n<li><strong>0.5% are classified as binders =&gt; 0.005 * 600k =&gt; 3k</strong></li>\n<li><strong>0.538% are classified as binders in training data ( both test and train have similar distribution )</strong>]</li>\n</ul>",
  "messages": [
    {
      "id": "2737578",
      "postDate": "04/05/2024 20:02:57",
      "content": "<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein</p>\n</blockquote>\n<table>\n<thead>\n<tr>\n<th>protein_name</th>\n<th>train counts</th>\n<th>test count</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>HSA</td>\n<td>98,415,610</td>\n<td>557,895</td>\n</tr>\n<tr>\n<td>BRD4</td>\n<td>98,415,610</td>\n<td>558,859</td>\n</tr>\n<tr>\n<td>sEH</td>\n<td>98,415,610</td>\n<td>558,142</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>As we have 98M training examples per protein</li>\n</ul>\n<p>So, validation examples per protein is 200k so 600k samples ( total test samples = 878,022 )</p>\n<table>\n<thead>\n<tr>\n<th>Public/Private</th>\n<th>Samples</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Public 35%</td>\n<td>307,308</td>\n</tr>\n<tr>\n<td>Private 65%</td>\n<td>570,714</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<table>\n<thead>\n<tr>\n<th>Building Blocks</th>\n<th>Train</th>\n<th>Test</th>\n<th>Common</th>\n<th>Test New</th>\n<th>Train Extra</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>buildingblock1_smiles</td>\n<td>271</td>\n<td>340</td>\n<td>271</td>\n<td>69</td>\n<td>0</td>\n</tr>\n<tr>\n<td>buildingblock2_smiles</td>\n<td>693</td>\n<td>1,139</td>\n<td>693</td>\n<td>446</td>\n<td>0</td>\n</tr>\n<tr>\n<td>buildingblock3_smiles</td>\n<td>871</td>\n<td>1,388</td>\n<td>870</td>\n<td>517</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<table>\n<thead>\n<tr>\n<th>Molecules/building blocks</th>\n<th>Train  (total samples)</th>\n<th>Test (total samples)</th>\n<th>Common with Train</th>\n<th>Common with Test</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Molecules</td>\n<td>98,415,610</td>\n<td>878,022</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>3 building blocks concat</td>\n<td>98,415,610</td>\n<td>878,022</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>3 building blocks combinations</td>\n<td>98,415,610</td>\n<td>878,022</td>\n<td>98,415,068 =&gt; <strong>99%</strong></td>\n<td>508,983 =&gt; <strong>58%</strong></td>\n</tr>\n<tr>\n<td><strong>- no common molecule smiels between train and test datasets.</strong></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>- 58% of the building blocks of test set are in train's building blocks</strong></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. <strong>To test generalizability, the test set contains building blocks that are not in the training set.</strong> These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).</p>\n</blockquote>\n<ul>\n<li><strong>all building blocks are new ( not part of train ) =&gt; all new molecules</strong></li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103\" target=\"_blank\">Answers - Discussion - deleted topic</a></p>\n<blockquote>\n  <p>This is a good catch and something we weren't super clear on. <strong>The test set is made from the combination of several different splitting strategies.</strong> That statement was only meant to describe one of the strategies: the <strong>bb-split</strong>. for this split we hold out certain building blocks. But the overall test set also includes molecules from a <strong>scaffold split</strong> and a <strong>random split</strong>, hence the overlap. by <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> </p>\n  <ul>\n  <li><strong>bb-split</strong> =&gt; new molecules</li>\n  <li><strong>scaffold split</strong> =&gt; based on scaffold of molecules diversity ( nice article about scaffold split -&gt; <a href=\"https://www.oloren.ai/blog/scaff-split\" target=\"_blank\">https://www.oloren.ai/blog/scaff-split</a> )</li>\n  <li><strong>random split</strong> =&gt; -- </li>\n  </ul>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491415#2737565\" target=\"_blank\">Answers - Discussion - CV Strategy</a></p>\n<blockquote>\n  <p>I think splitting on all 3 building blocks concatenated together is (basically) the same as splitting on the molecule.<br>\n  Here are a few resources to help think about ways to split your data:<br>\n  <a href=\"https://deepchem.readthedocs.io/en/latest/api_reference/splitters.html\" target=\"_blank\">https://deepchem.readthedocs.io/en/latest/api_reference/splitters.html</a><br>\n  <a href=\"https://tdcommons.ai/functions/data_split/\" target=\"_blank\">https://tdcommons.ai/functions/data_split/</a><br>\n  <a href=\"https://lifesci.dgl.ai/api/utils.splitters.html\" target=\"_blank\">https://lifesci.dgl.ai/api/utils.splitters.html</a> </p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: <strong>roughly 0.5% of examples are classified as binders</strong>; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).</p>\n</blockquote>\n<ul>\n<li><strong>0.5% are classified as binders =&gt; 0.005 * 600k =&gt; 3k</strong></li>\n<li><strong>0.538% are classified as binders in training data ( both test and train have similar distribution )</strong>]</li>\n</ul>",
      "rawMarkdown": "> All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein\n\n\n| protein_name | train counts | test count |\n| --- | --- | --- | \n| HSA | 98,415,610 | 557,895 |\n| BRD4 | 98,415,610    | 558,859 |\n| sEH | 98,415,610  | 558,142 |\n- As we have 98M training examples per protein\n\n\nSo, validation examples per protein is 200k so 600k samples ( total test samples = 878,022 )\n| Public/Private | Samples |\n| --- | --- |\n| Public 35%   |  307,308  |\n| Private 65%  |   570,714 |\n\n---\n\n| Building Blocks | Train | Test | Common | Test New | Train Extra | \n| --- | --- | --- | --- | --- | --- | \n| buildingblock1_smiles | 271 | 340 | 271 | 69 |  0 | \n| buildingblock2_smiles | 693 | 1,139  | 693 | 446 | 0 |\n| buildingblock3_smiles | 871 | 1,388 | 870 | 517 | 1 |\n\n---\n\n| Molecules/building blocks | Train  (total samples) | Test (total samples) | Common with Train | Common with Test |\n| --- | --- | --- | --- |\n| Molecules |  98,415,610 |  878,022 | 0 | 0 |\n| 3 building blocks concat | 98,415,610 | 878,022 | 0 | 0 |\n| 3 building blocks combinations |  98,415,610 | 878,022 |  98,415,068 => **99%**  | 508,983 => **58%** |\n**- no common molecule smiels between train and test datasets.**\n**- 58% of the building blocks of test set are in train's building blocks** \n\n\n\n> All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. **To test generalizability, the test set contains building blocks that are not in the training set.** These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).\n\n- **all building blocks are new ( not part of train ) => all new molecules**\n\n[Answers - Discussion - deleted topic](https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103)\n> This is a good catch and something we weren't super clear on. **The test set is made from the combination of several different splitting strategies.** That statement was only meant to describe one of the strategies: the **bb-split**. for this split we hold out certain building blocks. But the overall test set also includes molecules from a **scaffold split** and a **random split**, hence the overlap. by @andrewdblevins \n- **bb-split** => new molecules\n- **scaffold split** => based on scaffold of molecules diversity ( nice article about scaffold split -> [https://www.oloren.ai/blog/scaff-split](https://www.oloren.ai/blog/scaff-split) )\n- **random split** => -- \n\n[Answers - Discussion - CV Strategy](https://www.kaggle.com/competitions/leash-BELKA/discussion/491415#2737565)\n> I think splitting on all 3 building blocks concatenated together is (basically) the same as splitting on the molecule.\nHere are a few resources to help think about ways to split your data:\nhttps://deepchem.readthedocs.io/en/latest/api_reference/splitters.html\nhttps://tdcommons.ai/functions/data_split/\nhttps://lifesci.dgl.ai/api/utils.splitters.html \n\n---\n\n> All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: **roughly 0.5% of examples are classified as binders**; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).\n\n- **0.5% are classified as binders => 0.005 * 600k => 3k**\n- **0.538% are classified as binders in training data ( both test and train have similar distribution )**]",
      "votes": null
    },
    {
      "id": "2763080",
      "postDate": "04/20/2024 09:10:41",
      "content": "<p>So we have both the priveta and public test in advance right?<br>\nSo we can predict every test case without time and resource constraint?<br>\nVery unusual… Why did they do this?<br>\nAlso, how do we have 870K test samples? Isn't it 200k*3 = 600k?<br>\nWhat does 'validation set' refer to? There's no validation.csv and only test.csv and train.csv.</p>",
      "rawMarkdown": "So we have both the priveta and public test in advance right?\nSo we can predict every test case without time and resource constraint?\nVery unusual... Why did they do this?\nAlso, how do we have 870K test samples? Isn't it 200k*3 = 600k?\nWhat does 'validation set' refer to? There's no validation.csv and only test.csv and train.csv.",
      "votes": null
    },
    {
      "id": "2763090",
      "postDate": "04/20/2024 09:19:56",
      "content": "<blockquote>\n  <p>So we have both the priveta and public test in advance right?</p>\n</blockquote>\n<p>Yes, no hidden test set</p>\n<hr>\n<blockquote>\n  <p>So we can predict every test case without time and resource constraint?</p>\n</blockquote>\n<p>Yes</p>\n<hr>\n<blockquote>\n  <p>Very unusual… Why did they do this?</p>\n</blockquote>\n<p>Because it is not code competition as test set is large maybe</p>\n<hr>\n<blockquote>\n  <p>Also, how do we have 870K test samples? Isn't it 200k*3 = 600k? What does 'validation set' refer to? </p>\n</blockquote>\n<p>There's no validation.csv and only test.csv and train.csv.</p>\n<hr>\n<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, <strong>200K validation examples per protein, and 360K test molecules per protein</strong>. To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).</p>\n</blockquote>\n<ul>\n<li>I felt, validation set maybe about public LB is my guess</li>\n</ul>",
      "rawMarkdown": "> So we have both the priveta and public test in advance right?\n\nYes, no hidden test set\n\n---\n\n> So we can predict every test case without time and resource constraint?\n\nYes\n\n---\n\n> Very unusual… Why did they do this?\n\nBecause it is not code competition as test set is large maybe\n\n---\n\n> Also, how do we have 870K test samples? Isn't it 200k*3 = 600k? What does 'validation set' refer to? \n\nThere's no validation.csv and only test.csv and train.csv.\n\n---\n\n> All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, **200K validation examples per protein, and 360K test molecules per protein**. To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).\n\n- I felt, validation set maybe about public LB is my guess",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2763080,
      "author_name": "gyulamaloveczky4",
      "author_url": "",
      "post_date": "04/20/2024 09:10:41",
      "content": "<p>So we have both the priveta and public test in advance right?<br>\nSo we can predict every test case without time and resource constraint?<br>\nVery unusual… Why did they do this?<br>\nAlso, how do we have 870K test samples? Isn't it 200k*3 = 600k?<br>\nWhat does 'validation set' refer to? There's no validation.csv and only test.csv and train.csv.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2763090,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "04/20/2024 09:19:56",
          "content": "<blockquote>\n  <p>So we have both the priveta and public test in advance right?</p>\n</blockquote>\n<p>Yes, no hidden test set</p>\n<hr>\n<blockquote>\n  <p>So we can predict every test case without time and resource constraint?</p>\n</blockquote>\n<p>Yes</p>\n<hr>\n<blockquote>\n  <p>Very unusual… Why did they do this?</p>\n</blockquote>\n<p>Because it is not code competition as test set is large maybe</p>\n<hr>\n<blockquote>\n  <p>Also, how do we have 870K test samples? Isn't it 200k*3 = 600k? What does 'validation set' refer to? </p>\n</blockquote>\n<p>There's no validation.csv and only test.csv and train.csv.</p>\n<hr>\n<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, <strong>200K validation examples per protein, and 360K test molecules per protein</strong>. To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).</p>\n</blockquote>\n<ul>\n<li>I felt, validation set maybe about public LB is my guess</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2737578": "> All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein\n\n\n| protein_name | train counts | test count |\n| --- | --- | --- | \n| HSA | 98,415,610 | 557,895 |\n| BRD4 | 98,415,610    | 558,859 |\n| sEH | 98,415,610  | 558,142 |\n- As we have 98M training examples per protein\n\n\nSo, validation examples per protein is 200k so 600k samples ( total test samples = 878,022 )\n| Public/Private | Samples |\n| --- | --- |\n| Public 35%   |  307,308  |\n| Private 65%  |   570,714 |\n\n---\n\n| Building Blocks | Train | Test | Common | Test New | Train Extra | \n| --- | --- | --- | --- | --- | --- | \n| buildingblock1_smiles | 271 | 340 | 271 | 69 |  0 | \n| buildingblock2_smiles | 693 | 1,139  | 693 | 446 | 0 |\n| buildingblock3_smiles | 871 | 1,388 | 870 | 517 | 1 |\n\n---\n\n| Molecules/building blocks | Train  (total samples) | Test (total samples) | Common with Train | Common with Test |\n| --- | --- | --- | --- |\n| Molecules |  98,415,610 |  878,022 | 0 | 0 |\n| 3 building blocks concat | 98,415,610 | 878,022 | 0 | 0 |\n| 3 building blocks combinations |  98,415,610 | 878,022 |  98,415,068 => **99%**  | 508,983 => **58%** |\n**- no common molecule smiels between train and test datasets.**\n**- 58% of the building blocks of test set are in train's building blocks** \n\n\n\n> All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. **To test generalizability, the test set contains building blocks that are not in the training set.** These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).\n\n- **all building blocks are new ( not part of train ) => all new molecules**\n\n[Answers - Discussion - deleted topic](https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103)\n> This is a good catch and something we weren't super clear on. **The test set is made from the combination of several different splitting strategies.** That statement was only meant to describe one of the strategies: the **bb-split**. for this split we hold out certain building blocks. But the overall test set also includes molecules from a **scaffold split** and a **random split**, hence the overlap. by @andrewdblevins \n- **bb-split** => new molecules\n- **scaffold split** => based on scaffold of molecules diversity ( nice article about scaffold split -> [https://www.oloren.ai/blog/scaff-split](https://www.oloren.ai/blog/scaff-split) )\n- **random split** => -- \n\n[Answers - Discussion - CV Strategy](https://www.kaggle.com/competitions/leash-BELKA/discussion/491415#2737565)\n> I think splitting on all 3 building blocks concatenated together is (basically) the same as splitting on the molecule.\nHere are a few resources to help think about ways to split your data:\nhttps://deepchem.readthedocs.io/en/latest/api_reference/splitters.html\nhttps://tdcommons.ai/functions/data_split/\nhttps://lifesci.dgl.ai/api/utils.splitters.html \n\n---\n\n> All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: **roughly 0.5% of examples are classified as binders**; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).\n\n- **0.5% are classified as binders => 0.005 * 600k => 3k**\n- **0.538% are classified as binders in training data ( both test and train have similar distribution )**]",
    "2763080": "So we have both the priveta and public test in advance right?\nSo we can predict every test case without time and resource constraint?\nVery unusual... Why did they do this?\nAlso, how do we have 870K test samples? Isn't it 200k*3 = 600k?\nWhat does 'validation set' refer to? There's no validation.csv and only test.csv and train.csv.",
    "2763090": "> So we have both the priveta and public test in advance right?\n\nYes, no hidden test set\n\n---\n\n> So we can predict every test case without time and resource constraint?\n\nYes\n\n---\n\n> Very unusual… Why did they do this?\n\nBecause it is not code competition as test set is large maybe\n\n---\n\n> Also, how do we have 870K test samples? Isn't it 200k*3 = 600k? What does 'validation set' refer to? \n\nThere's no validation.csv and only test.csv and train.csv.\n\n---\n\n> All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, **200K validation examples per protein, and 360K test molecules per protein**. To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).\n\n- I felt, validation set maybe about public LB is my guess"
  },
  "source": "meta"
}