{
  "id": 492126,
  "title": "Are the Train and Test building blocks really not overlapping?",
  "url": "/competitions/leash-BELKA/discussion/492126",
  "author_name": "",
  "post_date": "2024-04-08T17:56:46.838698700Z",
  "votes": 23,
  "comment_count": 19,
  "views": 0,
  "content": "<p><strong>Upd:</strong> This question arose because I had memorized the first wording of the test set description:</p>\n<blockquote>\n  <p>e.g., for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets</p>\n</blockquote>\n<p>It was silently corrected to:</p>\n<blockquote>\n  <p>the test set contains building blocks that are not in the training set</p>\n</blockquote>\n<p>(explained in the <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103\" target=\"_blank\">deleted topic</a>)<br>\nIf you read the first one, it's still confusing, so the test set actually contains ALL the smiles from the train building blocks, plus additional new smiles.</p>\n<hr>\n<p>In <a href=\"https://www.kaggle.com/code/antoninadolgorukova/belka-reading-and-quick-stats?scriptVersionId=171033835\" target=\"_blank\">this notebook</a> I have extracted all the unique building blocks from the train set as follows:</p>\n<pre><code>train_meta  open_dataset\ncon  dbConnectduckdbduckdb\ntrain  arrowto_duckdbtrain_meta table_name   con  con\n\ntrain_bb_smiles  uniqueas.data.tablerbind\n    dbGetQuerycon \n    dbGetQuerycon \n    dbGetQuerycon \n    \nnrowtrain_bb_smiles\n\n</code></pre>\n<p>And from the test set as follows:</p>\n<pre><code>test_dt  fread\ntest_bb_smiles  unique\n  test_dt uniquebuildingblock1_smiles\n  test_dt uniquebuildingblock2_smiles\n  test_dt uniquebuildingblock3_smiles\n\ntest_bb_smiles\n\n</code></pre>\n<p>Then I did a check to see if there is any overlap:</p>\n<pre><code>nrowtrain_bb_smilessmiles  test_bb_smiles\n\n</code></pre>\n<p><strong>So it seems 1145 building blocks SMILES from the train set are also present in the test set.</strong><br>\nBut the description says:</p>\n<blockquote>\n  <p>To test generalizability, the test set contains building blocks that are not in the training set.</p>\n</blockquote>\n<p>Would be happy if someone could correct or explain…</p>",
  "messages": [
    {
      "id": "2742048",
      "postDate": "04/08/2024 17:56:46",
      "content": "<p><strong>Upd:</strong> This question arose because I had memorized the first wording of the test set description:</p>\n<blockquote>\n  <p>e.g., for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets</p>\n</blockquote>\n<p>It was silently corrected to:</p>\n<blockquote>\n  <p>the test set contains building blocks that are not in the training set</p>\n</blockquote>\n<p>(explained in the <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103\" target=\"_blank\">deleted topic</a>)<br>\nIf you read the first one, it's still confusing, so the test set actually contains ALL the smiles from the train building blocks, plus additional new smiles.</p>\n<hr>\n<p>In <a href=\"https://www.kaggle.com/code/antoninadolgorukova/belka-reading-and-quick-stats?scriptVersionId=171033835\" target=\"_blank\">this notebook</a> I have extracted all the unique building blocks from the train set as follows:</p>\n<pre><code>train_meta  open_dataset\ncon  dbConnectduckdbduckdb\ntrain  arrowto_duckdbtrain_meta table_name   con  con\n\ntrain_bb_smiles  uniqueas.data.tablerbind\n    dbGetQuerycon \n    dbGetQuerycon \n    dbGetQuerycon \n    \nnrowtrain_bb_smiles\n\n</code></pre>\n<p>And from the test set as follows:</p>\n<pre><code>test_dt  fread\ntest_bb_smiles  unique\n  test_dt uniquebuildingblock1_smiles\n  test_dt uniquebuildingblock2_smiles\n  test_dt uniquebuildingblock3_smiles\n\ntest_bb_smiles\n\n</code></pre>\n<p>Then I did a check to see if there is any overlap:</p>\n<pre><code>nrowtrain_bb_smilessmiles  test_bb_smiles\n\n</code></pre>\n<p><strong>So it seems 1145 building blocks SMILES from the train set are also present in the test set.</strong><br>\nBut the description says:</p>\n<blockquote>\n  <p>To test generalizability, the test set contains building blocks that are not in the training set.</p>\n</blockquote>\n<p>Would be happy if someone could correct or explain…</p>",
      "rawMarkdown": "**Upd:** This question arose because I had memorized the first wording of the test set description:\n>e.g., for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets\n\nIt was silently corrected to:\n>the test set contains building blocks that are not in the training set\n\n(explained in the [deleted topic](https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103))\nIf you read the first one, it's still confusing, so the test set actually contains ALL the smiles from the train building blocks, plus additional new smiles.\n****\nIn [this notebook](https://www.kaggle.com/code/antoninadolgorukova/belka-reading-and-quick-stats?scriptVersionId=171033835) I have extracted all the unique building blocks from the train set as follows:\n\n```r\ntrain_meta <- open_dataset('/kaggle/input/leash-BELKA/train.parquet')\ncon <- dbConnect(duckdb::duckdb())\ntrain <- arrow::to_duckdb(train_meta, table_name = \"train\", con = con)\n\ntrain_bb_smiles <- unique(as.data.table(rbind(\n    dbGetQuery(con, \"SELECT DISTINCT(buildingblock1_smiles) AS smiles FROM train\"),\n    dbGetQuery(con, \"SELECT DISTINCT(buildingblock2_smiles) AS smiles FROM train\"),\n    dbGetQuery(con, \"SELECT DISTINCT(buildingblock3_smiles) AS smiles FROM train\")\n    )))\nnrow(train_bb_smiles)\n# 1145\n```\n\nAnd from the test set as follows:\n\n```r\ntest_dt <- fread(\"/kaggle/input/leash-BELKA/test.csv\")\ntest_bb_smiles <- unique(c(\n  test_dt[, unique(buildingblock1_smiles)],\n  test_dt[, unique(buildingblock2_smiles)],\n  test_dt[, unique(buildingblock3_smiles)]\n))\nlength(test_bb_smiles)\n#2110\n```\n\nThen I did a check to see if there is any overlap:\n\n```r\nnrow(train_bb_smiles[smiles %in% test_bb_smiles])\n# 1145\n```\n**So it seems 1145 building blocks SMILES from the train set are also present in the test set.**\nBut the description says:\n\n>To test generalizability, the test set contains building blocks that are not in the training set.\n\nWould be happy if someone could correct or explain...",
      "votes": null
    },
    {
      "id": "2742063",
      "postDate": "04/08/2024 18:12:07",
      "content": "<p><a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> as per below query, no common when compare with exact combinations of buildingblock1_smiles, buildingblock2_smiles, buildingblock3_smiles or molecule_smiles</p>\n<pre><code> con.query().df()\n</code></pre>\n<hr>\n<blockquote>\n  <p><strong>Total all_buildingblock_smiles_count =&gt; 0</strong></p>\n</blockquote>\n<hr>\n<pre><code>con.query().df()\n</code></pre>\n<hr>\n<blockquote>\n  <p><strong>Total all_molecule_smiles_count =&gt; 0</strong></p>\n</blockquote>\n<hr>\n<pre><code>con.query().df()\n</code></pre>\n<hr>\n<blockquote>\n  <p><strong>Total other 3 building block smiles of train in test  =&gt; 369,039 out of 878,022 molecules</strong></p>\n</blockquote>\n<hr>\n<pre><code>con.query().df()\n</code></pre>\n<hr>\n<blockquote>\n  <p><strong>Total any one of 3 building block of train in test  =&gt; 508,983 out of 878,022 molecules</strong></p>\n</blockquote>\n<table>\n<thead>\n<tr>\n<th>Building Blocks</th>\n<th>Train</th>\n<th>Test</th>\n<th>Common</th>\n<th>Test New</th>\n<th>Train Extra</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>buildingblock1_smiles</td>\n<td>271</td>\n<td>340</td>\n<td>271</td>\n<td>69</td>\n<td>0</td>\n</tr>\n<tr>\n<td>buildingblock2_smiles</td>\n<td>693</td>\n<td>1,139</td>\n<td>693</td>\n<td>446</td>\n<td>0</td>\n</tr>\n<tr>\n<td>buildingblock3_smiles</td>\n<td>871</td>\n<td>1,388</td>\n<td>870</td>\n<td>517</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p>Am i missing anything here? I did similar analysis with all sql queries with train, test and compare both - Notebook - 🔍📈📊🧬 EDA SMILES 🧬📊📈🔍</p>",
      "rawMarkdown": "antoninadolgorukova as per below query, no common when compare with exact combinations of buildingblock1_smiles, buildingblock2_smiles, buildingblock3_smiles or molecule_smiles\n\n```python\n con.query(f\"\"\"(SELECT count(*) as all_buildingblock_smiles_count FROM (\n    SELECT distinct buildingblock1_smiles,buildingblock2_smiles,buildingblock3_smiles FROM parquet_scan('{test_path}')\n    INTERSECT\n    SELECT distinct buildingblock1_smiles,buildingblock2_smiles,buildingblock3_smiles FROM parquet_scan('{train_path}')\n    ) as t)\"\"\").df()\n```\n\n---\n\n> **Total all_buildingblock_smiles_count => 0**\n\n---\n\n```python\ncon.query(f\"\"\"(SELECT count(*) as all_molecule_smiles_count FROM (\n    SELECT distinct molecule_smiles FROM parquet_scan('{test_path}')\n    INTERSECT\n    SELECT distinct molecule_smiles FROM parquet_scan('{train_path}')\n    ) as t)\"\"\").df()\n```\n\n---\n\n> **Total all_molecule_smiles_count => 0**\n\n---\n\n```python\ncon.query(f\"\"\"(SELECT count(*) as new_building_blocks_count FROM (\n    SELECT distinct molecule_smiles FROM parquet_scan('{test_path}') \n    WHERE buildingblock1_smiles in (SELECT distinct buildingblock1_smiles FROM parquet_scan('{train_path}'))\n    AND buildingblock2_smiles in (SELECT distinct buildingblock2_smiles FROM parquet_scan('{train_path}'))\n    AND buildingblock3_smiles in (SELECT distinct buildingblock3_smiles FROM parquet_scan('{train_path}'))\n    ) as t)\"\"\").df()\n```\n\n---\n\n> **Total other 3 building block smiles of train in test  => 369,039 out of 878,022 molecules**\n\n---\n\n```python\ncon.query(f\"\"\"(SELECT count(*) as new_building_blocks_count FROM (\n    SELECT distinct molecule_smiles FROM parquet_scan('{test_path}') \n    WHERE buildingblock1_smiles not in (SELECT distinct buildingblock1_smiles FROM parquet_scan('{train_path}'))\n    OR buildingblock2_smiles not in (SELECT distinct buildingblock2_smiles FROM parquet_scan('{train_path}'))\n    OR buildingblock3_smiles not in (SELECT distinct buildingblock3_smiles FROM parquet_scan('{train_path}'))\n    ) as t)\"\"\").df()\n```\n\n---\n\n> **Total any one of 3 building block of train in test  => 508,983 out of 878,022 molecules**\n\n| Building Blocks | Train | Test | Common | Test New | Train Extra | \n| --- | --- | --- | --- | --- | --- | \n| buildingblock1_smiles | 271 | 340 | 271 | 69 |  0 | \n| buildingblock2_smiles | 693 | 1,139  | 693 | 446 | 0 |\n| buildingblock3_smiles | 871 | 1,388 | 870 | 517 | 1 |\n\n---\n\nAm i missing anything here? I did similar analysis with all sql queries with train, test and compare both - Notebook - 🔍📈📊🧬 EDA SMILES 🧬📊📈🔍",
      "votes": null
    },
    {
      "id": "2742139",
      "postDate": "04/08/2024 18:54:49",
      "content": "<p>Yes, great work! Thanks. I just compared all the unique BB smiles from the test with all the unique BB smiles from the train. They do indeed overlap, the test set contains 2110 unique BB smiles, 1145 of them are in the train set (and the train set contains only those 1145), and 965 are new.</p>\n<p>It seems I misinterpreted the wording of the description, the test set does contain BB smiles from the train, it just also contains new BBs too. Anyway, I would love to hear confirmation from the organizers. <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> </p>",
      "rawMarkdown": "Yes, great work! Thanks. I just compared all the unique BB smiles from the test with all the unique BB smiles from the train. They do indeed overlap, the test set contains 2110 unique BB smiles, 1145 of them are in the train set (and the train set contains only those 1145), and 965 are new.\n\nIt seems I misinterpreted the wording of the description, the test set does contain BB smiles from the train, it just also contains new BBs too. Anyway, I would love to hear confirmation from the organizers. @andrewdblevins",
      "votes": null
    },
    {
      "id": "2742171",
      "postDate": "04/08/2024 19:14:35",
      "content": "<p>See host confirmation <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "See host confirmation [here](https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103).",
      "votes": null
    },
    {
      "id": "2742181",
      "postDate": "04/08/2024 19:20:11",
      "content": "<p>Just about 1/3 of the test set share zero building blocks with the train set. The other 2/3 almost all(except 15) share 3 building blocks with the train set.</p>\n<p>So really there are two problems we are trying to solve. Making train and validation sets with no shared blocks gives validation scores of around 0.04. Whereas training and validating on shared blocks gives scores around 0.45 on my first attempts. </p>",
      "rawMarkdown": "Just about 1/3 of the test set share zero building blocks with the train set. The other 2/3 almost all(except 15) share 3 building blocks with the train set.\n\nSo really there are two problems we are trying to solve. Making train and validation sets with no shared blocks gives validation scores of around 0.04. Whereas training and validating on shared blocks gives scores around 0.45 on my first attempts.",
      "votes": null
    },
    {
      "id": "2742184",
      "postDate": "04/08/2024 19:23:20",
      "content": "<p>Great, thanks! Sorry that the topic was deleted. I think others like me also read the description before the correction and memorized this phrase:</p>\n<blockquote>\n  <p>e.g., for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets</p>\n</blockquote>\n<p>I could not find it when I was writing my question, and it confuses so I decided to ask anyway.</p>",
      "rawMarkdown": "Great, thanks! Sorry that the topic was deleted. I think others like me also read the description before the correction and memorized this phrase:\n\n>e.g., for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets\n\nI could not find it when I was writing my question, and it confuses so I decided to ask anyway.",
      "votes": null
    },
    {
      "id": "2742201",
      "postDate": "04/08/2024 19:45:22",
      "content": "<p>Yea it was a good topic but the one who opened it asked for votes…it's an important reminder not to do it haha</p>",
      "rawMarkdown": "Yea it was a good topic but the one who opened it asked for votes...it's an important reminder not to do it haha",
      "votes": null
    },
    {
      "id": "2742684",
      "postDate": "04/09/2024 03:00:22",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc32a341ab9c9eeeb7d16251d783190df%2FSelection_013.png?generation=1712631427893319&amp;alt=media\">  </p>\n<p>the diagram is for building block1.   <br>\nThe horizontal axis is block id (i.e. each value is a block type).  <br>\nHence there are about 271 unique  building block1 in train.  <br>\nHence there are about 341 unique  building block1 in test.  </p>\n<p>the vertical axis is the occurrence frequency of the block type.  </p>\n<hr>\n<p>hint: the test distribution may leak the binding target.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc32a341ab9c9eeeb7d16251d783190df%2FSelection_013.png?generation=1712631427893319&alt=media)  \n\nthe diagram is for building block1.   \nThe horizontal axis is block id (i.e. each value is a block type).  \nHence there are about 271 unique  building block1 in train.  \nHence there are about 341 unique  building block1 in test.  \n\nthe vertical axis is the occurrence frequency of the block type.  \n\n\n---\n\nhint: the test distribution may leak the binding target.",
      "votes": null
    },
    {
      "id": "2745446",
      "postDate": "04/10/2024 16:29:09",
      "content": "<p>I think you are right \"So really there are two problems we are trying to solve\".</p>",
      "rawMarkdown": "I think you are right \"So really there are two problems we are trying to solve\".",
      "votes": null
    },
    {
      "id": "2751073",
      "postDate": "04/14/2024 02:52:26",
      "content": "<p>I'll post my analysis here, it's pretty interesting.</p>\n<p>Credit to <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> 's <a href=\"https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration/notebook\" target=\"_blank\">awesome scaffolding notebook</a>, which enabled this breakdown.</p>\n<p>I found three categories, and looks like all three categories are completely non-overlapping in BBs (building blocks) from other categories:</p>\n<ul>\n<li>500,000 non-triazine AND all new BBs (29.9% of test dataset)</li>\n<li>67,779 triazine will all new BBs (5.8% of triazines; 4.0% of total test)</li>\n<li>1,107,117 triazines with BBs all included in train dataset (94.2% of triazines; 66.1% of all test)</li>\n</ul>\n<p>1,674,896 TOTAL TEST</p>\n<p>So non-overlapping CV: it only directly corresponds to just 4.0% of test dataset!</p>\n<p>Another 30%, the non-triazines, may or may not see a LB score that approximately aligns with the non-overlapping CV score. But they would be \"scaffolding hopping\" and might turn out to be a lot harder to predict than merely predicting a triazine with unseen BBs. Guess we'll see.</p>\n<p>And 66%, 2/3rds of the data, should align well with a more normal CV, though the test vs train still might need more analysis in terms of similarity scores or clustering or embeddings.</p>",
      "rawMarkdown": "I'll post my analysis here, it's pretty interesting.\n\nCredit to @chemdatafarmer 's [awesome scaffolding notebook](https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration/notebook), which enabled this breakdown.\n\nI found three categories, and looks like all three categories are completely non-overlapping in BBs (building blocks) from other categories:\n\n- 500,000 non-triazine AND all new BBs (29.9% of test dataset)\n- 67,779 triazine will all new BBs (5.8% of triazines; 4.0% of total test)\n- 1,107,117 triazines with BBs all included in train dataset (94.2% of triazines; 66.1% of all test)\n\n1,674,896 TOTAL TEST\n\nSo non-overlapping CV: it only directly corresponds to just 4.0% of test dataset!\n\nAnother 30%, the non-triazines, may or may not see a LB score that approximately aligns with the non-overlapping CV score. But they would be \"scaffolding hopping\" and might turn out to be a lot harder to predict than merely predicting a triazine with unseen BBs. Guess we'll see.\n\nAnd 66%, 2/3rds of the data, should align well with a more normal CV, though the test vs train still might need more analysis in terms of similarity scores or clustering or embeddings.",
      "votes": null
    },
    {
      "id": "2754093",
      "postDate": "04/15/2024 19:41:33",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">@devinanzelmo</a>. Great work. I just want to know  if your CV calculated over the entire dataset (the 98M molecule smiles)</p>",
      "rawMarkdown": "Hello @devinanzelmo. Great work. I just want to know  if your CV calculated over the entire dataset (the 98M molecule smiles)",
      "votes": null
    },
    {
      "id": "2754154",
      "postDate": "04/15/2024 21:07:40",
      "content": "<p>No it was on a much smaller subsample. The trick to getting scores matching the leaderboard is to subsample the data so the class balance is the same as the public leaderboard. This means there should be 1/125 as many positive examples as negative.</p>",
      "rawMarkdown": "No it was on a much smaller subsample. The trick to getting scores matching the leaderboard is to subsample the data so the class balance is the same as the public leaderboard. This means there should be 1/125 as many positive examples as negative.",
      "votes": null
    },
    {
      "id": "2754215",
      "postDate": "04/15/2024 22:17:22",
      "content": "<p>How do you know that it is 1/125? Is it some kind of clever probing? </p>",
      "rawMarkdown": "How do you know that it is 1/125? Is it some kind of clever probing?",
      "votes": null
    },
    {
      "id": "2754251",
      "postDate": "04/15/2024 23:03:18",
      "content": "<p>\" The trick to getting scores matching the leaderboard is to subsample the data so the class balance is the same as the public leaderboard.\"</p>\n<p>but there is a risk: the private dataset is assumed to have the same distribution as the public one.</p>\n<hr>\n<p>hint: not an issue since all test dataset is open (unlike code competition)</p>",
      "rawMarkdown": "\" The trick to getting scores matching the leaderboard is to subsample the data so the class balance is the same as the public leaderboard.\"\n\nbut there is a risk: the private dataset is assumed to have the same distribution as the public one.\n\n---\n\nhint: not an issue since all test dataset is open (unlike code competition)",
      "votes": null
    },
    {
      "id": "2754259",
      "postDate": "04/15/2024 23:22:30",
      "content": "<p>The score for a random/all ones/all zeros submission is 0.008, the way the metric works this means its about 1/125th positives.  At least this is what it looks like from playing with the metric. </p>",
      "rawMarkdown": "The score for a random/all ones/all zeros submission is 0.008, the way the metric works this means its about 1/125th positives.  At least this is what it looks like from playing with the metric.",
      "votes": null
    },
    {
      "id": "2754364",
      "postDate": "04/16/2024 01:45:58",
      "content": "<p>I think 0.008 is the default score, not a real one. When submitting all zeros, you should except the score to be 0/0 but it gives 0.008</p>",
      "rawMarkdown": "I think 0.008 is the default score, not a real one. When submitting all zeros, you should except the score to be 0/0 but it gives 0.008",
      "votes": null
    },
    {
      "id": "2754465",
      "postDate": "04/16/2024 04:21:44",
      "content": "<p>try it on your train set. submit a all zero (or other constant)will not give you zero. you can work out the maths too. <br>\nthis is the art of probing. test dataset is open. by submitting magic value to magic subgroup of test data, you can guess the test distribution.</p>",
      "rawMarkdown": "try it on your train set. submit a all zero (or other constant)will not give you zero. you can work out the maths too. \nthis is the art of probing. test dataset is open. by submitting magic value to magic subgroup of test data, you can guess the test distribution.",
      "votes": null
    },
    {
      "id": "2758962",
      "postDate": "04/18/2024 12:12:04",
      "content": "<p>I note that the information provided under the <a href=\"https://www.kaggle.com/competitions/leash-BELKA/data\" target=\"_blank\">Data tab of this competition</a> indicates that the positive rate is ~ 1 in 200 in both test and training sets. Verbatim, it says (bold highlighting is my emphasis):</p>\n<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. To test generalizability, the test set contains building blocks that are not in the training set. <strong>These datasets are very imbalanced: roughly 0.5% of examples are classified as binders</strong>; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).</p>\n</blockquote>",
      "rawMarkdown": "I note that the information provided under the [Data tab of this competition](https://www.kaggle.com/competitions/leash-BELKA/data) indicates that the positive rate is ~ 1 in 200 in both test and training sets. Verbatim, it says (bold highlighting is my emphasis):\n\n>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. To test generalizability, the test set contains building blocks that are not in the training set. **These datasets are very imbalanced: roughly 0.5% of examples are classified as binders**; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).",
      "votes": null
    },
    {
      "id": "2759026",
      "postDate": "04/18/2024 12:53:15",
      "content": "<p>If it was 1 in 200 then the score for all zeros would be 0.005. So there appear to be a few more positives in the public leaderboard data then stated. This could mean there are less then 1 in 200 for the private test set, but we can't be sure.  For the train data the fraction of positives is 0.0053 which is much closer to the stated value of the organizers then the public lb. </p>",
      "rawMarkdown": "If it was 1 in 200 then the score for all zeros would be 0.005. So there appear to be a few more positives in the public leaderboard data then stated. This could mean there are less then 1 in 200 for the private test set, but we can't be sure.  For the train data the fraction of positives is 0.0053 which is much closer to the stated value of the organizers then the public lb.",
      "votes": null
    },
    {
      "id": "2759134",
      "postDate": "04/18/2024 13:56:07",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">@devinanzelmo</a>, especially for the insightful comment about a possible difference of proportions in private vs public test sets. Since Kaggle LB scores are truncated, not rounded, I think the best interpretation is that the proportion of positives within the public test set is between 0.008 and 0.009, or between approximately 1 in 111 and 1 in 125.</p>",
      "rawMarkdown": "Thanks @devinanzelmo, especially for the insightful comment about a possible difference of proportions in private vs public test sets. Since Kaggle LB scores are truncated, not rounded, I think the best interpretation is that the proportion of positives within the public test set is between 0.008 and 0.009, or between approximately 1 in 111 and 1 in 125.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2742063,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "04/08/2024 18:12:07",
      "content": "<p><a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> as per below query, no common when compare with exact combinations of buildingblock1_smiles, buildingblock2_smiles, buildingblock3_smiles or molecule_smiles</p>\n<pre><code> con.query().df()\n</code></pre>\n<hr>\n<blockquote>\n  <p><strong>Total all_buildingblock_smiles_count =&gt; 0</strong></p>\n</blockquote>\n<hr>\n<pre><code>con.query().df()\n</code></pre>\n<hr>\n<blockquote>\n  <p><strong>Total all_molecule_smiles_count =&gt; 0</strong></p>\n</blockquote>\n<hr>\n<pre><code>con.query().df()\n</code></pre>\n<hr>\n<blockquote>\n  <p><strong>Total other 3 building block smiles of train in test  =&gt; 369,039 out of 878,022 molecules</strong></p>\n</blockquote>\n<hr>\n<pre><code>con.query().df()\n</code></pre>\n<hr>\n<blockquote>\n  <p><strong>Total any one of 3 building block of train in test  =&gt; 508,983 out of 878,022 molecules</strong></p>\n</blockquote>\n<table>\n<thead>\n<tr>\n<th>Building Blocks</th>\n<th>Train</th>\n<th>Test</th>\n<th>Common</th>\n<th>Test New</th>\n<th>Train Extra</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>buildingblock1_smiles</td>\n<td>271</td>\n<td>340</td>\n<td>271</td>\n<td>69</td>\n<td>0</td>\n</tr>\n<tr>\n<td>buildingblock2_smiles</td>\n<td>693</td>\n<td>1,139</td>\n<td>693</td>\n<td>446</td>\n<td>0</td>\n</tr>\n<tr>\n<td>buildingblock3_smiles</td>\n<td>871</td>\n<td>1,388</td>\n<td>870</td>\n<td>517</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p>Am i missing anything here? I did similar analysis with all sql queries with train, test and compare both - Notebook - 🔍📈📊🧬 EDA SMILES 🧬📊📈🔍</p>",
      "votes": null,
      "replies": [
        {
          "id": 2742139,
          "author_name": "antoninadolgorukova",
          "author_url": "",
          "post_date": "04/08/2024 18:54:49",
          "content": "<p>Yes, great work! Thanks. I just compared all the unique BB smiles from the test with all the unique BB smiles from the train. They do indeed overlap, the test set contains 2110 unique BB smiles, 1145 of them are in the train set (and the train set contains only those 1145), and 965 are new.</p>\n<p>It seems I misinterpreted the wording of the description, the test set does contain BB smiles from the train, it just also contains new BBs too. Anyway, I would love to hear confirmation from the organizers. <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> </p>",
          "votes": null,
          "replies": [
            {
              "id": 2742171,
              "author_name": "shlomoron",
              "author_url": "",
              "post_date": "04/08/2024 19:14:35",
              "content": "<p>See host confirmation <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103\" target=\"_blank\">here</a>.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2742184,
                  "author_name": "antoninadolgorukova",
                  "author_url": "",
                  "post_date": "04/08/2024 19:23:20",
                  "content": "<p>Great, thanks! Sorry that the topic was deleted. I think others like me also read the description before the correction and memorized this phrase:</p>\n<blockquote>\n  <p>e.g., for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets</p>\n</blockquote>\n<p>I could not find it when I was writing my question, and it confuses so I decided to ask anyway.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2742201,
                      "author_name": "shlomoron",
                      "author_url": "",
                      "post_date": "04/08/2024 19:45:22",
                      "content": "<p>Yea it was a good topic but the one who opened it asked for votes…it's an important reminder not to do it haha</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2742181,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "04/08/2024 19:20:11",
      "content": "<p>Just about 1/3 of the test set share zero building blocks with the train set. The other 2/3 almost all(except 15) share 3 building blocks with the train set.</p>\n<p>So really there are two problems we are trying to solve. Making train and validation sets with no shared blocks gives validation scores of around 0.04. Whereas training and validating on shared blocks gives scores around 0.45 on my first attempts. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2745446,
          "author_name": "ricopue",
          "author_url": "",
          "post_date": "04/10/2024 16:29:09",
          "content": "<p>I think you are right \"So really there are two problems we are trying to solve\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2754093,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "04/15/2024 19:41:33",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">@devinanzelmo</a>. Great work. I just want to know  if your CV calculated over the entire dataset (the 98M molecule smiles)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2754154,
          "author_name": "devinanzelmo",
          "author_url": "",
          "post_date": "04/15/2024 21:07:40",
          "content": "<p>No it was on a much smaller subsample. The trick to getting scores matching the leaderboard is to subsample the data so the class balance is the same as the public leaderboard. This means there should be 1/125 as many positive examples as negative.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2754215,
              "author_name": "shlomoron",
              "author_url": "",
              "post_date": "04/15/2024 22:17:22",
              "content": "<p>How do you know that it is 1/125? Is it some kind of clever probing? </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2754251,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "04/15/2024 23:03:18",
              "content": "<p>\" The trick to getting scores matching the leaderboard is to subsample the data so the class balance is the same as the public leaderboard.\"</p>\n<p>but there is a risk: the private dataset is assumed to have the same distribution as the public one.</p>\n<hr>\n<p>hint: not an issue since all test dataset is open (unlike code competition)</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2754259,
              "author_name": "devinanzelmo",
              "author_url": "",
              "post_date": "04/15/2024 23:22:30",
              "content": "<p>The score for a random/all ones/all zeros submission is 0.008, the way the metric works this means its about 1/125th positives.  At least this is what it looks like from playing with the metric. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2754364,
                  "author_name": "luispintoc",
                  "author_url": "",
                  "post_date": "04/16/2024 01:45:58",
                  "content": "<p>I think 0.008 is the default score, not a real one. When submitting all zeros, you should except the score to be 0/0 but it gives 0.008</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2754465,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "04/16/2024 04:21:44",
                      "content": "<p>try it on your train set. submit a all zero (or other constant)will not give you zero. you can work out the maths too. <br>\nthis is the art of probing. test dataset is open. by submitting magic value to magic subgroup of test data, you can guess the test distribution.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2758962,
                          "author_name": "jbomitchell",
                          "author_url": "",
                          "post_date": "04/18/2024 12:12:04",
                          "content": "<p>I note that the information provided under the <a href=\"https://www.kaggle.com/competitions/leash-BELKA/data\" target=\"_blank\">Data tab of this competition</a> indicates that the positive rate is ~ 1 in 200 in both test and training sets. Verbatim, it says (bold highlighting is my emphasis):</p>\n<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. To test generalizability, the test set contains building blocks that are not in the training set. <strong>These datasets are very imbalanced: roughly 0.5% of examples are classified as binders</strong>; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).</p>\n</blockquote>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 2759026,
              "author_name": "devinanzelmo",
              "author_url": "",
              "post_date": "04/18/2024 12:53:15",
              "content": "<p>If it was 1 in 200 then the score for all zeros would be 0.005. So there appear to be a few more positives in the public leaderboard data then stated. This could mean there are less then 1 in 200 for the private test set, but we can't be sure.  For the train data the fraction of positives is 0.0053 which is much closer to the stated value of the organizers then the public lb. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2759134,
                  "author_name": "jbomitchell",
                  "author_url": "",
                  "post_date": "04/18/2024 13:56:07",
                  "content": "<p>Thanks <a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">@devinanzelmo</a>, especially for the insightful comment about a possible difference of proportions in private vs public test sets. Since Kaggle LB scores are truncated, not rounded, I think the best interpretation is that the proportion of positives within the public test set is between 0.008 and 0.009, or between approximately 1 in 111 and 1 in 125.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2742684,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/09/2024 03:00:22",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc32a341ab9c9eeeb7d16251d783190df%2FSelection_013.png?generation=1712631427893319&amp;alt=media\">  </p>\n<p>the diagram is for building block1.   <br>\nThe horizontal axis is block id (i.e. each value is a block type).  <br>\nHence there are about 271 unique  building block1 in train.  <br>\nHence there are about 341 unique  building block1 in test.  </p>\n<p>the vertical axis is the occurrence frequency of the block type.  </p>\n<hr>\n<p>hint: the test distribution may leak the binding target.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2751073,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "04/14/2024 02:52:26",
      "content": "<p>I'll post my analysis here, it's pretty interesting.</p>\n<p>Credit to <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> 's <a href=\"https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration/notebook\" target=\"_blank\">awesome scaffolding notebook</a>, which enabled this breakdown.</p>\n<p>I found three categories, and looks like all three categories are completely non-overlapping in BBs (building blocks) from other categories:</p>\n<ul>\n<li>500,000 non-triazine AND all new BBs (29.9% of test dataset)</li>\n<li>67,779 triazine will all new BBs (5.8% of triazines; 4.0% of total test)</li>\n<li>1,107,117 triazines with BBs all included in train dataset (94.2% of triazines; 66.1% of all test)</li>\n</ul>\n<p>1,674,896 TOTAL TEST</p>\n<p>So non-overlapping CV: it only directly corresponds to just 4.0% of test dataset!</p>\n<p>Another 30%, the non-triazines, may or may not see a LB score that approximately aligns with the non-overlapping CV score. But they would be \"scaffolding hopping\" and might turn out to be a lot harder to predict than merely predicting a triazine with unseen BBs. Guess we'll see.</p>\n<p>And 66%, 2/3rds of the data, should align well with a more normal CV, though the test vs train still might need more analysis in terms of similarity scores or clustering or embeddings.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2742048": "**Upd:** This question arose because I had memorized the first wording of the test set description:\n>e.g., for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets\n\nIt was silently corrected to:\n>the test set contains building blocks that are not in the training set\n\n(explained in the [deleted topic](https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103))\nIf you read the first one, it's still confusing, so the test set actually contains ALL the smiles from the train building blocks, plus additional new smiles.\n****\nIn [this notebook](https://www.kaggle.com/code/antoninadolgorukova/belka-reading-and-quick-stats?scriptVersionId=171033835) I have extracted all the unique building blocks from the train set as follows:\n\n```r\ntrain_meta <- open_dataset('/kaggle/input/leash-BELKA/train.parquet')\ncon <- dbConnect(duckdb::duckdb())\ntrain <- arrow::to_duckdb(train_meta, table_name = \"train\", con = con)\n\ntrain_bb_smiles <- unique(as.data.table(rbind(\n    dbGetQuery(con, \"SELECT DISTINCT(buildingblock1_smiles) AS smiles FROM train\"),\n    dbGetQuery(con, \"SELECT DISTINCT(buildingblock2_smiles) AS smiles FROM train\"),\n    dbGetQuery(con, \"SELECT DISTINCT(buildingblock3_smiles) AS smiles FROM train\")\n    )))\nnrow(train_bb_smiles)\n# 1145\n```\n\nAnd from the test set as follows:\n\n```r\ntest_dt <- fread(\"/kaggle/input/leash-BELKA/test.csv\")\ntest_bb_smiles <- unique(c(\n  test_dt[, unique(buildingblock1_smiles)],\n  test_dt[, unique(buildingblock2_smiles)],\n  test_dt[, unique(buildingblock3_smiles)]\n))\nlength(test_bb_smiles)\n#2110\n```\n\nThen I did a check to see if there is any overlap:\n\n```r\nnrow(train_bb_smiles[smiles %in% test_bb_smiles])\n# 1145\n```\n**So it seems 1145 building blocks SMILES from the train set are also present in the test set.**\nBut the description says:\n\n>To test generalizability, the test set contains building blocks that are not in the training set.\n\nWould be happy if someone could correct or explain...",
    "2742063": "antoninadolgorukova as per below query, no common when compare with exact combinations of buildingblock1_smiles, buildingblock2_smiles, buildingblock3_smiles or molecule_smiles\n\n```python\n con.query(f\"\"\"(SELECT count(*) as all_buildingblock_smiles_count FROM (\n    SELECT distinct buildingblock1_smiles,buildingblock2_smiles,buildingblock3_smiles FROM parquet_scan('{test_path}')\n    INTERSECT\n    SELECT distinct buildingblock1_smiles,buildingblock2_smiles,buildingblock3_smiles FROM parquet_scan('{train_path}')\n    ) as t)\"\"\").df()\n```\n\n---\n\n> **Total all_buildingblock_smiles_count => 0**\n\n---\n\n```python\ncon.query(f\"\"\"(SELECT count(*) as all_molecule_smiles_count FROM (\n    SELECT distinct molecule_smiles FROM parquet_scan('{test_path}')\n    INTERSECT\n    SELECT distinct molecule_smiles FROM parquet_scan('{train_path}')\n    ) as t)\"\"\").df()\n```\n\n---\n\n> **Total all_molecule_smiles_count => 0**\n\n---\n\n```python\ncon.query(f\"\"\"(SELECT count(*) as new_building_blocks_count FROM (\n    SELECT distinct molecule_smiles FROM parquet_scan('{test_path}') \n    WHERE buildingblock1_smiles in (SELECT distinct buildingblock1_smiles FROM parquet_scan('{train_path}'))\n    AND buildingblock2_smiles in (SELECT distinct buildingblock2_smiles FROM parquet_scan('{train_path}'))\n    AND buildingblock3_smiles in (SELECT distinct buildingblock3_smiles FROM parquet_scan('{train_path}'))\n    ) as t)\"\"\").df()\n```\n\n---\n\n> **Total other 3 building block smiles of train in test  => 369,039 out of 878,022 molecules**\n\n---\n\n```python\ncon.query(f\"\"\"(SELECT count(*) as new_building_blocks_count FROM (\n    SELECT distinct molecule_smiles FROM parquet_scan('{test_path}') \n    WHERE buildingblock1_smiles not in (SELECT distinct buildingblock1_smiles FROM parquet_scan('{train_path}'))\n    OR buildingblock2_smiles not in (SELECT distinct buildingblock2_smiles FROM parquet_scan('{train_path}'))\n    OR buildingblock3_smiles not in (SELECT distinct buildingblock3_smiles FROM parquet_scan('{train_path}'))\n    ) as t)\"\"\").df()\n```\n\n---\n\n> **Total any one of 3 building block of train in test  => 508,983 out of 878,022 molecules**\n\n| Building Blocks | Train | Test | Common | Test New | Train Extra | \n| --- | --- | --- | --- | --- | --- | \n| buildingblock1_smiles | 271 | 340 | 271 | 69 |  0 | \n| buildingblock2_smiles | 693 | 1,139  | 693 | 446 | 0 |\n| buildingblock3_smiles | 871 | 1,388 | 870 | 517 | 1 |\n\n---\n\nAm i missing anything here? I did similar analysis with all sql queries with train, test and compare both - Notebook - 🔍📈📊🧬 EDA SMILES 🧬📊📈🔍",
    "2742139": "Yes, great work! Thanks. I just compared all the unique BB smiles from the test with all the unique BB smiles from the train. They do indeed overlap, the test set contains 2110 unique BB smiles, 1145 of them are in the train set (and the train set contains only those 1145), and 965 are new.\n\nIt seems I misinterpreted the wording of the description, the test set does contain BB smiles from the train, it just also contains new BBs too. Anyway, I would love to hear confirmation from the organizers. @andrewdblevins",
    "2742171": "See host confirmation [here](https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103).",
    "2742181": "Just about 1/3 of the test set share zero building blocks with the train set. The other 2/3 almost all(except 15) share 3 building blocks with the train set.\n\nSo really there are two problems we are trying to solve. Making train and validation sets with no shared blocks gives validation scores of around 0.04. Whereas training and validating on shared blocks gives scores around 0.45 on my first attempts.",
    "2742184": "Great, thanks! Sorry that the topic was deleted. I think others like me also read the description before the correction and memorized this phrase:\n\n>e.g., for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets\n\nI could not find it when I was writing my question, and it confuses so I decided to ask anyway.",
    "2742201": "Yea it was a good topic but the one who opened it asked for votes...it's an important reminder not to do it haha",
    "2742684": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc32a341ab9c9eeeb7d16251d783190df%2FSelection_013.png?generation=1712631427893319&alt=media)  \n\nthe diagram is for building block1.   \nThe horizontal axis is block id (i.e. each value is a block type).  \nHence there are about 271 unique  building block1 in train.  \nHence there are about 341 unique  building block1 in test.  \n\nthe vertical axis is the occurrence frequency of the block type.  \n\n\n---\n\nhint: the test distribution may leak the binding target.",
    "2745446": "I think you are right \"So really there are two problems we are trying to solve\".",
    "2751073": "I'll post my analysis here, it's pretty interesting.\n\nCredit to @chemdatafarmer 's [awesome scaffolding notebook](https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration/notebook), which enabled this breakdown.\n\nI found three categories, and looks like all three categories are completely non-overlapping in BBs (building blocks) from other categories:\n\n- 500,000 non-triazine AND all new BBs (29.9% of test dataset)\n- 67,779 triazine will all new BBs (5.8% of triazines; 4.0% of total test)\n- 1,107,117 triazines with BBs all included in train dataset (94.2% of triazines; 66.1% of all test)\n\n1,674,896 TOTAL TEST\n\nSo non-overlapping CV: it only directly corresponds to just 4.0% of test dataset!\n\nAnother 30%, the non-triazines, may or may not see a LB score that approximately aligns with the non-overlapping CV score. But they would be \"scaffolding hopping\" and might turn out to be a lot harder to predict than merely predicting a triazine with unseen BBs. Guess we'll see.\n\nAnd 66%, 2/3rds of the data, should align well with a more normal CV, though the test vs train still might need more analysis in terms of similarity scores or clustering or embeddings.",
    "2754093": "Hello @devinanzelmo. Great work. I just want to know  if your CV calculated over the entire dataset (the 98M molecule smiles)",
    "2754154": "No it was on a much smaller subsample. The trick to getting scores matching the leaderboard is to subsample the data so the class balance is the same as the public leaderboard. This means there should be 1/125 as many positive examples as negative.",
    "2754215": "How do you know that it is 1/125? Is it some kind of clever probing?",
    "2754251": "\" The trick to getting scores matching the leaderboard is to subsample the data so the class balance is the same as the public leaderboard.\"\n\nbut there is a risk: the private dataset is assumed to have the same distribution as the public one.\n\n---\n\nhint: not an issue since all test dataset is open (unlike code competition)",
    "2754259": "The score for a random/all ones/all zeros submission is 0.008, the way the metric works this means its about 1/125th positives.  At least this is what it looks like from playing with the metric.",
    "2754364": "I think 0.008 is the default score, not a real one. When submitting all zeros, you should except the score to be 0/0 but it gives 0.008",
    "2754465": "try it on your train set. submit a all zero (or other constant)will not give you zero. you can work out the maths too. \nthis is the art of probing. test dataset is open. by submitting magic value to magic subgroup of test data, you can guess the test distribution.",
    "2758962": "I note that the information provided under the [Data tab of this competition](https://www.kaggle.com/competitions/leash-BELKA/data) indicates that the positive rate is ~ 1 in 200 in both test and training sets. Verbatim, it says (bold highlighting is my emphasis):\n\n>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. To test generalizability, the test set contains building blocks that are not in the training set. **These datasets are very imbalanced: roughly 0.5% of examples are classified as binders**; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).",
    "2759026": "If it was 1 in 200 then the score for all zeros would be 0.005. So there appear to be a few more positives in the public leaderboard data then stated. This could mean there are less then 1 in 200 for the private test set, but we can't be sure.  For the train data the fraction of positives is 0.0053 which is much closer to the stated value of the organizers then the public lb.",
    "2759134": "Thanks @devinanzelmo, especially for the insightful comment about a possible difference of proportions in private vs public test sets. Since Kaggle LB scores are truncated, not rounded, I think the best interpretation is that the proportion of positives within the public test set is between 0.008 and 0.009, or between approximately 1 in 111 and 1 in 125."
  },
  "source": "meta"
}