{
  "id": 496576,
  "title": "The groupings and permutations of the competition data, in terms of building blocks",
  "url": "/competitions/leash-BELKA/discussion/496576",
  "author_name": "Robert Hatch",
  "post_date": "2024-04-21T17:10:42.764000",
  "votes": 75,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Abbreviations: buildingblock1: BB1. Etc.</p>\n<h1>Train</h1>\n<ul>\n<li>271 BB1</li>\n<li>691 BB2/3</li>\n<li>2 BB2 only with the 181 below (Interesting. Why?)</li>\n<li>181 BB3 only, paired with any BB2, just never paired with each other. (Interesting. Why?)</li>\n</ul>\n<p>All combinations of BB2 and BB3:<br>\n181x693 for all combinations that include the 181 and/or the 2.<br>\n691x692 / 2 for all combinations of the 691 with each other. The same BB is allowed and indeed happens exactly 691 times. For any x,y pair, it will always be listed in one order, not the other order. (Which is where the divide by 2 comes in).</p>\n<p>The interesting and nice thing is the math checks out! Total of ALL possible combinations:<br>\n271x((181x693) + (691x692/2)) = 98,784,649</p>\n<p>But wait, there's only 98.4 million in train!</p>\n<p>Well:<br>\n98784649 - 98415610 = 369039</p>\n<p>Those 369039, times 3 protein targets, equals: 1,107,117 test rows. Which is exactly the number of test rows with shared building blocks from train in the test dataset.</p>\n<h1>Test</h1>\n<ul>\n<li>1,107,117 rows (369,039 molecules) with shared BBs, from above calculations.</li>\n<li>33,813 rows (11,271 molecules) from: 17x((2x34)+(34x35/2)) for triazine core, non-shared BBs group1 with: 17 BB1, 34 BB2/3, 2 BB3 only.</li>\n<li>33,966 rows (11,322 molecules) from: 17x(36x37/2) for triazine core, non-shared BBs group2 with: 17BB1, 36BB2/3.<ul>\n<li>Note: again, permutation numbers exactly match the 67,779 rows and 22593 molecules with triazine core and non-shared BBs in test dataset.</li></ul></li>\n<li>500,000 rows (486390 molecules): non-triazine core (credit <a href=\"https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration/notebook\" target=\"_blank\">this notebook</a>), non-shared BBs: 36 BB1, 377 BB2 only, 446 BB3 only. No obvious reason or pattern for the number of molecules in test, it looks at first few glances like a random selection of 500,000 from all possible proteins and BB permutations: 3x36x377x446. Only 13k molecules sampled two or 3 times seemed low to me at first glance, if random, but actually the math checks out, ranging from 0 to 1/18 chance if every prior pick was a new molecule. So about 1/36 * 500,000 ~= 13k, matching observed data.</li>\n</ul>\n<p>Note that, in terms of a rigorous experiment testing setting, it would NOT make any sense to have that final 500,000 rows in both \"validation\" (aka public LB) and \"test\" (aka private LB), and the hosts wrote:</p>\n<blockquote>\n  <p>200K validation examples per protein, and 360K test molecules per protein</p>\n</blockquote>\n<p>So by far the most likely possibility is:</p>\n<h3>Public LB</h3>\n<ul>\n<li>50% of shared BBs: 184,519 per protein.</li>\n<li>group 1 OR group 2: 11,271 per protein.</li>\n</ul>\n<p>Rounded to nearest 10K equals 200k \"validation\" per protein</p>\n<h3>Private LB</h3>\n<ul>\n<li>50% of shared BBs: 184,520 per protein</li>\n<li>group 1 OR group 2: 11,322 per protein.</li>\n<li>The non-triazine core group: 166,667 per protein on average.</li>\n</ul>\n<p>Rounded to nearest 10K: 360K \"test\" per protein</p>",
  "messages": [
    {
      "id": 2766380,
      "postDate": "2024-04-21T17:10:42.763Z",
      "content": "<p>Abbreviations: buildingblock1: BB1. Etc.</p>\n<h1>Train</h1>\n<ul>\n<li>271 BB1</li>\n<li>691 BB2/3</li>\n<li>2 BB2 only with the 181 below (Interesting. Why?)</li>\n<li>181 BB3 only, paired with any BB2, just never paired with each other. (Interesting. Why?)</li>\n</ul>\n<p>All combinations of BB2 and BB3:<br>\n181x693 for all combinations that include the 181 and/or the 2.<br>\n691x692 / 2 for all combinations of the 691 with each other. The same BB is allowed and indeed happens exactly 691 times. For any x,y pair, it will always be listed in one order, not the other order. (Which is where the divide by 2 comes in).</p>\n<p>The interesting and nice thing is the math checks out! Total of ALL possible combinations:<br>\n271x((181x693) + (691x692/2)) = 98,784,649</p>\n<p>But wait, there's only 98.4 million in train!</p>\n<p>Well:<br>\n98784649 - 98415610 = 369039</p>\n<p>Those 369039, times 3 protein targets, equals: 1,107,117 test rows. Which is exactly the number of test rows with shared building blocks from train in the test dataset.</p>\n<h1>Test</h1>\n<ul>\n<li>1,107,117 rows (369,039 molecules) with shared BBs, from above calculations.</li>\n<li>33,813 rows (11,271 molecules) from: 17x((2x34)+(34x35/2)) for triazine core, non-shared BBs group1 with: 17 BB1, 34 BB2/3, 2 BB3 only.</li>\n<li>33,966 rows (11,322 molecules) from: 17x(36x37/2) for triazine core, non-shared BBs group2 with: 17BB1, 36BB2/3.<ul>\n<li>Note: again, permutation numbers exactly match the 67,779 rows and 22593 molecules with triazine core and non-shared BBs in test dataset.</li></ul></li>\n<li>500,000 rows (486390 molecules): non-triazine core (credit <a href=\"https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration/notebook\" target=\"_blank\">this notebook</a>), non-shared BBs: 36 BB1, 377 BB2 only, 446 BB3 only. No obvious reason or pattern for the number of molecules in test, it looks at first few glances like a random selection of 500,000 from all possible proteins and BB permutations: 3x36x377x446. Only 13k molecules sampled two or 3 times seemed low to me at first glance, if random, but actually the math checks out, ranging from 0 to 1/18 chance if every prior pick was a new molecule. So about 1/36 * 500,000 ~= 13k, matching observed data.</li>\n</ul>\n<p>Note that, in terms of a rigorous experiment testing setting, it would NOT make any sense to have that final 500,000 rows in both \"validation\" (aka public LB) and \"test\" (aka private LB), and the hosts wrote:</p>\n<blockquote>\n  <p>200K validation examples per protein, and 360K test molecules per protein</p>\n</blockquote>\n<p>So by far the most likely possibility is:</p>\n<h3>Public LB</h3>\n<ul>\n<li>50% of shared BBs: 184,519 per protein.</li>\n<li>group 1 OR group 2: 11,271 per protein.</li>\n</ul>\n<p>Rounded to nearest 10K equals 200k \"validation\" per protein</p>\n<h3>Private LB</h3>\n<ul>\n<li>50% of shared BBs: 184,520 per protein</li>\n<li>group 1 OR group 2: 11,322 per protein.</li>\n<li>The non-triazine core group: 166,667 per protein on average.</li>\n</ul>\n<p>Rounded to nearest 10K: 360K \"test\" per protein</p>",
      "rawMarkdown": "Abbreviations: buildingblock1: BB1. Etc.\n# Train\n* 271 BB1\n* 691 BB2/3\n* 2 BB2 only with the 181 below (Interesting. Why?)\n* 181 BB3 only, paired with any BB2, just never paired with each other. (Interesting. Why?)\n\nAll combinations of BB2 and BB3:\n181x693 for all combinations that include the 181 and/or the 2.\n691x692 / 2 for all combinations of the 691 with each other. The same BB is allowed and indeed happens exactly 691 times. For any x,y pair, it will always be listed in one order, not the other order. (Which is where the divide by 2 comes in).\n\nThe interesting and nice thing is the math checks out! Total of ALL possible combinations:\n271x((181x693) + (691x692/2)) = 98,784,649\n\nBut wait, there's only 98.4 million in train!\n\nWell:\n98784649 - 98415610 = 369039\n\nThose 369039, times 3 protein targets, equals: 1,107,117 test rows. Which is exactly the number of test rows with shared building blocks from train in the test dataset.\n\n# Test\n* 1,107,117 rows (369,039 molecules) with shared BBs, from above calculations.\n* 33,813 rows (11,271 molecules) from: 17x((2x34)+(34x35/2)) for triazine core, non-shared BBs group1 with: 17 BB1, 34 BB2/3, 2 BB3 only.\n* 33,966 rows (11,322 molecules) from: 17x(36x37/2) for triazine core, non-shared BBs group2 with: 17BB1, 36BB2/3.\n  * Note: again, permutation numbers exactly match the 67,779 rows and 22593 molecules with triazine core and non-shared BBs in test dataset.\n* 500,000 rows (486390 molecules): non-triazine core (credit [this notebook](https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration/notebook)), non-shared BBs: 36 BB1, 377 BB2 only, 446 BB3 only. No obvious reason or pattern for the number of molecules in test, it looks at first few glances like a random selection of 500,000 from all possible proteins and BB permutations: 3x36x377x446. Only 13k molecules sampled two or 3 times seemed low to me at first glance, if random, but actually the math checks out, ranging from 0 to 1/18 chance if every prior pick was a new molecule. So about 1/36 * 500,000 ~= 13k, matching observed data.\n\nNote that, in terms of a rigorous experiment testing setting, it would NOT make any sense to have that final 500,000 rows in both \"validation\" (aka public LB) and \"test\" (aka private LB), and the hosts wrote:\n> 200K validation examples per protein, and 360K test molecules per protein\n\nSo by far the most likely possibility is:\n\n### Public LB\n* 50% of shared BBs: 184,519 per protein.\n* group 1 OR group 2: 11,271 per protein.\n\nRounded to nearest 10K equals 200k \"validation\" per protein\n\n### Private LB\n* 50% of shared BBs: 184,520 per protein\n* group 1 OR group 2: 11,322 per protein.\n* The non-triazine core group: 166,667 per protein on average.\n\nRounded to nearest 10K: 360K \"test\" per protein",
      "votes": 73
    },
    {
      "id": 2781137,
      "postDate": "2024-04-28T15:37:53.557Z",
      "content": "<p>to put into picture</p>\n<p>FOR TRAIN:<br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5db8dd160ffadd6ca30b43c2eabb8cf8%2FSelection_056.png?generation=1714318669316671&amp;alt=media\"></p>\n<p>FOR TEST:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F37d94487c09461f130fb48af3fa4ee2f%2FSelection_059.png?generation=1714355321671144&amp;alt=media\"></p>",
      "rawMarkdown": "to put into picture\n\nFOR TRAIN:\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5db8dd160ffadd6ca30b43c2eabb8cf8%2FSelection_056.png?generation=1714318669316671&alt=media)\n\nFOR TEST:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F37d94487c09461f130fb48af3fa4ee2f%2FSelection_059.png?generation=1714355321671144&alt=media)\n",
      "votes": 12,
      "replies": [
        {
          "id": 2781745,
          "postDate": "2024-04-29T02:00:05.310Z",
          "content": "<p>my guess on public and private split</p>\n<p>i train a classifier on the train share blocks</p>\n<ul>\n<li>train and validation  ap are almost similar</li>\n<li>public ap should be similar for the share block</li>\n</ul>\n<p>From experiments, applying classifier on non share block has almost zero detection</p>\n<pre><code>= r*validation + (very small value from non r= % of in public test = about %\n\n can estimated similarity of non </code></pre>",
          "rawMarkdown": "my guess on public and private split\n\ni train a classifier on the train share blocks\n- train and validation  ap are almost similar\n- public ap should be similar for the share block\n\nFrom experiments, applying classifier on non share block has almost zero detection\n\n```\nLB = r*validation + (very small value from non share blocks)\nr= % of share blocks in public test = about 87%\n\n\"very small value\" can be estimated by similarity of share and non share blocks\n```"
        }
      ]
    },
    {
      "id": 2794083,
      "postDate": "2024-05-05T06:08:30.703Z",
      "content": "<p>i read a paper an understand what the host is trying to do:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcc932ce8c13a7f5066be711d8da671cb%2FSelection_066.png?generation=1714889299525183&amp;alt=media\"></p>",
      "rawMarkdown": "i read a paper an understand what the host is trying to do:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcc932ce8c13a7f5066be711d8da671cb%2FSelection_066.png?generation=1714889299525183&alt=media)",
      "votes": 6,
      "replies": [
        {
          "id": 2794603,
          "postDate": "2024-05-05T12:01:39.937Z",
          "content": "<p>forget to put link to paper</p>\n<p>Machine learning on DNA-encoded library count data using an uncertainty-aware probabilistic loss<br>\nfunction<br>\n<a href=\"https://arxiv.org/pdf/2108.12471\" target=\"_blank\">https://arxiv.org/pdf/2108.12471</a><br>\n<a href=\"https://github.com/coleygroup/del_qsar\" target=\"_blank\">https://github.com/coleygroup/del_qsar</a></p>\n<p>related: Training Machine Learning-based QSAR models with Conformal Prediction on Experimental<br>\nData from DNA-Encoded Chemical Libraries<br>\n<a href=\"https://www.diva-portal.org/smash/get/diva2:1575162/FULLTEXT01.pdf\" target=\"_blank\">https://www.diva-portal.org/smash/get/diva2:1575162/FULLTEXT01.pdf</a></p>",
          "rawMarkdown": "forget to put link to paper\n\nMachine learning on DNA-encoded library count data using an uncertainty-aware probabilistic loss\nfunction\nhttps://arxiv.org/pdf/2108.12471\nhttps://github.com/coleygroup/del_qsar\n\nrelated: Training Machine Learning-based QSAR models with Conformal Prediction on Experimental\nData from DNA-Encoded Chemical Libraries\nhttps://www.diva-portal.org/smash/get/diva2:1575162/FULLTEXT01.pdf",
          "votes": 5,
          "replies": [
            {
              "id": 2796580,
              "postDate": "2024-05-06T09:57:33.530Z",
              "content": "<p>Indeed, conformal prediction has a long history of use by some of innovative companies like Astra Zeneca for drug discovery. </p>\n<p>More on all things conformal prediction -&gt; <a href=\"https://github.com/valeman/awesome-conformal-prediction\" target=\"_blank\">https://github.com/valeman/awesome-conformal-prediction</a></p>",
              "rawMarkdown": "Indeed, conformal prediction has a long history of use by some of innovative companies like Astra Zeneca for drug discovery. \n\nMore on all things conformal prediction -> https://github.com/valeman/awesome-conformal-prediction",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2794572,
      "postDate": "2024-05-05T11:36:51.430Z",
      "content": "<p>As BBs are share between BB2 and BB3, there is an interesting twist related to symmetry of the triazines:</p>\n<p>The synthesis procedure DNA codes for BB2 and BB3 separately, so if the reagent sets share BB reagents then there should be two synthesis routes leading to the same molecule if the BB2 and BB3 are different. I'll check if this is the case or if syntheses yielding the same final molecule have been combined in the hit scoring (which I doubt given the statistics in this contribution).<br>\nIn the first case, it might be interesting to judge how reliable the hit calling (or der sequential synthesis procedure) is in the data as [BBa,BBb,BBc] should hit the same target combination as [BBa,BBc,BBb].</p>",
      "rawMarkdown": "As BBs are share between BB2 and BB3, there is an interesting twist related to symmetry of the triazines:\n\nThe synthesis procedure DNA codes for BB2 and BB3 separately, so if the reagent sets share BB reagents then there should be two synthesis routes leading to the same molecule if the BB2 and BB3 are different. I'll check if this is the case or if syntheses yielding the same final molecule have been combined in the hit scoring (which I doubt given the statistics in this contribution).\nIn the first case, it might be interesting to judge how reliable the hit calling (or der sequential synthesis procedure) is in the data as [BBa,BBb,BBc] should hit the same target combination as [BBa,BBc,BBb].",
      "votes": 1,
      "replies": [
        {
          "id": 2794846,
          "postDate": "2024-05-05T14:33:59.657Z",
          "content": "<p>For the 691 BBs shared between BB2 and BB3 in train, if you pick any two of these and look for all matching results in train, they are only listed one of the two possible ways in train dataset. So this seems to be accounted for. </p>",
          "rawMarkdown": "For the 691 BBs shared between BB2 and BB3 in train, if you pick any two of these and look for all matching results in train, they are only listed one of the two possible ways in train dataset. So this seems to be accounted for. ",
          "votes": 2,
          "replies": [
            {
              "id": 2794936,
              "postDate": "2024-05-05T15:26:33.403Z",
              "content": "<p>Thanks for letting me know. I was about to check myself.</p>",
              "rawMarkdown": "Thanks for letting me know. I was about to check myself.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2808308,
      "postDate": "2024-05-12T06:17:32.510Z",
      "content": "<p>I admit, I don't quite fully understand the breakdown. What does group1 and group2 mean exactly?</p>",
      "rawMarkdown": "I admit, I don't quite fully understand the breakdown. What does group1 and group2 mean exactly?",
      "replies": [
        {
          "id": 2809205,
          "postDate": "2024-05-12T15:34:23.490Z",
          "content": "<p>For instance this quote?</p>\n<blockquote>\n  <p>group 1 OR group 2: 11,271 per protein.</p>\n</blockquote>\n<p>There are some \"nonshare\" building blocks in test, that DO have the triazine core. I called this subset group 1 <em>and</em> group 2, because when you look closer there are two subsets that don't share any overlap with each other either. The groupings simply mean building block groupings. Grouped so that we can simply say this:</p>\n<p>Basically:</p>\n<ul>\n<li>Shared group: huge group that includes all train data. No overlap with any other group's building blocks</li>\n<li>Triazine nonshare group 1. No overlap with any other group's building blocks</li>\n<li>Triazine nonshare group 2. No overlap with any other group's building blocks</li>\n<li>Non-triazine group. No overlap with any other group's building blocks.</li>\n</ul>",
          "rawMarkdown": "For instance this quote?\n\n> group 1 OR group 2: 11,271 per protein.\n\nThere are some \"nonshare\" building blocks in test, that DO have the triazine core. I called this subset group 1 *and* group 2, because when you look closer there are two subsets that don't share any overlap with each other either. The groupings simply mean building block groupings. Grouped so that we can simply say this:\n\nBasically:\n* Shared group: huge group that includes all train data. No overlap with any other group's building blocks\n* Triazine nonshare group 1. No overlap with any other group's building blocks\n* Triazine nonshare group 2. No overlap with any other group's building blocks\n* Non-triazine group. No overlap with any other group's building blocks.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2831242,
      "postDate": "2024-05-23T15:22:40.247Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2781137,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-28T15:37:53.557000",
      "content": "<p>to put into picture</p>\n<p>FOR TRAIN:<br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5db8dd160ffadd6ca30b43c2eabb8cf8%2FSelection_056.png?generation=1714318669316671&amp;alt=media\"></p>\n<p>FOR TEST:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F37d94487c09461f130fb48af3fa4ee2f%2FSelection_059.png?generation=1714355321671144&amp;alt=media\"></p>",
      "votes": 12,
      "replies": [
        {
          "id": 2781745,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-04-29T02:00:05.310000",
          "content": "<p>my guess on public and private split</p>\n<p>i train a classifier on the train share blocks</p>\n<ul>\n<li>train and validation  ap are almost similar</li>\n<li>public ap should be similar for the share block</li>\n</ul>\n<p>From experiments, applying classifier on non share block has almost zero detection</p>\n<pre><code>= r*validation + (very small value from non r= % of in public test = about %\n\n can estimated similarity of non </code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2794083,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-05T06:08:30.703000",
      "content": "<p>i read a paper an understand what the host is trying to do:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcc932ce8c13a7f5066be711d8da671cb%2FSelection_066.png?generation=1714889299525183&amp;alt=media\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 2794603,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-05-05T12:01:39.937000",
          "content": "<p>forget to put link to paper</p>\n<p>Machine learning on DNA-encoded library count data using an uncertainty-aware probabilistic loss<br>\nfunction<br>\n<a href=\"https://arxiv.org/pdf/2108.12471\" target=\"_blank\">https://arxiv.org/pdf/2108.12471</a><br>\n<a href=\"https://github.com/coleygroup/del_qsar\" target=\"_blank\">https://github.com/coleygroup/del_qsar</a></p>\n<p>related: Training Machine Learning-based QSAR models with Conformal Prediction on Experimental<br>\nData from DNA-Encoded Chemical Libraries<br>\n<a href=\"https://www.diva-portal.org/smash/get/diva2:1575162/FULLTEXT01.pdf\" target=\"_blank\">https://www.diva-portal.org/smash/get/diva2:1575162/FULLTEXT01.pdf</a></p>",
          "votes": 5,
          "replies": [
            {
              "id": 2796580,
              "author_name": "predict_addict",
              "author_url": "",
              "post_date": "2024-05-06T09:57:33.530000",
              "content": "<p>Indeed, conformal prediction has a long history of use by some of innovative companies like Astra Zeneca for drug discovery. </p>\n<p>More on all things conformal prediction -&gt; <a href=\"https://github.com/valeman/awesome-conformal-prediction\" target=\"_blank\">https://github.com/valeman/awesome-conformal-prediction</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2794572,
      "author_name": "Bernhard Rohde",
      "author_url": "",
      "post_date": "2024-05-05T11:36:51.430000",
      "content": "<p>As BBs are share between BB2 and BB3, there is an interesting twist related to symmetry of the triazines:</p>\n<p>The synthesis procedure DNA codes for BB2 and BB3 separately, so if the reagent sets share BB reagents then there should be two synthesis routes leading to the same molecule if the BB2 and BB3 are different. I'll check if this is the case or if syntheses yielding the same final molecule have been combined in the hit scoring (which I doubt given the statistics in this contribution).<br>\nIn the first case, it might be interesting to judge how reliable the hit calling (or der sequential synthesis procedure) is in the data as [BBa,BBb,BBc] should hit the same target combination as [BBa,BBc,BBb].</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2794846,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2024-05-05T14:33:59.657000",
          "content": "<p>For the 691 BBs shared between BB2 and BB3 in train, if you pick any two of these and look for all matching results in train, they are only listed one of the two possible ways in train dataset. So this seems to be accounted for. </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2794936,
              "author_name": "Bernhard Rohde",
              "author_url": "",
              "post_date": "2024-05-05T15:26:33.403000",
              "content": "<p>Thanks for letting me know. I was about to check myself.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2808308,
      "author_name": "Stephen Lee",
      "author_url": "",
      "post_date": "2024-05-12T06:17:32.510000",
      "content": "<p>I admit, I don't quite fully understand the breakdown. What does group1 and group2 mean exactly?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2809205,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2024-05-12T15:34:23.490000",
          "content": "<p>For instance this quote?</p>\n<blockquote>\n  <p>group 1 OR group 2: 11,271 per protein.</p>\n</blockquote>\n<p>There are some \"nonshare\" building blocks in test, that DO have the triazine core. I called this subset group 1 <em>and</em> group 2, because when you look closer there are two subsets that don't share any overlap with each other either. The groupings simply mean building block groupings. Grouped so that we can simply say this:</p>\n<p>Basically:</p>\n<ul>\n<li>Shared group: huge group that includes all train data. No overlap with any other group's building blocks</li>\n<li>Triazine nonshare group 1. No overlap with any other group's building blocks</li>\n<li>Triazine nonshare group 2. No overlap with any other group's building blocks</li>\n<li>Non-triazine group. No overlap with any other group's building blocks.</li>\n</ul>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2831242,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-23T15:22:40.247000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2766380": "Abbreviations: buildingblock1: BB1. Etc.\n# Train\n* 271 BB1\n* 691 BB2/3\n* 2 BB2 only with the 181 below (Interesting. Why?)\n* 181 BB3 only, paired with any BB2, just never paired with each other. (Interesting. Why?)\n\nAll combinations of BB2 and BB3:\n181x693 for all combinations that include the 181 and/or the 2.\n691x692 / 2 for all combinations of the 691 with each other. The same BB is allowed and indeed happens exactly 691 times. For any x,y pair, it will always be listed in one order, not the other order. (Which is where the divide by 2 comes in).\n\nThe interesting and nice thing is the math checks out! Total of ALL possible combinations:\n271x((181x693) + (691x692/2)) = 98,784,649\n\nBut wait, there's only 98.4 million in train!\n\nWell:\n98784649 - 98415610 = 369039\n\nThose 369039, times 3 protein targets, equals: 1,107,117 test rows. Which is exactly the number of test rows with shared building blocks from train in the test dataset.\n\n# Test\n* 1,107,117 rows (369,039 molecules) with shared BBs, from above calculations.\n* 33,813 rows (11,271 molecules) from: 17x((2x34)+(34x35/2)) for triazine core, non-shared BBs group1 with: 17 BB1, 34 BB2/3, 2 BB3 only.\n* 33,966 rows (11,322 molecules) from: 17x(36x37/2) for triazine core, non-shared BBs group2 with: 17BB1, 36BB2/3.\n  * Note: again, permutation numbers exactly match the 67,779 rows and 22593 molecules with triazine core and non-shared BBs in test dataset.\n* 500,000 rows (486390 molecules): non-triazine core (credit [this notebook](https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration/notebook)), non-shared BBs: 36 BB1, 377 BB2 only, 446 BB3 only. No obvious reason or pattern for the number of molecules in test, it looks at first few glances like a random selection of 500,000 from all possible proteins and BB permutations: 3x36x377x446. Only 13k molecules sampled two or 3 times seemed low to me at first glance, if random, but actually the math checks out, ranging from 0 to 1/18 chance if every prior pick was a new molecule. So about 1/36 * 500,000 ~= 13k, matching observed data.\n\nNote that, in terms of a rigorous experiment testing setting, it would NOT make any sense to have that final 500,000 rows in both \"validation\" (aka public LB) and \"test\" (aka private LB), and the hosts wrote:\n> 200K validation examples per protein, and 360K test molecules per protein\n\nSo by far the most likely possibility is:\n\n### Public LB\n* 50% of shared BBs: 184,519 per protein.\n* group 1 OR group 2: 11,271 per protein.\n\nRounded to nearest 10K equals 200k \"validation\" per protein\n\n### Private LB\n* 50% of shared BBs: 184,520 per protein\n* group 1 OR group 2: 11,322 per protein.\n* The non-triazine core group: 166,667 per protein on average.\n\nRounded to nearest 10K: 360K \"test\" per protein",
    "2781137": "to put into picture\n\nFOR TRAIN:\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5db8dd160ffadd6ca30b43c2eabb8cf8%2FSelection_056.png?generation=1714318669316671&alt=media)\n\nFOR TEST:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F37d94487c09461f130fb48af3fa4ee2f%2FSelection_059.png?generation=1714355321671144&alt=media)\n",
    "2794083": "i read a paper an understand what the host is trying to do:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcc932ce8c13a7f5066be711d8da671cb%2FSelection_066.png?generation=1714889299525183&alt=media)",
    "2794572": "As BBs are share between BB2 and BB3, there is an interesting twist related to symmetry of the triazines:\n\nThe synthesis procedure DNA codes for BB2 and BB3 separately, so if the reagent sets share BB reagents then there should be two synthesis routes leading to the same molecule if the BB2 and BB3 are different. I'll check if this is the case or if syntheses yielding the same final molecule have been combined in the hit scoring (which I doubt given the statistics in this contribution).\nIn the first case, it might be interesting to judge how reliable the hit calling (or der sequential synthesis procedure) is in the data as [BBa,BBb,BBc] should hit the same target combination as [BBa,BBc,BBb].",
    "2808308": "I admit, I don't quite fully understand the breakdown. What does group1 and group2 mean exactly?",
    "2831242": ""
  }
}