{
  "id": 509015,
  "title": "Local validation for molecules with nonshared BBs",
  "url": "/competitions/leash-BELKA/discussion/509015",
  "author_name": "",
  "post_date": "2024-05-31T19:32:11.060020900Z",
  "votes": 11,
  "comment_count": 9,
  "views": 0,
  "content": "<p>We faced with absence of correlation between local validation and LB scores for each target (nonshared BBs; shared BBs correlate perfectly).<br>\nLocal scores are &lt;0.1 for BRD4 and HSA, and LB scores for submit with only BRD4 or HSA for nonshared BBs be like 0.09 and 0.05. So actual AP on LB for both targets should be ~0.4 to get such scores in average with 0 submission for other targets and shared BBs (0 or any constant submission get score 0.023 - it is ratio of positive samples). <br>\n5/6 * 0.023 + 1/6 * x = 0.09, so x = 0.425 - quite far from any validation score in our experiments or in published results in <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/505985\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/505985</a> and other topics. </p>",
  "messages": [
    {
      "id": "2848105",
      "postDate": "05/31/2024 19:32:11",
      "content": "<p>We faced with absence of correlation between local validation and LB scores for each target (nonshared BBs; shared BBs correlate perfectly).<br>\nLocal scores are &lt;0.1 for BRD4 and HSA, and LB scores for submit with only BRD4 or HSA for nonshared BBs be like 0.09 and 0.05. So actual AP on LB for both targets should be ~0.4 to get such scores in average with 0 submission for other targets and shared BBs (0 or any constant submission get score 0.023 - it is ratio of positive samples). <br>\n5/6 * 0.023 + 1/6 * x = 0.09, so x = 0.425 - quite far from any validation score in our experiments or in published results in <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/505985\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/505985</a> and other topics. </p>",
      "rawMarkdown": "We faced with absence of correlation between local validation and LB scores for each target (nonshared BBs; shared BBs correlate perfectly).\nLocal scores are <0.1 for BRD4 and HSA, and LB scores for submit with only BRD4 or HSA for nonshared BBs be like 0.09 and 0.05. So actual AP on LB for both targets should be ~0.4 to get such scores in average with 0 submission for other targets and shared BBs (0 or any constant submission get score 0.023 - it is ratio of positive samples). \n5/6 * 0.023 + 1/6 * x = 0.09, so x = 0.425 - quite far from any validation score in our experiments or in published results in https://www.kaggle.com/competitions/leash-BELKA/discussion/505985 and other topics.",
      "votes": null
    },
    {
      "id": "2848360",
      "postDate": "06/01/2024 02:08:48",
      "content": "<p>\"quite far from any validation score in our experiments or in published results in https://www.kaggle.com/competitions/leash-BELKA/discussion/505985\"</p>\n<p>it just mean my public score for nonshare bb is very high, probably around 0.2 to 0.3.<br>\n(i.e. public lb = share + nonshare = (0.62 + 0.22)/2)<br>\nThis could be \"false results\" because number of nonshare molecules in public test is very small. (it is easier to get good nonshare results after \"many trials\")</p>\n<p>\"Local scores are &lt;0.1 for BRD4 and HSA, a …\"<br>\nif you check your training log carefully, you may see fluctuation. don't look at the final metric. check the whole loss curve vs training iteration. (there may be overfitting for nonshare and maybe earlier models generalised better)</p>\n<p>lastly, since there is domain shift, you need NOT submit one models for both share and nonshare.<br>\nyou can made separate models for them.</p>",
      "rawMarkdown": "\"quite far from any validation score in our experiments or in published results in https://www.kaggle.com/competitions/leash-BELKA/discussion/505985\"\n\nit just mean my public score for nonshare bb is very high, probably around 0.2 to 0.3.\n(i.e. public lb = share + nonshare = (0.62 + 0.22)/2)\nThis could be \"false results\" because number of nonshare molecules in public test is very small. (it is easier to get good nonshare results after \"many trials\")\n\n\"Local scores are <0.1 for BRD4 and HSA, a ...\"\nif you check your training log carefully, you may see fluctuation. don't look at the final metric. check the whole loss curve vs training iteration. (there may be overfitting for nonshare and maybe earlier models generalised better)\n\nlastly, since there is domain shift, you need NOT submit one models for both share and nonshare.\nyou can made separate models for them.",
      "votes": null
    },
    {
      "id": "2848577",
      "postDate": "06/01/2024 05:05:00",
      "content": "<p>It is true that we can get \"spurious results\" with a small nonshared public test size, cuz small changes in model predictions or changes of class balance in the training data can lead to relatively large changes in the LB MAP score. So bettter to not look at LB score, agree. But the private test for this group (nonshared triazines) is similar in size and accounts for 1/3 of the total score. This means that neither the public nor the private LB scores for the nonshared part can accurately reflect the true performance of the model on a larger, more diverse data set.</p>\n<p>So in my understanding, even if we had a representative non-shared part with a sufficient sample size for training and local validation, and found a model that generalizes well, we might get randomly biased scores on the public/private tests due to its small size and specific classes balance for each target. </p>",
      "rawMarkdown": "It is true that we can get \"spurious results\" with a small nonshared public test size, cuz small changes in model predictions or changes of class balance in the training data can lead to relatively large changes in the LB MAP score. So bettter to not look at LB score, agree. But the private test for this group (nonshared triazines) is similar in size and accounts for 1/3 of the total score. This means that neither the public nor the private LB scores for the nonshared part can accurately reflect the true performance of the model on a larger, more diverse data set.\n\nSo in my understanding, even if we had a representative non-shared part with a sufficient sample size for training and local validation, and found a model that generalizes well, we might get randomly biased scores on the public/private tests due to its small size and specific classes balance for each target.",
      "votes": null
    },
    {
      "id": "2848598",
      "postDate": "06/01/2024 05:15:33",
      "content": "<p>\"But the private test for this group (nonshared triazines) is similar in size and accounts for 1/3 of the total score. This means that neither the public nor the private LB scores for the nonshared part can accurately reflect the true performance of the model on a larger, more diverse data set. …\"</p>\n<p>maybe there will be large shakeup … if we just use pure machine learning.</p>\n<hr>\n<p>\" and found a model that generalizes well, we might get randomly biased scores on the public/private tests due to its small size and specific classes balance to each target. \"</p>\n<p>This is really depends on the hidden private data. If the hosts are \"kind\", they would have selected the hidden test molecules that are \"somehow similar\"  to the train. But if they really want to do a \"stress test\", the train and test could be something very different. We need a chemist to look at the train and test data to tell us what is the expected generalisation.</p>\n<hr>\n<p>but then, computation drug discovery has been there quite a while. i wonder how did the large drug company discover hits? I believe that those expensive computation simulation should help.</p>",
      "rawMarkdown": "\"But the private test for this group (nonshared triazines) is similar in size and accounts for 1/3 of the total score. This means that neither the public nor the private LB scores for the nonshared part can accurately reflect the true performance of the model on a larger, more diverse data set. ...\"\n\nmaybe there will be large shakeup ... if we just use pure machine learning.\n\n---\n\n\" and found a model that generalizes well, we might get randomly biased scores on the public/private tests due to its small size and specific classes balance to each target. \"\n\nThis is really depends on the hidden private data. If the hosts are \"kind\", they would have selected the hidden test molecules that are \"somehow similar\"  to the train. But if they really want to do a \"stress test\", the train and test could be something very different. We need a chemist to look at the train and test data to tell us what is the expected generalisation.\n\n\n---\n\nbut then, computation drug discovery has been there quite a while. i wonder how did the large drug company discover hits? I believe that those expensive computation simulation should help.",
      "votes": null
    },
    {
      "id": "2848679",
      "postDate": "06/01/2024 05:43:21",
      "content": "<blockquote>\n  <p>large drug company</p>\n</blockquote>\n<p>Big pharma do perform enrichment, i.e. identification of topN most promising molecules among compounds library for further experimental investigation. In general, drug discovery is not about binary classification.</p>",
      "rawMarkdown": ">large drug company\n\nBig pharma do perform enrichment, i.e. identification of topN most promising molecules among compounds library for further experimental investigation. In general, drug discovery is not about binary classification.",
      "votes": null
    },
    {
      "id": "2848728",
      "postDate": "06/01/2024 06:28:11",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F682731%2Ffb19668d3452bab6a823351973f6ade6%2FBRD4.PNG?generation=1717222942717364&amp;alt=media\"><br>\nHere is an example of training curves for BRD4. The same fluctuations range is well reproduced between models, runs, features and validation sets.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F682731%2Ffb19668d3452bab6a823351973f6ade6%2FBRD4.PNG?generation=1717222942717364&alt=media)\nHere is an example of training curves for BRD4. The same fluctuations range is well reproduced between models, runs, features and validation sets.",
      "votes": null
    },
    {
      "id": "2848820",
      "postDate": "06/01/2024 08:26:18",
      "content": "<p>note that the hit fraction of the test is not the same as train. we only know train has about 0.005, test is 0.023-0.005=0.018.</p>\n<p>you have to do you experiments with hit fractions 0.018 for  22593 validation molecules.<br>\nfurther, here are about 35 unique nonshare test block1, each has about 666 molecules.</p>\n<p>in my experiment, depends on how i sample the validation molecules, validation AP can go between 0.050 to 0.150.</p>\n<pre><code>len(test_df)\n\n\n(( test_df.triazine) &amp;(test_df.   \n((~test_df.triazine) &amp;(test_df.   \n(( test_df.triazine) &amp;(test_df.nonshare)).mean() \n((~test_df.triazine) &amp;(test_df.nonshare)).mean() \n\ndf = test df for nonshare \n\ndf., \n\nlen(df.\nlen(df.\nlen(df.\n\n\n* = \nthey only sample  (is this ? are the already enrichments?)\n</code></pre>",
      "rawMarkdown": "note that the hit fraction of the test is not the same as train. we only know train has about 0.005, test is 0.023-0.005=0.018.\n\nyou have to do you experiments with hit fractions 0.018 for  22593 validation molecules.\nfurther, here are about 35 unique nonshare test block1, each has about 666 molecules.\n\nin my experiment, depends on how i sample the validation molecules, validation AP can go between 0.050 to 0.150.\n\n\n```\nlen(test_df)\n878022\n'''\n(( test_df.triazine) &(test_df.share)).mean()    # 0.4203072360373658\n((~test_df.triazine) &(test_df.share)).mean()    # 0\n(( test_df.triazine) &(test_df.nonshare)).mean() # 0.025731701483561915 #22593\n((~test_df.triazine) &(test_df.nonshare)).mean() # 0.5539610624790723\n\ndf = test df for nonshare #22593\n\ndf.buildingblock1_id.value_counts()\n666, 663\n\nlen(df.buildingblock1_id.unique())\n34\nlen(df.buildingblock2_id.unique())\n70\nlen(df.buildingblock3_id.unique())\n72\n\n\n70*72 = 5040\nbut they only sample 666 (is this random? or are the BB already enrichments?)\n```",
      "votes": null
    },
    {
      "id": "2848833",
      "postDate": "06/01/2024 08:33:20",
      "content": "<blockquote>\n  <p>test is 0.023-0.005=0.018</p>\n</blockquote>\n<p>Test is 0.023*2-0.005=0.041! 0.023 is a mean ratio for shared and nonhared, so (0.041+0.005)/2=0.023 (if we assume test ratio for shared BBs equals train ratio).</p>",
      "rawMarkdown": ">test is 0.023-0.005=0.018\n\nTest is 0.023*2-0.005=0.041! 0.023 is a mean ratio for shared and nonhared, so (0.041+0.005)/2=0.023 (if we assume test ratio for shared BBs equals train ratio).",
      "votes": null
    },
    {
      "id": "2848837",
      "postDate": "06/01/2024 08:34:26",
      "content": "<blockquote>\n  <p>but they only sample 666 (is this random?)</p>\n</blockquote>\n<p>I'm pretty sure that non-random selection is a key. Differences between scores are too big to explain it by variation in positive samples fraction.</p>",
      "rawMarkdown": ">but they only sample 666 (is this random?)\n\nI'm pretty sure that non-random selection is a key. Differences between scores are too big to explain it by variation in positive samples fraction.",
      "votes": null
    },
    {
      "id": "2848840",
      "postDate": "06/01/2024 08:38:52",
      "content": "<p>results are very different for different validation splits. the last 3 cols in the validation are non share<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F67dbeab3e65f1e60509e03622df21fd3%2FSelection_169.png?generation=1717231130275116&amp;alt=media\"></p>\n<p>in other experiments, SEH could be higher, etc</p>",
      "rawMarkdown": "results are very different for different validation splits. the last 3 cols in the validation are non share\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F67dbeab3e65f1e60509e03622df21fd3%2FSelection_169.png?generation=1717231130275116&alt=media)\n\nin other experiments, SEH could be higher, etc",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2848360,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/01/2024 02:08:48",
      "content": "<p>\"quite far from any validation score in our experiments or in published results in https://www.kaggle.com/competitions/leash-BELKA/discussion/505985\"</p>\n<p>it just mean my public score for nonshare bb is very high, probably around 0.2 to 0.3.<br>\n(i.e. public lb = share + nonshare = (0.62 + 0.22)/2)<br>\nThis could be \"false results\" because number of nonshare molecules in public test is very small. (it is easier to get good nonshare results after \"many trials\")</p>\n<p>\"Local scores are &lt;0.1 for BRD4 and HSA, a …\"<br>\nif you check your training log carefully, you may see fluctuation. don't look at the final metric. check the whole loss curve vs training iteration. (there may be overfitting for nonshare and maybe earlier models generalised better)</p>\n<p>lastly, since there is domain shift, you need NOT submit one models for both share and nonshare.<br>\nyou can made separate models for them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2848577,
          "author_name": "antoninadolgorukova",
          "author_url": "",
          "post_date": "06/01/2024 05:05:00",
          "content": "<p>It is true that we can get \"spurious results\" with a small nonshared public test size, cuz small changes in model predictions or changes of class balance in the training data can lead to relatively large changes in the LB MAP score. So bettter to not look at LB score, agree. But the private test for this group (nonshared triazines) is similar in size and accounts for 1/3 of the total score. This means that neither the public nor the private LB scores for the nonshared part can accurately reflect the true performance of the model on a larger, more diverse data set.</p>\n<p>So in my understanding, even if we had a representative non-shared part with a sufficient sample size for training and local validation, and found a model that generalizes well, we might get randomly biased scores on the public/private tests due to its small size and specific classes balance for each target. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2848598,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "06/01/2024 05:15:33",
              "content": "<p>\"But the private test for this group (nonshared triazines) is similar in size and accounts for 1/3 of the total score. This means that neither the public nor the private LB scores for the nonshared part can accurately reflect the true performance of the model on a larger, more diverse data set. …\"</p>\n<p>maybe there will be large shakeup … if we just use pure machine learning.</p>\n<hr>\n<p>\" and found a model that generalizes well, we might get randomly biased scores on the public/private tests due to its small size and specific classes balance to each target. \"</p>\n<p>This is really depends on the hidden private data. If the hosts are \"kind\", they would have selected the hidden test molecules that are \"somehow similar\"  to the train. But if they really want to do a \"stress test\", the train and test could be something very different. We need a chemist to look at the train and test data to tell us what is the expected generalisation.</p>\n<hr>\n<p>but then, computation drug discovery has been there quite a while. i wonder how did the large drug company discover hits? I believe that those expensive computation simulation should help.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2848679,
                  "author_name": "ogurtsov",
                  "author_url": "",
                  "post_date": "06/01/2024 05:43:21",
                  "content": "<blockquote>\n  <p>large drug company</p>\n</blockquote>\n<p>Big pharma do perform enrichment, i.e. identification of topN most promising molecules among compounds library for further experimental investigation. In general, drug discovery is not about binary classification.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2848728,
          "author_name": "ogurtsov",
          "author_url": "",
          "post_date": "06/01/2024 06:28:11",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F682731%2Ffb19668d3452bab6a823351973f6ade6%2FBRD4.PNG?generation=1717222942717364&amp;alt=media\"><br>\nHere is an example of training curves for BRD4. The same fluctuations range is well reproduced between models, runs, features and validation sets.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2848820,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "06/01/2024 08:26:18",
              "content": "<p>note that the hit fraction of the test is not the same as train. we only know train has about 0.005, test is 0.023-0.005=0.018.</p>\n<p>you have to do you experiments with hit fractions 0.018 for  22593 validation molecules.<br>\nfurther, here are about 35 unique nonshare test block1, each has about 666 molecules.</p>\n<p>in my experiment, depends on how i sample the validation molecules, validation AP can go between 0.050 to 0.150.</p>\n<pre><code>len(test_df)\n\n\n(( test_df.triazine) &amp;(test_df.   \n((~test_df.triazine) &amp;(test_df.   \n(( test_df.triazine) &amp;(test_df.nonshare)).mean() \n((~test_df.triazine) &amp;(test_df.nonshare)).mean() \n\ndf = test df for nonshare \n\ndf., \n\nlen(df.\nlen(df.\nlen(df.\n\n\n* = \nthey only sample  (is this ? are the already enrichments?)\n</code></pre>",
              "votes": null,
              "replies": [
                {
                  "id": 2848833,
                  "author_name": "ogurtsov",
                  "author_url": "",
                  "post_date": "06/01/2024 08:33:20",
                  "content": "<blockquote>\n  <p>test is 0.023-0.005=0.018</p>\n</blockquote>\n<p>Test is 0.023*2-0.005=0.041! 0.023 is a mean ratio for shared and nonhared, so (0.041+0.005)/2=0.023 (if we assume test ratio for shared BBs equals train ratio).</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2848837,
                  "author_name": "ogurtsov",
                  "author_url": "",
                  "post_date": "06/01/2024 08:34:26",
                  "content": "<blockquote>\n  <p>but they only sample 666 (is this random?)</p>\n</blockquote>\n<p>I'm pretty sure that non-random selection is a key. Differences between scores are too big to explain it by variation in positive samples fraction.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2848840,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "06/01/2024 08:38:52",
                  "content": "<p>results are very different for different validation splits. the last 3 cols in the validation are non share<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F67dbeab3e65f1e60509e03622df21fd3%2FSelection_169.png?generation=1717231130275116&amp;alt=media\"></p>\n<p>in other experiments, SEH could be higher, etc</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2848105": "We faced with absence of correlation between local validation and LB scores for each target (nonshared BBs; shared BBs correlate perfectly).\nLocal scores are <0.1 for BRD4 and HSA, and LB scores for submit with only BRD4 or HSA for nonshared BBs be like 0.09 and 0.05. So actual AP on LB for both targets should be ~0.4 to get such scores in average with 0 submission for other targets and shared BBs (0 or any constant submission get score 0.023 - it is ratio of positive samples). \n5/6 * 0.023 + 1/6 * x = 0.09, so x = 0.425 - quite far from any validation score in our experiments or in published results in https://www.kaggle.com/competitions/leash-BELKA/discussion/505985 and other topics.",
    "2848360": "\"quite far from any validation score in our experiments or in published results in https://www.kaggle.com/competitions/leash-BELKA/discussion/505985\"\n\nit just mean my public score for nonshare bb is very high, probably around 0.2 to 0.3.\n(i.e. public lb = share + nonshare = (0.62 + 0.22)/2)\nThis could be \"false results\" because number of nonshare molecules in public test is very small. (it is easier to get good nonshare results after \"many trials\")\n\n\"Local scores are <0.1 for BRD4 and HSA, a ...\"\nif you check your training log carefully, you may see fluctuation. don't look at the final metric. check the whole loss curve vs training iteration. (there may be overfitting for nonshare and maybe earlier models generalised better)\n\nlastly, since there is domain shift, you need NOT submit one models for both share and nonshare.\nyou can made separate models for them.",
    "2848577": "It is true that we can get \"spurious results\" with a small nonshared public test size, cuz small changes in model predictions or changes of class balance in the training data can lead to relatively large changes in the LB MAP score. So bettter to not look at LB score, agree. But the private test for this group (nonshared triazines) is similar in size and accounts for 1/3 of the total score. This means that neither the public nor the private LB scores for the nonshared part can accurately reflect the true performance of the model on a larger, more diverse data set.\n\nSo in my understanding, even if we had a representative non-shared part with a sufficient sample size for training and local validation, and found a model that generalizes well, we might get randomly biased scores on the public/private tests due to its small size and specific classes balance for each target.",
    "2848598": "\"But the private test for this group (nonshared triazines) is similar in size and accounts for 1/3 of the total score. This means that neither the public nor the private LB scores for the nonshared part can accurately reflect the true performance of the model on a larger, more diverse data set. ...\"\n\nmaybe there will be large shakeup ... if we just use pure machine learning.\n\n---\n\n\" and found a model that generalizes well, we might get randomly biased scores on the public/private tests due to its small size and specific classes balance to each target. \"\n\nThis is really depends on the hidden private data. If the hosts are \"kind\", they would have selected the hidden test molecules that are \"somehow similar\"  to the train. But if they really want to do a \"stress test\", the train and test could be something very different. We need a chemist to look at the train and test data to tell us what is the expected generalisation.\n\n\n---\n\nbut then, computation drug discovery has been there quite a while. i wonder how did the large drug company discover hits? I believe that those expensive computation simulation should help.",
    "2848679": ">large drug company\n\nBig pharma do perform enrichment, i.e. identification of topN most promising molecules among compounds library for further experimental investigation. In general, drug discovery is not about binary classification.",
    "2848728": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F682731%2Ffb19668d3452bab6a823351973f6ade6%2FBRD4.PNG?generation=1717222942717364&alt=media)\nHere is an example of training curves for BRD4. The same fluctuations range is well reproduced between models, runs, features and validation sets.",
    "2848820": "note that the hit fraction of the test is not the same as train. we only know train has about 0.005, test is 0.023-0.005=0.018.\n\nyou have to do you experiments with hit fractions 0.018 for  22593 validation molecules.\nfurther, here are about 35 unique nonshare test block1, each has about 666 molecules.\n\nin my experiment, depends on how i sample the validation molecules, validation AP can go between 0.050 to 0.150.\n\n\n```\nlen(test_df)\n878022\n'''\n(( test_df.triazine) &(test_df.share)).mean()    # 0.4203072360373658\n((~test_df.triazine) &(test_df.share)).mean()    # 0\n(( test_df.triazine) &(test_df.nonshare)).mean() # 0.025731701483561915 #22593\n((~test_df.triazine) &(test_df.nonshare)).mean() # 0.5539610624790723\n\ndf = test df for nonshare #22593\n\ndf.buildingblock1_id.value_counts()\n666, 663\n\nlen(df.buildingblock1_id.unique())\n34\nlen(df.buildingblock2_id.unique())\n70\nlen(df.buildingblock3_id.unique())\n72\n\n\n70*72 = 5040\nbut they only sample 666 (is this random? or are the BB already enrichments?)\n```",
    "2848833": ">test is 0.023-0.005=0.018\n\nTest is 0.023*2-0.005=0.041! 0.023 is a mean ratio for shared and nonhared, so (0.041+0.005)/2=0.023 (if we assume test ratio for shared BBs equals train ratio).",
    "2848837": ">but they only sample 666 (is this random?)\n\nI'm pretty sure that non-random selection is a key. Differences between scores are too big to explain it by variation in positive samples fraction.",
    "2848840": "results are very different for different validation splits. the last 3 cols in the validation are non share\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F67dbeab3e65f1e60509e03622df21fd3%2FSelection_169.png?generation=1717231130275116&alt=media)\n\nin other experiments, SEH could be higher, etc"
  },
  "source": "meta"
}