{
  "id": 501140,
  "title": "[maybe obsolete] are you overfitting model to sharing blocks?",
  "url": "/competitions/leash-BELKA/discussion/501140",
  "author_name": "hengck23",
  "post_date": "2024-05-08T07:51:55.955000",
  "votes": 9,
  "comment_count": 20,
  "views": 0,
  "content": "<p><strong>UPDATED!! the scoring metric has been updated. so some strategy and discussion may not be valid anymore !!!</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F18c0f3c299addf7c9ae36692d2545c19%2FSelection_080.png?generation=1715154492163145&amp;alt=media\"></p>\n<p>i made a third submission:</p>\n<pre><code> =model2_submit_df.copy()\nreplaced.loc[nonshare,]=model1_submit_df.loc[nonshare,]\n\n\nlbscore = .\n</code></pre>\n<p>implications:</p>\n<ul>\n<li>nonsharing block indeed contribute to public LB score</li>\n<li>be care in validation and interpretation of local and public scores</li>\n</ul>",
  "messages": [
    {
      "id": 2800473,
      "postDate": "2024-05-08T07:51:55.957Z",
      "content": "<p><strong>UPDATED!! the scoring metric has been updated. so some strategy and discussion may not be valid anymore !!!</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F18c0f3c299addf7c9ae36692d2545c19%2FSelection_080.png?generation=1715154492163145&amp;alt=media\"></p>\n<p>i made a third submission:</p>\n<pre><code> =model2_submit_df.copy()\nreplaced.loc[nonshare,]=model1_submit_df.loc[nonshare,]\n\n\nlbscore = .\n</code></pre>\n<p>implications:</p>\n<ul>\n<li>nonsharing block indeed contribute to public LB score</li>\n<li>be care in validation and interpretation of local and public scores</li>\n</ul>",
      "rawMarkdown": "**UPDATED!! the scoring metric has been updated. so some strategy and discussion may not be valid anymore !!!**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F18c0f3c299addf7c9ae36692d2545c19%2FSelection_080.png?generation=1715154492163145&alt=media)\n\ni made a third submission:\n```\nreplaced =model2_submit_df.copy()\nreplaced.loc[nonshare,'binds']=model1_submit_df.loc[nonshare,'binds']\n# this is actually not very correct because of different calibration of the two models\n\nlbscore = 0.612\n```\n\nimplications:\n- nonsharing block indeed contribute to public LB score\n- be care in validation and interpretation of local and public scores",
      "votes": 9
    },
    {
      "id": 2802312,
      "postDate": "2024-05-09T01:56:57.627Z",
      "content": "<p>maybe very good news?</p>\n<p>umap of ecfp fingerprint of building blocks<br>\nblue : share bb<br>\nred : nonshare bb<br>\ngrey: external bb (external data from my other post)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fab4bcf5987371c81571a1b677ee11ee4%2FSelection_082.png?generation=1715219732483573&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fad5c52f7a81c1c28a1e7f5ed93b90d87%2FSelection_083.png?generation=1715219744081586&amp;alt=media\"></p>",
      "rawMarkdown": "maybe very good news?\n\numap of ecfp fingerprint of building blocks\nblue : share bb\nred : nonshare bb\ngrey: external bb (external data from my other post)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fab4bcf5987371c81571a1b677ee11ee4%2FSelection_082.png?generation=1715219732483573&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fad5c52f7a81c1c28a1e7f5ed93b90d87%2FSelection_083.png?generation=1715219744081586&alt=media)",
      "votes": 3,
      "replies": [
        {
          "id": 2802386,
          "postDate": "2024-05-09T02:50:42.140Z",
          "content": "<p>The external data is the sEH data you posted about weeks(?) ago? </p>",
          "rawMarkdown": "The external data is the sEH data you posted about weeks(?) ago? ",
          "replies": [
            {
              "id": 2802419,
              "postDate": "2024-05-09T03:32:53.933Z",
              "content": "<p>yes. if it gives binding affinity as readout values (not 0,1 label).<br>\nif you have ideas how to use that, please let me know.</p>\n<p>right now i can only think of ranking loss for external data</p>\n<p>update: i figure it out. there is overlap between blue and gray  that can used to transform targets of different dataset</p>",
              "rawMarkdown": "yes. if it gives binding affinity as readout values (not 0,1 label).\nif you have ideas how to use that, please let me know.\n\nright now i can only think of ranking loss for external data\n\nupdate: i figure it out. there is overlap between blue and gray  that can used to transform targets of different dataset"
            },
            {
              "id": 2802425,
              "postDate": "2024-05-09T03:35:23.703Z",
              "content": "<p>\"The 133M small molecule library used here, AMA014, was provided by AlphaMa\"</p>\n<p>it is probably easy to find compounds similar to AMA014</p>",
              "rawMarkdown": "\"The 133M small molecule library used here, AMA014, was provided by AlphaMa\"\n\nit is probably easy to find compounds similar to AMA014"
            },
            {
              "id": 2803330,
              "postDate": "2024-05-09T12:33:12.993Z",
              "content": "<p>You can convert to binary with all non-zero label = 1. That's actually pretty equivalent to the train data of this competition, the external data is just giving more info that we don't have access to with the competition hosted data. </p>",
              "rawMarkdown": "You can convert to binary with all non-zero label = 1. That's actually pretty equivalent to the train data of this competition, the external data is just giving more info that we don't have access to with the competition hosted data. "
            }
          ]
        }
      ]
    },
    {
      "id": 2812021,
      "postDate": "2024-05-14T04:20:55.247Z",
      "content": "<p>i have new results and prove that my observations are correct:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0e4305dd56d5c61c6561cbccb5ad9617%2FSelection_112.png?generation=1715660177707043&amp;alt=media\"></p>\n<ul>\n<li>if you overfit share blocks, nonshare blocks can get worse</li>\n<li>CV/LB of share block correlates very strong.</li>\n</ul>\n<p>if you perform worse than other kagglers who you using the model architecture and same CV, it could be because of the nonshare part. hence those top in the leaderboard has better nonshare block score then you (either via training or artificial adjusting). I think top 20 kagglers have now about same share lb score about 0.595.</p>\n<p>i think the top peforming cv of share is about 0.73 (one model)</p>",
      "rawMarkdown": "i have new results and prove that my observations are correct:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0e4305dd56d5c61c6561cbccb5ad9617%2FSelection_112.png?generation=1715660177707043&alt=media)\n\n- if you overfit share blocks, nonshare blocks can get worse\n- CV/LB of share block correlates very strong.\n\nif you perform worse than other kagglers who you using the model architecture and same CV, it could be because of the nonshare part. hence those top in the leaderboard has better nonshare block score then you (either via training or artificial adjusting). I think top 20 kagglers have now about same share lb score about 0.595.\n\ni think the top peforming cv of share is about 0.73 (one model)",
      "votes": 1
    },
    {
      "id": 2800527,
      "postDate": "2024-05-08T08:08:06.530Z",
      "content": "<p>while it is difficult to find external data that has pos labels, i think it is easier to find those with neg labels. can anyone advised?</p>",
      "rawMarkdown": "while it is difficult to find external data that has pos labels, i think it is easier to find those with neg labels. can anyone advised?",
      "votes": 1
    },
    {
      "id": 2816163,
      "postDate": "2024-05-16T07:33:53.090Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F481be0a0f5cf01848f5f1f9506254401%2FSelection_126.png?generation=1715844792449133&amp;alt=media\"></p>\n<p>YES!!! training with external dataset improves  nonshare LB, but by a little</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F481be0a0f5cf01848f5f1f9506254401%2FSelection_126.png?generation=1715844792449133&alt=media)\n\nYES!!! training with external dataset improves  nonshare LB, but by a little",
      "votes": 2
    },
    {
      "id": 2803679,
      "postDate": "2024-05-09T15:36:23.593Z",
      "content": "<p>at first i though it is the question if we can generalised, i am wrong.<br>\nit is whether we can detect at all !!!<br>\n(or maybe external data protein segment is different, or experiment setting, etc … let's wait for the lb score)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F13b1bde7ec5435bcb59b8356a8012df8%2FSelection_092.png?generation=1715309251814154&amp;alt=media\"></p>\n<p>i have to change my strategy</p>",
      "rawMarkdown": "at first i though it is the question if we can generalised, i am wrong.\nit is whether we can detect at all !!!\n(or maybe external data protein segment is different, or experiment setting, etc ... let's wait for the lb score)\n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F13b1bde7ec5435bcb59b8356a8012df8%2FSelection_092.png?generation=1715309251814154&alt=media)\n\ni have to change my strategy",
      "votes": 2,
      "replies": [
        {
          "id": 2803782,
          "postDate": "2024-05-09T16:05:22.747Z",
          "content": "<p>That's your score matrix for train on one dataset, predict on both? Neither dataset is able to predict the other hardly at all, it appears? That was just with xgb on ECFP?</p>\n<p>Yikes that's low though</p>",
          "rawMarkdown": "That's your score matrix for train on one dataset, predict on both? Neither dataset is able to predict the other hardly at all, it appears? That was just with xgb on ECFP?\n\nYikes that's low though",
          "votes": 1,
          "replies": [
            {
              "id": 2803789,
              "postDate": "2024-05-09T16:05:55.900Z",
              "content": "<p>I wish it surprised me more, lol. </p>",
              "rawMarkdown": "I wish it surprised me more, lol. ",
              "votes": 1
            },
            {
              "id": 2803801,
              "postDate": "2024-05-09T16:08:53.423Z",
              "content": "<p>\" Neither dataset is able to predict the other hardly at all, it appears? \"<br>\nThat is true.</p>\n<p>the lb score confirms it.<br>\nif that is true then private score will be  ….</p>",
              "rawMarkdown": "\" Neither dataset is able to predict the other hardly at all, it appears? \"\nThat is true.\n\nthe lb score confirms it.\nif that is true then private score will be  ....",
              "votes": 1
            },
            {
              "id": 2803810,
              "postDate": "2024-05-09T16:17:05.557Z",
              "content": "<p>Also even if your score had been high, this could've been true about the other two proteins which may not have external data. Each protein may behave differently in terms of generalization, is what I would expect. </p>\n<p>The state of the art seems more about getting the best hits within your top 10% or 1% of guesses. But that can be achieved and still get a bad MAP, I believe. End of competition reveal will be interesting, for sure!</p>",
              "rawMarkdown": "Also even if your score had been high, this could've been true about the other two proteins which may not have external data. Each protein may behave differently in terms of generalization, is what I would expect. \n\nThe state of the art seems more about getting the best hits within your top 10% or 1% of guesses. But that can be achieved and still get a bad MAP, I believe. End of competition reveal will be interesting, for sure!",
              "votes": 1
            },
            {
              "id": 2803811,
              "postDate": "2024-05-09T16:18:21.357Z",
              "content": "<p>there are some catch:</p>\n<ul>\n<li>need to probe what is the max value of LB(nonshare). we do not know if 0.011, 0.012 is high value or not.</li>\n<li>there is a possible use of train data=kaggle+external. i also have a 2nd set of external data on SEH</li>\n<li>need to carry out experiment on self-supervised learning of test data soon.</li>\n<li>someone need to submit docking results soon.</li>\n<li>need to try pretrain model + linear model? (aka few shot)</li>\n</ul>",
              "rawMarkdown": "there are some catch:\n- need to probe what is the max value of LB(nonshare). we do not know if 0.011, 0.012 is high value or not.\n- there is a possible use of train data=kaggle+external. i also have a 2nd set of external data on SEH\n- need to carry out experiment on self-supervised learning of test data soon.\n- someone need to submit docking results soon.\n- need to try pretrain model + linear model? (aka few shot)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2812070,
      "postDate": "2024-05-14T05:13:00.320Z",
      "content": "<p>Hello, I am Rajiv from Mumbai. I am not able to open the xl input as it is very large. Can you help? Connect on <a href=\"mailto:rajivkjs@gmail.com\">rajivkjs@gmail.com</a></p>",
      "rawMarkdown": "Hello, I am Rajiv from Mumbai. I am not able to open the xl input as it is very large. Can you help? Connect on rajivkjs@gmail.com"
    },
    {
      "id": 2801196,
      "postDate": "2024-05-08T14:36:34.490Z",
      "content": "<p>By sharing blocks, you mean building blocks that are in common between the training set and the test set?  And nonsharing means building blocks that are unique to the test set?  </p>",
      "rawMarkdown": "By sharing blocks, you mean building blocks that are in common between the training set and the test set?  And nonsharing means building blocks that are unique to the test set?  ",
      "replies": [
        {
          "id": 2801228,
          "postDate": "2024-05-08T14:50:15.843Z",
          "content": "<p>yes, you are correct</p>",
          "rawMarkdown": "yes, you are correct",
          "votes": 1
        }
      ]
    },
    {
      "id": 2800488,
      "postDate": "2024-05-08T07:56:51.167Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> yes</p>",
      "rawMarkdown": "@hengck23 yes"
    },
    {
      "id": 2816887,
      "postDate": "2024-05-16T15:23:11.633Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2814672,
      "postDate": "2024-05-15T13:36:42.667Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2802312,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-09T01:56:57.627000",
      "content": "<p>maybe very good news?</p>\n<p>umap of ecfp fingerprint of building blocks<br>\nblue : share bb<br>\nred : nonshare bb<br>\ngrey: external bb (external data from my other post)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fab4bcf5987371c81571a1b677ee11ee4%2FSelection_082.png?generation=1715219732483573&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fad5c52f7a81c1c28a1e7f5ed93b90d87%2FSelection_083.png?generation=1715219744081586&amp;alt=media\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2802386,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2024-05-09T02:50:42.140000",
          "content": "<p>The external data is the sEH data you posted about weeks(?) ago? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2802419,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-09T03:32:53.933000",
              "content": "<p>yes. if it gives binding affinity as readout values (not 0,1 label).<br>\nif you have ideas how to use that, please let me know.</p>\n<p>right now i can only think of ranking loss for external data</p>\n<p>update: i figure it out. there is overlap between blue and gray  that can used to transform targets of different dataset</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2802425,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-09T03:35:23.703000",
              "content": "<p>\"The 133M small molecule library used here, AMA014, was provided by AlphaMa\"</p>\n<p>it is probably easy to find compounds similar to AMA014</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2803330,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-05-09T12:33:12.993000",
              "content": "<p>You can convert to binary with all non-zero label = 1. That's actually pretty equivalent to the train data of this competition, the external data is just giving more info that we don't have access to with the competition hosted data. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2812021,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-14T04:20:55.247000",
      "content": "<p>i have new results and prove that my observations are correct:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0e4305dd56d5c61c6561cbccb5ad9617%2FSelection_112.png?generation=1715660177707043&amp;alt=media\"></p>\n<ul>\n<li>if you overfit share blocks, nonshare blocks can get worse</li>\n<li>CV/LB of share block correlates very strong.</li>\n</ul>\n<p>if you perform worse than other kagglers who you using the model architecture and same CV, it could be because of the nonshare part. hence those top in the leaderboard has better nonshare block score then you (either via training or artificial adjusting). I think top 20 kagglers have now about same share lb score about 0.595.</p>\n<p>i think the top peforming cv of share is about 0.73 (one model)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2800527,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-08T08:08:06.530000",
      "content": "<p>while it is difficult to find external data that has pos labels, i think it is easier to find those with neg labels. can anyone advised?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2816163,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-16T07:33:53.090000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F481be0a0f5cf01848f5f1f9506254401%2FSelection_126.png?generation=1715844792449133&amp;alt=media\"></p>\n<p>YES!!! training with external dataset improves  nonshare LB, but by a little</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2803679,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-09T15:36:23.593000",
      "content": "<p>at first i though it is the question if we can generalised, i am wrong.<br>\nit is whether we can detect at all !!!<br>\n(or maybe external data protein segment is different, or experiment setting, etc … let's wait for the lb score)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F13b1bde7ec5435bcb59b8356a8012df8%2FSelection_092.png?generation=1715309251814154&amp;alt=media\"></p>\n<p>i have to change my strategy</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2803782,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2024-05-09T16:05:22.747000",
          "content": "<p>That's your score matrix for train on one dataset, predict on both? Neither dataset is able to predict the other hardly at all, it appears? That was just with xgb on ECFP?</p>\n<p>Yikes that's low though</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2803789,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-05-09T16:05:55.900000",
              "content": "<p>I wish it surprised me more, lol. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2803801,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-09T16:08:53.423000",
              "content": "<p>\" Neither dataset is able to predict the other hardly at all, it appears? \"<br>\nThat is true.</p>\n<p>the lb score confirms it.<br>\nif that is true then private score will be  ….</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2803810,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-05-09T16:17:05.557000",
              "content": "<p>Also even if your score had been high, this could've been true about the other two proteins which may not have external data. Each protein may behave differently in terms of generalization, is what I would expect. </p>\n<p>The state of the art seems more about getting the best hits within your top 10% or 1% of guesses. But that can be achieved and still get a bad MAP, I believe. End of competition reveal will be interesting, for sure!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2803811,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-09T16:18:21.357000",
              "content": "<p>there are some catch:</p>\n<ul>\n<li>need to probe what is the max value of LB(nonshare). we do not know if 0.011, 0.012 is high value or not.</li>\n<li>there is a possible use of train data=kaggle+external. i also have a 2nd set of external data on SEH</li>\n<li>need to carry out experiment on self-supervised learning of test data soon.</li>\n<li>someone need to submit docking results soon.</li>\n<li>need to try pretrain model + linear model? (aka few shot)</li>\n</ul>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2812070,
      "author_name": "Prof. Rajiv Iyer ",
      "author_url": "",
      "post_date": "2024-05-14T05:13:00.320000",
      "content": "<p>Hello, I am Rajiv from Mumbai. I am not able to open the xl input as it is very large. Can you help? Connect on <a href=\"mailto:rajivkjs@gmail.com\">rajivkjs@gmail.com</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2801196,
      "author_name": "KirkDCO",
      "author_url": "",
      "post_date": "2024-05-08T14:36:34.490000",
      "content": "<p>By sharing blocks, you mean building blocks that are in common between the training set and the test set?  And nonsharing means building blocks that are unique to the test set?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2801228,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-05-08T14:50:15.843000",
          "content": "<p>yes, you are correct</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2800488,
      "author_name": "Shubham Kale",
      "author_url": "",
      "post_date": "2024-05-08T07:56:51.167000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> yes</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2816887,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-16T15:23:11.633000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2814672,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-15T13:36:42.667000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2800473": "**UPDATED!! the scoring metric has been updated. so some strategy and discussion may not be valid anymore !!!**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F18c0f3c299addf7c9ae36692d2545c19%2FSelection_080.png?generation=1715154492163145&alt=media)\n\ni made a third submission:\n```\nreplaced =model2_submit_df.copy()\nreplaced.loc[nonshare,'binds']=model1_submit_df.loc[nonshare,'binds']\n# this is actually not very correct because of different calibration of the two models\n\nlbscore = 0.612\n```\n\nimplications:\n- nonsharing block indeed contribute to public LB score\n- be care in validation and interpretation of local and public scores",
    "2802312": "maybe very good news?\n\numap of ecfp fingerprint of building blocks\nblue : share bb\nred : nonshare bb\ngrey: external bb (external data from my other post)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fab4bcf5987371c81571a1b677ee11ee4%2FSelection_082.png?generation=1715219732483573&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fad5c52f7a81c1c28a1e7f5ed93b90d87%2FSelection_083.png?generation=1715219744081586&alt=media)",
    "2812021": "i have new results and prove that my observations are correct:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0e4305dd56d5c61c6561cbccb5ad9617%2FSelection_112.png?generation=1715660177707043&alt=media)\n\n- if you overfit share blocks, nonshare blocks can get worse\n- CV/LB of share block correlates very strong.\n\nif you perform worse than other kagglers who you using the model architecture and same CV, it could be because of the nonshare part. hence those top in the leaderboard has better nonshare block score then you (either via training or artificial adjusting). I think top 20 kagglers have now about same share lb score about 0.595.\n\ni think the top peforming cv of share is about 0.73 (one model)",
    "2800527": "while it is difficult to find external data that has pos labels, i think it is easier to find those with neg labels. can anyone advised?",
    "2816163": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F481be0a0f5cf01848f5f1f9506254401%2FSelection_126.png?generation=1715844792449133&alt=media)\n\nYES!!! training with external dataset improves  nonshare LB, but by a little",
    "2803679": "at first i though it is the question if we can generalised, i am wrong.\nit is whether we can detect at all !!!\n(or maybe external data protein segment is different, or experiment setting, etc ... let's wait for the lb score)\n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F13b1bde7ec5435bcb59b8356a8012df8%2FSelection_092.png?generation=1715309251814154&alt=media)\n\ni have to change my strategy",
    "2812070": "Hello, I am Rajiv from Mumbai. I am not able to open the xl input as it is very large. Can you help? Connect on rajivkjs@gmail.com",
    "2801196": "By sharing blocks, you mean building blocks that are in common between the training set and the test set?  And nonsharing means building blocks that are unique to the test set?  ",
    "2800488": "@hengck23 yes",
    "2816887": "",
    "2814672": ""
  }
}