{
  "id": 174589,
  "title": "Estimate the Shake",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174589",
  "author_name": "Qishen Ha",
  "post_date": "2020-08-14T08:00:10.489000",
  "votes": 56,
  "comment_count": 30,
  "views": 0,
  "content": "<h1>Description</h1>\n<p>We all know that we have an extremely unstable public LB, with around 78 pos samples and around 3295 neg samples.</p>\n<p>Look at the distribution of test data and oof, I guess that positive and negative samples in private LB have similar proportions to the public one. That's around 178 pos samples and around 7687 neg samples.</p>\n<p>So I did a simulation base on this hypothesis to estimate the shake.</p>\n<p>We use a <code>auc = 0.9525</code> oof file to run following code</p>\n<h1>Code</h1>\n<pre><code>aucs1 = []\nfor i in tqdm(range(10000)):\n    df_this = pd.concat([df2020[df2020.target == 1].sample(78), df2020[df2020.target == 0].sample(3295)])\n    aucs1.append(roc_auc_score(df_this.target, df_this.pred))\n\naucs2 = []\nfor i in tqdm(range(10000)):\n    df_this = pd.concat([df2020[df2020.target == 1].sample(178), df2020[df2020.target == 0].sample(7687)])\n    aucs2.append(roc_auc_score(df_this.target, df_this.pred))\n\nsns.distplot(aucs1, bins=100)\nsns.distplot(aucs2, bins=100)\n</code></pre>\n<h1>Output</h1>\n<h1><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2Feb6f6780f3783d30a9b10340eb8657da%2Fimage%20(7).png?generation=1597391568832528&amp;alt=media\" alt=\"\"></h1>\n<h1>Conclusions</h1>\n<ul>\n<li>If your CV is 0.95, while your public LB is only 0.92 (or is &gt; 0.97), it is not weird.</li>\n<li>If your CV is 0.95, you still have a certain probability that finished with a lower ranking than teams with CV 0.94</li>\n<li>Anyone who has a public LB &gt; 0.92 have a chance to shake up to gold zone at the end.</li>\n</ul>",
  "messages": [
    {
      "id": 970109,
      "postDate": "2020-08-14T08:00:10.490Z",
      "content": "<h1>Description</h1>\n<p>We all know that we have an extremely unstable public LB, with around 78 pos samples and around 3295 neg samples.</p>\n<p>Look at the distribution of test data and oof, I guess that positive and negative samples in private LB have similar proportions to the public one. That's around 178 pos samples and around 7687 neg samples.</p>\n<p>So I did a simulation base on this hypothesis to estimate the shake.</p>\n<p>We use a <code>auc = 0.9525</code> oof file to run following code</p>\n<h1>Code</h1>\n<pre><code>aucs1 = []\nfor i in tqdm(range(10000)):\n    df_this = pd.concat([df2020[df2020.target == 1].sample(78), df2020[df2020.target == 0].sample(3295)])\n    aucs1.append(roc_auc_score(df_this.target, df_this.pred))\n\naucs2 = []\nfor i in tqdm(range(10000)):\n    df_this = pd.concat([df2020[df2020.target == 1].sample(178), df2020[df2020.target == 0].sample(7687)])\n    aucs2.append(roc_auc_score(df_this.target, df_this.pred))\n\nsns.distplot(aucs1, bins=100)\nsns.distplot(aucs2, bins=100)\n</code></pre>\n<h1>Output</h1>\n<h1><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2Feb6f6780f3783d30a9b10340eb8657da%2Fimage%20(7).png?generation=1597391568832528&amp;alt=media\" alt=\"\"></h1>\n<h1>Conclusions</h1>\n<ul>\n<li>If your CV is 0.95, while your public LB is only 0.92 (or is &gt; 0.97), it is not weird.</li>\n<li>If your CV is 0.95, you still have a certain probability that finished with a lower ranking than teams with CV 0.94</li>\n<li>Anyone who has a public LB &gt; 0.92 have a chance to shake up to gold zone at the end.</li>\n</ul>",
      "rawMarkdown": "# Description\n\nWe all know that we have an extremely unstable public LB, with around 78 pos samples and around 3295 neg samples.\n\nLook at the distribution of test data and oof, I guess that positive and negative samples in private LB have similar proportions to the public one. That's around 178 pos samples and around 7687 neg samples.\n\nSo I did a simulation base on this hypothesis to estimate the shake.\n\nWe use a `auc = 0.9525` oof file to run following code\n\n# Code\n\n```\naucs1 = []\nfor i in tqdm(range(10000)):\n    df_this = pd.concat([df2020[df2020.target == 1].sample(78), df2020[df2020.target == 0].sample(3295)])\n    aucs1.append(roc_auc_score(df_this.target, df_this.pred))\n    \naucs2 = []\nfor i in tqdm(range(10000)):\n    df_this = pd.concat([df2020[df2020.target == 1].sample(178), df2020[df2020.target == 0].sample(7687)])\n    aucs2.append(roc_auc_score(df_this.target, df_this.pred))\n\nsns.distplot(aucs1, bins=100)\nsns.distplot(aucs2, bins=100)\n```\n\n# Output\n\n# ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2Feb6f6780f3783d30a9b10340eb8657da%2Fimage%20(7).png?generation=1597391568832528&alt=media)\n\n# Conclusions\n\n* If your CV is 0.95, while your public LB is only 0.92 (or is > 0.97), it is not weird.\n* If your CV is 0.95, you still have a certain probability that finished with a lower ranking than teams with CV 0.94\n* Anyone who has a public LB > 0.92 have a chance to shake up to gold zone at the end.",
      "votes": 55
    },
    {
      "id": 970954,
      "postDate": "2020-08-15T03:24:54.843Z",
      "content": "<p>I implemented <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> idea below. I simulated 1000 public private leaderboards. Then I took my 45 offline models (from 2 months of work) and computed their public and private LB scores. Next I computed the 45 differences between simulated private and public score for each of 1000 simulations and subtracted the mean each simulation. Below is a plot of the 45,000 (mean standardized) shakeup differences. The standard deviation is 0.0113. That means that 75% of teams with <strong>honest models</strong> will have their leaderboard score change plus minus 0.0113 (relatively compared with others) (from public to private) which is plus minus 700 positions. Big Shakeup! (23.5% of teams will see plus or minus between 0.0113 and 0.025! A public LB 0.945 can turn into private LB 0.97 with 1 in 200 chance!).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F8dec330bedb1237f5171b8c259c9badd%2Fnew.png?generation=1597473209592889&amp;alt=media\" alt=\"\"></p>\n<p>Notes: I used 45 diverse models OOF. These are every backbone, every image resolution, every augmentation, every learning schedule, every preprocess, every everything! These are 2 months of diverse models with strong CV LB.</p>\n<p>\"Honest\" means that your model is not the result of probing or ensembling using LB score. This means, that you produced your model offline either single model, or used CV to ensemble. The code to produce the above plot is</p>\n<pre><code>diff = []\nfor j in range(1000):\n\n    # PUBLIC LB\n    x = pd.concat([oof2[oof2.target == 1].sample(78), oof2[oof2.target == 0].sample(3295)])\n    # PRIVATE LB\n    oof3 = oof2.loc[~oof2.index.isin(x.index)]\n    y = pd.concat([oof3[oof3.target == 1].sample(178), oof3[oof3.target == 0].sample(7687)])\n\n   # COMPUTE SHAKEUP FOR 45 DIVERSE MODELS\n   temp = []\n   for k in range(45):\n       # COMPUTE PUBLIC LB FOR MODEL K\n       idx = x.index.values\n       public = roc_auc_score( oof[0].loc[oof[0].index.isin(idx),'target'], \n               oof[k].loc[oof[k].index.isin(idx),'pred'] )\n       # COMPUTE PRIVATE LB FOR MODEL K\n       idx = y.index.values\n       private = roc_auc_score( oof[0].loc[oof[0].index.isin(idx),'target'], \n               oof[k].loc[oof[k].index.isin(idx),'pred'] )\n       # SAVE SHAKEUP\n       temp.append(private-public)\n    temp = np.array(temp); temp -= np.mean(temp)\n    diff += list(temp)\n</code></pre>",
      "rawMarkdown": "I implemented @philippsinger idea below. I simulated 1000 public private leaderboards. Then I took my 45 offline models (from 2 months of work) and computed their public and private LB scores. Next I computed the 45 differences between simulated private and public score for each of 1000 simulations and subtracted the mean each simulation. Below is a plot of the 45,000 (mean standardized) shakeup differences. The standard deviation is 0.0113. That means that 75% of teams with **honest models** will have their leaderboard score change plus minus 0.0113 (relatively compared with others) (from public to private) which is plus minus 700 positions. Big Shakeup! (23.5% of teams will see plus or minus between 0.0113 and 0.025! A public LB 0.945 can turn into private LB 0.97 with 1 in 200 chance!).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F8dec330bedb1237f5171b8c259c9badd%2Fnew.png?generation=1597473209592889&alt=media)\n\nNotes: I used 45 diverse models OOF. These are every backbone, every image resolution, every augmentation, every learning schedule, every preprocess, every everything! These are 2 months of diverse models with strong CV LB.\n\n\"Honest\" means that your model is not the result of probing or ensembling using LB score. This means, that you produced your model offline either single model, or used CV to ensemble. The code to produce the above plot is\n\n    diff = []\n    for j in range(1000):\n    \n        # PUBLIC LB\n        x = pd.concat([oof2[oof2.target == 1].sample(78), oof2[oof2.target == 0].sample(3295)])\n        # PRIVATE LB\n        oof3 = oof2.loc[~oof2.index.isin(x.index)]\n        y = pd.concat([oof3[oof3.target == 1].sample(178), oof3[oof3.target == 0].sample(7687)])\n\n       # COMPUTE SHAKEUP FOR 45 DIVERSE MODELS\n       temp = []\n       for k in range(45):\n           # COMPUTE PUBLIC LB FOR MODEL K\n           idx = x.index.values\n           public = roc_auc_score( oof[0].loc[oof[0].index.isin(idx),'target'], \n                   oof[k].loc[oof[k].index.isin(idx),'pred'] )\n           # COMPUTE PRIVATE LB FOR MODEL K\n           idx = y.index.values\n           private = roc_auc_score( oof[0].loc[oof[0].index.isin(idx),'target'], \n                   oof[k].loc[oof[k].index.isin(idx),'pred'] )\n           # SAVE SHAKEUP\n           temp.append(private-public)\n        temp = np.array(temp); temp -= np.mean(temp)\n        diff += list(temp)",
      "votes": 16,
      "replies": [
        {
          "id": 970976,
          "postDate": "2020-08-15T04:04:18.523Z",
          "content": "<p>does that mean that my cv .96 model with .92lb score has a chance lol</p>",
          "rawMarkdown": "does that mean that my cv .96 model with .92lb score has a chance lol",
          "votes": 1
        },
        {
          "id": 971041,
          "postDate": "2020-08-15T06:00:55.763Z",
          "content": "<p>Yup. It could be 0.97 on private leaderboard. But relatively speaking, it can only become \"0.945\". (Relative means, that when your score goes up other teams' scores goes up. So at most, you'll gain 0.025 above the mean increase).</p>",
          "rawMarkdown": "Yup. It could be 0.97 on private leaderboard. But relatively speaking, it can only become \"0.945\". (Relative means, that when your score goes up other teams' scores goes up. So at most, you'll gain 0.025 above the mean increase).",
          "votes": 1
        },
        {
          "id": 971048,
          "postDate": "2020-08-15T06:12:54.560Z",
          "content": "<p>So, this is going to be another M5. Nobody is safe.</p>",
          "rawMarkdown": "So, this is going to be another M5. Nobody is safe.",
          "votes": 1
        },
        {
          "id": 971071,
          "postDate": "2020-08-15T06:42:39.777Z",
          "content": "<p>UPDATE: I updated my model to be relative. For example, if all 45 models increase exactly 0.02 for one simulated public private LB. And all models decrease exactly 0.01 for another simulated public private LB. Then there will be no shakeup because everyone increased or decreased together. </p>\n<p>There is only a shakeup when given a public private LB, some teams AUC score goes up while others go down. To update my simulation to be relative, i subtracted the mean from each of 1000 simulated public private LB.</p>",
          "rawMarkdown": "UPDATE: I updated my model to be relative. For example, if all 45 models increase exactly 0.02 for one simulated public private LB. And all models decrease exactly 0.01 for another simulated public private LB. Then there will be no shakeup because everyone increased or decreased together. \n\nThere is only a shakeup when given a public private LB, some teams AUC score goes up while others go down. To update my simulation to be relative, i subtracted the mean from each of 1000 simulated public private LB.",
          "votes": 1
        },
        {
          "id": 971090,
          "postDate": "2020-08-15T07:04:44.203Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Hi Chris, thanks again for your insightful contribution. Just to ask you, what's your take on blending subs in this comp --&gt; some of which achieves a high LB score.</p>",
          "rawMarkdown": "@cdeotte Hi Chris, thanks again for your insightful contribution. Just to ask you, what's your take on blending subs in this comp --> some of which achieves a high LB score."
        },
        {
          "id": 971142,
          "postDate": "2020-08-15T08:00:34.670Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true,
          "replies": [
            {
              "id": 971161,
              "postDate": "2020-08-15T08:21:31.297Z",
              "content": "<p><a href=\"https://www.kaggle.com/bawaguava\" target=\"_blank\">@bawaguava</a> Thanks for the heads up: What if the cv and lb score are very close to each other; for example 0.9560 and 0.9580? What does this signify? Also, what if the CV is way higher than lb by 0.02?</p>",
              "rawMarkdown": "@bawaguava Thanks for the heads up: What if the cv and lb score are very close to each other; for example 0.9560 and 0.9580? What does this signify? Also, what if the CV is way higher than lb by 0.02?"
            },
            {
              "id": 971168,
              "postDate": "2020-08-15T08:32:56.583Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 971169,
              "postDate": "2020-08-15T08:35:49.027Z",
              "content": "<p><a href=\"https://www.kaggle.com/bawaguava\" target=\"_blank\">@bawaguava</a> Thanks a lot for the insight, so a similar cv and lb score is a good indication of your model's generalization.</p>",
              "rawMarkdown": "@bawaguava Thanks a lot for the insight, so a similar cv and lb score is a good indication of your model's generalization."
            }
          ]
        },
        {
          "id": 971170,
          "postDate": "2020-08-15T08:39:25.720Z",
          "content": "<p>I used my 9 models ensemble oof, simulated 10k delta between public and private LB and got the <code>std = 0.010633</code>, which is almost the same as your single model std.</p>",
          "rawMarkdown": "I used my 9 models ensemble oof, simulated 10k delta between public and private LB and got the `std = 0.010633`, which is almost the same as your single model std.\n\n",
          "votes": 3
        },
        {
          "id": 971193,
          "postDate": "2020-08-15T09:19:34.230Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Great. <br>\nBut what if the public/private was not splited proportionlly ? We're going to shake more 😂</p>",
          "rawMarkdown": "@cdeotte Great. \nBut what if the public/private was not splited proportionlly ? We're going to shake more 😂",
          "votes": 1
        },
        {
          "id": 971215,
          "postDate": "2020-08-15T09:40:00.267Z",
          "content": "<p>Yes, this is assuming equal split, and even then the shakeup potential looks high.</p>",
          "rawMarkdown": "Yes, this is assuming equal split, and even then the shakeup potential looks high.",
          "votes": 1
        },
        {
          "id": 971353,
          "postDate": "2020-08-15T12:40:41.713Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> now in order to take into account the fact that LB shows best score of all submissions: shouldn't you compare the max LB score for the 1000 LBs to each actual score in the 1000 Private LB in order to have a better idea of the shakup to come? (this would be shake up assuming a 1000 submissions for 45 participants?) </p>",
          "rawMarkdown": "@cdeotte now in order to take into account the fact that LB shows best score of all submissions: shouldn't you compare the max LB score for the 1000 LBs to each actual score in the 1000 Private LB in order to have a better idea of the shakup to come? (this would be shake up assuming a 1000 submissions for 45 participants?) "
        },
        {
          "id": 971881,
          "postDate": "2020-08-16T03:29:19.257Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 970125,
      "postDate": "2020-08-14T08:16:31.933Z",
      "content": "<p>I think this type of shakeup analysis is in general not fully correct. It just simulates the difference in AUC scores across sub-samples, which is not too surprising as there are easier and harder subsamples. Now you could say that public LB might be an easier subsample, and private LB is a harder subsample (with a lower std across subsamples). But now theoretically everyone would move according to the difficulty of the subsample.</p>\n<p>The better way is to take k models, and evaluate the std on same subsamples.</p>",
      "rawMarkdown": "I think this type of shakeup analysis is in general not fully correct. It just simulates the difference in AUC scores across sub-samples, which is not too surprising as there are easier and harder subsamples. Now you could say that public LB might be an easier subsample, and private LB is a harder subsample (with a lower std across subsamples). But now theoretically everyone would move according to the difficulty of the subsample.\n\nThe better way is to take k models, and evaluate the std on same subsamples.",
      "votes": 8,
      "replies": [
        {
          "id": 970157,
          "postDate": "2020-08-14T08:56:48.503Z",
          "content": "<p>I agree with you by part, IMO no shakeup analysis is fully correct, its all based on some hypothesis.<br>\nLike what you said public LB might be an easier subsample, and private LB is a harder subsample, is also a kind of hypothesis.<br>\nHere I just assume that private LB have a similar distribution with public LB and training set, but didn't make any hypothesis on the difficulty of the public/private subset.</p>",
          "rawMarkdown": "I agree with you by part, IMO no shakeup analysis is fully correct, its all based on some hypothesis.\nLike what you said public LB might be an easier subsample, and private LB is a harder subsample, is also a kind of hypothesis.\nHere I just assume that private LB have a similar distribution with public LB and training set, but didn't make any hypothesis on the difficulty of the public/private subset.",
          "votes": 2
        },
        {
          "id": 970180,
          "postDate": "2020-08-14T09:11:54.970Z",
          "content": "<p>But if this hypothesis is true, how would this lead to a shakeup?</p>",
          "rawMarkdown": "But if this hypothesis is true, how would this lead to a shakeup?"
        },
        {
          "id": 970240,
          "postDate": "2020-08-14T09:39:54.320Z",
          "content": "<p>I believe it can be caused by random fluctuations due to the lack of positive samples in the test set.</p>",
          "rawMarkdown": "I believe it can be caused by random fluctuations due to the lack of positive samples in the test set.",
          "votes": -1
        },
        {
          "id": 971087,
          "postDate": "2020-08-15T06:58:17.657Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Good point. I think i implemented what you suggest above. We want the variance not the shift.</p>\n<p>If all teams increase by 0.01 going from public to private then there is no shakeup. But if some teams increase and other teams decrease, then there is a shakeup.</p>\n<p>To simulate this, i make a public and private leaderboard. Then I compute scores for my 45 offline models. Next I compute all the differences between private and public and subtract the mean. All we have left is the variance. </p>\n<p>I repeat this 1000 times, so we have 45,000 numbers representing variance. Lastly I plot these 45,000 numbers.</p>",
          "rawMarkdown": "@philippsinger Good point. I think i implemented what you suggest above. We want the variance not the shift.\n\nIf all teams increase by 0.01 going from public to private then there is no shakeup. But if some teams increase and other teams decrease, then there is a shakeup.\n\nTo simulate this, i make a public and private leaderboard. Then I compute scores for my 45 offline models. Next I compute all the differences between private and public and subtract the mean. All we have left is the variance. \n\nI repeat this 1000 times, so we have 45,000 numbers representing variance. Lastly I plot these 45,000 numbers.",
          "votes": 2
        },
        {
          "id": 971106,
          "postDate": "2020-08-15T07:24:38.580Z",
          "content": "<p>Awesome <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! Thanks!</p>",
          "rawMarkdown": "Awesome @cdeotte! Thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 971846,
      "postDate": "2020-08-16T02:02:52.693Z",
      "content": "<p>Achive that CV was impossible for me over the last days of this competition, but i'm not sad, i'm just need to learn a lot more. Meanwhile, i'm very excited to see the <strong>Private Leaderboard</strong>. This competition It's been fun, and i learned more in 2 months than 1 year by myself. Good to luck to everyone.</p>",
      "rawMarkdown": "Achive that CV was impossible for me over the last days of this competition, but i'm not sad, i'm just need to learn a lot more. Meanwhile, i'm very excited to see the **Private Leaderboard**. This competition It's been fun, and i learned more in 2 months than 1 year by myself. Good to luck to everyone.",
      "votes": 1
    },
    {
      "id": 971866,
      "postDate": "2020-08-16T02:51:41.680Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F385905%2Fde962ed75ded1fcf1e6548d85f4a9c66%2Finbox_1424766_ca6ef42ae6cea851c090d0a8df65d50d_index.jpeg?generation=1597546286594466&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F385905%2Fde962ed75ded1fcf1e6548d85f4a9c66%2Finbox_1424766_ca6ef42ae6cea851c090d0a8df65d50d_index.jpeg?generation=1597546286594466&alt=media)",
      "votes": 2
    },
    {
      "id": 970219,
      "postDate": "2020-08-14T09:31:21.077Z",
      "content": "<p>Really exited to see the leaderboard  shake up. 🤩😬</p>",
      "rawMarkdown": "Really exited to see the leaderboard  shake up. 🤩😬",
      "votes": 1,
      "replies": [
        {
          "id": 970810,
          "postDate": "2020-08-14T20:24:10.847Z",
          "content": "<p>why you'll feel the excitement? </p>",
          "rawMarkdown": "why you'll feel the excitement? "
        },
        {
          "id": 971323,
          "postDate": "2020-08-15T11:46:38.633Z",
          "content": "<p><a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> Hey said 'exited' not 'excited'.😂</p>",
          "rawMarkdown": "@ipythonx Hey said 'exited' not 'excited'.😂",
          "votes": 2
        }
      ]
    },
    {
      "id": 971513,
      "postDate": "2020-08-15T15:59:57.023Z",
      "content": "<h1><img src=\"https://i.imgur.com/ZyFqvq9.png\" alt=\"https://i.imgur.com/ZyFqvq9.png\"></h1>\n<p>My Cv Status is somewhat strange<br>\nstandard deviation is too small compared to yours<br>\nPublic LB, std is 0.008<br>\nFor Private LB, std is 0.005</p>\n<p>assuming my current cv status is appropriate, I don't have a lottery!</p>",
      "rawMarkdown": "#![https://i.imgur.com/ZyFqvq9.png](https://i.imgur.com/ZyFqvq9.png)\n\nMy Cv Status is somewhat strange\nstandard deviation is too small compared to yours\nPublic LB, std is 0.008\nFor Private LB, std is 0.005\n\nassuming my current cv status is appropriate, I don't have a lottery!",
      "votes": -1,
      "replies": [
        {
          "id": 971853,
          "postDate": "2020-08-16T02:33:53.793Z",
          "content": "<p>We are saying <code>np.std(auc1-auc2)</code> here. But not np.std(auc1) or np.std(auc2)</p>",
          "rawMarkdown": "We are saying `np.std(auc1-auc2)` here. But not np.std(auc1) or np.std(auc2)",
          "votes": 4
        },
        {
          "id": 971953,
          "postDate": "2020-08-16T05:44:53.953Z",
          "content": "<p>oh i see. i will check np.std(auc1-auc2) right now!</p>\n<p>====<br>\nmy np.std(auc1-auc2) is around 0.0099. i have a lottery too.</p>",
          "rawMarkdown": "oh i see. i will check np.std(auc1-auc2) right now!\n\n====\nmy np.std(auc1-auc2) is around 0.0099. i have a lottery too."
        }
      ]
    },
    {
      "id": 972487,
      "postDate": "2020-08-16T15:26:50.050Z",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 970954,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-15T03:24:54.843000",
      "content": "<p>I implemented <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> idea below. I simulated 1000 public private leaderboards. Then I took my 45 offline models (from 2 months of work) and computed their public and private LB scores. Next I computed the 45 differences between simulated private and public score for each of 1000 simulations and subtracted the mean each simulation. Below is a plot of the 45,000 (mean standardized) shakeup differences. The standard deviation is 0.0113. That means that 75% of teams with <strong>honest models</strong> will have their leaderboard score change plus minus 0.0113 (relatively compared with others) (from public to private) which is plus minus 700 positions. Big Shakeup! (23.5% of teams will see plus or minus between 0.0113 and 0.025! A public LB 0.945 can turn into private LB 0.97 with 1 in 200 chance!).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F8dec330bedb1237f5171b8c259c9badd%2Fnew.png?generation=1597473209592889&amp;alt=media\" alt=\"\"></p>\n<p>Notes: I used 45 diverse models OOF. These are every backbone, every image resolution, every augmentation, every learning schedule, every preprocess, every everything! These are 2 months of diverse models with strong CV LB.</p>\n<p>\"Honest\" means that your model is not the result of probing or ensembling using LB score. This means, that you produced your model offline either single model, or used CV to ensemble. The code to produce the above plot is</p>\n<pre><code>diff = []\nfor j in range(1000):\n\n    # PUBLIC LB\n    x = pd.concat([oof2[oof2.target == 1].sample(78), oof2[oof2.target == 0].sample(3295)])\n    # PRIVATE LB\n    oof3 = oof2.loc[~oof2.index.isin(x.index)]\n    y = pd.concat([oof3[oof3.target == 1].sample(178), oof3[oof3.target == 0].sample(7687)])\n\n   # COMPUTE SHAKEUP FOR 45 DIVERSE MODELS\n   temp = []\n   for k in range(45):\n       # COMPUTE PUBLIC LB FOR MODEL K\n       idx = x.index.values\n       public = roc_auc_score( oof[0].loc[oof[0].index.isin(idx),'target'], \n               oof[k].loc[oof[k].index.isin(idx),'pred'] )\n       # COMPUTE PRIVATE LB FOR MODEL K\n       idx = y.index.values\n       private = roc_auc_score( oof[0].loc[oof[0].index.isin(idx),'target'], \n               oof[k].loc[oof[k].index.isin(idx),'pred'] )\n       # SAVE SHAKEUP\n       temp.append(private-public)\n    temp = np.array(temp); temp -= np.mean(temp)\n    diff += list(temp)\n</code></pre>",
      "votes": 16,
      "replies": [
        {
          "id": 970976,
          "author_name": "isaac in",
          "author_url": "",
          "post_date": "2020-08-15T04:04:18.523000",
          "content": "<p>does that mean that my cv .96 model with .92lb score has a chance lol</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 971041,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-15T06:00:55.763000",
          "content": "<p>Yup. It could be 0.97 on private leaderboard. But relatively speaking, it can only become \"0.945\". (Relative means, that when your score goes up other teams' scores goes up. So at most, you'll gain 0.025 above the mean increase).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 971048,
          "author_name": "Kumar Shubham",
          "author_url": "",
          "post_date": "2020-08-15T06:12:54.560000",
          "content": "<p>So, this is going to be another M5. Nobody is safe.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 971071,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-15T06:42:39.777000",
          "content": "<p>UPDATE: I updated my model to be relative. For example, if all 45 models increase exactly 0.02 for one simulated public private LB. And all models decrease exactly 0.01 for another simulated public private LB. Then there will be no shakeup because everyone increased or decreased together. </p>\n<p>There is only a shakeup when given a public private LB, some teams AUC score goes up while others go down. To update my simulation to be relative, i subtracted the mean from each of 1000 simulated public private LB.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 971090,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2020-08-15T07:04:44.203000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Hi Chris, thanks again for your insightful contribution. Just to ask you, what's your take on blending subs in this comp --&gt; some of which achieves a high LB score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 971142,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-15T08:00:34.670000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 971161,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-15T08:21:31.297000",
              "content": "<p><a href=\"https://www.kaggle.com/bawaguava\" target=\"_blank\">@bawaguava</a> Thanks for the heads up: What if the cv and lb score are very close to each other; for example 0.9560 and 0.9580? What does this signify? Also, what if the CV is way higher than lb by 0.02?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 971168,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-08-15T08:32:56.583000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 971169,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-15T08:35:49.027000",
              "content": "<p><a href=\"https://www.kaggle.com/bawaguava\" target=\"_blank\">@bawaguava</a> Thanks a lot for the insight, so a similar cv and lb score is a good indication of your model's generalization.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 971170,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-08-15T08:39:25.720000",
          "content": "<p>I used my 9 models ensemble oof, simulated 10k delta between public and private LB and got the <code>std = 0.010633</code>, which is almost the same as your single model std.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 971193,
          "author_name": "Changyi",
          "author_url": "",
          "post_date": "2020-08-15T09:19:34.230000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Great. <br>\nBut what if the public/private was not splited proportionlly ? We're going to shake more 😂</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 971215,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-15T09:40:00.267000",
          "content": "<p>Yes, this is assuming equal split, and even then the shakeup potential looks high.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 971353,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-08-15T12:40:41.713000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> now in order to take into account the fact that LB shows best score of all submissions: shouldn't you compare the max LB score for the 1000 LBs to each actual score in the 1000 Private LB in order to have a better idea of the shakup to come? (this would be shake up assuming a 1000 submissions for 45 participants?) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 971881,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-16T03:29:19.257000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 970125,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-08-14T08:16:31.933000",
      "content": "<p>I think this type of shakeup analysis is in general not fully correct. It just simulates the difference in AUC scores across sub-samples, which is not too surprising as there are easier and harder subsamples. Now you could say that public LB might be an easier subsample, and private LB is a harder subsample (with a lower std across subsamples). But now theoretically everyone would move according to the difficulty of the subsample.</p>\n<p>The better way is to take k models, and evaluate the std on same subsamples.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 970157,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-08-14T08:56:48.503000",
          "content": "<p>I agree with you by part, IMO no shakeup analysis is fully correct, its all based on some hypothesis.<br>\nLike what you said public LB might be an easier subsample, and private LB is a harder subsample, is also a kind of hypothesis.<br>\nHere I just assume that private LB have a similar distribution with public LB and training set, but didn't make any hypothesis on the difficulty of the public/private subset.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 970180,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-14T09:11:54.970000",
          "content": "<p>But if this hypothesis is true, how would this lead to a shakeup?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 970240,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-08-14T09:39:54.320000",
          "content": "<p>I believe it can be caused by random fluctuations due to the lack of positive samples in the test set.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 971087,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-15T06:58:17.657000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Good point. I think i implemented what you suggest above. We want the variance not the shift.</p>\n<p>If all teams increase by 0.01 going from public to private then there is no shakeup. But if some teams increase and other teams decrease, then there is a shakeup.</p>\n<p>To simulate this, i make a public and private leaderboard. Then I compute scores for my 45 offline models. Next I compute all the differences between private and public and subtract the mean. All we have left is the variance. </p>\n<p>I repeat this 1000 times, so we have 45,000 numbers representing variance. Lastly I plot these 45,000 numbers.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 971106,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-15T07:24:38.580000",
          "content": "<p>Awesome <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! Thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 971846,
      "author_name": "Hiram Coria 🧬",
      "author_url": "",
      "post_date": "2020-08-16T02:02:52.693000",
      "content": "<p>Achive that CV was impossible for me over the last days of this competition, but i'm not sad, i'm just need to learn a lot more. Meanwhile, i'm very excited to see the <strong>Private Leaderboard</strong>. This competition It's been fun, and i learned more in 2 months than 1 year by myself. Good to luck to everyone.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 971866,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2020-08-16T02:51:41.680000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F385905%2Fde962ed75ded1fcf1e6548d85f4a9c66%2Finbox_1424766_ca6ef42ae6cea851c090d0a8df65d50d_index.jpeg?generation=1597546286594466&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 970219,
      "author_name": "Mo Fahad",
      "author_url": "",
      "post_date": "2020-08-14T09:31:21.077000",
      "content": "<p>Really exited to see the leaderboard  shake up. 🤩😬</p>",
      "votes": 1,
      "replies": [
        {
          "id": 970810,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-08-14T20:24:10.847000",
          "content": "<p>why you'll feel the excitement? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 971323,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2020-08-15T11:46:38.633000",
          "content": "<p><a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> Hey said 'exited' not 'excited'.😂</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 971513,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2020-08-15T15:59:57.023000",
      "content": "<h1><img src=\"https://i.imgur.com/ZyFqvq9.png\" alt=\"https://i.imgur.com/ZyFqvq9.png\"></h1>\n<p>My Cv Status is somewhat strange<br>\nstandard deviation is too small compared to yours<br>\nPublic LB, std is 0.008<br>\nFor Private LB, std is 0.005</p>\n<p>assuming my current cv status is appropriate, I don't have a lottery!</p>",
      "votes": -1,
      "replies": [
        {
          "id": 971853,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-08-16T02:33:53.793000",
          "content": "<p>We are saying <code>np.std(auc1-auc2)</code> here. But not np.std(auc1) or np.std(auc2)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 971953,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2020-08-16T05:44:53.953000",
          "content": "<p>oh i see. i will check np.std(auc1-auc2) right now!</p>\n<p>====<br>\nmy np.std(auc1-auc2) is around 0.0099. i have a lottery too.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 972487,
      "author_name": "Brandon Nova",
      "author_url": "",
      "post_date": "2020-08-16T15:26:50.050000",
      "content": "<p>thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "970109": "# Description\n\nWe all know that we have an extremely unstable public LB, with around 78 pos samples and around 3295 neg samples.\n\nLook at the distribution of test data and oof, I guess that positive and negative samples in private LB have similar proportions to the public one. That's around 178 pos samples and around 7687 neg samples.\n\nSo I did a simulation base on this hypothesis to estimate the shake.\n\nWe use a `auc = 0.9525` oof file to run following code\n\n# Code\n\n```\naucs1 = []\nfor i in tqdm(range(10000)):\n    df_this = pd.concat([df2020[df2020.target == 1].sample(78), df2020[df2020.target == 0].sample(3295)])\n    aucs1.append(roc_auc_score(df_this.target, df_this.pred))\n    \naucs2 = []\nfor i in tqdm(range(10000)):\n    df_this = pd.concat([df2020[df2020.target == 1].sample(178), df2020[df2020.target == 0].sample(7687)])\n    aucs2.append(roc_auc_score(df_this.target, df_this.pred))\n\nsns.distplot(aucs1, bins=100)\nsns.distplot(aucs2, bins=100)\n```\n\n# Output\n\n# ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2Feb6f6780f3783d30a9b10340eb8657da%2Fimage%20(7).png?generation=1597391568832528&alt=media)\n\n# Conclusions\n\n* If your CV is 0.95, while your public LB is only 0.92 (or is > 0.97), it is not weird.\n* If your CV is 0.95, you still have a certain probability that finished with a lower ranking than teams with CV 0.94\n* Anyone who has a public LB > 0.92 have a chance to shake up to gold zone at the end.",
    "970954": "I implemented @philippsinger idea below. I simulated 1000 public private leaderboards. Then I took my 45 offline models (from 2 months of work) and computed their public and private LB scores. Next I computed the 45 differences between simulated private and public score for each of 1000 simulations and subtracted the mean each simulation. Below is a plot of the 45,000 (mean standardized) shakeup differences. The standard deviation is 0.0113. That means that 75% of teams with **honest models** will have their leaderboard score change plus minus 0.0113 (relatively compared with others) (from public to private) which is plus minus 700 positions. Big Shakeup! (23.5% of teams will see plus or minus between 0.0113 and 0.025! A public LB 0.945 can turn into private LB 0.97 with 1 in 200 chance!).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F8dec330bedb1237f5171b8c259c9badd%2Fnew.png?generation=1597473209592889&alt=media)\n\nNotes: I used 45 diverse models OOF. These are every backbone, every image resolution, every augmentation, every learning schedule, every preprocess, every everything! These are 2 months of diverse models with strong CV LB.\n\n\"Honest\" means that your model is not the result of probing or ensembling using LB score. This means, that you produced your model offline either single model, or used CV to ensemble. The code to produce the above plot is\n\n    diff = []\n    for j in range(1000):\n    \n        # PUBLIC LB\n        x = pd.concat([oof2[oof2.target == 1].sample(78), oof2[oof2.target == 0].sample(3295)])\n        # PRIVATE LB\n        oof3 = oof2.loc[~oof2.index.isin(x.index)]\n        y = pd.concat([oof3[oof3.target == 1].sample(178), oof3[oof3.target == 0].sample(7687)])\n\n       # COMPUTE SHAKEUP FOR 45 DIVERSE MODELS\n       temp = []\n       for k in range(45):\n           # COMPUTE PUBLIC LB FOR MODEL K\n           idx = x.index.values\n           public = roc_auc_score( oof[0].loc[oof[0].index.isin(idx),'target'], \n                   oof[k].loc[oof[k].index.isin(idx),'pred'] )\n           # COMPUTE PRIVATE LB FOR MODEL K\n           idx = y.index.values\n           private = roc_auc_score( oof[0].loc[oof[0].index.isin(idx),'target'], \n                   oof[k].loc[oof[k].index.isin(idx),'pred'] )\n           # SAVE SHAKEUP\n           temp.append(private-public)\n        temp = np.array(temp); temp -= np.mean(temp)\n        diff += list(temp)",
    "970125": "I think this type of shakeup analysis is in general not fully correct. It just simulates the difference in AUC scores across sub-samples, which is not too surprising as there are easier and harder subsamples. Now you could say that public LB might be an easier subsample, and private LB is a harder subsample (with a lower std across subsamples). But now theoretically everyone would move according to the difficulty of the subsample.\n\nThe better way is to take k models, and evaluate the std on same subsamples.",
    "971846": "Achive that CV was impossible for me over the last days of this competition, but i'm not sad, i'm just need to learn a lot more. Meanwhile, i'm very excited to see the **Private Leaderboard**. This competition It's been fun, and i learned more in 2 months than 1 year by myself. Good to luck to everyone.",
    "971866": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F385905%2Fde962ed75ded1fcf1e6548d85f4a9c66%2Finbox_1424766_ca6ef42ae6cea851c090d0a8df65d50d_index.jpeg?generation=1597546286594466&alt=media)",
    "970219": "Really exited to see the leaderboard  shake up. 🤩😬",
    "971513": "#![https://i.imgur.com/ZyFqvq9.png](https://i.imgur.com/ZyFqvq9.png)\n\nMy Cv Status is somewhat strange\nstandard deviation is too small compared to yours\nPublic LB, std is 0.008\nFor Private LB, std is 0.005\n\nassuming my current cv status is appropriate, I don't have a lottery!",
    "972487": "thanks for sharing!"
  }
}