{
  "id": 106844,
  "title": "Time to think about LB shake",
  "url": "/competitions/aptos2019-blindness-detection/discussion/106844",
  "author_name": "",
  "post_date": "2019-08-31T07:02:55.808730700Z",
  "votes": 12,
  "comment_count": 42,
  "views": 0,
  "content": "<p>What your thought about LB shake ?</p>\n\n<p>I guess shake is moderate.</p>\n\n<p>Reason: <br>\n- grade distribution will be different between public and private\n- QWK is unstable evaluation metric</p>",
  "messages": [
    {
      "id": "614176",
      "postDate": "08/31/2019 07:02:55",
      "content": "<p>What your thought about LB shake ?</p>\n\n<p>I guess shake is moderate.</p>\n\n<p>Reason: <br>\n- grade distribution will be different between public and private\n- QWK is unstable evaluation metric</p>",
      "rawMarkdown": "What your thought about LB shake ?\n\nI guess shake is moderate.\n\nReason:  \n- grade distribution will be different between public and private\n- QWK is unstable evaluation metric",
      "votes": null
    },
    {
      "id": "614191",
      "postDate": "08/31/2019 07:31:43",
      "content": "<blockquote>\n  <p>grade distribution will be different between public and private</p>\n</blockquote>\n\n<p>How can you be so sure about this?</p>",
      "rawMarkdown": "&gt;grade distribution will be different between public and private\n\n How can you be so sure about this?",
      "votes": null
    },
    {
      "id": "614198",
      "postDate": "08/31/2019 07:41:58",
      "content": "<p>A Difficult question. <br>\nBut if we use common sense, public distribution would be abnormal. <br>\nI suppose the number of private test dataset is enough large to reflect usual grade distribution(less severity).  </p>\n\n<p><strong>update 1</strong>\nBelow figure was posted by <a href=\"/puremath86\">@puremath86</a> and shows the predicted public distribution. <br>\nWe can easily see how this distribution is strange compared with the usual distribution. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F838971%2Fa86219e1c1eca9d598aaf69adb8bb66a%2FCapture2.PNG?generation=1561952985345841&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>update 2</strong>\nMy best single model (LB = .830) distribution is as following.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fdece607277ac88541c04b34f81666e21%2Fimage.png?generation=1567331246673536&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "A Difficult question.  \nBut if we use common sense, public distribution would be abnormal.  \nI suppose the number of private test dataset is enough large to reflect usual grade distribution(less severity).  \n\n**update 1**\nBelow figure was posted by @puremath86 and shows the predicted public distribution.    \nWe can easily see how this distribution is strange compared with the usual distribution.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F838971%2Fa86219e1c1eca9d598aaf69adb8bb66a%2FCapture2.PNG?generation=1561952985345841&amp;alt=media)\n\n**update 2**\nMy best single model (LB = .830) distribution is as following.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fdece607277ac88541c04b34f81666e21%2Fimage.png?generation=1567331246673536&amp;alt=media)",
      "votes": null
    },
    {
      "id": "614201",
      "postDate": "08/31/2019 07:47:17",
      "content": "<p>yeah, a big problem...</p>",
      "rawMarkdown": "yeah, a big problem...",
      "votes": null
    },
    {
      "id": "614221",
      "postDate": "08/31/2019 08:21:06",
      "content": "<p>Tbh, I don't know what the private test set is gonna be like. Just an additional information:</p>\n\n<p>Normalized distribution, class wise 2015 competition.</p>\n\n<p>2015's train set:\n0    0.734783\n2    0.150658\n1    0.069550\n3    0.024853\n4    0.020156</p>\n\n<p>2015 public test set:\n0    0.745461\n2    0.144783\n1    0.066019\n4    0.022006\n3    0.021731</p>\n\n<p>2015 private test set:\n0    0.735950\n2    0.147223\n1    0.071291\n3    0.022897\n4    0.022639</p>",
      "rawMarkdown": "Tbh, I don't know what the private test set is gonna be like. Just an additional information:\n\nNormalized distribution, class wise 2015 competition.\n\n2015's train set:\n0    0.734783\n2    0.150658\n1    0.069550\n3    0.024853\n4    0.020156\n\n2015 public test set:\n0    0.745461\n2    0.144783\n1    0.066019\n4    0.022006\n3    0.021731\n\n2015 private test set:\n0    0.735950\n2    0.147223\n1    0.071291\n3    0.022897\n4    0.022639",
      "votes": null
    },
    {
      "id": "614373",
      "postDate": "08/31/2019 12:07:42",
      "content": "<p>I expect teams with an LB score of 0.81 or higher to shift about ± 0.03 〜 0.04.</p>",
      "rawMarkdown": "I expect teams with an LB score of 0.81 or higher to shift about ± 0.03 〜 0.04.",
      "votes": null
    },
    {
      "id": "614385",
      "postDate": "08/31/2019 12:21:53",
      "content": "<p>is this a gut feeling or something else?</p>",
      "rawMarkdown": "is this a gut feeling or something else?",
      "votes": null
    },
    {
      "id": "614402",
      "postDate": "08/31/2019 12:53:21",
      "content": "<p>Yes. It is just a feeling based on my trial. The truth is only God knows.</p>",
      "rawMarkdown": "Yes. It is just a feeling based on my trial. The truth is only God knows.",
      "votes": null
    },
    {
      "id": "614969",
      "postDate": "09/01/2019 09:57:26",
      "content": "<p><a href=\"/hirune924\">@hirune924</a> \nI almost stand by you. <br>\nIn my simulation, score will fluctuate by at least 1% within about 50% confidential interval. <br>\nIt's not surprising if shake will be equal to such a high degree shift (3 ~ 4%).  </p>",
      "rawMarkdown": "hirune924 \nI almost stand by you.  \nIn my simulation, score will fluctuate by at least 1% within about 50% confidential interval.  \nIt's not surprising if shake will be equal to such a high degree shift (3 ~ 4%).",
      "votes": null
    },
    {
      "id": "614994",
      "postDate": "09/01/2019 10:35:04",
      "content": "<p><a href=\"/maxwell110\">@maxwell110</a> I see your plots. I kinda get the same distribution for my best test submission, but that doesn't tell us anything about the private test set, right? Just curious, has anyone seen any Kaggle competition where there was a major difference in public and private test set? </p>\n\n<p>A point to remember, QWK shouldn't be trusted, two quite different predictions can have same QWK score. For example: one of my submissions which scores 0.829 on LB:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F71d8f80636b50eceac15423406a6966e%2Fdist.png?generation=1567334053540493&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "maxwell110 I see your plots. I kinda get the same distribution for my best test submission, but that doesn't tell us anything about the private test set, right? Just curious, has anyone seen any Kaggle competition where there was a major difference in public and private test set? \n\nA point to remember, QWK shouldn't be trusted, two quite different predictions can have same QWK score. For example: one of my submissions which scores 0.829 on LB:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F71d8f80636b50eceac15423406a6966e%2Fdist.png?generation=1567334053540493&amp;alt=media)",
      "votes": null
    },
    {
      "id": "614999",
      "postDate": "09/01/2019 10:43:46",
      "content": "<p>This is almost the first competition I've worked on until the last, so I'm not sure about Shake.</p>\n\n<p>I think that there is little difference in class distribution between public and private.  (my hope)\nHowever, in serious cases, I expect that the number of data is small and there is a variety of image findings. I'm afraid that this may cause Shake in the high score zone.</p>\n\n<p>Considering my observations (metrics vs losses, CV vs LB, etc.), I suspect that teams with high LB (e. g. above 0.83) are overfitting.  But this may be a biased analysis by wish. 😎 </p>",
      "rawMarkdown": "This is almost the first competition I've worked on until the last, so I'm not sure about Shake.\n\nI think that there is little difference in class distribution between public and private.  (my hope)\nHowever, in serious cases, I expect that the number of data is small and there is a variety of image findings. I'm afraid that this may cause Shake in the high score zone.\n\nConsidering my observations (metrics vs losses, CV vs LB, etc.), I suspect that teams with high LB (e. g. above 0.83) are overfitting.  But this may be a biased analysis by wish. 😎",
      "votes": null
    },
    {
      "id": "615013",
      "postDate": "09/01/2019 11:04:40",
      "content": "<p><a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> There are many competitions where private and public are totally different. One extreme case is “VSB powerline prediction”, where shakeup was like an earthquake there. I fell from 100 to 5xx . Top3 (or top5) also fell like 100s places. (But my model was not designed to be robust at that time, so that was my bad)</p>\n\n<p>In this competition, however, I agree with <a href=\"/maxwell110\">@maxwell110</a> ; I think due to the nature and size of dataset, the private test set should have similar distribution to training set of 2015/2019, but may be not like public test set.</p>",
      "rawMarkdown": "rishabhiitbhu There are many competitions where private and public are totally different. One extreme case is “VSB powerline prediction”, where shakeup was like an earthquake there. I fell from 100 to 5xx . Top3 (or top5) also fell like 100s places. (But my model was not designed to be robust at that time, so that was my bad)\n\nIn this competition, however, I agree with @maxwell110 ; I think due to the nature and size of dataset, the private test set should have similar distribution to training set of 2015/2019, but may be not like public test set.",
      "votes": null
    },
    {
      "id": "615014",
      "postDate": "09/01/2019 11:05:26",
      "content": "<p><a href=\"/octpath0302\">@octpath0302</a> \nAlmost the same opinion.\nGood Luck to you!</p>",
      "rawMarkdown": "octpath0302 \nAlmost the same opinion.\nGood Luck to you!",
      "votes": null
    },
    {
      "id": "615035",
      "postDate": "09/01/2019 11:32:52",
      "content": "<p>who knows? all we can do is trust our cv?</p>",
      "rawMarkdown": "who knows? all we can do is trust our cv?",
      "votes": null
    },
    {
      "id": "615110",
      "postDate": "09/01/2019 13:31:39",
      "content": "<p>And if we removed all 0s in 2015 train, 2015 test, 2019 train. We can find the ratios between class 1, 2, 3, 4 of these three datasets are similar. Maybe that's the natural rule.</p>",
      "rawMarkdown": "And if we removed all 0s in 2015 train, 2015 test, 2019 train. We can find the ratios between class 1, 2, 3, 4 of these three datasets are similar. Maybe that's the natural rule.",
      "votes": null
    },
    {
      "id": "615628",
      "postDate": "09/02/2019 07:34:24",
      "content": "<p>Anybody, who has higher score than my team is overfitting :-D</p>",
      "rawMarkdown": "Anybody, who has higher score than my team is overfitting :-D",
      "votes": null
    },
    {
      "id": "615630",
      "postDate": "09/02/2019 07:36:52",
      "content": "<p>I measured high variance (0.007), just by changing random seed</p>",
      "rawMarkdown": "I measured high variance (0.007), just by changing random seed",
      "votes": null
    },
    {
      "id": "615935",
      "postDate": "09/02/2019 14:18:47",
      "content": "<p><a href=\"/nemethpeti\">@nemethpeti</a> True, this variance is prominent when using TTA.</p>",
      "rawMarkdown": "nemethpeti True, this variance is prominent when using TTA.",
      "votes": null
    },
    {
      "id": "615950",
      "postDate": "09/02/2019 14:33:58",
      "content": "<p>Yeah... I think they will be a shake the question is how bad. Whats intresting about this compettioin that you can  achieve LB score more than 0.81 easily using difrrent variations of training data (atleast 3 or 4 diffrent ways  e.g just sing 2019 data, 2019 data with some combination from 2015 and etc). All of this methods will give you diffrent CV. The question is what should you trust ? =) </p>",
      "rawMarkdown": "Yeah... I think they will be a shake the question is how bad. Whats intresting about this compettioin that you can  achieve LB score more than 0.81 easily using difrrent variations of training data (atleast 3 or 4 diffrent ways  e.g just sing 2019 data, 2019 data with some combination from 2015 and etc). All of this methods will give you diffrent CV. The question is what should you trust ? =)",
      "votes": null
    },
    {
      "id": "615978",
      "postDate": "09/02/2019 14:59:40",
      "content": "<p>agree with you, I don't know what should be trusted...😂 </p>",
      "rawMarkdown": "agree with you, I don't know what should be trusted...😂",
      "votes": null
    },
    {
      "id": "615988",
      "postDate": "09/02/2019 15:09:06",
      "content": "<p>At the end I think I am going to roll dice and select submission =) </p>\n\n<p>Edit Here is my stratagy =) \n1) go <a href=\"https://www.google.com/search?q=dice+roller\">here</a>. Change the number of sides to number of total submisions \n2) roll the dice \n3) select submision file corresponding to the number of the dice =) </p>",
      "rawMarkdown": "At the end I think I am going to roll dice and select submission =) \n\nEdit Here is my stratagy =) \n1) go [here](https://www.google.com/search?q=dice+roller). Change the number of sides to number of total submisions \n2) roll the dice \n3) select submision file corresponding to the number of the dice =)",
      "votes": null
    },
    {
      "id": "616021",
      "postDate": "09/02/2019 15:56:39",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> Haha, thanks for the advice. When almost all models are reaching val qwk of 0.90+ and public test is crazy like this, your strategy seems legit.</p>",
      "rawMarkdown": "drhabib Haha, thanks for the advice. When almost all models are reaching val qwk of 0.90+ and public test is crazy like this, your strategy seems legit.",
      "votes": null
    },
    {
      "id": "616510",
      "postDate": "09/03/2019 07:40:15",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> Can you share 1-2 variations of training data to achieve &gt; 0.81? </p>",
      "rawMarkdown": "drhabib Can you share 1-2 variations of training data to achieve &gt; 0.81?",
      "votes": null
    },
    {
      "id": "617043",
      "postDate": "09/03/2019 17:38:06",
      "content": "<p><img src=\"https://sun9-44.userapi.com/c854532/v854532836/dd1fa/E_ziPIlBdRA.jpg\" alt=\"\"></p>",
      "rawMarkdown": "![](https://sun9-44.userapi.com/c854532/v854532836/dd1fa/E_ziPIlBdRA.jpg)",
      "votes": null
    },
    {
      "id": "617240",
      "postDate": "09/03/2019 23:33:53",
      "content": "<p>The small public sample size itself could lead to huge shake.</p>\n\n<p>I run the following experiment using 2015 data:</p>\n\n<ol>\n<li><p>Sampled ~15,000 data from the test set to simulate the complete test set in this competition. </p></li>\n<li><p>Sampling ~ 2,000 data from the 15,000 data to simulate the public test set using different random seeds. Do this for ~1,000 times. </p></li>\n</ol>\n\n<p>I am choosing 2,000 and 15,000 because we have 1,928 images in the public and around 13,000 images in the private for this competition. The 'public qwk' distribution is following. While the 'true qwk' over 15,000 data is around 0.835, the qwk on 'public set' could vary from 0.79 to 0.87, depending on different the random seeds used for sampling. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Ffdaebf5d6877e6dc3baad98ff8a4f89c%2Ffigure1.png?generation=1567553360059542&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The small public sample size itself could lead to huge shake.\n\nI run the following experiment using 2015 data:\n\n1. Sampled ~15,000 data from the test set to simulate the complete test set in this competition. \n\n2. Sampling ~ 2,000 data from the 15,000 data to simulate the public test set using different random seeds. Do this for ~1,000 times. \n\nI am choosing 2,000 and 15,000 because we have 1,928 images in the public and around 13,000 images in the private for this competition. The 'public qwk' distribution is following. While the 'true qwk' over 15,000 data is around 0.835, the qwk on 'public set' could vary from 0.79 to 0.87, depending on different the random seeds used for sampling. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Ffdaebf5d6877e6dc3baad98ff8a4f89c%2Ffigure1.png?generation=1567553360059542&amp;alt=media)",
      "votes": null
    },
    {
      "id": "617242",
      "postDate": "09/03/2019 23:40:02",
      "content": "<p>Gary is gonna get 1st place.</p>",
      "rawMarkdown": "Gary is gonna get 1st place.",
      "votes": null
    },
    {
      "id": "617255",
      "postDate": "09/04/2019 00:15:44",
      "content": "<p>Can't really tell is this wholehearted blessing or poisonous milk. </p>",
      "rawMarkdown": "Can't really tell is this wholehearted blessing or poisonous milk.",
      "votes": null
    },
    {
      "id": "617335",
      "postDate": "09/04/2019 02:56:27",
      "content": "<p>Take a look at past 10 shake-up visualization here : \n<a href=\"https://www.kaggle.com/pednoi/visualize-the-shakeups-of-10-recent-competitions\">https://www.kaggle.com/pednoi/visualize-the-shakeups-of-10-recent-competitions</a></p>",
      "rawMarkdown": "Take a look at past 10 shake-up visualization here : \nhttps://www.kaggle.com/pednoi/visualize-the-shakeups-of-10-recent-competitions",
      "votes": null
    },
    {
      "id": "617378",
      "postDate": "09/04/2019 04:13:55",
      "content": "<p>Good analysis！</p>",
      "rawMarkdown": "Good analysis！",
      "votes": null
    },
    {
      "id": "617505",
      "postDate": "09/04/2019 07:38:47",
      "content": "<p>So the shake will be huge. Everyone that got his LB &gt; 0.81 have a chance to grab a gold medal.</p>",
      "rawMarkdown": "So the shake will be huge. Everyone that got his LB &gt; 0.81 have a chance to grab a gold medal.",
      "votes": null
    },
    {
      "id": "617960",
      "postDate": "09/04/2019 16:59:17",
      "content": "<p>Just following up regarding the sample size. If we follow the 2015 public-private ratio (~10k public， ~43k private), the distribution of 'public qwk' (see below) is much more stable than the one I showed previous.   </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Fd0a4697eaeea968318705141f6f2cb70%2Fqwk_distribution_sample10000.png?generation=1567616351714122&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Just following up regarding the sample size. If we follow the 2015 public-private ratio (~10k public， ~43k private), the distribution of 'public qwk' (see below) is much more stable than the one I showed previous.   \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Fd0a4697eaeea968318705141f6f2cb70%2Fqwk_distribution_sample10000.png?generation=1567616351714122&amp;alt=media)",
      "votes": null
    },
    {
      "id": "618121",
      "postDate": "09/04/2019 21:42:37",
      "content": "<p>Really impressive analysis !  💯 </p>",
      "rawMarkdown": "Really impressive analysis !  💯",
      "votes": null
    },
    {
      "id": "618334",
      "postDate": "09/05/2019 06:06:30",
      "content": "<p>Now time to pray.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F9728832ceea692a1b06ace2ecdac2a96%2Fpray.jpg?generation=1567663581744594&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Now time to pray.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F9728832ceea692a1b06ace2ecdac2a96%2Fpray.jpg?generation=1567663581744594&amp;alt=media)",
      "votes": null
    },
    {
      "id": "619450",
      "postDate": "09/06/2019 08:00:17",
      "content": "<p>i dont know why kaggle cant select the max of all the scores...i dont think that will require more CPUs and resources..\nSince every one is blind to private test so using max is very efficient strategy..\nwhat if top rankers lower  score submission gets a higher score..</p>",
      "rawMarkdown": "i dont know why kaggle cant select the max of all the scores...i dont think that will require more CPUs and resources..\nSince every one is blind to private test so using max is very efficient strategy..\nwhat if top rankers lower  score submission gets a higher score..",
      "votes": null
    },
    {
      "id": "619632",
      "postDate": "09/06/2019 11:29:58",
      "content": "<p>Part of the challenge in any competition is creating a model that is generalisable. If Kaggle just selected your max then no doubt in many cases we would just get a model that happens to perform a little better on this particular test set, which isn't what we want.</p>",
      "rawMarkdown": "Part of the challenge in any competition is creating a model that is generalisable. If Kaggle just selected your max then no doubt in many cases we would just get a model that happens to perform a little better on this particular test set, which isn't what we want.",
      "votes": null
    },
    {
      "id": "619636",
      "postDate": "09/06/2019 11:39:25",
      "content": "<p>The main challenge is to select the right submissions for private dataset. In real world you also have to select your models you put into production. This is a skill on itself, and more often than not one of the hardest ones.</p>",
      "rawMarkdown": "The main challenge is to select the right submissions for private dataset. In real world you also have to select your models you put into production. This is a skill on itself, and more often than not one of the hardest ones.",
      "votes": null
    },
    {
      "id": "619870",
      "postDate": "09/06/2019 17:17:53",
      "content": "<p>But in real world we have more information regarding the unknown data, how they are collected, distribution, etc. </p>\n\n<p>In all the competitions, we do not know how the public and private test set are separated, we can only pray they are divided randomly. </p>\n\n<p>For this competition, I am extremely interested to know how they divide public and private. It is very hard for me to believe they are separated randomly, since the diabetic retinopathy grading distribution should not vary a lot in different regions. </p>",
      "rawMarkdown": "But in real world we have more information regarding the unknown data, how they are collected, distribution, etc. \n\nIn all the competitions, we do not know how the public and private test set are separated, we can only pray they are divided randomly. \n\nFor this competition, I am extremely interested to know how they divide public and private. It is very hard for me to believe they are separated randomly, since the diabetic retinopathy grading distribution should not vary a lot in different regions.",
      "votes": null
    },
    {
      "id": "620102",
      "postDate": "09/07/2019 03:19:18",
      "content": "<p>Ha ha, great advice, and I'll take it tomorrow👀 </p>",
      "rawMarkdown": "Ha ha, great advice, and I'll take it tomorrow👀",
      "votes": null
    },
    {
      "id": "620244",
      "postDate": "09/07/2019 07:49:26",
      "content": "<p><a href=\"/naivelamb\">@naivelamb</a> Fair point!</p>",
      "rawMarkdown": "naivelamb Fair point!",
      "votes": null
    },
    {
      "id": "620659",
      "postDate": "09/07/2019 19:56:10",
      "content": "<p>very nice</p>",
      "rawMarkdown": "very nice",
      "votes": null
    },
    {
      "id": "620721",
      "postDate": "09/07/2019 22:44:32",
      "content": "<p>some how i have always found in all submission that unless 2 category count hitting 1200+ mark score was always below .80 ... with 1 and 3s  between 130 and 160... giving 80+ score..</p>",
      "rawMarkdown": "some how i have always found in all submission that unless 2 category count hitting 1200+ mark score was always below .80 ... with 1 and 3s  between 130 and 160... giving 80+ score..",
      "votes": null
    },
    {
      "id": "620855",
      "postDate": "09/08/2019 02:24:36",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F1a74fc74a500408034e77a6abe09c095%2Fasdfasdf.jpg?generation=1567909407471950&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F1a74fc74a500408034e77a6abe09c095%2Fasdfasdf.jpg?generation=1567909407471950&amp;alt=media)",
      "votes": null
    },
    {
      "id": "620958",
      "postDate": "09/08/2019 05:14:47",
      "content": "<p>A final milk shake after conpleting mixing...I gained 50 ,could have been 90 \n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107946\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107946</a></p>",
      "rawMarkdown": "A final milk shake after conpleting mixing...I gained 50 ,could have been 90 \nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107946",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 614191,
      "author_name": "rishabhiitbhu",
      "author_url": "",
      "post_date": "08/31/2019 07:31:43",
      "content": "<blockquote>\n  <p>grade distribution will be different between public and private</p>\n</blockquote>\n\n<p>How can you be so sure about this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 614198,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "08/31/2019 07:41:58",
          "content": "<p>A Difficult question. <br>\nBut if we use common sense, public distribution would be abnormal. <br>\nI suppose the number of private test dataset is enough large to reflect usual grade distribution(less severity).  </p>\n\n<p><strong>update 1</strong>\nBelow figure was posted by <a href=\"/puremath86\">@puremath86</a> and shows the predicted public distribution. <br>\nWe can easily see how this distribution is strange compared with the usual distribution. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F838971%2Fa86219e1c1eca9d598aaf69adb8bb66a%2FCapture2.PNG?generation=1561952985345841&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>update 2</strong>\nMy best single model (LB = .830) distribution is as following.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fdece607277ac88541c04b34f81666e21%2Fimage.png?generation=1567331246673536&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 614221,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "08/31/2019 08:21:06",
          "content": "<p>Tbh, I don't know what the private test set is gonna be like. Just an additional information:</p>\n\n<p>Normalized distribution, class wise 2015 competition.</p>\n\n<p>2015's train set:\n0    0.734783\n2    0.150658\n1    0.069550\n3    0.024853\n4    0.020156</p>\n\n<p>2015 public test set:\n0    0.745461\n2    0.144783\n1    0.066019\n4    0.022006\n3    0.021731</p>\n\n<p>2015 private test set:\n0    0.735950\n2    0.147223\n1    0.071291\n3    0.022897\n4    0.022639</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 614994,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "09/01/2019 10:35:04",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> I see your plots. I kinda get the same distribution for my best test submission, but that doesn't tell us anything about the private test set, right? Just curious, has anyone seen any Kaggle competition where there was a major difference in public and private test set? </p>\n\n<p>A point to remember, QWK shouldn't be trusted, two quite different predictions can have same QWK score. For example: one of my submissions which scores 0.829 on LB:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F71d8f80636b50eceac15423406a6966e%2Fdist.png?generation=1567334053540493&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615013,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "09/01/2019 11:04:40",
          "content": "<p><a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> There are many competitions where private and public are totally different. One extreme case is “VSB powerline prediction”, where shakeup was like an earthquake there. I fell from 100 to 5xx . Top3 (or top5) also fell like 100s places. (But my model was not designed to be robust at that time, so that was my bad)</p>\n\n<p>In this competition, however, I agree with <a href=\"/maxwell110\">@maxwell110</a> ; I think due to the nature and size of dataset, the private test set should have similar distribution to training set of 2015/2019, but may be not like public test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615035,
          "author_name": "garybios",
          "author_url": "",
          "post_date": "09/01/2019 11:32:52",
          "content": "<p>who knows? all we can do is trust our cv?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615110,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "09/01/2019 13:31:39",
          "content": "<p>And if we removed all 0s in 2015 train, 2015 test, 2019 train. We can find the ratios between class 1, 2, 3, 4 of these three datasets are similar. Maybe that's the natural rule.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620721,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/07/2019 22:44:32",
          "content": "<p>some how i have always found in all submission that unless 2 category count hitting 1200+ mark score was always below .80 ... with 1 and 3s  between 130 and 160... giving 80+ score..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 614201,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "08/31/2019 07:47:17",
      "content": "<p>yeah, a big problem...</p>",
      "votes": null,
      "replies": [
        {
          "id": 617242,
          "author_name": "manyfoldcv",
          "author_url": "",
          "post_date": "09/03/2019 23:40:02",
          "content": "<p>Gary is gonna get 1st place.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617255,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "09/04/2019 00:15:44",
          "content": "<p>Can't really tell is this wholehearted blessing or poisonous milk. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 614373,
      "author_name": "hirune924",
      "author_url": "",
      "post_date": "08/31/2019 12:07:42",
      "content": "<p>I expect teams with an LB score of 0.81 or higher to shift about ± 0.03 〜 0.04.</p>",
      "votes": null,
      "replies": [
        {
          "id": 614385,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "08/31/2019 12:21:53",
          "content": "<p>is this a gut feeling or something else?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 614402,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "08/31/2019 12:53:21",
          "content": "<p>Yes. It is just a feeling based on my trial. The truth is only God knows.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 614969,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "09/01/2019 09:57:26",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> \nI almost stand by you. <br>\nIn my simulation, score will fluctuate by at least 1% within about 50% confidential interval. <br>\nIt's not surprising if shake will be equal to such a high degree shift (3 ~ 4%).  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615630,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "09/02/2019 07:36:52",
          "content": "<p>I measured high variance (0.007), just by changing random seed</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615935,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "09/02/2019 14:18:47",
          "content": "<p><a href=\"/nemethpeti\">@nemethpeti</a> True, this variance is prominent when using TTA.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 614999,
      "author_name": "octpath0302",
      "author_url": "",
      "post_date": "09/01/2019 10:43:46",
      "content": "<p>This is almost the first competition I've worked on until the last, so I'm not sure about Shake.</p>\n\n<p>I think that there is little difference in class distribution between public and private.  (my hope)\nHowever, in serious cases, I expect that the number of data is small and there is a variety of image findings. I'm afraid that this may cause Shake in the high score zone.</p>\n\n<p>Considering my observations (metrics vs losses, CV vs LB, etc.), I suspect that teams with high LB (e. g. above 0.83) are overfitting.  But this may be a biased analysis by wish. 😎 </p>",
      "votes": null,
      "replies": [
        {
          "id": 615014,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "09/01/2019 11:05:26",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a> \nAlmost the same opinion.\nGood Luck to you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615628,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "09/02/2019 07:34:24",
          "content": "<p>Anybody, who has higher score than my team is overfitting :-D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 615950,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "09/02/2019 14:33:58",
      "content": "<p>Yeah... I think they will be a shake the question is how bad. Whats intresting about this compettioin that you can  achieve LB score more than 0.81 easily using difrrent variations of training data (atleast 3 or 4 diffrent ways  e.g just sing 2019 data, 2019 data with some combination from 2015 and etc). All of this methods will give you diffrent CV. The question is what should you trust ? =) </p>",
      "votes": null,
      "replies": [
        {
          "id": 615978,
          "author_name": "garybios",
          "author_url": "",
          "post_date": "09/02/2019 14:59:40",
          "content": "<p>agree with you, I don't know what should be trusted...😂 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615988,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "09/02/2019 15:09:06",
          "content": "<p>At the end I think I am going to roll dice and select submission =) </p>\n\n<p>Edit Here is my stratagy =) \n1) go <a href=\"https://www.google.com/search?q=dice+roller\">here</a>. Change the number of sides to number of total submisions \n2) roll the dice \n3) select submision file corresponding to the number of the dice =) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 616021,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "09/02/2019 15:56:39",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Haha, thanks for the advice. When almost all models are reaching val qwk of 0.90+ and public test is crazy like this, your strategy seems legit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 616510,
          "author_name": "ankit001",
          "author_url": "",
          "post_date": "09/03/2019 07:40:15",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Can you share 1-2 variations of training data to achieve &gt; 0.81? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 619450,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/06/2019 08:00:17",
          "content": "<p>i dont know why kaggle cant select the max of all the scores...i dont think that will require more CPUs and resources..\nSince every one is blind to private test so using max is very efficient strategy..\nwhat if top rankers lower  score submission gets a higher score..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 619632,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "09/06/2019 11:29:58",
          "content": "<p>Part of the challenge in any competition is creating a model that is generalisable. If Kaggle just selected your max then no doubt in many cases we would just get a model that happens to perform a little better on this particular test set, which isn't what we want.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 619636,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "09/06/2019 11:39:25",
          "content": "<p>The main challenge is to select the right submissions for private dataset. In real world you also have to select your models you put into production. This is a skill on itself, and more often than not one of the hardest ones.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 619870,
          "author_name": "naivelamb",
          "author_url": "",
          "post_date": "09/06/2019 17:17:53",
          "content": "<p>But in real world we have more information regarding the unknown data, how they are collected, distribution, etc. </p>\n\n<p>In all the competitions, we do not know how the public and private test set are separated, we can only pray they are divided randomly. </p>\n\n<p>For this competition, I am extremely interested to know how they divide public and private. It is very hard for me to believe they are separated randomly, since the diabetic retinopathy grading distribution should not vary a lot in different regions. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620102,
          "author_name": "homoalways",
          "author_url": "",
          "post_date": "09/07/2019 03:19:18",
          "content": "<p>Ha ha, great advice, and I'll take it tomorrow👀 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620244,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "09/07/2019 07:49:26",
          "content": "<p><a href=\"/naivelamb\">@naivelamb</a> Fair point!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 617043,
      "author_name": "kupchanski",
      "author_url": "",
      "post_date": "09/03/2019 17:38:06",
      "content": "<p><img src=\"https://sun9-44.userapi.com/c854532/v854532836/dd1fa/E_ziPIlBdRA.jpg\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 618334,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "09/05/2019 06:06:30",
          "content": "<p>Now time to pray.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F9728832ceea692a1b06ace2ecdac2a96%2Fpray.jpg?generation=1567663581744594&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620855,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "09/08/2019 02:24:36",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F1a74fc74a500408034e77a6abe09c095%2Fasdfasdf.jpg?generation=1567909407471950&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620958,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/08/2019 05:14:47",
          "content": "<p>A final milk shake after conpleting mixing...I gained 50 ,could have been 90 \n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107946\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107946</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 617240,
      "author_name": "naivelamb",
      "author_url": "",
      "post_date": "09/03/2019 23:33:53",
      "content": "<p>The small public sample size itself could lead to huge shake.</p>\n\n<p>I run the following experiment using 2015 data:</p>\n\n<ol>\n<li><p>Sampled ~15,000 data from the test set to simulate the complete test set in this competition. </p></li>\n<li><p>Sampling ~ 2,000 data from the 15,000 data to simulate the public test set using different random seeds. Do this for ~1,000 times. </p></li>\n</ol>\n\n<p>I am choosing 2,000 and 15,000 because we have 1,928 images in the public and around 13,000 images in the private for this competition. The 'public qwk' distribution is following. While the 'true qwk' over 15,000 data is around 0.835, the qwk on 'public set' could vary from 0.79 to 0.87, depending on different the random seeds used for sampling. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Ffdaebf5d6877e6dc3baad98ff8a4f89c%2Ffigure1.png?generation=1567553360059542&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 617378,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "09/04/2019 04:13:55",
          "content": "<p>Good analysis！</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617505,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "09/04/2019 07:38:47",
          "content": "<p>So the shake will be huge. Everyone that got his LB &gt; 0.81 have a chance to grab a gold medal.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617960,
          "author_name": "naivelamb",
          "author_url": "",
          "post_date": "09/04/2019 16:59:17",
          "content": "<p>Just following up regarding the sample size. If we follow the 2015 public-private ratio (~10k public， ~43k private), the distribution of 'public qwk' (see below) is much more stable than the one I showed previous.   </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Fd0a4697eaeea968318705141f6f2cb70%2Fqwk_distribution_sample10000.png?generation=1567616351714122&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 618121,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "09/04/2019 21:42:37",
          "content": "<p>Really impressive analysis !  💯 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 617335,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "09/04/2019 02:56:27",
      "content": "<p>Take a look at past 10 shake-up visualization here : \n<a href=\"https://www.kaggle.com/pednoi/visualize-the-shakeups-of-10-recent-competitions\">https://www.kaggle.com/pednoi/visualize-the-shakeups-of-10-recent-competitions</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 620659,
      "author_name": "harrymi",
      "author_url": "",
      "post_date": "09/07/2019 19:56:10",
      "content": "<p>very nice</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "614176": "What your thought about LB shake ?\n\nI guess shake is moderate.\n\nReason:  \n- grade distribution will be different between public and private\n- QWK is unstable evaluation metric",
    "614191": "&gt;grade distribution will be different between public and private\n\n How can you be so sure about this?",
    "614198": "A Difficult question.  \nBut if we use common sense, public distribution would be abnormal.  \nI suppose the number of private test dataset is enough large to reflect usual grade distribution(less severity).  \n\n**update 1**\nBelow figure was posted by @puremath86 and shows the predicted public distribution.    \nWe can easily see how this distribution is strange compared with the usual distribution.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F838971%2Fa86219e1c1eca9d598aaf69adb8bb66a%2FCapture2.PNG?generation=1561952985345841&amp;alt=media)\n\n**update 2**\nMy best single model (LB = .830) distribution is as following.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fdece607277ac88541c04b34f81666e21%2Fimage.png?generation=1567331246673536&amp;alt=media)",
    "614201": "yeah, a big problem...",
    "614221": "Tbh, I don't know what the private test set is gonna be like. Just an additional information:\n\nNormalized distribution, class wise 2015 competition.\n\n2015's train set:\n0    0.734783\n2    0.150658\n1    0.069550\n3    0.024853\n4    0.020156\n\n2015 public test set:\n0    0.745461\n2    0.144783\n1    0.066019\n4    0.022006\n3    0.021731\n\n2015 private test set:\n0    0.735950\n2    0.147223\n1    0.071291\n3    0.022897\n4    0.022639",
    "614373": "I expect teams with an LB score of 0.81 or higher to shift about ± 0.03 〜 0.04.",
    "614385": "is this a gut feeling or something else?",
    "614402": "Yes. It is just a feeling based on my trial. The truth is only God knows.",
    "614969": "hirune924 \nI almost stand by you.  \nIn my simulation, score will fluctuate by at least 1% within about 50% confidential interval.  \nIt's not surprising if shake will be equal to such a high degree shift (3 ~ 4%).",
    "614994": "maxwell110 I see your plots. I kinda get the same distribution for my best test submission, but that doesn't tell us anything about the private test set, right? Just curious, has anyone seen any Kaggle competition where there was a major difference in public and private test set? \n\nA point to remember, QWK shouldn't be trusted, two quite different predictions can have same QWK score. For example: one of my submissions which scores 0.829 on LB:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F71d8f80636b50eceac15423406a6966e%2Fdist.png?generation=1567334053540493&amp;alt=media)",
    "614999": "This is almost the first competition I've worked on until the last, so I'm not sure about Shake.\n\nI think that there is little difference in class distribution between public and private.  (my hope)\nHowever, in serious cases, I expect that the number of data is small and there is a variety of image findings. I'm afraid that this may cause Shake in the high score zone.\n\nConsidering my observations (metrics vs losses, CV vs LB, etc.), I suspect that teams with high LB (e. g. above 0.83) are overfitting.  But this may be a biased analysis by wish. 😎",
    "615013": "rishabhiitbhu There are many competitions where private and public are totally different. One extreme case is “VSB powerline prediction”, where shakeup was like an earthquake there. I fell from 100 to 5xx . Top3 (or top5) also fell like 100s places. (But my model was not designed to be robust at that time, so that was my bad)\n\nIn this competition, however, I agree with @maxwell110 ; I think due to the nature and size of dataset, the private test set should have similar distribution to training set of 2015/2019, but may be not like public test set.",
    "615014": "octpath0302 \nAlmost the same opinion.\nGood Luck to you!",
    "615035": "who knows? all we can do is trust our cv?",
    "615110": "And if we removed all 0s in 2015 train, 2015 test, 2019 train. We can find the ratios between class 1, 2, 3, 4 of these three datasets are similar. Maybe that's the natural rule.",
    "615628": "Anybody, who has higher score than my team is overfitting :-D",
    "615630": "I measured high variance (0.007), just by changing random seed",
    "615935": "nemethpeti True, this variance is prominent when using TTA.",
    "615950": "Yeah... I think they will be a shake the question is how bad. Whats intresting about this compettioin that you can  achieve LB score more than 0.81 easily using difrrent variations of training data (atleast 3 or 4 diffrent ways  e.g just sing 2019 data, 2019 data with some combination from 2015 and etc). All of this methods will give you diffrent CV. The question is what should you trust ? =)",
    "615978": "agree with you, I don't know what should be trusted...😂",
    "615988": "At the end I think I am going to roll dice and select submission =) \n\nEdit Here is my stratagy =) \n1) go [here](https://www.google.com/search?q=dice+roller). Change the number of sides to number of total submisions \n2) roll the dice \n3) select submision file corresponding to the number of the dice =)",
    "616021": "drhabib Haha, thanks for the advice. When almost all models are reaching val qwk of 0.90+ and public test is crazy like this, your strategy seems legit.",
    "616510": "drhabib Can you share 1-2 variations of training data to achieve &gt; 0.81?",
    "617043": "![](https://sun9-44.userapi.com/c854532/v854532836/dd1fa/E_ziPIlBdRA.jpg)",
    "617240": "The small public sample size itself could lead to huge shake.\n\nI run the following experiment using 2015 data:\n\n1. Sampled ~15,000 data from the test set to simulate the complete test set in this competition. \n\n2. Sampling ~ 2,000 data from the 15,000 data to simulate the public test set using different random seeds. Do this for ~1,000 times. \n\nI am choosing 2,000 and 15,000 because we have 1,928 images in the public and around 13,000 images in the private for this competition. The 'public qwk' distribution is following. While the 'true qwk' over 15,000 data is around 0.835, the qwk on 'public set' could vary from 0.79 to 0.87, depending on different the random seeds used for sampling. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Ffdaebf5d6877e6dc3baad98ff8a4f89c%2Ffigure1.png?generation=1567553360059542&amp;alt=media)",
    "617242": "Gary is gonna get 1st place.",
    "617255": "Can't really tell is this wholehearted blessing or poisonous milk.",
    "617335": "Take a look at past 10 shake-up visualization here : \nhttps://www.kaggle.com/pednoi/visualize-the-shakeups-of-10-recent-competitions",
    "617378": "Good analysis！",
    "617505": "So the shake will be huge. Everyone that got his LB &gt; 0.81 have a chance to grab a gold medal.",
    "617960": "Just following up regarding the sample size. If we follow the 2015 public-private ratio (~10k public， ~43k private), the distribution of 'public qwk' (see below) is much more stable than the one I showed previous.   \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Fd0a4697eaeea968318705141f6f2cb70%2Fqwk_distribution_sample10000.png?generation=1567616351714122&amp;alt=media)",
    "618121": "Really impressive analysis !  💯",
    "618334": "Now time to pray.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F9728832ceea692a1b06ace2ecdac2a96%2Fpray.jpg?generation=1567663581744594&amp;alt=media)",
    "619450": "i dont know why kaggle cant select the max of all the scores...i dont think that will require more CPUs and resources..\nSince every one is blind to private test so using max is very efficient strategy..\nwhat if top rankers lower  score submission gets a higher score..",
    "619632": "Part of the challenge in any competition is creating a model that is generalisable. If Kaggle just selected your max then no doubt in many cases we would just get a model that happens to perform a little better on this particular test set, which isn't what we want.",
    "619636": "The main challenge is to select the right submissions for private dataset. In real world you also have to select your models you put into production. This is a skill on itself, and more often than not one of the hardest ones.",
    "619870": "But in real world we have more information regarding the unknown data, how they are collected, distribution, etc. \n\nIn all the competitions, we do not know how the public and private test set are separated, we can only pray they are divided randomly. \n\nFor this competition, I am extremely interested to know how they divide public and private. It is very hard for me to believe they are separated randomly, since the diabetic retinopathy grading distribution should not vary a lot in different regions.",
    "620102": "Ha ha, great advice, and I'll take it tomorrow👀",
    "620244": "naivelamb Fair point!",
    "620659": "very nice",
    "620721": "some how i have always found in all submission that unless 2 category count hitting 1200+ mark score was always below .80 ... with 1 and 3s  between 130 and 160... giving 80+ score..",
    "620855": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F1a74fc74a500408034e77a6abe09c095%2Fasdfasdf.jpg?generation=1567909407471950&amp;alt=media)",
    "620958": "A final milk shake after conpleting mixing...I gained 50 ,could have been 90 \nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107946"
  },
  "source": "meta"
}