{
  "id": 217328,
  "title": "A shake-up estimation",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/217328",
  "author_name": "",
  "post_date": "2021-02-06T11:05:57.548479300Z",
  "votes": 38,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I did an experiment to guess how much a shake-up is going to be  <br>\nFrom my CV 0.904 OOF, I randomly split the oof into 31% public and 69% private<br>\nWith 100,000 iterations, I plotted these graphs</p>\n<p><img src=\"https://i.imgur.com/ppBRLwZ.png\" alt=\"\"></p>\n<p><img src=\"https://i.imgur.com/5e0RSS9.png\" alt=\"\"></p>\n<p>I think the shake-up will not be so large in terms of leaderboard scores<br>\n75% of  (Public LB - Private LB) are in the range of  [-0.0028, 0.003]<br>\n90 % cases of (Public LB - Private LB) are in the range of [-0.0055, 0.0056]</p>\n<p>Trusting in LB might not be a bad choice in this competition…</p>",
  "messages": [
    {
      "id": "1188585",
      "postDate": "02/06/2021 11:05:57",
      "content": "<p>I did an experiment to guess how much a shake-up is going to be  <br>\nFrom my CV 0.904 OOF, I randomly split the oof into 31% public and 69% private<br>\nWith 100,000 iterations, I plotted these graphs</p>\n<p><img src=\"https://i.imgur.com/ppBRLwZ.png\" alt=\"\"></p>\n<p><img src=\"https://i.imgur.com/5e0RSS9.png\" alt=\"\"></p>\n<p>I think the shake-up will not be so large in terms of leaderboard scores<br>\n75% of  (Public LB - Private LB) are in the range of  [-0.0028, 0.003]<br>\n90 % cases of (Public LB - Private LB) are in the range of [-0.0055, 0.0056]</p>\n<p>Trusting in LB might not be a bad choice in this competition…</p>",
      "rawMarkdown": "I did an experiment to guess how much a shake-up is going to be  \nFrom my CV 0.904 OOF, I randomly split the oof into 31% public and 69% private\nWith 100,000 iterations, I plotted these graphs\n\n![](https://i.imgur.com/ppBRLwZ.png)\n\n![](https://i.imgur.com/5e0RSS9.png)\n\nI think the shake-up will not be so large in terms of leaderboard scores\n75% of  (Public LB - Private LB) are in the range of  [-0.0028, 0.003]\n90 % cases of (Public LB - Private LB) are in the range of [-0.0055, 0.0056]\n\nTrusting in LB might not be a bad choice in this competition...",
      "votes": null
    },
    {
      "id": "1188698",
      "postDate": "02/06/2021 13:04:35",
      "content": "<p>Excellent work!<br>\nMy previous guess is [+-0.005, +-0.009].</p>\n<p>According to your experiments, it is even a smaller shake then I expect.</p>",
      "rawMarkdown": "Excellent work!\nMy previous guess is [+-0.005, +-0.009].\n\nAccording to your experiments, it is even a smaller shake then I expect.",
      "votes": null
    },
    {
      "id": "1188789",
      "postDate": "02/06/2021 14:24:36",
      "content": "<p>I don't know, half the people here are saying, the shake will be huge, only rely on your cv and the other half says it won't be that bad. </p>",
      "rawMarkdown": "I don't know, half the people here are saying, the shake will be huge, only rely on your cv and the other half says it won't be that bad.",
      "votes": null
    },
    {
      "id": "1188860",
      "postDate": "02/06/2021 14:47:29",
      "content": "<p>The shake will not that bad in terms of score. 2-3 previous leaks were all within +- 0.003.<br>\nHowever, too much people are now with the same score, so the shake of the rank might be huge. </p>",
      "rawMarkdown": "The shake will not that bad in terms of score. 2-3 previous leaks were all within +- 0.003.\nHowever, too much people are now with the same score, so the shake of the rank might be huge.",
      "votes": null
    },
    {
      "id": "1188892",
      "postDate": "02/06/2021 15:04:22",
      "content": "<p>Let me give my opinion cautiously about the private leaderboard rank<br>\nI think the private leaderboard rank will not be volatile like the private leaderboard score which I mentioned above<br>\nIt seems like that a lot of competitors lack diversity.<br>\nThe Bronze medal range is stuck at 0.903<br>\nWith some additions to public notebooks, some competitors can improve further to LB 0.903, 0.904, 0.905  more, and we are stuck in that so narrow range.<br>\nEven though competitors just resort to LB and fail, it will result in falling together.</p>\n<p>Those who creatively develop a strong classifier can luckily shake-up but it will be few</p>",
      "rawMarkdown": "Let me give my opinion cautiously about the private leaderboard rank\nI think the private leaderboard rank will not be volatile like the private leaderboard score which I mentioned above\nIt seems like that a lot of competitors lack diversity.\nThe Bronze medal range is stuck at 0.903\nWith some additions to public notebooks, some competitors can improve further to LB 0.903, 0.904, 0.905  more, and we are stuck in that so narrow range.\nEven though competitors just resort to LB and fail, it will result in falling together.\n\nThose who creatively develop a strong classifier can luckily shake-up but it will be few",
      "votes": null
    },
    {
      "id": "1188932",
      "postDate": "02/06/2021 15:34:38",
      "content": "<p>Good analysis! The number of test images is roughly 15,000. So I think it will be better that you randomly <strong>sample and split</strong> the oof into 31% and 69%, for an expectation.</p>",
      "rawMarkdown": "Good analysis! The number of test images is roughly 15,000. So I think it will be better that you randomly **sample and split** the oof into 31% and 69%, for an expectation.",
      "votes": null
    },
    {
      "id": "1189581",
      "postDate": "02/07/2021 05:55:42",
      "content": "<p>Thank you for the good analysis!<br>\nActually, the Public score is more or less too fits the public noisy test data, so I think the Public-Private range will shift to the positive side of your analysis results.</p>",
      "rawMarkdown": "Thank you for the good analysis!\nActually, the Public score is more or less too fits the public noisy test data, so I think the Public-Private range will shift to the positive side of your analysis results.",
      "votes": null
    },
    {
      "id": "1190782",
      "postDate": "02/08/2021 02:14:00",
      "content": "<p>i think cv is unreliable, it got at least 3-4% noisy labels. 👀</p>",
      "rawMarkdown": "i think cv is unreliable, it got at least 3-4% noisy labels. 👀",
      "votes": null
    },
    {
      "id": "1195083",
      "postDate": "02/10/2021 14:37:30",
      "content": "<p>I feel the shake up will have major impact in .903 to .905 range as most lies in it. So someone in 100 may goes into 200 and vice versa.</p>",
      "rawMarkdown": "I feel the shake up will have major impact in .903 to .905 range as most lies in it. So someone in 100 may goes into 200 and vice versa.",
      "votes": null
    },
    {
      "id": "1195161",
      "postDate": "02/10/2021 15:38:53",
      "content": "<p>In terms of the leaderboard score, I agree.<br>\nthere is a 30% chance that a public score .903 turns to .905 and vice versa.</p>",
      "rawMarkdown": "In terms of the leaderboard score, I agree.\nthere is a 30% chance that a public score .903 turns to .905 and vice versa.",
      "votes": null
    },
    {
      "id": "1196918",
      "postDate": "02/11/2021 18:56:22",
      "content": "<p>I think shake up will have impact every where. Even at the top. </p>",
      "rawMarkdown": "I think shake up will have impact every where. Even at the top.",
      "votes": null
    },
    {
      "id": "1196968",
      "postDate": "02/11/2021 19:52:29",
      "content": "<p>Totally agree, I can think of the few points related to the shake up.<br>\nFirstly some might be using one or two folds for submission based on maximum cv score out of 5 folds (in case 5 folds are used) which might give a little boost in lb but might not generalise with remaining 69% test data. </p>\n<p>Secondly how noisy and similar/different is 69% of the test data with 31% data is not known. </p>\n<p>Thirdly models used might be over fitting 31% of the test data because everyone wants to rise in lb  but may or may not work as accurately with the remaining test data.</p>",
      "rawMarkdown": "Totally agree, I can think of the few points related to the shake up.\nFirstly some might be using one or two folds for submission based on maximum cv score out of 5 folds (in case 5 folds are used) which might give a little boost in lb but might not generalise with remaining 69% test data. \n\nSecondly how noisy and similar/different is 69% of the test data with 31% data is not known. \n\nThirdly models used might be over fitting 31% of the test data because everyone wants to rise in lb  but may or may not work as accurately with the remaining test data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1188698,
      "author_name": "zlanan",
      "author_url": "",
      "post_date": "02/06/2021 13:04:35",
      "content": "<p>Excellent work!<br>\nMy previous guess is [+-0.005, +-0.009].</p>\n<p>According to your experiments, it is even a smaller shake then I expect.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1188789,
      "author_name": "alexanderriedel",
      "author_url": "",
      "post_date": "02/06/2021 14:24:36",
      "content": "<p>I don't know, half the people here are saying, the shake will be huge, only rely on your cv and the other half says it won't be that bad. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1188860,
          "author_name": "zlanan",
          "author_url": "",
          "post_date": "02/06/2021 14:47:29",
          "content": "<p>The shake will not that bad in terms of score. 2-3 previous leaks were all within +- 0.003.<br>\nHowever, too much people are now with the same score, so the shake of the rank might be huge. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1188892,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "02/06/2021 15:04:22",
          "content": "<p>Let me give my opinion cautiously about the private leaderboard rank<br>\nI think the private leaderboard rank will not be volatile like the private leaderboard score which I mentioned above<br>\nIt seems like that a lot of competitors lack diversity.<br>\nThe Bronze medal range is stuck at 0.903<br>\nWith some additions to public notebooks, some competitors can improve further to LB 0.903, 0.904, 0.905  more, and we are stuck in that so narrow range.<br>\nEven though competitors just resort to LB and fail, it will result in falling together.</p>\n<p>Those who creatively develop a strong classifier can luckily shake-up but it will be few</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1190782,
          "author_name": "raininbox",
          "author_url": "",
          "post_date": "02/08/2021 02:14:00",
          "content": "<p>i think cv is unreliable, it got at least 3-4% noisy labels. 👀</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1188932,
      "author_name": "yukia18",
      "author_url": "",
      "post_date": "02/06/2021 15:34:38",
      "content": "<p>Good analysis! The number of test images is roughly 15,000. So I think it will be better that you randomly <strong>sample and split</strong> the oof into 31% and 69%, for an expectation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1189581,
      "author_name": "sunakuzira",
      "author_url": "",
      "post_date": "02/07/2021 05:55:42",
      "content": "<p>Thank you for the good analysis!<br>\nActually, the Public score is more or less too fits the public noisy test data, so I think the Public-Private range will shift to the positive side of your analysis results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1195083,
      "author_name": "vickygoyal",
      "author_url": "",
      "post_date": "02/10/2021 14:37:30",
      "content": "<p>I feel the shake up will have major impact in .903 to .905 range as most lies in it. So someone in 100 may goes into 200 and vice versa.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1195161,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "02/10/2021 15:38:53",
          "content": "<p>In terms of the leaderboard score, I agree.<br>\nthere is a 30% chance that a public score .903 turns to .905 and vice versa.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1196918,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "02/11/2021 18:56:22",
          "content": "<p>I think shake up will have impact every where. Even at the top. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1196968,
          "author_name": "vickygoyal",
          "author_url": "",
          "post_date": "02/11/2021 19:52:29",
          "content": "<p>Totally agree, I can think of the few points related to the shake up.<br>\nFirstly some might be using one or two folds for submission based on maximum cv score out of 5 folds (in case 5 folds are used) which might give a little boost in lb but might not generalise with remaining 69% test data. </p>\n<p>Secondly how noisy and similar/different is 69% of the test data with 31% data is not known. </p>\n<p>Thirdly models used might be over fitting 31% of the test data because everyone wants to rise in lb  but may or may not work as accurately with the remaining test data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1188585": "I did an experiment to guess how much a shake-up is going to be  \nFrom my CV 0.904 OOF, I randomly split the oof into 31% public and 69% private\nWith 100,000 iterations, I plotted these graphs\n\n![](https://i.imgur.com/ppBRLwZ.png)\n\n![](https://i.imgur.com/5e0RSS9.png)\n\nI think the shake-up will not be so large in terms of leaderboard scores\n75% of  (Public LB - Private LB) are in the range of  [-0.0028, 0.003]\n90 % cases of (Public LB - Private LB) are in the range of [-0.0055, 0.0056]\n\nTrusting in LB might not be a bad choice in this competition...",
    "1188698": "Excellent work!\nMy previous guess is [+-0.005, +-0.009].\n\nAccording to your experiments, it is even a smaller shake then I expect.",
    "1188789": "I don't know, half the people here are saying, the shake will be huge, only rely on your cv and the other half says it won't be that bad.",
    "1188860": "The shake will not that bad in terms of score. 2-3 previous leaks were all within +- 0.003.\nHowever, too much people are now with the same score, so the shake of the rank might be huge.",
    "1188892": "Let me give my opinion cautiously about the private leaderboard rank\nI think the private leaderboard rank will not be volatile like the private leaderboard score which I mentioned above\nIt seems like that a lot of competitors lack diversity.\nThe Bronze medal range is stuck at 0.903\nWith some additions to public notebooks, some competitors can improve further to LB 0.903, 0.904, 0.905  more, and we are stuck in that so narrow range.\nEven though competitors just resort to LB and fail, it will result in falling together.\n\nThose who creatively develop a strong classifier can luckily shake-up but it will be few",
    "1188932": "Good analysis! The number of test images is roughly 15,000. So I think it will be better that you randomly **sample and split** the oof into 31% and 69%, for an expectation.",
    "1189581": "Thank you for the good analysis!\nActually, the Public score is more or less too fits the public noisy test data, so I think the Public-Private range will shift to the positive side of your analysis results.",
    "1190782": "i think cv is unreliable, it got at least 3-4% noisy labels. 👀",
    "1195083": "I feel the shake up will have major impact in .903 to .905 range as most lies in it. So someone in 100 may goes into 200 and vice versa.",
    "1195161": "In terms of the leaderboard score, I agree.\nthere is a 30% chance that a public score .903 turns to .905 and vice versa.",
    "1196918": "I think shake up will have impact every where. Even at the top.",
    "1196968": "Totally agree, I can think of the few points related to the shake up.\nFirstly some might be using one or two folds for submission based on maximum cv score out of 5 folds (in case 5 folds are used) which might give a little boost in lb but might not generalise with remaining 69% test data. \n\nSecondly how noisy and similar/different is 69% of the test data with 31% data is not known. \n\nThirdly models used might be over fitting 31% of the test data because everyone wants to rise in lb  but may or may not work as accurately with the remaining test data."
  },
  "source": "meta"
}