{
  "id": 173438,
  "title": "Should we expect a shake-up?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173438",
  "author_name": "",
  "post_date": "2020-08-09T08:12:36.053711300Z",
  "votes": 12,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hello! How do you think, should we expect a shake-up in this competition? And if so, why?</p>\n\n<p>I personally expect to shake-up just cause a few following reasons:\n* Relation of the public and private dataset (30% and 70%), such a big relation difference between two datasets usually causes big shake-ups.\n* A lot of overfitted participants at the public LB\n* Many people just submitted CSV-files from the public notebooks, you can see a lot of people with the same score.\n* The Featured Prediction Competition is a cause of shake-up itself. This type of competition is much more susceptible to the shake-up than any Research Code Competition. The best example is M5 competitions with their 1000-2000 places numerous shake-ups.\n* Complexity and ambiguity of the issue. The task to predict a probability is one more factor of a shake-up. It's not just classification with softmax, it's a sigmoid task. So it's pretty possible to get overfitted randomly tuning CSV-files.</p>\n\n<p>UPD:\n* A huge class imbalance in the training dataset.</p>\n\n<p>I wish you all a successful shake-up.</p>",
  "messages": [
    {
      "id": "963701",
      "postDate": "08/09/2020 08:12:36",
      "content": "<p>Hello! How do you think, should we expect a shake-up in this competition? And if so, why?</p>\n\n<p>I personally expect to shake-up just cause a few following reasons:\n* Relation of the public and private dataset (30% and 70%), such a big relation difference between two datasets usually causes big shake-ups.\n* A lot of overfitted participants at the public LB\n* Many people just submitted CSV-files from the public notebooks, you can see a lot of people with the same score.\n* The Featured Prediction Competition is a cause of shake-up itself. This type of competition is much more susceptible to the shake-up than any Research Code Competition. The best example is M5 competitions with their 1000-2000 places numerous shake-ups.\n* Complexity and ambiguity of the issue. The task to predict a probability is one more factor of a shake-up. It's not just classification with softmax, it's a sigmoid task. So it's pretty possible to get overfitted randomly tuning CSV-files.</p>\n\n<p>UPD:\n* A huge class imbalance in the training dataset.</p>\n\n<p>I wish you all a successful shake-up.</p>",
      "rawMarkdown": "Hello! How do you think, should we expect a shake-up in this competition? And if so, why?\n\nI personally expect to shake-up just cause a few following reasons:\n* Relation of the public and private dataset (30% and 70%), such a big relation difference between two datasets usually causes big shake-ups.\n* A lot of overfitted participants at the public LB\n* Many people just submitted CSV-files from the public notebooks, you can see a lot of people with the same score.\n* The Featured Prediction Competition is a cause of shake-up itself. This type of competition is much more susceptible to the shake-up than any Research Code Competition. The best example is M5 competitions with their 1000-2000 places numerous shake-ups.\n* Complexity and ambiguity of the issue. The task to predict a probability is one more factor of a shake-up. It's not just classification with softmax, it's a sigmoid task. So it's pretty possible to get overfitted randomly tuning CSV-files.\n\nUPD:\n* A huge class imbalance in the training dataset.\n\nI wish you all a successful shake-up.",
      "votes": null
    },
    {
      "id": "963724",
      "postDate": "08/09/2020 08:32:09",
      "content": "<p>There are only 78 ones in public test. So, the role of luck/overfitting increases a lot. So, I expect a huge shakeup!! There is such a class imbalance, I feel even if you have good CV, you are not safe in this scenario. </p>",
      "rawMarkdown": "There are only 78 ones in public test. So, the role of luck/overfitting increases a lot. So, I expect a huge shakeup!! There is such a class imbalance, I feel even if you have good CV, you are not safe in this scenario.",
      "votes": null
    },
    {
      "id": "963775",
      "postDate": "08/09/2020 09:16:40",
      "content": "<p>All reasons you provide are legit for expecting a strong shakeup. </p>\n\n<p>At the same time, the test data distribution is somewhat different from the train data distribution, so it is very difficult to say how much of public LB overfitting is indeed overfitting and does not relate to private LB. If distribution discrepancies between train/test are much higher than those between public/private test, it can make trusting LB a good idea and reduce the magnitude of shakeup. </p>\n\n<p>Will see soon!</p>",
      "rawMarkdown": "All reasons you provide are legit for expecting a strong shakeup. \n\nAt the same time, the test data distribution is somewhat different from the train data distribution, so it is very difficult to say how much of public LB overfitting is indeed overfitting and does not relate to private LB. If distribution discrepancies between train/test are much higher than those between public/private test, it can make trusting LB a good idea and reduce the magnitude of shakeup. \n\nWill see soon!",
      "votes": null
    },
    {
      "id": "963810",
      "postDate": "08/09/2020 10:07:55",
      "content": "<p>In CV we trust...</p>",
      "rawMarkdown": "In CV we trust...",
      "votes": null
    },
    {
      "id": "963816",
      "postDate": "08/09/2020 10:16:16",
      "content": "<p>Even if there is distribution discrepancies between train/test, do you think 78 ones in public LB is enough to make generalisation decisions (for private lb)?</p>",
      "rawMarkdown": "Even if there is distribution discrepancies between train/test, do you think 78 ones in public LB is enough to make generalisation decisions (for private lb)?",
      "votes": null
    },
    {
      "id": "963821",
      "postDate": "08/09/2020 10:23:41",
      "content": "<p>Of course it is not enough. What I am saying is that we should not ignore public LB because chasing CV can lead to overfitting the train set. We know there are some distribution discrepancies. They will definitely play some role, but this effect might be very small in the end compared to the overfitting consequences.</p>",
      "rawMarkdown": "Of course it is not enough. What I am saying is that we should not ignore public LB because chasing CV can lead to overfitting the train set. We know there are some distribution discrepancies. They will definitely play some role, but this effect might be very small in the end compared to the overfitting consequences.",
      "votes": null
    },
    {
      "id": "963824",
      "postDate": "08/09/2020 10:28:10",
      "content": "<p><a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> what do you mean with distribution discrepancies? target distribution?</p>",
      "rawMarkdown": "kozodoi what do you mean with distribution discrepancies? target distribution?",
      "votes": null
    },
    {
      "id": "963836",
      "postDate": "08/09/2020 10:46:05",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I think he may be referring to this discussion here <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028</a>. There are certain images in test which are significantly different from train.</p>",
      "rawMarkdown": "philippsinger I think he may be referring to this discussion here https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028. There are certain images in test which are significantly different from train.",
      "votes": null
    },
    {
      "id": "963844",
      "postDate": "08/09/2020 10:55:55",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> yes, I was mainly referring to the discussion topic that <a href=\"/ks2019\">@ks2019</a> mentioned above.</p>",
      "rawMarkdown": "philippsinger yes, I was mainly referring to the discussion topic that @ks2019 mentioned above.",
      "votes": null
    },
    {
      "id": "964095",
      "postDate": "08/09/2020 15:27:42",
      "content": "<p>Definitely yes.</p>",
      "rawMarkdown": "Definitely yes.",
      "votes": null
    },
    {
      "id": "964134",
      "postDate": "08/09/2020 16:15:23",
      "content": "<p>I totally agree with you. I have chris's stratified CV 0.9614 but LB just got 95.10\nI think this implies discrepancies between the train sets and the test sets </p>",
      "rawMarkdown": "I totally agree with you. I have chris's stratified CV 0.9614 but LB just got 95.10\nI think this implies discrepancies between the train sets and the test sets",
      "votes": null
    },
    {
      "id": "965751",
      "postDate": "08/10/2020 20:17:52",
      "content": "<p>I wonder if SMOTE will help with the class imbalance? </p>",
      "rawMarkdown": "I wonder if SMOTE will help with the class imbalance?",
      "votes": null
    },
    {
      "id": "965853",
      "postDate": "08/10/2020 22:53:34",
      "content": "<p>Have you ever seen a case where it helped?  In the only cases I saw people where doing SMOTE before fold split. That way thee was massive target leak to validation fold.  When SMOTE was used properly, i.e. on each training fold then there was no upside.</p>\n<p>I'd be happy to be prove wrong.</p>",
      "rawMarkdown": "Have you ever seen a case where it helped?  In the only cases I saw people where doing SMOTE before fold split. That way thee was massive target leak to validation fold.  When SMOTE was used properly, i.e. on each training fold then there was no upside.\n\nI'd be happy to be prove wrong.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 964095,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "08/09/2020 15:27:42",
      "content": "<p>Definitely yes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 963724,
      "author_name": "ks2019",
      "author_url": "",
      "post_date": "08/09/2020 08:32:09",
      "content": "<p>There are only 78 ones in public test. So, the role of luck/overfitting increases a lot. So, I expect a huge shakeup!! There is such a class imbalance, I feel even if you have good CV, you are not safe in this scenario. </p>",
      "votes": null,
      "replies": [
        {
          "id": 964134,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "08/09/2020 16:15:23",
          "content": "<p>I totally agree with you. I have chris's stratified CV 0.9614 but LB just got 95.10\nI think this implies discrepancies between the train sets and the test sets </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 963775,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "08/09/2020 09:16:40",
      "content": "<p>All reasons you provide are legit for expecting a strong shakeup. </p>\n\n<p>At the same time, the test data distribution is somewhat different from the train data distribution, so it is very difficult to say how much of public LB overfitting is indeed overfitting and does not relate to private LB. If distribution discrepancies between train/test are much higher than those between public/private test, it can make trusting LB a good idea and reduce the magnitude of shakeup. </p>\n\n<p>Will see soon!</p>",
      "votes": null,
      "replies": [
        {
          "id": 963816,
          "author_name": "ks2019",
          "author_url": "",
          "post_date": "08/09/2020 10:16:16",
          "content": "<p>Even if there is distribution discrepancies between train/test, do you think 78 ones in public LB is enough to make generalisation decisions (for private lb)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 963821,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "08/09/2020 10:23:41",
          "content": "<p>Of course it is not enough. What I am saying is that we should not ignore public LB because chasing CV can lead to overfitting the train set. We know there are some distribution discrepancies. They will definitely play some role, but this effect might be very small in the end compared to the overfitting consequences.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 963824,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/09/2020 10:28:10",
          "content": "<p><a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> what do you mean with distribution discrepancies? target distribution?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 963836,
          "author_name": "ks2019",
          "author_url": "",
          "post_date": "08/09/2020 10:46:05",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I think he may be referring to this discussion here <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028</a>. There are certain images in test which are significantly different from train.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 963844,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "08/09/2020 10:55:55",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> yes, I was mainly referring to the discussion topic that <a href=\"/ks2019\">@ks2019</a> mentioned above.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 963810,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "08/09/2020 10:07:55",
      "content": "<p>In CV we trust...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 965751,
      "author_name": "srikanthpotukuchi",
      "author_url": "",
      "post_date": "08/10/2020 20:17:52",
      "content": "<p>I wonder if SMOTE will help with the class imbalance? </p>",
      "votes": null,
      "replies": [
        {
          "id": 965853,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2020 22:53:34",
          "content": "<p>Have you ever seen a case where it helped?  In the only cases I saw people where doing SMOTE before fold split. That way thee was massive target leak to validation fold.  When SMOTE was used properly, i.e. on each training fold then there was no upside.</p>\n<p>I'd be happy to be prove wrong.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "963701": "Hello! How do you think, should we expect a shake-up in this competition? And if so, why?\n\nI personally expect to shake-up just cause a few following reasons:\n* Relation of the public and private dataset (30% and 70%), such a big relation difference between two datasets usually causes big shake-ups.\n* A lot of overfitted participants at the public LB\n* Many people just submitted CSV-files from the public notebooks, you can see a lot of people with the same score.\n* The Featured Prediction Competition is a cause of shake-up itself. This type of competition is much more susceptible to the shake-up than any Research Code Competition. The best example is M5 competitions with their 1000-2000 places numerous shake-ups.\n* Complexity and ambiguity of the issue. The task to predict a probability is one more factor of a shake-up. It's not just classification with softmax, it's a sigmoid task. So it's pretty possible to get overfitted randomly tuning CSV-files.\n\nUPD:\n* A huge class imbalance in the training dataset.\n\nI wish you all a successful shake-up.",
    "963724": "There are only 78 ones in public test. So, the role of luck/overfitting increases a lot. So, I expect a huge shakeup!! There is such a class imbalance, I feel even if you have good CV, you are not safe in this scenario.",
    "963775": "All reasons you provide are legit for expecting a strong shakeup. \n\nAt the same time, the test data distribution is somewhat different from the train data distribution, so it is very difficult to say how much of public LB overfitting is indeed overfitting and does not relate to private LB. If distribution discrepancies between train/test are much higher than those between public/private test, it can make trusting LB a good idea and reduce the magnitude of shakeup. \n\nWill see soon!",
    "963810": "In CV we trust...",
    "963816": "Even if there is distribution discrepancies between train/test, do you think 78 ones in public LB is enough to make generalisation decisions (for private lb)?",
    "963821": "Of course it is not enough. What I am saying is that we should not ignore public LB because chasing CV can lead to overfitting the train set. We know there are some distribution discrepancies. They will definitely play some role, but this effect might be very small in the end compared to the overfitting consequences.",
    "963824": "kozodoi what do you mean with distribution discrepancies? target distribution?",
    "963836": "philippsinger I think he may be referring to this discussion here https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028. There are certain images in test which are significantly different from train.",
    "963844": "philippsinger yes, I was mainly referring to the discussion topic that @ks2019 mentioned above.",
    "964095": "Definitely yes.",
    "964134": "I totally agree with you. I have chris's stratified CV 0.9614 but LB just got 95.10\nI think this implies discrepancies between the train sets and the test sets",
    "965751": "I wonder if SMOTE will help with the class imbalance?",
    "965853": "Have you ever seen a case where it helped?  In the only cases I saw people where doing SMOTE before fold split. That way thee was massive target leak to validation fold.  When SMOTE was used properly, i.e. on each training fold then there was no upside.\n\nI'd be happy to be prove wrong."
  },
  "source": "meta"
}