{
  "id": 198198,
  "title": "Is K fold Cross Validation helps for this imbalance problem ? ",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/198198",
  "author_name": "",
  "post_date": "2020-11-20T08:01:00.685072500Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Cross Validation Strategy for this problem , Kindly suggest other cross validation strategy better suits for this problem </p>",
  "messages": [
    {
      "id": "1084619",
      "postDate": "11/20/2020 08:01:00",
      "content": "<p>Cross Validation Strategy for this problem , Kindly suggest other cross validation strategy better suits for this problem </p>",
      "rawMarkdown": "Cross Validation Strategy for this problem , Kindly suggest other cross validation strategy better suits for this problem",
      "votes": null
    },
    {
      "id": "1084633",
      "postDate": "11/20/2020 08:18:36",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/praveengovi\" target=\"_blank\">@praveengovi</a>,</p>\n<p>In general one could look at using <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html\" target=\"_blank\">StratifiedKFold</a>, where the folds are made by preserving the percentage of samples for each class, or one could take a look at the <a href=\"https://github.com/scikit-learn-contrib/imbalanced-learn\" target=\"_blank\">imbalanced-learn</a> package, which implements various under and over re-sampling techniques. However, that said I am not familiar with this competition, and these techniques <em>may</em> not be the most appropriate ones.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @praveengovi,\n\nIn general one could look at using [StratifiedKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html), where the folds are made by preserving the percentage of samples for each class, or one could take a look at the [imbalanced-learn](https://github.com/scikit-learn-contrib/imbalanced-learn) package, which implements various under and over re-sampling techniques. However, that said I am not familiar with this competition, and these techniques *may* not be the most appropriate ones.\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1137700",
      "postDate": "01/04/2021 06:59:39",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> and <a href=\"https://www.kaggle.com/praveengovi\" target=\"_blank\">@praveengovi</a> for bringing this topic up. I looked at other discussions regarding the same,  but I still have doubts how to settle on cv strategy. Using StratifiedKFold would create a validation set which is same distribution as training set. But ideally, we would want the validation set to be same distribution of the test set (which is always unknown). So, how to figure out if we have robust validation strategy?</p>",
      "rawMarkdown": "Thanks @carlmcbrideellis and @praveengovi for bringing this topic up. I looked at other discussions regarding the same,  but I still have doubts how to settle on cv strategy. Using StratifiedKFold would create a validation set which is same distribution as training set. But ideally, we would want the validation set to be same distribution of the test set (which is always unknown). So, how to figure out if we have robust validation strategy?",
      "votes": null
    },
    {
      "id": "1137727",
      "postDate": "01/04/2021 07:17:17",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/suryajrrafl\" target=\"_blank\">@suryajrrafl</a>,</p>\n<p>That is a very good question; one would hope that the hidden test data has the same distribution as ones training data. The way to check this is once your local CV is up and running, then make a submission to the competition and see if the leaderboard score is the 'same' as your local CV score. However, the trick is not to abuse the leaderboard score by making too many tests, as this leads to overfitting!</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @suryajrrafl,\n\nThat is a very good question; one would hope that the hidden test data has the same distribution as ones training data. The way to check this is once your local CV is up and running, then make a submission to the competition and see if the leaderboard score is the 'same' as your local CV score. However, the trick is not to abuse the leaderboard score by making too many tests, as this leads to overfitting!\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1138433",
      "postDate": "01/04/2021 17:26:47",
      "content": "<p>There are several discussion posts where folks have described the test distribution based on LB probing that they have performed.  In the real world you might not be able to determine the distribution - on Kaggle you can beat the system by probing the LB.  </p>\n<p>The only issue with probing is that you only see your score for a part of the test set.  For example, if you create a model that always submits class as 0 than you get a score of ,048 - indicating that 4.8% of the test set REPORTED to you is class 0.  </p>",
      "rawMarkdown": "There are several discussion posts where folks have described the test distribution based on LB probing that they have performed.  In the real world you might not be able to determine the distribution - on Kaggle you can beat the system by probing the LB.  \n\nThe only issue with probing is that you only see your score for a part of the test set.  For example, if you create a model that always submits class as 0 than you get a score of ,048 - indicating that 4.8% of the test set REPORTED to you is class 0.",
      "votes": null
    },
    {
      "id": "1138851",
      "postDate": "01/05/2021 02:39:39",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> for the headstart. I'll try to check the posts you mentioned. </p>",
      "rawMarkdown": "Thanks @pcjimmmy for the headstart. I'll try to check the posts you mentioned.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1084633,
      "author_name": "carlmcbrideellis",
      "author_url": "",
      "post_date": "11/20/2020 08:18:36",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/praveengovi\" target=\"_blank\">@praveengovi</a>,</p>\n<p>In general one could look at using <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html\" target=\"_blank\">StratifiedKFold</a>, where the folds are made by preserving the percentage of samples for each class, or one could take a look at the <a href=\"https://github.com/scikit-learn-contrib/imbalanced-learn\" target=\"_blank\">imbalanced-learn</a> package, which implements various under and over re-sampling techniques. However, that said I am not familiar with this competition, and these techniques <em>may</em> not be the most appropriate ones.</p>\n<p>All the best,<br>\ncarl</p>",
      "votes": null,
      "replies": [
        {
          "id": 1137700,
          "author_name": "suryajrrafl",
          "author_url": "",
          "post_date": "01/04/2021 06:59:39",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> and <a href=\"https://www.kaggle.com/praveengovi\" target=\"_blank\">@praveengovi</a> for bringing this topic up. I looked at other discussions regarding the same,  but I still have doubts how to settle on cv strategy. Using StratifiedKFold would create a validation set which is same distribution as training set. But ideally, we would want the validation set to be same distribution of the test set (which is always unknown). So, how to figure out if we have robust validation strategy?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137727,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "01/04/2021 07:17:17",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/suryajrrafl\" target=\"_blank\">@suryajrrafl</a>,</p>\n<p>That is a very good question; one would hope that the hidden test data has the same distribution as ones training data. The way to check this is once your local CV is up and running, then make a submission to the competition and see if the leaderboard score is the 'same' as your local CV score. However, the trick is not to abuse the leaderboard score by making too many tests, as this leads to overfitting!</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1138433,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/04/2021 17:26:47",
          "content": "<p>There are several discussion posts where folks have described the test distribution based on LB probing that they have performed.  In the real world you might not be able to determine the distribution - on Kaggle you can beat the system by probing the LB.  </p>\n<p>The only issue with probing is that you only see your score for a part of the test set.  For example, if you create a model that always submits class as 0 than you get a score of ,048 - indicating that 4.8% of the test set REPORTED to you is class 0.  </p>",
          "votes": null,
          "replies": [
            {
              "id": 1138851,
              "author_name": "suryajrrafl",
              "author_url": "",
              "post_date": "01/05/2021 02:39:39",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> for the headstart. I'll try to check the posts you mentioned. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1084619": "Cross Validation Strategy for this problem , Kindly suggest other cross validation strategy better suits for this problem",
    "1084633": "Dear @praveengovi,\n\nIn general one could look at using [StratifiedKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html), where the folds are made by preserving the percentage of samples for each class, or one could take a look at the [imbalanced-learn](https://github.com/scikit-learn-contrib/imbalanced-learn) package, which implements various under and over re-sampling techniques. However, that said I am not familiar with this competition, and these techniques *may* not be the most appropriate ones.\n\nAll the best,\ncarl",
    "1137700": "Thanks @carlmcbrideellis and @praveengovi for bringing this topic up. I looked at other discussions regarding the same,  but I still have doubts how to settle on cv strategy. Using StratifiedKFold would create a validation set which is same distribution as training set. But ideally, we would want the validation set to be same distribution of the test set (which is always unknown). So, how to figure out if we have robust validation strategy?",
    "1137727": "Dear @suryajrrafl,\n\nThat is a very good question; one would hope that the hidden test data has the same distribution as ones training data. The way to check this is once your local CV is up and running, then make a submission to the competition and see if the leaderboard score is the 'same' as your local CV score. However, the trick is not to abuse the leaderboard score by making too many tests, as this leads to overfitting!\n\nAll the best,\ncarl",
    "1138433": "There are several discussion posts where folks have described the test distribution based on LB probing that they have performed.  In the real world you might not be able to determine the distribution - on Kaggle you can beat the system by probing the LB.  \n\nThe only issue with probing is that you only see your score for a part of the test set.  For example, if you create a model that always submits class as 0 than you get a score of ,048 - indicating that 4.8% of the test set REPORTED to you is class 0.",
    "1138851": "Thanks @pcjimmmy for the headstart. I'll try to check the posts you mentioned."
  },
  "source": "meta"
}