{
  "id": 157499,
  "title": "How to set up a good CV strategy ?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/157499",
  "author_name": "",
  "post_date": "2020-06-10T22:07:30.139420100Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hey everyone this is my second competition on kaggle and I see a lot of people talking about the importance of having a good cv strategy (not only on this competition). So my understanding is to use K-Fold CV or Stratified K-Fold CV (like this competition because data is imbalanced and we want to have the same proportions) and we eventually calculate the total loss of the K-Folds. As for predictions (and submitting to LB) we predict on each time we train and test on one fold then average results.  If I'm wrong or forgetting something correct me and If I'm right isn't this a slow way to experiment and iterate ? Any tips on the matter would be extremely appreciated thanks ! </p>",
  "messages": [
    {
      "id": "881284",
      "postDate": "06/10/2020 22:07:30",
      "content": "<p>Hey everyone this is my second competition on kaggle and I see a lot of people talking about the importance of having a good cv strategy (not only on this competition). So my understanding is to use K-Fold CV or Stratified K-Fold CV (like this competition because data is imbalanced and we want to have the same proportions) and we eventually calculate the total loss of the K-Folds. As for predictions (and submitting to LB) we predict on each time we train and test on one fold then average results.  If I'm wrong or forgetting something correct me and If I'm right isn't this a slow way to experiment and iterate ? Any tips on the matter would be extremely appreciated thanks ! </p>",
      "rawMarkdown": "Hey everyone this is my second competition on kaggle and I see a lot of people talking about the importance of having a good cv strategy (not only on this competition). So my understanding is to use K-Fold CV or Stratified K-Fold CV (like this competition because data is imbalanced and we want to have the same proportions) and we eventually calculate the total loss of the K-Folds. As for predictions (and submitting to LB) we predict on each time we train and test on one fold then average results.  If I'm wrong or forgetting something correct me and If I'm right isn't this a slow way to experiment and iterate ? Any tips on the matter would be extremely appreciated thanks !",
      "votes": null
    },
    {
      "id": "881629",
      "postDate": "06/11/2020 08:06:49",
      "content": "<p>From my limited experience training models on the competition data, regular stratified/group KFold is not enough. </p>\n\n<p>The reason is even if you use as little as 3 folds, each of them would contain only a couple of hundreds of positive examples, which is clearly too few to consider the score on this subset to be reliable. But what one could do instead?</p>\n\n<p>Some ideas off the top of my head:</p>\n\n<ul>\n<li>When comparing different models, compare not only the mean score across folds but also variance. A rule of thumb I often use is to consider a new model better than the old one if its mean score is at least 1 standard deviation far from the mean score of the old model. To estimate the variance more properly, you might want to use more fold or use multiple KFold splits (repeated KFold). A more rigorous way of comparing models would be to use statistical tests. You might benefit from reading <a href=\"https://medium.com/@vktech/practitioners-guide-to-statistical-tests-ed2d580ef04f\">this article</a>. One of the authors <a href=\"/daniel89\">@daniel89</a> is actually a gold medal winner of two past Kaggle competitions.</li>\n<li>Increase the validation test size. Use external data for training and use more of the competition data for validation. This would result in a slightly worse performance of your models, but the scores should be much less shaky.</li>\n</ul>",
      "rawMarkdown": "From my limited experience training models on the competition data, regular stratified/group KFold is not enough. \n\nThe reason is even if you use as little as 3 folds, each of them would contain only a couple of hundreds of positive examples, which is clearly too few to consider the score on this subset to be reliable. But what one could do instead?\n\nSome ideas off the top of my head:\n\n* When comparing different models, compare not only the mean score across folds but also variance. A rule of thumb I often use is to consider a new model better than the old one if its mean score is at least 1 standard deviation far from the mean score of the old model. To estimate the variance more properly, you might want to use more fold or use multiple KFold splits (repeated KFold). A more rigorous way of comparing models would be to use statistical tests. You might benefit from reading [this article](https://medium.com/@vktech/practitioners-guide-to-statistical-tests-ed2d580ef04f). One of the authors @daniel89 is actually a gold medal winner of two past Kaggle competitions.\n* Increase the validation test size. Use external data for training and use more of the competition data for validation. This would result in a slightly worse performance of your models, but the scores should be much less shaky.",
      "votes": null
    },
    {
      "id": "881997",
      "postDate": "06/11/2020 14:19:52",
      "content": "<p>-</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "883490",
      "postDate": "06/12/2020 17:20:47",
      "content": "<ul>\n<li>Use GroupKFold, data has some duplicates, that share unique patient_id.</li>\n<li>Save some data for test and do not tune hyper parameters on it.</li>\n<li>Try oversampling methods to increase number of positive cases while training. In last year ISIC Pneumothorax best solution was to start with 80% positive cases and then slowly decrease distribution of samples to original one.</li>\n</ul>",
      "rawMarkdown": "Use GroupKFold, data has some duplicates, that share unique patient_id.\n- Save some data for test and do not tune hyper parameters on it.\n- Try oversampling methods to increase number of positive cases while training. In last year ISIC Pneumothorax best solution was to start with 80% positive cases and then slowly decrease distribution of samples to original one.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 881629,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "06/11/2020 08:06:49",
      "content": "<p>From my limited experience training models on the competition data, regular stratified/group KFold is not enough. </p>\n\n<p>The reason is even if you use as little as 3 folds, each of them would contain only a couple of hundreds of positive examples, which is clearly too few to consider the score on this subset to be reliable. But what one could do instead?</p>\n\n<p>Some ideas off the top of my head:</p>\n\n<ul>\n<li>When comparing different models, compare not only the mean score across folds but also variance. A rule of thumb I often use is to consider a new model better than the old one if its mean score is at least 1 standard deviation far from the mean score of the old model. To estimate the variance more properly, you might want to use more fold or use multiple KFold splits (repeated KFold). A more rigorous way of comparing models would be to use statistical tests. You might benefit from reading <a href=\"https://medium.com/@vktech/practitioners-guide-to-statistical-tests-ed2d580ef04f\">this article</a>. One of the authors <a href=\"/daniel89\">@daniel89</a> is actually a gold medal winner of two past Kaggle competitions.</li>\n<li>Increase the validation test size. Use external data for training and use more of the competition data for validation. This would result in a slightly worse performance of your models, but the scores should be much less shaky.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 881997,
      "author_name": "luuuth",
      "author_url": "",
      "post_date": "06/11/2020 14:19:52",
      "content": "<p>-</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 883490,
      "author_name": "zakajd",
      "author_url": "",
      "post_date": "06/12/2020 17:20:47",
      "content": "<ul>\n<li>Use GroupKFold, data has some duplicates, that share unique patient_id.</li>\n<li>Save some data for test and do not tune hyper parameters on it.</li>\n<li>Try oversampling methods to increase number of positive cases while training. In last year ISIC Pneumothorax best solution was to start with 80% positive cases and then slowly decrease distribution of samples to original one.</li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "881284": "Hey everyone this is my second competition on kaggle and I see a lot of people talking about the importance of having a good cv strategy (not only on this competition). So my understanding is to use K-Fold CV or Stratified K-Fold CV (like this competition because data is imbalanced and we want to have the same proportions) and we eventually calculate the total loss of the K-Folds. As for predictions (and submitting to LB) we predict on each time we train and test on one fold then average results.  If I'm wrong or forgetting something correct me and If I'm right isn't this a slow way to experiment and iterate ? Any tips on the matter would be extremely appreciated thanks !",
    "881629": "From my limited experience training models on the competition data, regular stratified/group KFold is not enough. \n\nThe reason is even if you use as little as 3 folds, each of them would contain only a couple of hundreds of positive examples, which is clearly too few to consider the score on this subset to be reliable. But what one could do instead?\n\nSome ideas off the top of my head:\n\n* When comparing different models, compare not only the mean score across folds but also variance. A rule of thumb I often use is to consider a new model better than the old one if its mean score is at least 1 standard deviation far from the mean score of the old model. To estimate the variance more properly, you might want to use more fold or use multiple KFold splits (repeated KFold). A more rigorous way of comparing models would be to use statistical tests. You might benefit from reading [this article](https://medium.com/@vktech/practitioners-guide-to-statistical-tests-ed2d580ef04f). One of the authors @daniel89 is actually a gold medal winner of two past Kaggle competitions.\n* Increase the validation test size. Use external data for training and use more of the competition data for validation. This would result in a slightly worse performance of your models, but the scores should be much less shaky.",
    "881997": "",
    "883490": "Use GroupKFold, data has some duplicates, that share unique patient_id.\n- Save some data for test and do not tune hyper parameters on it.\n- Try oversampling methods to increase number of positive cases while training. In last year ISIC Pneumothorax best solution was to start with 80% positive cases and then slowly decrease distribution of samples to original one."
  },
  "source": "meta"
}