{
  "id": 122849,
  "title": "Multi-label stratified Kfold",
  "url": "/competitions/bengaliai-cv19/discussion/122849",
  "author_name": "",
  "post_date": "2019-12-23T08:19:58.098072200Z",
  "votes": 54,
  "comment_count": 14,
  "views": 0,
  "content": "<p>In previous multi-label classification competitions like imet, top participants use <strong>iterative stratification</strong> (<a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a>) to split the dataset (for example: <a href=\"https://www.kaggle.com/c/imet-2019-fgvc6/discussion/94687\">https://www.kaggle.com/c/imet-2019-fgvc6/discussion/94687</a>)</p>\n\n<p>Therefore, here I wrote a simple code to show how to use this method, feel free to take a look at my kernel: <a href=\"https://www.kaggle.com/yiheng/iterative-stratification\">https://www.kaggle.com/yiheng/iterative-stratification</a></p>\n\n<p>It would be very helpful to correct my mistakes, thanks! </p>\n\n<p>Welcome to post the comparisons between it and the original kfold.</p>",
  "messages": [
    {
      "id": "701205",
      "postDate": "12/23/2019 08:19:58",
      "content": "<p>In previous multi-label classification competitions like imet, top participants use <strong>iterative stratification</strong> (<a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a>) to split the dataset (for example: <a href=\"https://www.kaggle.com/c/imet-2019-fgvc6/discussion/94687\">https://www.kaggle.com/c/imet-2019-fgvc6/discussion/94687</a>)</p>\n\n<p>Therefore, here I wrote a simple code to show how to use this method, feel free to take a look at my kernel: <a href=\"https://www.kaggle.com/yiheng/iterative-stratification\">https://www.kaggle.com/yiheng/iterative-stratification</a></p>\n\n<p>It would be very helpful to correct my mistakes, thanks! </p>\n\n<p>Welcome to post the comparisons between it and the original kfold.</p>",
      "rawMarkdown": "In previous multi-label classification competitions like imet, top participants use **iterative stratification** (https://github.com/trent-b/iterative-stratification) to split the dataset (for example: https://www.kaggle.com/c/imet-2019-fgvc6/discussion/94687)\n\nTherefore, here I wrote a simple code to show how to use this method, feel free to take a look at my kernel: https://www.kaggle.com/yiheng/iterative-stratification\n\nIt would be very helpful to correct my mistakes, thanks! \n\nWelcome to post the comparisons between it and the original kfold.",
      "votes": null
    },
    {
      "id": "751209",
      "postDate": "02/20/2020 04:23:06",
      "content": "<p>If you are using my kfold file (the output of my kernel) for training, here are some of my submission results for reference.</p>\n\n<p>Training set: fold != 0\nValidation set: fold == 0\nLocal Scores:\n0.9921,  0.9926, 0.9949, 0.9965, 0.9968\nLB Scores:\n0.9834, 0.9839, 0.9859, 0.9880, 0.9884</p>",
      "rawMarkdown": "If you are using my kfold file (the output of my kernel) for training, here are some of my submission results for reference.\n\nTraining set: fold != 0\nValidation set: fold == 0\nLocal Scores:\n0.9921,  0.9926, 0.9949, 0.9965, 0.9968\nLB Scores:\n0.9834, 0.9839, 0.9859, 0.9880, 0.9884",
      "votes": null
    },
    {
      "id": "751267",
      "postDate": "02/20/2020 04:36:56",
      "content": "<p><a href=\"/yiheng\">@yiheng</a> Hi Venn, May I ask how many epochs did you train?😃 </p>",
      "rawMarkdown": "yiheng Hi Venn, May I ask how many epochs did you train?😃",
      "votes": null
    },
    {
      "id": "751288",
      "postDate": "02/20/2020 04:49:58",
      "content": "<p>~150 epochs : )</p>",
      "rawMarkdown": "~150 epochs : )",
      "votes": null
    },
    {
      "id": "751293",
      "postDate": "02/20/2020 04:54:41",
      "content": "<p>ok thanks a lot, My cv is hard to break 0.980 at 60 epoch, maybe I need try more epoch</p>",
      "rawMarkdown": "ok thanks a lot, My cv is hard to break 0.980 at 60 epoch, maybe I need try more epoch",
      "votes": null
    },
    {
      "id": "751467",
      "postDate": "02/20/2020 08:00:08",
      "content": "<p><a href=\"/yiheng\">@yiheng</a> , have you tried full KFold yet?. I am curious about the variance between folds</p>",
      "rawMarkdown": "yiheng , have you tried full KFold yet?. I am curious about the variance between folds",
      "votes": null
    },
    {
      "id": "751591",
      "postDate": "02/20/2020 10:11:09",
      "content": "<p>In my experience, some folds have better CV(and LB) scores</p>",
      "rawMarkdown": "In my experience, some folds have better CV(and LB) scores",
      "votes": null
    },
    {
      "id": "751818",
      "postDate": "02/20/2020 14:47:46",
      "content": "<p>Nope, cuz training all folds take too much times : )</p>",
      "rawMarkdown": "Nope, cuz training all folds take too much times : )",
      "votes": null
    },
    {
      "id": "751822",
      "postDate": "02/20/2020 14:48:54",
      "content": "<p>Thanks for your notification : ) Let me reach to 0.998+ score first</p>",
      "rawMarkdown": "Thanks for your notification : ) Let me reach to 0.998+ score first",
      "votes": null
    },
    {
      "id": "752314",
      "postDate": "02/20/2020 21:38:43",
      "content": "<p>Is shuffleStratifiedSplit same thing? I'm using it since it can divide train/val to 2 folds with different sizes.</p>",
      "rawMarkdown": "Is shuffleStratifiedSplit same thing? I'm using it since it can divide train/val to 2 folds with different sizes.",
      "votes": null
    },
    {
      "id": "752975",
      "postDate": "02/21/2020 15:33:50",
      "content": "<p><a href=\"/yiheng\">@yiheng</a> , If it's not too much to ask, can you let us know what the scores (Local CV) are at the 25th, 50th, 100th epoch for any one of the folds ? Thanks ! </p>",
      "rawMarkdown": "yiheng , If it's not too much to ask, can you let us know what the scores (Local CV) are at the 25th, 50th, 100th epoch for any one of the folds ? Thanks !",
      "votes": null
    },
    {
      "id": "753058",
      "postDate": "02/21/2020 17:06:48",
      "content": "<p>I have a general question regarding the methodology on how to use the Kfold for cross validation. If I understand correctly, let's say there are 5 fold, you train 5 models, each with one of the fold as validation data and the other 4 as training data. And then you look at the average accuracy on the 5 models.</p>\n\n<p>My question is, for submission, are you using a new model that is trained on all the data? In this final model training, how do you determine which epoch to use? And how do you decide on the learning rate reduction etc? Because when training the model with the full data set, the loss reduction could be different from CV. In this case, would you still choose the same epoch number as you found in the CV that performs the best?</p>\n\n<p>I think my main question is when training the model with the full training data set, what should be changed (or not changed at all) v.s. the CV training pipeline, and how to choose which epoch to use when submitting (choose the one that has the lowest loss in the final training, or choose the epoch number that has the best performance in the CV)?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "I have a general question regarding the methodology on how to use the Kfold for cross validation. If I understand correctly, let's say there are 5 fold, you train 5 models, each with one of the fold as validation data and the other 4 as training data. And then you look at the average accuracy on the 5 models.\n\nMy question is, for submission, are you using a new model that is trained on all the data? In this final model training, how do you determine which epoch to use? And how do you decide on the learning rate reduction etc? Because when training the model with the full data set, the loss reduction could be different from CV. In this case, would you still choose the same epoch number as you found in the CV that performs the best?\n\nI think my main question is when training the model with the full training data set, what should be changed (or not changed at all) v.s. the CV training pipeline, and how to choose which epoch to use when submitting (choose the one that has the lowest loss in the final training, or choose the epoch number that has the best performance in the CV)?\n\nThanks!",
      "votes": null
    },
    {
      "id": "753299",
      "postDate": "02/22/2020 02:00:29",
      "content": "<p>Generally, I will not train a new model on all data. 5 models each trained on 80% data will be used and their predictions can be averaged.</p>",
      "rawMarkdown": "Generally, I will not train a new model on all data. 5 models each trained on 80% data will be used and their predictions can be averaged.",
      "votes": null
    },
    {
      "id": "753300",
      "postDate": "02/22/2020 02:01:39",
      "content": "<p>actually it depends on your lr schedule. </p>",
      "rawMarkdown": "actually it depends on your lr schedule.",
      "votes": null
    },
    {
      "id": "753843",
      "postDate": "02/22/2020 17:57:09",
      "content": "<p>I see. Thanks!</p>",
      "rawMarkdown": "I see. Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 751209,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "02/20/2020 04:23:06",
      "content": "<p>If you are using my kfold file (the output of my kernel) for training, here are some of my submission results for reference.</p>\n\n<p>Training set: fold != 0\nValidation set: fold == 0\nLocal Scores:\n0.9921,  0.9926, 0.9949, 0.9965, 0.9968\nLB Scores:\n0.9834, 0.9839, 0.9859, 0.9880, 0.9884</p>",
      "votes": null,
      "replies": [
        {
          "id": 751267,
          "author_name": "hesene",
          "author_url": "",
          "post_date": "02/20/2020 04:36:56",
          "content": "<p><a href=\"/yiheng\">@yiheng</a> Hi Venn, May I ask how many epochs did you train?😃 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751288,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "02/20/2020 04:49:58",
          "content": "<p>~150 epochs : )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751293,
          "author_name": "hesene",
          "author_url": "",
          "post_date": "02/20/2020 04:54:41",
          "content": "<p>ok thanks a lot, My cv is hard to break 0.980 at 60 epoch, maybe I need try more epoch</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751467,
          "author_name": "backaggle",
          "author_url": "",
          "post_date": "02/20/2020 08:00:08",
          "content": "<p><a href=\"/yiheng\">@yiheng</a> , have you tried full KFold yet?. I am curious about the variance between folds</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751591,
          "author_name": "bibek777",
          "author_url": "",
          "post_date": "02/20/2020 10:11:09",
          "content": "<p>In my experience, some folds have better CV(and LB) scores</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751818,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "02/20/2020 14:47:46",
          "content": "<p>Nope, cuz training all folds take too much times : )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751822,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "02/20/2020 14:48:54",
          "content": "<p>Thanks for your notification : ) Let me reach to 0.998+ score first</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 752975,
          "author_name": "sachinprabhu",
          "author_url": "",
          "post_date": "02/21/2020 15:33:50",
          "content": "<p><a href=\"/yiheng\">@yiheng</a> , If it's not too much to ask, can you let us know what the scores (Local CV) are at the 25th, 50th, 100th epoch for any one of the folds ? Thanks ! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 753300,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "02/22/2020 02:01:39",
          "content": "<p>actually it depends on your lr schedule. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 752314,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "02/20/2020 21:38:43",
      "content": "<p>Is shuffleStratifiedSplit same thing? I'm using it since it can divide train/val to 2 folds with different sizes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 753058,
      "author_name": "axiostpc",
      "author_url": "",
      "post_date": "02/21/2020 17:06:48",
      "content": "<p>I have a general question regarding the methodology on how to use the Kfold for cross validation. If I understand correctly, let's say there are 5 fold, you train 5 models, each with one of the fold as validation data and the other 4 as training data. And then you look at the average accuracy on the 5 models.</p>\n\n<p>My question is, for submission, are you using a new model that is trained on all the data? In this final model training, how do you determine which epoch to use? And how do you decide on the learning rate reduction etc? Because when training the model with the full data set, the loss reduction could be different from CV. In this case, would you still choose the same epoch number as you found in the CV that performs the best?</p>\n\n<p>I think my main question is when training the model with the full training data set, what should be changed (or not changed at all) v.s. the CV training pipeline, and how to choose which epoch to use when submitting (choose the one that has the lowest loss in the final training, or choose the epoch number that has the best performance in the CV)?</p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 753299,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "02/22/2020 02:00:29",
          "content": "<p>Generally, I will not train a new model on all data. 5 models each trained on 80% data will be used and their predictions can be averaged.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 753843,
          "author_name": "axiostpc",
          "author_url": "",
          "post_date": "02/22/2020 17:57:09",
          "content": "<p>I see. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "701205": "In previous multi-label classification competitions like imet, top participants use **iterative stratification** (https://github.com/trent-b/iterative-stratification) to split the dataset (for example: https://www.kaggle.com/c/imet-2019-fgvc6/discussion/94687)\n\nTherefore, here I wrote a simple code to show how to use this method, feel free to take a look at my kernel: https://www.kaggle.com/yiheng/iterative-stratification\n\nIt would be very helpful to correct my mistakes, thanks! \n\nWelcome to post the comparisons between it and the original kfold.",
    "751209": "If you are using my kfold file (the output of my kernel) for training, here are some of my submission results for reference.\n\nTraining set: fold != 0\nValidation set: fold == 0\nLocal Scores:\n0.9921,  0.9926, 0.9949, 0.9965, 0.9968\nLB Scores:\n0.9834, 0.9839, 0.9859, 0.9880, 0.9884",
    "751267": "yiheng Hi Venn, May I ask how many epochs did you train?😃",
    "751288": "~150 epochs : )",
    "751293": "ok thanks a lot, My cv is hard to break 0.980 at 60 epoch, maybe I need try more epoch",
    "751467": "yiheng , have you tried full KFold yet?. I am curious about the variance between folds",
    "751591": "In my experience, some folds have better CV(and LB) scores",
    "751818": "Nope, cuz training all folds take too much times : )",
    "751822": "Thanks for your notification : ) Let me reach to 0.998+ score first",
    "752314": "Is shuffleStratifiedSplit same thing? I'm using it since it can divide train/val to 2 folds with different sizes.",
    "752975": "yiheng , If it's not too much to ask, can you let us know what the scores (Local CV) are at the 25th, 50th, 100th epoch for any one of the folds ? Thanks !",
    "753058": "I have a general question regarding the methodology on how to use the Kfold for cross validation. If I understand correctly, let's say there are 5 fold, you train 5 models, each with one of the fold as validation data and the other 4 as training data. And then you look at the average accuracy on the 5 models.\n\nMy question is, for submission, are you using a new model that is trained on all the data? In this final model training, how do you determine which epoch to use? And how do you decide on the learning rate reduction etc? Because when training the model with the full data set, the loss reduction could be different from CV. In this case, would you still choose the same epoch number as you found in the CV that performs the best?\n\nI think my main question is when training the model with the full training data set, what should be changed (or not changed at all) v.s. the CV training pipeline, and how to choose which epoch to use when submitting (choose the one that has the lowest loss in the final training, or choose the epoch number that has the best performance in the CV)?\n\nThanks!",
    "753299": "Generally, I will not train a new model on all data. 5 models each trained on 80% data will be used and their predictions can be averaged.",
    "753300": "actually it depends on your lr schedule.",
    "753843": "I see. Thanks!"
  },
  "source": "meta"
}