{
  "id": 155345,
  "title": "Using Stratified GroupKFold for CV",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/155345",
  "author_name": "",
  "post_date": "2020-06-01T10:29:55.576993200Z",
  "votes": 8,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Since the data is severely imbalanced, we should implement Stratified GroupKFold, but its not implemented in the existing libraries; if one is interested, you can refer to the <a href=\"https://www.kaggle.com/jakubwasikowski/stratified-group-k-fold-cross-validation\">link</a> here for this implementation.</p>\n\n<p>The reason I think GroupKFold might be better than normal KFold is as follows: Quoting from scikit-learn's website, it says:</p>\n\n<p>The Independent and identically distributed (i.i.d.) assumption is broken if the underlying generative process yield groups of dependent samples.</p>\n\n<p>Such a grouping of data is domain specific. An example would be when there is medical data collected from multiple patients, with multiple samples taken from each patient. And such data is likely to be dependent on the individual group. In our example, the patient id for each sample will be its group identifier.</p>\n\n<p>In this case we would like to know if a model trained on a particular set of groups generalizes well to the unseen groups. To measure this, we need to ensure that all the samples in the validation fold come from groups that are not represented at all in the paired training fold.</p>\n\n<p>There are 2056 unique <code>patient_ids</code> out of a whooping 33126 rows. This means that many patients have multiple images. And as I mentioned earlier, when we do KFold splitting, let's say there are 20 images for patient 1, we may have patient 1's data/images in both the training and validation set. For example, in the splitting process, there are 15 images of patient 1 in the training set, and there are 5 images in the validation set; then this may not be ideal since the model has already seen 15 of the images for patient 1 and can easily remember features that are unique to patient 1, and therefore predict well in the validation set for the same patient 1. Therefore, the normal K-Fold cross validation method may give over optimistic results and fail to generalize well to more unseen images.</p>\n\n<p>So I made a trial <a href=\"https://www.kaggle.com/reighns/groupkfold-efficientbnet-trial\">notebook</a> to test out the idea of using GroupKFold. I strongly recommend using Stratified GroupKFold as our cv environment since the data is quite imbalanced. I will update the notebook if I manage to implement a stratified version. What do you guys think?</p>",
  "messages": [
    {
      "id": "869865",
      "postDate": "06/01/2020 10:29:55",
      "content": "<p>Since the data is severely imbalanced, we should implement Stratified GroupKFold, but its not implemented in the existing libraries; if one is interested, you can refer to the <a href=\"https://www.kaggle.com/jakubwasikowski/stratified-group-k-fold-cross-validation\">link</a> here for this implementation.</p>\n\n<p>The reason I think GroupKFold might be better than normal KFold is as follows: Quoting from scikit-learn's website, it says:</p>\n\n<p>The Independent and identically distributed (i.i.d.) assumption is broken if the underlying generative process yield groups of dependent samples.</p>\n\n<p>Such a grouping of data is domain specific. An example would be when there is medical data collected from multiple patients, with multiple samples taken from each patient. And such data is likely to be dependent on the individual group. In our example, the patient id for each sample will be its group identifier.</p>\n\n<p>In this case we would like to know if a model trained on a particular set of groups generalizes well to the unseen groups. To measure this, we need to ensure that all the samples in the validation fold come from groups that are not represented at all in the paired training fold.</p>\n\n<p>There are 2056 unique <code>patient_ids</code> out of a whooping 33126 rows. This means that many patients have multiple images. And as I mentioned earlier, when we do KFold splitting, let's say there are 20 images for patient 1, we may have patient 1's data/images in both the training and validation set. For example, in the splitting process, there are 15 images of patient 1 in the training set, and there are 5 images in the validation set; then this may not be ideal since the model has already seen 15 of the images for patient 1 and can easily remember features that are unique to patient 1, and therefore predict well in the validation set for the same patient 1. Therefore, the normal K-Fold cross validation method may give over optimistic results and fail to generalize well to more unseen images.</p>\n\n<p>So I made a trial <a href=\"https://www.kaggle.com/reighns/groupkfold-efficientbnet-trial\">notebook</a> to test out the idea of using GroupKFold. I strongly recommend using Stratified GroupKFold as our cv environment since the data is quite imbalanced. I will update the notebook if I manage to implement a stratified version. What do you guys think?</p>",
      "rawMarkdown": "Since the data is severely imbalanced, we should implement Stratified GroupKFold, but its not implemented in the existing libraries; if one is interested, you can refer to the [link](https://www.kaggle.com/jakubwasikowski/stratified-group-k-fold-cross-validation) here for this implementation.\n\nThe reason I think GroupKFold might be better than normal KFold is as follows: Quoting from scikit-learn's website, it says:\n\nThe Independent and identically distributed (i.i.d.) assumption is broken if the underlying generative process yield groups of dependent samples.\n\nSuch a grouping of data is domain specific. An example would be when there is medical data collected from multiple patients, with multiple samples taken from each patient. And such data is likely to be dependent on the individual group. In our example, the patient id for each sample will be its group identifier.\n\nIn this case we would like to know if a model trained on a particular set of groups generalizes well to the unseen groups. To measure this, we need to ensure that all the samples in the validation fold come from groups that are not represented at all in the paired training fold.\n\nThere are 2056 unique `patient_ids` out of a whooping 33126 rows. This means that many patients have multiple images. And as I mentioned earlier, when we do KFold splitting, let's say there are 20 images for patient 1, we may have patient 1's data/images in both the training and validation set. For example, in the splitting process, there are 15 images of patient 1 in the training set, and there are 5 images in the validation set; then this may not be ideal since the model has already seen 15 of the images for patient 1 and can easily remember features that are unique to patient 1, and therefore predict well in the validation set for the same patient 1. Therefore, the normal K-Fold cross validation method may give over optimistic results and fail to generalize well to more unseen images.\n\nSo I made a trial [notebook](https://www.kaggle.com/reighns/groupkfold-efficientbnet-trial) to test out the idea of using GroupKFold. I strongly recommend using Stratified GroupKFold as our cv environment since the data is quite imbalanced. I will update the notebook if I manage to implement a stratified version. What do you guys think?",
      "votes": null
    },
    {
      "id": "869881",
      "postDate": "06/01/2020 10:49:43",
      "content": "<p>Hey,thats very useful!!\nThanks for sharing👍 </p>",
      "rawMarkdown": "Hey,thats very useful!!\nThanks for sharing👍",
      "votes": null
    },
    {
      "id": "869936",
      "postDate": "06/01/2020 11:41:13",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "869956",
      "postDate": "06/01/2020 11:57:07",
      "content": "<p>I tried it yesterday (Patient_ids grouped and stratified K-fold) , but got worst results (compared to simple Stratified K-fold)</p>\n\n<p>Tell us about your experiments. </p>",
      "rawMarkdown": "I tried it yesterday (Patient_ids grouped and stratified K-fold) , but got worst results (compared to simple Stratified K-fold)\n\nTell us about your experiments.",
      "votes": null
    },
    {
      "id": "870004",
      "postDate": "06/01/2020 12:26:49",
      "content": "<p>My first trial got me 0.888 on public LB. Ideally, stratified GroupKFold might make more sense, however I have yet to find a good way to code this out. I shared a link on the implementation of Stratified GroupKFold, but I have not implemented it yet haha. <a href=\"/serigne\">@serigne</a> </p>",
      "rawMarkdown": "My first trial got me 0.888 on public LB. Ideally, stratified GroupKFold might make more sense, however I have yet to find a good way to code this out. I shared a link on the implementation of Stratified GroupKFold, but I have not implemented it yet haha. @serigne",
      "votes": null
    },
    {
      "id": "918643",
      "postDate": "07/07/2020 11:31:09",
      "content": "<p><a href=\"/reighns\">@reighns</a>  I am using GroupKFold and observed that it also considering the percentage of target in each fold along with specified group as well. Am I getting it by chance or GroupKFold ensures stratification as well?</p>\n\n<p>Here are the results of 5-folds:\n<code>\n4    6626\n3    6625\n2    6625\n1    6625\n0    6625\n</code></p>\n\n<p>and in each fold, no. of positive and negative samples are almost same but have a little variation in each fold,</p>\n\n<p><code>\n0    6511\n1     114\n</code></p>\n\n<p>It seems like GroupKFold ensuring stratification by target as well. Is it doing by chance or what you say about it ?</p>",
      "rawMarkdown": "reighns  I am using GroupKFold and observed that it also considering the percentage of target in each fold along with specified group as well. Am I getting it by chance or GroupKFold ensures stratification as well?\n\nHere are the results of 5-folds:\n```\n4    6626\n3    6625\n2    6625\n1    6625\n0    6625\n```\n\nand in each fold, no. of positive and negative samples are almost same but have a little variation in each fold,\n\n```\n0    6511\n1     114\n```\n\nIt seems like GroupKFold ensuring stratification by target as well. Is it doing by chance or what you say about it ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 869881,
      "author_name": "proxyy",
      "author_url": "",
      "post_date": "06/01/2020 10:49:43",
      "content": "<p>Hey,thats very useful!!\nThanks for sharing👍 </p>",
      "votes": null,
      "replies": [
        {
          "id": 869936,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "06/01/2020 11:41:13",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 869956,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "06/01/2020 11:57:07",
      "content": "<p>I tried it yesterday (Patient_ids grouped and stratified K-fold) , but got worst results (compared to simple Stratified K-fold)</p>\n\n<p>Tell us about your experiments. </p>",
      "votes": null,
      "replies": [
        {
          "id": 870004,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "06/01/2020 12:26:49",
          "content": "<p>My first trial got me 0.888 on public LB. Ideally, stratified GroupKFold might make more sense, however I have yet to find a good way to code this out. I shared a link on the implementation of Stratified GroupKFold, but I have not implemented it yet haha. <a href=\"/serigne\">@serigne</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 918643,
      "author_name": "abdurrehman245",
      "author_url": "",
      "post_date": "07/07/2020 11:31:09",
      "content": "<p><a href=\"/reighns\">@reighns</a>  I am using GroupKFold and observed that it also considering the percentage of target in each fold along with specified group as well. Am I getting it by chance or GroupKFold ensures stratification as well?</p>\n\n<p>Here are the results of 5-folds:\n<code>\n4    6626\n3    6625\n2    6625\n1    6625\n0    6625\n</code></p>\n\n<p>and in each fold, no. of positive and negative samples are almost same but have a little variation in each fold,</p>\n\n<p><code>\n0    6511\n1     114\n</code></p>\n\n<p>It seems like GroupKFold ensuring stratification by target as well. Is it doing by chance or what you say about it ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "869865": "Since the data is severely imbalanced, we should implement Stratified GroupKFold, but its not implemented in the existing libraries; if one is interested, you can refer to the [link](https://www.kaggle.com/jakubwasikowski/stratified-group-k-fold-cross-validation) here for this implementation.\n\nThe reason I think GroupKFold might be better than normal KFold is as follows: Quoting from scikit-learn's website, it says:\n\nThe Independent and identically distributed (i.i.d.) assumption is broken if the underlying generative process yield groups of dependent samples.\n\nSuch a grouping of data is domain specific. An example would be when there is medical data collected from multiple patients, with multiple samples taken from each patient. And such data is likely to be dependent on the individual group. In our example, the patient id for each sample will be its group identifier.\n\nIn this case we would like to know if a model trained on a particular set of groups generalizes well to the unseen groups. To measure this, we need to ensure that all the samples in the validation fold come from groups that are not represented at all in the paired training fold.\n\nThere are 2056 unique `patient_ids` out of a whooping 33126 rows. This means that many patients have multiple images. And as I mentioned earlier, when we do KFold splitting, let's say there are 20 images for patient 1, we may have patient 1's data/images in both the training and validation set. For example, in the splitting process, there are 15 images of patient 1 in the training set, and there are 5 images in the validation set; then this may not be ideal since the model has already seen 15 of the images for patient 1 and can easily remember features that are unique to patient 1, and therefore predict well in the validation set for the same patient 1. Therefore, the normal K-Fold cross validation method may give over optimistic results and fail to generalize well to more unseen images.\n\nSo I made a trial [notebook](https://www.kaggle.com/reighns/groupkfold-efficientbnet-trial) to test out the idea of using GroupKFold. I strongly recommend using Stratified GroupKFold as our cv environment since the data is quite imbalanced. I will update the notebook if I manage to implement a stratified version. What do you guys think?",
    "869881": "Hey,thats very useful!!\nThanks for sharing👍",
    "869936": "Thanks!",
    "869956": "I tried it yesterday (Patient_ids grouped and stratified K-fold) , but got worst results (compared to simple Stratified K-fold)\n\nTell us about your experiments.",
    "870004": "My first trial got me 0.888 on public LB. Ideally, stratified GroupKFold might make more sense, however I have yet to find a good way to code this out. I shared a link on the implementation of Stratified GroupKFold, but I have not implemented it yet haha. @serigne",
    "918643": "reighns  I am using GroupKFold and observed that it also considering the percentage of target in each fold along with specified group as well. Am I getting it by chance or GroupKFold ensures stratification as well?\n\nHere are the results of 5-folds:\n```\n4    6626\n3    6625\n2    6625\n1    6625\n0    6625\n```\n\nand in each fold, no. of positive and negative samples are almost same but have a little variation in each fold,\n\n```\n0    6511\n1     114\n```\n\nIt seems like GroupKFold ensuring stratification by target as well. Is it doing by chance or what you say about it ?"
  },
  "source": "meta"
}