{
  "id": 156002,
  "title": "Implementation of Stratified Group K-Folds",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/156002",
  "author_name": "Alexey Pronin",
  "post_date": "2020-06-04T02:29:19.299000",
  "votes": 16,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I would like to continue the topic started in <a href=\"https://www.kaggle.com/reighns/groupkfold-efficientbnet-trial\">this kernel</a> by <a href=\"https://www.kaggle.com/reighns\">reigHns</a>. </p>\n\n<p>One of the challenges that we are facing in this competition is finding the best way to implement a K-fold cross-validation. On one hand, we have a very imbalanced dataset, so some sort of stratification seems to be in order here. On the other hand, we know that some of the patients present in the dataset are represented by multiple images and it is probably a good idea not to mix images corresponding to the same <code>patient_id</code> in our training and validation sets. The latter can be achieved by using sklearn Group K-Fold split. Unfortunately, sklearn does not have an implementation of Stratified Group K-Folds, so we have to make one from scratch. Fortunately, one such <a href=\"https://www.kaggle.com/jakubwasikowski/stratified-group-k-fold-cross-validation\">implementation of Stratified Group K-Folds</a> was already developed in the PetFinder competition, so we can just take it and use it here. And it seems to be working pretty well as can be seen from my public kernel:</p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/siim-stratified-groupkfold\">SIIM Stratified GroupKFold 5-folds</a> </p>\n\n<p>-- the 5-fold split has almost the same percentages of 0's and 1's in each fold and the unique <code>patient_id</code> values do not overlap between different folds. I can think of two possible ways how these folds can be utilized: we can use them to do cross-validation on jpeg files or we can make a new set of tfrecord files based on these folds. </p>",
  "messages": [
    {
      "id": 873306,
      "postDate": "2020-06-04T02:29:19.300Z",
      "content": "<p>I would like to continue the topic started in <a href=\"https://www.kaggle.com/reighns/groupkfold-efficientbnet-trial\">this kernel</a> by <a href=\"https://www.kaggle.com/reighns\">reigHns</a>. </p>\n\n<p>One of the challenges that we are facing in this competition is finding the best way to implement a K-fold cross-validation. On one hand, we have a very imbalanced dataset, so some sort of stratification seems to be in order here. On the other hand, we know that some of the patients present in the dataset are represented by multiple images and it is probably a good idea not to mix images corresponding to the same <code>patient_id</code> in our training and validation sets. The latter can be achieved by using sklearn Group K-Fold split. Unfortunately, sklearn does not have an implementation of Stratified Group K-Folds, so we have to make one from scratch. Fortunately, one such <a href=\"https://www.kaggle.com/jakubwasikowski/stratified-group-k-fold-cross-validation\">implementation of Stratified Group K-Folds</a> was already developed in the PetFinder competition, so we can just take it and use it here. And it seems to be working pretty well as can be seen from my public kernel:</p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/siim-stratified-groupkfold\">SIIM Stratified GroupKFold 5-folds</a> </p>\n\n<p>-- the 5-fold split has almost the same percentages of 0's and 1's in each fold and the unique <code>patient_id</code> values do not overlap between different folds. I can think of two possible ways how these folds can be utilized: we can use them to do cross-validation on jpeg files or we can make a new set of tfrecord files based on these folds. </p>",
      "rawMarkdown": "I would like to continue the topic started in [this kernel](https://www.kaggle.com/reighns/groupkfold-efficientbnet-trial) by [reigHns](https://www.kaggle.com/reighns). \n\nOne of the challenges that we are facing in this competition is finding the best way to implement a K-fold cross-validation. On one hand, we have a very imbalanced dataset, so some sort of stratification seems to be in order here. On the other hand, we know that some of the patients present in the dataset are represented by multiple images and it is probably a good idea not to mix images corresponding to the same `patient_id` in our training and validation sets. The latter can be achieved by using sklearn Group K-Fold split. Unfortunately, sklearn does not have an implementation of Stratified Group K-Folds, so we have to make one from scratch. Fortunately, one such [implementation of Stratified Group K-Folds](https://www.kaggle.com/jakubwasikowski/stratified-group-k-fold-cross-validation) was already developed in the PetFinder competition, so we can just take it and use it here. And it seems to be working pretty well as can be seen from my public kernel:\n\n[SIIM Stratified GroupKFold 5-folds](https://www.kaggle.com/graf10a/siim-stratified-groupkfold) \n\n-- the 5-fold split has almost the same percentages of 0's and 1's in each fold and the unique `patient_id` values do not overlap between different folds. I can think of two possible ways how these folds can be utilized: we can use them to do cross-validation on jpeg files or we can make a new set of tfrecord files based on these folds. ",
      "votes": 14
    },
    {
      "id": 926847,
      "postDate": "2020-07-13T03:32:39.577Z",
      "content": "<p>Thanks <a href=\"/graf10a\">@graf10a</a> for your CV strategy. Now I'm using it as my baseline.\nMine CV AUC each folder is: 0.913/0.937/0.884/0.883/0.906\nOverall OOF AUC is 0.892; LB 0.929 (with external, 256x256, Eff-b0)\nDo you also observe the AUC difference in each folder?</p>",
      "rawMarkdown": "Thanks @graf10a for your CV strategy. Now I'm using it as my baseline.\nMine CV AUC each folder is: 0.913/0.937/0.884/0.883/0.906\nOverall OOF AUC is 0.892; LB 0.929 (with external, 256x256, Eff-b0)\nDo you also observe the AUC difference in each folder?\n",
      "votes": 1,
      "replies": [
        {
          "id": 926854,
          "postDate": "2020-07-13T03:42:57.530Z",
          "content": "<p><a href=\"/waylongo\">@waylongo</a> Yes, I do see a great deal of variability across the folds. Makes me very nervous about the private LB 😲 .</p>",
          "rawMarkdown": "@waylongo Yes, I do see a great deal of variability across the folds. Makes me very nervous about the private LB 😲 .",
          "votes": 1
        }
      ]
    },
    {
      "id": 892301,
      "postDate": "2020-06-18T19:29:50.857Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 892433,
          "postDate": "2020-06-18T21:43:28.580Z",
          "content": "<p>Thank you! It seems to be working reasonably well.</p>",
          "rawMarkdown": "Thank you! It seems to be working reasonably well.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 926847,
      "author_name": "Waylon Wu",
      "author_url": "",
      "post_date": "2020-07-13T03:32:39.577000",
      "content": "<p>Thanks <a href=\"/graf10a\">@graf10a</a> for your CV strategy. Now I'm using it as my baseline.\nMine CV AUC each folder is: 0.913/0.937/0.884/0.883/0.906\nOverall OOF AUC is 0.892; LB 0.929 (with external, 256x256, Eff-b0)\nDo you also observe the AUC difference in each folder?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 926854,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-07-13T03:42:57.530000",
          "content": "<p><a href=\"/waylongo\">@waylongo</a> Yes, I do see a great deal of variability across the folds. Makes me very nervous about the private LB 😲 .</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 892301,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-18T19:29:50.857000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 892433,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-06-18T21:43:28.580000",
          "content": "<p>Thank you! It seems to be working reasonably well.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "873306": "I would like to continue the topic started in [this kernel](https://www.kaggle.com/reighns/groupkfold-efficientbnet-trial) by [reigHns](https://www.kaggle.com/reighns). \n\nOne of the challenges that we are facing in this competition is finding the best way to implement a K-fold cross-validation. On one hand, we have a very imbalanced dataset, so some sort of stratification seems to be in order here. On the other hand, we know that some of the patients present in the dataset are represented by multiple images and it is probably a good idea not to mix images corresponding to the same `patient_id` in our training and validation sets. The latter can be achieved by using sklearn Group K-Fold split. Unfortunately, sklearn does not have an implementation of Stratified Group K-Folds, so we have to make one from scratch. Fortunately, one such [implementation of Stratified Group K-Folds](https://www.kaggle.com/jakubwasikowski/stratified-group-k-fold-cross-validation) was already developed in the PetFinder competition, so we can just take it and use it here. And it seems to be working pretty well as can be seen from my public kernel:\n\n[SIIM Stratified GroupKFold 5-folds](https://www.kaggle.com/graf10a/siim-stratified-groupkfold) \n\n-- the 5-fold split has almost the same percentages of 0's and 1's in each fold and the unique `patient_id` values do not overlap between different folds. I can think of two possible ways how these folds can be utilized: we can use them to do cross-validation on jpeg files or we can make a new set of tfrecord files based on these folds. ",
    "926847": "Thanks @graf10a for your CV strategy. Now I'm using it as my baseline.\nMine CV AUC each folder is: 0.913/0.937/0.884/0.883/0.906\nOverall OOF AUC is 0.892; LB 0.929 (with external, 256x256, Eff-b0)\nDo you also observe the AUC difference in each folder?\n",
    "892301": ""
  }
}