{
  "id": 306521,
  "title": "Stratified KFold v. Group KFold (aka. I'm a dummy)",
  "url": "/competitions/happy-whale-and-dolphin/discussion/306521",
  "author_name": "Darien Schettler",
  "post_date": "2022-02-09T18:39:32.255000",
  "votes": 58,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Hi all, I'm very used to seeing success shared (and sharing my own), however, failures and mistakes are just as common and sometimes just as important.</p>\n<p>I made a really dumb mistake early on in this competition and it yielded days of frustration. I wanted to share this mistake with you all so you know that not everyone is crushing it all the time. (I dunno, I might max out at crushing it 1% of the time). This is a pretty clear example of moving too fast and not taking the time to think through what you want to do and why. </p>\n<p>I'm not sure if anyone needs this, or wants to read a post like this. Maybe it just makes me look like a fool. However, I know I can often get overwhelmed by the absolutely astonishing brilliance of many people on Kaggle, LinkedIn, etc., and sometimes it feels like anything other than perfection amounts to failure.</p>\n<p>I just wanted to share something other people might be able to resonate with to remind them that the path to success is not a straight line… it's rebounding from failure after failure and never giving up.</p>\n<hr>\n<p><strong>My mistake:</strong></p>\n<ul>\n<li>I used <strong><code>GroupKFold</code></strong> instead of <strong><code>StratifiedKFold</code></strong> to create my train/val splits for training.</li>\n<li>I used <strong><code>species</code></strong> as the label in <strong><code>GroupKFold</code></strong> and <strong><code>individual_id</code></strong> as the group identifier.</li>\n<li>When using <strong><code>StratifiedKFold</code></strong>, you should be using <strong><code>individual_id</code></strong> as the label and there is no <strong><code>group</code></strong> argument to be provided.</li>\n</ul>\n<p><strong>Why is this wrong?</strong></p>\n<ul>\n<li><strong><code>GroupKFold</code></strong> will generate train/val datasets that keeps all entities within a given <strong><em>group</em></strong> together within the dataset. This means that, because I used <strong><code>individual_id</code></strong> to define groups, all the images/examples for a given individual would be contained within either the train or val dataset (but would never be found in both).</li>\n<li><strong><code>GroupKFold</code></strong> will then ensure that a similar distribution for the given label is found in both train and val respectively. i.e. if 5% of the train dataset is made up of <strong><code>pilot_whale</code></strong> than 5% of the val dataset will also be made up of <strong><code>pilot_whale</code></strong>.</li>\n<li><strong><code>StratifiedKFold</code></strong> on the other hand will simply generate train/val datasets that ensure a similar distribution for the given label. As mentioned previously, if we provide the label of <strong><code>individual_id</code></strong>, this will <strong><em>MIX BETWEEN THE DATASETS</em></strong> images/examples for a particular individual. It will attempt to make sure that for a given individual the same percentage of examples in the respective datasets will yield that individual. In reality this is hard to achieve as some individuals have very few images (if # of images is less than # of folds, obviously you can't split properly).</li>\n</ul>\n<p><strong>Why did I make this mistake?</strong></p>\n<ul>\n<li>For some reason I thought we had to keep the examples for a given individual within the respective dataset.</li>\n<li>This makes NO SENSE, when you consider the approach of training using a clustering algorithm or something similar (ArcFace)</li>\n<li>You need images of the same individual in both the train and the validation dataset, or else you can't possibly see if your model has learned how to distinguish between unique individuals.</li>\n</ul>\n<hr>\n<p>The simple analogy to explain what I did:</p>\n<p><strong>Imagine you are trying to classify cat v. dog</strong></p>\n<ul>\n<li>I essentially put all the cat photos into the train dataset and all of the dog photos into the val dataset (i.e. species used to group)</li>\n</ul>\n<hr>\n<p>😅😅😅😅So yeah! 😅😅😅😅</p>\n<hr>\n<p>At least I figured it out and got it all fixed! Cheers!</p>",
  "messages": [
    {
      "id": 1683344,
      "postDate": "2022-02-09T18:39:32.257Z",
      "content": "<p>Hi all, I'm very used to seeing success shared (and sharing my own), however, failures and mistakes are just as common and sometimes just as important.</p>\n<p>I made a really dumb mistake early on in this competition and it yielded days of frustration. I wanted to share this mistake with you all so you know that not everyone is crushing it all the time. (I dunno, I might max out at crushing it 1% of the time). This is a pretty clear example of moving too fast and not taking the time to think through what you want to do and why. </p>\n<p>I'm not sure if anyone needs this, or wants to read a post like this. Maybe it just makes me look like a fool. However, I know I can often get overwhelmed by the absolutely astonishing brilliance of many people on Kaggle, LinkedIn, etc., and sometimes it feels like anything other than perfection amounts to failure.</p>\n<p>I just wanted to share something other people might be able to resonate with to remind them that the path to success is not a straight line… it's rebounding from failure after failure and never giving up.</p>\n<hr>\n<p><strong>My mistake:</strong></p>\n<ul>\n<li>I used <strong><code>GroupKFold</code></strong> instead of <strong><code>StratifiedKFold</code></strong> to create my train/val splits for training.</li>\n<li>I used <strong><code>species</code></strong> as the label in <strong><code>GroupKFold</code></strong> and <strong><code>individual_id</code></strong> as the group identifier.</li>\n<li>When using <strong><code>StratifiedKFold</code></strong>, you should be using <strong><code>individual_id</code></strong> as the label and there is no <strong><code>group</code></strong> argument to be provided.</li>\n</ul>\n<p><strong>Why is this wrong?</strong></p>\n<ul>\n<li><strong><code>GroupKFold</code></strong> will generate train/val datasets that keeps all entities within a given <strong><em>group</em></strong> together within the dataset. This means that, because I used <strong><code>individual_id</code></strong> to define groups, all the images/examples for a given individual would be contained within either the train or val dataset (but would never be found in both).</li>\n<li><strong><code>GroupKFold</code></strong> will then ensure that a similar distribution for the given label is found in both train and val respectively. i.e. if 5% of the train dataset is made up of <strong><code>pilot_whale</code></strong> than 5% of the val dataset will also be made up of <strong><code>pilot_whale</code></strong>.</li>\n<li><strong><code>StratifiedKFold</code></strong> on the other hand will simply generate train/val datasets that ensure a similar distribution for the given label. As mentioned previously, if we provide the label of <strong><code>individual_id</code></strong>, this will <strong><em>MIX BETWEEN THE DATASETS</em></strong> images/examples for a particular individual. It will attempt to make sure that for a given individual the same percentage of examples in the respective datasets will yield that individual. In reality this is hard to achieve as some individuals have very few images (if # of images is less than # of folds, obviously you can't split properly).</li>\n</ul>\n<p><strong>Why did I make this mistake?</strong></p>\n<ul>\n<li>For some reason I thought we had to keep the examples for a given individual within the respective dataset.</li>\n<li>This makes NO SENSE, when you consider the approach of training using a clustering algorithm or something similar (ArcFace)</li>\n<li>You need images of the same individual in both the train and the validation dataset, or else you can't possibly see if your model has learned how to distinguish between unique individuals.</li>\n</ul>\n<hr>\n<p>The simple analogy to explain what I did:</p>\n<p><strong>Imagine you are trying to classify cat v. dog</strong></p>\n<ul>\n<li>I essentially put all the cat photos into the train dataset and all of the dog photos into the val dataset (i.e. species used to group)</li>\n</ul>\n<hr>\n<p>😅😅😅😅So yeah! 😅😅😅😅</p>\n<hr>\n<p>At least I figured it out and got it all fixed! Cheers!</p>",
      "rawMarkdown": "Hi all, I'm very used to seeing success shared (and sharing my own), however, failures and mistakes are just as common and sometimes just as important.\n\nI made a really dumb mistake early on in this competition and it yielded days of frustration. I wanted to share this mistake with you all so you know that not everyone is crushing it all the time. (I dunno, I might max out at crushing it 1% of the time). This is a pretty clear example of moving too fast and not taking the time to think through what you want to do and why. \n\nI'm not sure if anyone needs this, or wants to read a post like this. Maybe it just makes me look like a fool. However, I know I can often get overwhelmed by the absolutely astonishing brilliance of many people on Kaggle, LinkedIn, etc., and sometimes it feels like anything other than perfection amounts to failure.\n\nI just wanted to share something other people might be able to resonate with to remind them that the path to success is not a straight line... it's rebounding from failure after failure and never giving up.\n\n---\n\n**My mistake:**\n* I used **`GroupKFold`** instead of **`StratifiedKFold`** to create my train/val splits for training.\n* I used **`species`** as the label in **`GroupKFold`** and **`individual_id`** as the group identifier.\n* When using **`StratifiedKFold`**, you should be using **`individual_id`** as the label and there is no **`group`** argument to be provided.\n\n**Why is this wrong?**\n* **`GroupKFold`** will generate train/val datasets that keeps all entities within a given ***group*** together within the dataset. This means that, because I used **`individual_id`** to define groups, all the images/examples for a given individual would be contained within either the train or val dataset (but would never be found in both).\n* **`GroupKFold`** will then ensure that a similar distribution for the given label is found in both train and val respectively. i.e. if 5% of the train dataset is made up of **`pilot_whale`** than 5% of the val dataset will also be made up of **`pilot_whale`**.\n* **`StratifiedKFold`** on the other hand will simply generate train/val datasets that ensure a similar distribution for the given label. As mentioned previously, if we provide the label of **`individual_id`**, this will ***MIX BETWEEN THE DATASETS*** images/examples for a particular individual. It will attempt to make sure that for a given individual the same percentage of examples in the respective datasets will yield that individual. In reality this is hard to achieve as some individuals have very few images (if # of images is less than # of folds, obviously you can't split properly).\n\n**Why did I make this mistake?**\n* For some reason I thought we had to keep the examples for a given individual within the respective dataset.\n* This makes NO SENSE, when you consider the approach of training using a clustering algorithm or something similar (ArcFace)\n* You need images of the same individual in both the train and the validation dataset, or else you can't possibly see if your model has learned how to distinguish between unique individuals.\n\n---\n\nThe simple analogy to explain what I did:\n\n**Imagine you are trying to classify cat v. dog**\n* I essentially put all the cat photos into the train dataset and all of the dog photos into the val dataset (i.e. species used to group)\n\n---\n\n😅😅😅😅So yeah! 😅😅😅😅\n\n---\n\nAt least I figured it out and got it all fixed! Cheers!",
      "votes": 58
    },
    {
      "id": 1683380,
      "postDate": "2022-02-09T19:09:50.090Z",
      "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> everyone makes mistakes, but you overcame this one and thus learned something new! I'm sure you won't forget it.</p>\n<p>I have always used <code>StratifiedKFold</code>, never felt the need of using <code>GroupKFold</code>. Any examples come to mind so as to where <code>GroupKFold</code> would allow for better folds? Would love to know 😄</p>",
      "rawMarkdown": "@dschettler8845 everyone makes mistakes, but you overcame this one and thus learned something new! I'm sure you won't forget it.\n\nI have always used `StratifiedKFold`, never felt the need of using `GroupKFold`. Any examples come to mind so as to where `GroupKFold` would allow for better folds? Would love to know 😄",
      "votes": 1,
      "replies": [
        {
          "id": 1683418,
          "postDate": "2022-02-09T19:36:11.950Z",
          "content": "<p>I remember I used groupKfold in the shopee comp, generally use this if you think you will encounter a totally different class in the test data.</p>",
          "rawMarkdown": "I remember I used groupKfold in the shopee comp, generally use this if you think you will encounter a totally different class in the test data.",
          "votes": 1
        },
        {
          "id": 1683435,
          "postDate": "2022-02-09T19:53:55.080Z",
          "content": "<p>I see, haven't really had any experiences with this type of problem. Will keep in mind for the future, thanks</p>",
          "rawMarkdown": "I see, haven't really had any experiences with this type of problem. Will keep in mind for the future, thanks",
          "votes": 1
        },
        {
          "id": 1683462,
          "postDate": "2022-02-09T20:14:40.013Z",
          "content": "<p>I think normally you use GroupKFold to mitigate potential leakage.</p>\n<p>I'm not super confident on this… but I think in the COTS competition you might want to keep images from short video clips in the same split. This would prevent the model from being able to infer subsequent frames that have been leaked into the validation split.</p>\n<p>i.e. You would pass to GroupKFold an identifier (x), the class (y) and the video clip (group). Then the resulting dataset splits will have approximately even distributions of the relevant classes… but all the images for a given video clip will be kept together.  </p>",
          "rawMarkdown": "I think normally you use GroupKFold to mitigate potential leakage.\n\nI'm not super confident on this... but I think in the COTS competition you might want to keep images from short video clips in the same split. This would prevent the model from being able to infer subsequent frames that have been leaked into the validation split.\n\ni.e. You would pass to GroupKFold an identifier (x), the class (y) and the video clip (group). Then the resulting dataset splits will have approximately even distributions of the relevant classes... but all the images for a given video clip will be kept together.  ",
          "votes": 1
        },
        {
          "id": 1683464,
          "postDate": "2022-02-09T20:15:04.697Z",
          "content": "<p>Also, thanks for the support <a href=\"https://www.kaggle.com/asarvazyan\" target=\"_blank\">@asarvazyan</a> !</p>",
          "rawMarkdown": "Also, thanks for the support @asarvazyan !",
          "votes": 1
        },
        {
          "id": 1683476,
          "postDate": "2022-02-09T20:23:18.597Z",
          "content": "<p>Makes sense, thanks for the example!</p>",
          "rawMarkdown": "Makes sense, thanks for the example!"
        }
      ]
    },
    {
      "id": 1683615,
      "postDate": "2022-02-09T23:12:23.607Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> for sharing.<br>\nOne quick question that is bothering me: how do you actually set up a cross validation in a Computer Vision (or NLP) comp, where generally we have large datasets or models need a lot of time for training?</p>\n<p>Let me explain better. I know how CV works and I use it with no problems at all in tabular tasks. I generally split my dataset in 5 stratified (or group) folds, train, predict and save my OOF and predictions. This is generally done for later ensembling.<br>\nIn NLP or Com. Vision cross validation this entire process seems \"demanding\" to me. Is it ok to fine tune a model 5 times? Is it a standard procedure?</p>",
      "rawMarkdown": "Thanks @dschettler8845 for sharing.\nOne quick question that is bothering me: how do you actually set up a cross validation in a Computer Vision (or NLP) comp, where generally we have large datasets or models need a lot of time for training?\n\nLet me explain better. I know how CV works and I use it with no problems at all in tabular tasks. I generally split my dataset in 5 stratified (or group) folds, train, predict and save my OOF and predictions. This is generally done for later ensembling.\nIn NLP or Com. Vision cross validation this entire process seems \"demanding\" to me. Is it ok to fine tune a model 5 times? Is it a standard procedure?",
      "votes": 2,
      "replies": [
        {
          "id": 1683824,
          "postDate": "2022-02-10T04:27:07.200Z",
          "content": "<p>Yes, it is especially demanding in this competition as any fold wont have all the \"individual_ids\". Do you mean to ask that is it okay to train a model 5 times on 5 different folds? Yes. It is a must.</p>",
          "rawMarkdown": "Yes, it is especially demanding in this competition as any fold wont have all the \"individual_ids\". Do you mean to ask that is it okay to train a model 5 times on 5 different folds? Yes. It is a must.",
          "votes": 2
        },
        {
          "id": 1684083,
          "postDate": "2022-02-10T08:45:26.897Z",
          "content": "<p>Cross-validation is very popular on Kaggle but it's not a necessity. You don't see cross-validation on academic papers or benchmarks very often. Since those benchmarks are quite large datasets like this one, they mostly use train/val/test splits. A 70% (train), 10% (val), 20% (test) split is very popular on those benchmarks. You can use something similar if you have resource limitations. </p>",
          "rawMarkdown": "Cross-validation is very popular on Kaggle but it's not a necessity. You don't see cross-validation on academic papers or benchmarks very often. Since those benchmarks are quite large datasets like this one, they mostly use train/val/test splits. A 70% (train), 10% (val), 20% (test) split is very popular on those benchmarks. You can use something similar if you have resource limitations. ",
          "votes": 3
        },
        {
          "id": 1684089,
          "postDate": "2022-02-10T08:57:57.833Z",
          "content": "<p>Thanks for your answers <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> .</p>\n<p>The question arises precisely from the fact that in academic papers one does not often (perhaps almost never) see a CV on large datasets but here on Kaggle it is almost a standard practice.</p>\n<p>However, do you agree that the \"standard\" procedure is:</p>\n<ul>\n<li>k-fold creation</li>\n<li>training of k models</li>\n<li>prediction on OOF</li>\n<li>saving submission and OOF prediction</li>\n<li>possible ensembling</li>\n</ul>",
          "rawMarkdown": "Thanks for your answers @harshitsheoran @gunesevitan .\n\nThe question arises precisely from the fact that in academic papers one does not often (perhaps almost never) see a CV on large datasets but here on Kaggle it is almost a standard practice.\n\nHowever, do you agree that the \"standard\" procedure is:\n- k-fold creation\n- training of k models\n- prediction on OOF\n- saving submission and OOF prediction\n- possible ensembling",
          "votes": 3
        },
        {
          "id": 1698693,
          "postDate": "2022-02-20T15:29:36.250Z",
          "content": "<p>To add to the discussion, I have read about a good validation strategy depending on the dataset size <a href=\"https://sebastianraschka.com/pdf/slides/2021-04_czi.pdf\" target=\"_blank\">here</a>: </p>\n<p><a href=\"https://ibb.co/fkwkTQH\"><img src=\"https://i.ibb.co/YcscVbh/validation-strategy.png\" alt=\"validation-strategy\"></a></p>\n<p>The dataset in this competition is of \"medium\" size I would say, so cross-validation will be useful. If you have limited resources, I would suggest dropping it for now and once you have a good single model, move to cross-validation.</p>\n<p>I hope this helps.</p>",
          "rawMarkdown": "To add to the discussion, I have read about a good validation strategy depending on the dataset size [here](https://sebastianraschka.com/pdf/slides/2021-04_czi.pdf): \n\n<a href=\"https://ibb.co/fkwkTQH\"><img src=\"https://i.ibb.co/YcscVbh/validation-strategy.png\" alt=\"validation-strategy\" border=\"0\"></a>\n\n\nThe dataset in this competition is of \"medium\" size I would say, so cross-validation will be useful. If you have limited resources, I would suggest dropping it for now and once you have a good single model, move to cross-validation.\n\nI hope this helps.\n\n",
          "votes": 1
        },
        {
          "id": 1704969,
          "postDate": "2022-02-26T02:57:08.130Z",
          "content": "<p>Can I ask a question?<br>\nSo why rather than re-training the model with the whole dataset after making sure that CV returns a good result for that specific hyperparameters, we stick with model than are trained with less data? Are the ensemble of 5 models (with less training data each) more performant?<br>\nThank you!</p>",
          "rawMarkdown": "Can I ask a question?\nSo why rather than re-training the model with the whole dataset after making sure that CV returns a good result for that specific hyperparameters, we stick with model than are trained with less data? Are the ensemble of 5 models (with less training data each) more performant?\nThank you!"
        }
      ]
    },
    {
      "id": 1683416,
      "postDate": "2022-02-09T19:35:15.313Z",
      "content": "<p>Thanks for sharing, I think we should all share our failures, especially in the field of ML where bugs are very hard to find.</p>",
      "rawMarkdown": "Thanks for sharing, I think we should all share our failures, especially in the field of ML where bugs are very hard to find.",
      "votes": 2,
      "replies": [
        {
          "id": 1683457,
          "postDate": "2022-02-09T20:10:48.517Z",
          "content": "<p>Absolutely! </p>",
          "rawMarkdown": "Absolutely! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1684075,
      "postDate": "2022-02-10T08:37:48.123Z",
      "content": "<p>It's very counter-intuitive to use non-overlapping group splits here. I wonder how did you come up with that approach. :D</p>",
      "rawMarkdown": "It's very counter-intuitive to use non-overlapping group splits here. I wonder how did you come up with that approach. :D",
      "replies": [
        {
          "id": 1684549,
          "postDate": "2022-02-10T15:23:12.860Z",
          "content": "<p>I think I didn't fully grasp what GroupKFold was doing… I thought that it would ensure an even distribution of individual IDs maybe? But I didn't even really take the time to understand what that could mean. I won't make that mistake again.</p>\n<p>Totally a dummy maneuver!</p>",
          "rawMarkdown": "I think I didn't fully grasp what GroupKFold was doing... I thought that it would ensure an even distribution of individual IDs maybe? But I didn't even really take the time to understand what that could mean. I won't make that mistake again.\n\nTotally a dummy maneuver!",
          "votes": 2
        }
      ]
    },
    {
      "id": 1698488,
      "postDate": "2022-02-20T12:17:10.423Z",
      "content": "<p>Thanks for sharing this, it should always feel safe to share failures, even though I won't consider this a real failure since you haven't put this into production and the competition isn't over yet. 👌</p>",
      "rawMarkdown": "Thanks for sharing this, it should always feel safe to share failures, even though I won't consider this a real failure since you haven't put this into production and the competition isn't over yet. 👌"
    },
    {
      "id": 1684978,
      "postDate": "2022-02-10T22:53:19.580Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>, thanks for this. I am trying to get my head around the different cv schemes at the moment. I found <a href=\"https://scikit-learn.org/stable/auto_examples/model_selection/plot_cv_indices.html\" target=\"_blank\">this page</a> pretty helpful.</p>",
      "rawMarkdown": "Hi @dschettler8845, thanks for this. I am trying to get my head around the different cv schemes at the moment. I found [this page](https://scikit-learn.org/stable/auto_examples/model_selection/plot_cv_indices.html) pretty helpful."
    },
    {
      "id": 1684958,
      "postDate": "2022-02-10T22:18:51.477Z",
      "content": "<p>Well, that is an experience. Thanks for sharing <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> . In my opinion, knowing your mistake is as important as what you are doing right!.</p>",
      "rawMarkdown": "Well, that is an experience. Thanks for sharing @dschettler8845 . In my opinion, knowing your mistake is as important as what you are doing right!."
    },
    {
      "id": 1698030,
      "postDate": "2022-02-20T04:35:16.190Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1683380,
      "author_name": "asarvazyan",
      "author_url": "",
      "post_date": "2022-02-09T19:09:50.090000",
      "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> everyone makes mistakes, but you overcame this one and thus learned something new! I'm sure you won't forget it.</p>\n<p>I have always used <code>StratifiedKFold</code>, never felt the need of using <code>GroupKFold</code>. Any examples come to mind so as to where <code>GroupKFold</code> would allow for better folds? Would love to know 😄</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1683418,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2022-02-09T19:36:11.950000",
          "content": "<p>I remember I used groupKfold in the shopee comp, generally use this if you think you will encounter a totally different class in the test data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1683435,
          "author_name": "asarvazyan",
          "author_url": "",
          "post_date": "2022-02-09T19:53:55.080000",
          "content": "<p>I see, haven't really had any experiences with this type of problem. Will keep in mind for the future, thanks</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1683462,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2022-02-09T20:14:40.013000",
          "content": "<p>I think normally you use GroupKFold to mitigate potential leakage.</p>\n<p>I'm not super confident on this… but I think in the COTS competition you might want to keep images from short video clips in the same split. This would prevent the model from being able to infer subsequent frames that have been leaked into the validation split.</p>\n<p>i.e. You would pass to GroupKFold an identifier (x), the class (y) and the video clip (group). Then the resulting dataset splits will have approximately even distributions of the relevant classes… but all the images for a given video clip will be kept together.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1683464,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2022-02-09T20:15:04.697000",
          "content": "<p>Also, thanks for the support <a href=\"https://www.kaggle.com/asarvazyan\" target=\"_blank\">@asarvazyan</a> !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1683476,
          "author_name": "asarvazyan",
          "author_url": "",
          "post_date": "2022-02-09T20:23:18.597000",
          "content": "<p>Makes sense, thanks for the example!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1683615,
      "author_name": "Jacopo Repossi",
      "author_url": "",
      "post_date": "2022-02-09T23:12:23.607000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> for sharing.<br>\nOne quick question that is bothering me: how do you actually set up a cross validation in a Computer Vision (or NLP) comp, where generally we have large datasets or models need a lot of time for training?</p>\n<p>Let me explain better. I know how CV works and I use it with no problems at all in tabular tasks. I generally split my dataset in 5 stratified (or group) folds, train, predict and save my OOF and predictions. This is generally done for later ensembling.<br>\nIn NLP or Com. Vision cross validation this entire process seems \"demanding\" to me. Is it ok to fine tune a model 5 times? Is it a standard procedure?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1683824,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-02-10T04:27:07.200000",
          "content": "<p>Yes, it is especially demanding in this competition as any fold wont have all the \"individual_ids\". Do you mean to ask that is it okay to train a model 5 times on 5 different folds? Yes. It is a must.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1684083,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-02-10T08:45:26.897000",
          "content": "<p>Cross-validation is very popular on Kaggle but it's not a necessity. You don't see cross-validation on academic papers or benchmarks very often. Since those benchmarks are quite large datasets like this one, they mostly use train/val/test splits. A 70% (train), 10% (val), 20% (test) split is very popular on those benchmarks. You can use something similar if you have resource limitations. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1684089,
          "author_name": "Jacopo Repossi",
          "author_url": "",
          "post_date": "2022-02-10T08:57:57.833000",
          "content": "<p>Thanks for your answers <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> .</p>\n<p>The question arises precisely from the fact that in academic papers one does not often (perhaps almost never) see a CV on large datasets but here on Kaggle it is almost a standard practice.</p>\n<p>However, do you agree that the \"standard\" procedure is:</p>\n<ul>\n<li>k-fold creation</li>\n<li>training of k models</li>\n<li>prediction on OOF</li>\n<li>saving submission and OOF prediction</li>\n<li>possible ensembling</li>\n</ul>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1698693,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2022-02-20T15:29:36.250000",
          "content": "<p>To add to the discussion, I have read about a good validation strategy depending on the dataset size <a href=\"https://sebastianraschka.com/pdf/slides/2021-04_czi.pdf\" target=\"_blank\">here</a>: </p>\n<p><a href=\"https://ibb.co/fkwkTQH\"><img src=\"https://i.ibb.co/YcscVbh/validation-strategy.png\" alt=\"validation-strategy\"></a></p>\n<p>The dataset in this competition is of \"medium\" size I would say, so cross-validation will be useful. If you have limited resources, I would suggest dropping it for now and once you have a good single model, move to cross-validation.</p>\n<p>I hope this helps.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1704969,
          "author_name": "termanteus",
          "author_url": "",
          "post_date": "2022-02-26T02:57:08.130000",
          "content": "<p>Can I ask a question?<br>\nSo why rather than re-training the model with the whole dataset after making sure that CV returns a good result for that specific hyperparameters, we stick with model than are trained with less data? Are the ensemble of 5 models (with less training data each) more performant?<br>\nThank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1683416,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2022-02-09T19:35:15.313000",
      "content": "<p>Thanks for sharing, I think we should all share our failures, especially in the field of ML where bugs are very hard to find.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1683457,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2022-02-09T20:10:48.517000",
          "content": "<p>Absolutely! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1684075,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-02-10T08:37:48.123000",
      "content": "<p>It's very counter-intuitive to use non-overlapping group splits here. I wonder how did you come up with that approach. :D</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1684549,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2022-02-10T15:23:12.860000",
          "content": "<p>I think I didn't fully grasp what GroupKFold was doing… I thought that it would ensure an even distribution of individual IDs maybe? But I didn't even really take the time to understand what that could mean. I won't make that mistake again.</p>\n<p>Totally a dummy maneuver!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1698488,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2022-02-20T12:17:10.423000",
      "content": "<p>Thanks for sharing this, it should always feel safe to share failures, even though I won't consider this a real failure since you haven't put this into production and the competition isn't over yet. 👌</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1684978,
      "author_name": "Nick Potter",
      "author_url": "",
      "post_date": "2022-02-10T22:53:19.580000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>, thanks for this. I am trying to get my head around the different cv schemes at the moment. I found <a href=\"https://scikit-learn.org/stable/auto_examples/model_selection/plot_cv_indices.html\" target=\"_blank\">this page</a> pretty helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1684958,
      "author_name": "N. Peker Çelik",
      "author_url": "",
      "post_date": "2022-02-10T22:18:51.477000",
      "content": "<p>Well, that is an experience. Thanks for sharing <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> . In my opinion, knowing your mistake is as important as what you are doing right!.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1698030,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-20T04:35:16.190000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1683344": "Hi all, I'm very used to seeing success shared (and sharing my own), however, failures and mistakes are just as common and sometimes just as important.\n\nI made a really dumb mistake early on in this competition and it yielded days of frustration. I wanted to share this mistake with you all so you know that not everyone is crushing it all the time. (I dunno, I might max out at crushing it 1% of the time). This is a pretty clear example of moving too fast and not taking the time to think through what you want to do and why. \n\nI'm not sure if anyone needs this, or wants to read a post like this. Maybe it just makes me look like a fool. However, I know I can often get overwhelmed by the absolutely astonishing brilliance of many people on Kaggle, LinkedIn, etc., and sometimes it feels like anything other than perfection amounts to failure.\n\nI just wanted to share something other people might be able to resonate with to remind them that the path to success is not a straight line... it's rebounding from failure after failure and never giving up.\n\n---\n\n**My mistake:**\n* I used **`GroupKFold`** instead of **`StratifiedKFold`** to create my train/val splits for training.\n* I used **`species`** as the label in **`GroupKFold`** and **`individual_id`** as the group identifier.\n* When using **`StratifiedKFold`**, you should be using **`individual_id`** as the label and there is no **`group`** argument to be provided.\n\n**Why is this wrong?**\n* **`GroupKFold`** will generate train/val datasets that keeps all entities within a given ***group*** together within the dataset. This means that, because I used **`individual_id`** to define groups, all the images/examples for a given individual would be contained within either the train or val dataset (but would never be found in both).\n* **`GroupKFold`** will then ensure that a similar distribution for the given label is found in both train and val respectively. i.e. if 5% of the train dataset is made up of **`pilot_whale`** than 5% of the val dataset will also be made up of **`pilot_whale`**.\n* **`StratifiedKFold`** on the other hand will simply generate train/val datasets that ensure a similar distribution for the given label. As mentioned previously, if we provide the label of **`individual_id`**, this will ***MIX BETWEEN THE DATASETS*** images/examples for a particular individual. It will attempt to make sure that for a given individual the same percentage of examples in the respective datasets will yield that individual. In reality this is hard to achieve as some individuals have very few images (if # of images is less than # of folds, obviously you can't split properly).\n\n**Why did I make this mistake?**\n* For some reason I thought we had to keep the examples for a given individual within the respective dataset.\n* This makes NO SENSE, when you consider the approach of training using a clustering algorithm or something similar (ArcFace)\n* You need images of the same individual in both the train and the validation dataset, or else you can't possibly see if your model has learned how to distinguish between unique individuals.\n\n---\n\nThe simple analogy to explain what I did:\n\n**Imagine you are trying to classify cat v. dog**\n* I essentially put all the cat photos into the train dataset and all of the dog photos into the val dataset (i.e. species used to group)\n\n---\n\n😅😅😅😅So yeah! 😅😅😅😅\n\n---\n\nAt least I figured it out and got it all fixed! Cheers!",
    "1683380": "@dschettler8845 everyone makes mistakes, but you overcame this one and thus learned something new! I'm sure you won't forget it.\n\nI have always used `StratifiedKFold`, never felt the need of using `GroupKFold`. Any examples come to mind so as to where `GroupKFold` would allow for better folds? Would love to know 😄",
    "1683615": "Thanks @dschettler8845 for sharing.\nOne quick question that is bothering me: how do you actually set up a cross validation in a Computer Vision (or NLP) comp, where generally we have large datasets or models need a lot of time for training?\n\nLet me explain better. I know how CV works and I use it with no problems at all in tabular tasks. I generally split my dataset in 5 stratified (or group) folds, train, predict and save my OOF and predictions. This is generally done for later ensembling.\nIn NLP or Com. Vision cross validation this entire process seems \"demanding\" to me. Is it ok to fine tune a model 5 times? Is it a standard procedure?",
    "1683416": "Thanks for sharing, I think we should all share our failures, especially in the field of ML where bugs are very hard to find.",
    "1684075": "It's very counter-intuitive to use non-overlapping group splits here. I wonder how did you come up with that approach. :D",
    "1698488": "Thanks for sharing this, it should always feel safe to share failures, even though I won't consider this a real failure since you haven't put this into production and the competition isn't over yet. 👌",
    "1684978": "Hi @dschettler8845, thanks for this. I am trying to get my head around the different cv schemes at the moment. I found [this page](https://scikit-learn.org/stable/auto_examples/model_selection/plot_cv_indices.html) pretty helpful.",
    "1684958": "Well, that is an experience. Thanks for sharing @dschettler8845 . In my opinion, knowing your mistake is as important as what you are doing right!.",
    "1698030": ""
  }
}