{
  "id": 201911,
  "title": "(Cross-)validation strategy for this competition",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/201911",
  "author_name": "",
  "post_date": "2020-12-07T08:15:26.191603700Z",
  "votes": 16,
  "comment_count": 1,
  "views": 0,
  "content": "<p>We all know how important a good validation scheme is for exploring our modeling ideas and subsequently ensembling predictions. I have been wondering whether there is anything special we should consider for our CV scheme in this competition (it seems less obvious than with time series or tabular data with related records):</p>\n<ul>\n<li><strong>Simple k-fold (with k=5,6 or 7)</strong>: seems straightforward and I guess you'd need a reason to deviate from that.</li>\n<li><strong>Repeated k-fold</strong>: The other obvious choice besides a simple k-fold is to do a repeated k-fold (e.g. split into 5 folds twice in different ways) in order to not depend on the specific fold-splits / reduce noise a bit more. However, that may slow down iteration, I'm curious about people's experience in that respect (or does that depend on how powerful your local machine is?).</li>\n<li><strong>Stratified k-fold by label</strong>: Perhaps one should stratify the CV-splits by label - although <a href=\"https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy\" target=\"_blank\">when I looked at that</a> it does not look like you'd not get too low with the rarest category even without that. To some extent, we'd ideally split in the same way training and testing data were split for the competition and I suppose we do not know how that was done (if stratified, then you'd want to do stratified CV).</li>\n<li><strong>Stratified k-fold by image type</strong>: Perhaps one might want to stratify by certain types/clusters of images and perhaps there is more to that: the rarest type of image (those of plant roots - <a href=\"https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy\" target=\"_blank\">that's about 65 training images</a>) are pretty rare. If you want them present in all validation folds of your CV-split, you may have to stratify by image cluster (esp. for 7- or higher-fold CV, seems less of a problem with 5-fold).</li>\n<li><strong>Single train-test split</strong>: I does not seem to me that the dataset is so huge that this is either necessary (due to processing time) or a good idea (reliability of CV).</li>\n</ul>\n<p>Other things that crossed my mind was whether when I get additional images from somewhere, are people that have done a lot of vision competitions before typically splitting those among the folds (incl. using them as out-of-fold predictions)? Or do you use those always as part of the training for all folds? In some sense that might feel safer, as the out-of-fold validation would hopefully be more like the final private LB, if we only use images from this competition for that.</p>",
  "messages": [
    {
      "id": "1104788",
      "postDate": "12/07/2020 08:15:26",
      "content": "<p>We all know how important a good validation scheme is for exploring our modeling ideas and subsequently ensembling predictions. I have been wondering whether there is anything special we should consider for our CV scheme in this competition (it seems less obvious than with time series or tabular data with related records):</p>\n<ul>\n<li><strong>Simple k-fold (with k=5,6 or 7)</strong>: seems straightforward and I guess you'd need a reason to deviate from that.</li>\n<li><strong>Repeated k-fold</strong>: The other obvious choice besides a simple k-fold is to do a repeated k-fold (e.g. split into 5 folds twice in different ways) in order to not depend on the specific fold-splits / reduce noise a bit more. However, that may slow down iteration, I'm curious about people's experience in that respect (or does that depend on how powerful your local machine is?).</li>\n<li><strong>Stratified k-fold by label</strong>: Perhaps one should stratify the CV-splits by label - although <a href=\"https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy\" target=\"_blank\">when I looked at that</a> it does not look like you'd not get too low with the rarest category even without that. To some extent, we'd ideally split in the same way training and testing data were split for the competition and I suppose we do not know how that was done (if stratified, then you'd want to do stratified CV).</li>\n<li><strong>Stratified k-fold by image type</strong>: Perhaps one might want to stratify by certain types/clusters of images and perhaps there is more to that: the rarest type of image (those of plant roots - <a href=\"https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy\" target=\"_blank\">that's about 65 training images</a>) are pretty rare. If you want them present in all validation folds of your CV-split, you may have to stratify by image cluster (esp. for 7- or higher-fold CV, seems less of a problem with 5-fold).</li>\n<li><strong>Single train-test split</strong>: I does not seem to me that the dataset is so huge that this is either necessary (due to processing time) or a good idea (reliability of CV).</li>\n</ul>\n<p>Other things that crossed my mind was whether when I get additional images from somewhere, are people that have done a lot of vision competitions before typically splitting those among the folds (incl. using them as out-of-fold predictions)? Or do you use those always as part of the training for all folds? In some sense that might feel safer, as the out-of-fold validation would hopefully be more like the final private LB, if we only use images from this competition for that.</p>",
      "rawMarkdown": "We all know how important a good validation scheme is for exploring our modeling ideas and subsequently ensembling predictions. I have been wondering whether there is anything special we should consider for our CV scheme in this competition (it seems less obvious than with time series or tabular data with related records):\n- **Simple k-fold (with k=5,6 or 7)**: seems straightforward and I guess you'd need a reason to deviate from that.\n- **Repeated k-fold**: The other obvious choice besides a simple k-fold is to do a repeated k-fold (e.g. split into 5 folds twice in different ways) in order to not depend on the specific fold-splits / reduce noise a bit more. However, that may slow down iteration, I'm curious about people's experience in that respect (or does that depend on how powerful your local machine is?).\n- **Stratified k-fold by label**: Perhaps one should stratify the CV-splits by label - although [when I looked at that](https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy) it does not look like you'd not get too low with the rarest category even without that. To some extent, we'd ideally split in the same way training and testing data were split for the competition and I suppose we do not know how that was done (if stratified, then you'd want to do stratified CV).\n- **Stratified k-fold by image type**: Perhaps one might want to stratify by certain types/clusters of images and perhaps there is more to that: the rarest type of image (those of plant roots - [that's about 65 training images](https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy)) are pretty rare. If you want them present in all validation folds of your CV-split, you may have to stratify by image cluster (esp. for 7- or higher-fold CV, seems less of a problem with 5-fold).\n- **Single train-test split**: I does not seem to me that the dataset is so huge that this is either necessary (due to processing time) or a good idea (reliability of CV).\n\nOther things that crossed my mind was whether when I get additional images from somewhere, are people that have done a lot of vision competitions before typically splitting those among the folds (incl. using them as out-of-fold predictions)? Or do you use those always as part of the training for all folds? In some sense that might feel safer, as the out-of-fold validation would hopefully be more like the final private LB, if we only use images from this competition for that.",
      "votes": null
    },
    {
      "id": "1105067",
      "postDate": "12/07/2020 13:53:59",
      "content": "<p>I think it makes sense to use Stratified k-fold by label since the training data is highly imbalanced where 2 classes dominate the outcomes. (Class 3 and Class 4)</p>\n<p>Somebody posted that while submitting the final label as 3 the accuracy came out to be 61.5 which is in-line with the training set distribution, so I am assuming the private test set would also go by the same distribution of public test set.</p>",
      "rawMarkdown": "I think it makes sense to use Stratified k-fold by label since the training data is highly imbalanced where 2 classes dominate the outcomes. (Class 3 and Class 4)\n\nSomebody posted that while submitting the final label as 3 the accuracy came out to be 61.5 which is in-line with the training set distribution, so I am assuming the private test set would also go by the same distribution of public test set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1105067,
      "author_name": "harveenchadha",
      "author_url": "",
      "post_date": "12/07/2020 13:53:59",
      "content": "<p>I think it makes sense to use Stratified k-fold by label since the training data is highly imbalanced where 2 classes dominate the outcomes. (Class 3 and Class 4)</p>\n<p>Somebody posted that while submitting the final label as 3 the accuracy came out to be 61.5 which is in-line with the training set distribution, so I am assuming the private test set would also go by the same distribution of public test set.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1104788": "We all know how important a good validation scheme is for exploring our modeling ideas and subsequently ensembling predictions. I have been wondering whether there is anything special we should consider for our CV scheme in this competition (it seems less obvious than with time series or tabular data with related records):\n- **Simple k-fold (with k=5,6 or 7)**: seems straightforward and I guess you'd need a reason to deviate from that.\n- **Repeated k-fold**: The other obvious choice besides a simple k-fold is to do a repeated k-fold (e.g. split into 5 folds twice in different ways) in order to not depend on the specific fold-splits / reduce noise a bit more. However, that may slow down iteration, I'm curious about people's experience in that respect (or does that depend on how powerful your local machine is?).\n- **Stratified k-fold by label**: Perhaps one should stratify the CV-splits by label - although [when I looked at that](https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy) it does not look like you'd not get too low with the rarest category even without that. To some extent, we'd ideally split in the same way training and testing data were split for the competition and I suppose we do not know how that was done (if stratified, then you'd want to do stratified CV).\n- **Stratified k-fold by image type**: Perhaps one might want to stratify by certain types/clusters of images and perhaps there is more to that: the rarest type of image (those of plant roots - [that's about 65 training images](https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy)) are pretty rare. If you want them present in all validation folds of your CV-split, you may have to stratify by image cluster (esp. for 7- or higher-fold CV, seems less of a problem with 5-fold).\n- **Single train-test split**: I does not seem to me that the dataset is so huge that this is either necessary (due to processing time) or a good idea (reliability of CV).\n\nOther things that crossed my mind was whether when I get additional images from somewhere, are people that have done a lot of vision competitions before typically splitting those among the folds (incl. using them as out-of-fold predictions)? Or do you use those always as part of the training for all folds? In some sense that might feel safer, as the out-of-fold validation would hopefully be more like the final private LB, if we only use images from this competition for that.",
    "1105067": "I think it makes sense to use Stratified k-fold by label since the training data is highly imbalanced where 2 classes dominate the outcomes. (Class 3 and Class 4)\n\nSomebody posted that while submitting the final label as 3 the accuracy came out to be 61.5 which is in-line with the training set distribution, so I am assuming the private test set would also go by the same distribution of public test set."
  },
  "source": "meta"
}