{
  "id": 107649,
  "title": "Cross-Validation in Image Segmentation",
  "url": "/competitions/understanding_cloud_organization/discussion/107649",
  "author_name": "",
  "post_date": "2019-09-05T15:30:32.647473100Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I am quite new to image segmentation and have just got started within past few days. I have been going through some recent image segmentation competition leaderboard like  <a href=\"https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/leaderboard\">SIIM-ACR Pneumothorax Segmentation</a> and notice that a lot of people that scored well in top 50 in public leaderboard perform poorly in private leaderboard. I presume this is mainly due to their model being too overfitted toward the public test set. If I am not wrong, this can be prevented if the model has a reliable local validation, such as K-fold validation, to assess their models performance. If that's the case, I have some questions on the implementing cross validation in image segmentation</p>\n\n<ol>\n<li><p>Do you stratify each fold with certain criteria? <a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/69291\">In 1st Place TGS Competition Solution</a>, additional information on salt depth level is given and the folds is stratified by the depth level. In this case, do we stratify by the type of cloud, amount of mask, both, or just randomly created folds is good enough?</p></li>\n<li><p>Is the model overall performance determine by the average validation accuracy in each fold?</p></li>\n<li><p>How to predict test set with folds, do we need to use all 5 folds model's prediction on the test set and average their pixel wise probability , etc?</p></li>\n</ol>\n\n<p>Edit: Fixed link for LB</p>",
  "messages": [
    {
      "id": "618878",
      "postDate": "09/05/2019 15:30:32",
      "content": "<p>I am quite new to image segmentation and have just got started within past few days. I have been going through some recent image segmentation competition leaderboard like  <a href=\"https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/leaderboard\">SIIM-ACR Pneumothorax Segmentation</a> and notice that a lot of people that scored well in top 50 in public leaderboard perform poorly in private leaderboard. I presume this is mainly due to their model being too overfitted toward the public test set. If I am not wrong, this can be prevented if the model has a reliable local validation, such as K-fold validation, to assess their models performance. If that's the case, I have some questions on the implementing cross validation in image segmentation</p>\n\n<ol>\n<li><p>Do you stratify each fold with certain criteria? <a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/69291\">In 1st Place TGS Competition Solution</a>, additional information on salt depth level is given and the folds is stratified by the depth level. In this case, do we stratify by the type of cloud, amount of mask, both, or just randomly created folds is good enough?</p></li>\n<li><p>Is the model overall performance determine by the average validation accuracy in each fold?</p></li>\n<li><p>How to predict test set with folds, do we need to use all 5 folds model's prediction on the test set and average their pixel wise probability , etc?</p></li>\n</ol>\n\n<p>Edit: Fixed link for LB</p>",
      "rawMarkdown": "I am quite new to image segmentation and have just got started within past few days. I have been going through some recent image segmentation competition leaderboard like  [SIIM-ACR Pneumothorax Segmentation](https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/leaderboard) and notice that a lot of people that scored well in top 50 in public leaderboard perform poorly in private leaderboard. I presume this is mainly due to their model being too overfitted toward the public test set. If I am not wrong, this can be prevented if the model has a reliable local validation, such as K-fold validation, to assess their models performance. If that's the case, I have some questions on the implementing cross validation in image segmentation\n\n1. Do you stratify each fold with certain criteria? [In 1st Place TGS Competition Solution](https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/69291), additional information on salt depth level is given and the folds is stratified by the depth level. In this case, do we stratify by the type of cloud, amount of mask, both, or just randomly created folds is good enough?\n\n2. Is the model overall performance determine by the average validation accuracy in each fold?\n\n3. How to predict test set with folds, do we need to use all 5 folds model's prediction on the test set and average their pixel wise probability , etc?\n\n\nEdit: Fixed link for LB",
      "votes": null
    },
    {
      "id": "619020",
      "postDate": "09/05/2019 18:45:31",
      "content": "<p>You can stratify on attributes like number, size/area, and type of cloud classes. Does it help? Well it can’t hurt versus random. I usually average on pixel wise probs. Trust local CV not LB :)</p>",
      "rawMarkdown": "You can stratify on attributes like number, size/area, and type of cloud classes. Does it help? Well it can’t hurt versus random. I usually average on pixel wise probs. Trust local CV not LB :)",
      "votes": null
    },
    {
      "id": "619580",
      "postDate": "09/06/2019 10:09:25",
      "content": "<p>Pneumothorax competition was a 2-stage competition. So, the current comparison between Public and Private Leaderboards is not valid.</p>\n\n<ol>\n<li>I guess, it's always useful to use some kind of stratification instead of a random split. For this competition <a href=\"/robga\">@robga</a> mentioned a couple of possible options. Do not forget to place each image into a separate fold (as long as the train data contains 4 rows for each image).</li>\n<li>To get the local performance, the simplest way is to find the mean over all dice scores. However, a better approach is to look at the <code>mean(dice_scores) - std(dice_scores)</code>. It allows to take into account score deviation from one fold to another.</li>\n<li>Yes, for the test predictions you could find the average of mask probabilities over all folds. It could be a simple arithmetic mean, or geometric mean for each pixel. Also, you could combine the probabilities in some other ways: like, taking the maximum over all the folds.</li>\n</ol>",
      "rawMarkdown": "Pneumothorax competition was a 2-stage competition. So, the current comparison between Public and Private Leaderboards is not valid.\n\n1. I guess, it's always useful to use some kind of stratification instead of a random split. For this competition @robga mentioned a couple of possible options. Do not forget to place each image into a separate fold (as long as the train data contains 4 rows for each image).\n2. To get the local performance, the simplest way is to find the mean over all dice scores. However, a better approach is to look at the `mean(dice_scores) - std(dice_scores)`. It allows to take into account score deviation from one fold to another.\n3. Yes, for the test predictions you could find the average of mask probabilities over all folds. It could be a simple arithmetic mean, or geometric mean for each pixel. Also, you could combine the probabilities in some other ways: like, taking the maximum over all the folds.",
      "votes": null
    },
    {
      "id": "619690",
      "postDate": "09/06/2019 13:20:10",
      "content": "<p>Thanks! <code>mean(dice_scores) - std(dice_scores)</code> is a good criteria.</p>",
      "rawMarkdown": "Thanks! `mean(dice_scores) - std(dice_scores)` is a good criteria.",
      "votes": null
    },
    {
      "id": "647267",
      "postDate": "10/12/2019 10:08:40",
      "content": "<p>How is everyone dealing with k-fold cross-validation? My single model takes ~8 hours to train (30 epochs). If I were to take 5 folds it will take me ~40 hours. I have very limited GCP credits and am using a single P100 GPU.</p>",
      "rawMarkdown": "How is everyone dealing with k-fold cross-validation? My single model takes ~8 hours to train (30 epochs). If I were to take 5 folds it will take me ~40 hours. I have very limited GCP credits and am using a single P100 GPU.",
      "votes": null
    },
    {
      "id": "651669",
      "postDate": "10/17/2019 19:18:47",
      "content": "<p>What image size? And what model?</p>",
      "rawMarkdown": "What image size? And what model?",
      "votes": null
    },
    {
      "id": "651687",
      "postDate": "10/17/2019 19:51:41",
      "content": "<p>From the beginning, I am working with images resized to 350 x 525 px (required output size for masks). Earlier I was using a resnet50 encoded UNet with ~30 epochs. I am using efficientnet-b2 now, it is much much faster. </p>",
      "rawMarkdown": "From the beginning, I am working with images resized to 350 x 525 px (required output size for masks). Earlier I was using a resnet50 encoded UNet with ~30 epochs. I am using efficientnet-b2 now, it is much much faster.",
      "votes": null
    },
    {
      "id": "651692",
      "postDate": "10/17/2019 19:56:31",
      "content": "<p>Unfortunately image competitions require good GPUs. For now i'm using a similar size and efficientnetb3... taking around 4 hours/fold in a single V100</p>",
      "rawMarkdown": "Unfortunately image competitions require good GPUs. For now i'm using a similar size and efficientnetb3... taking around 4 hours/fold in a single V100",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 619020,
      "author_name": "robga",
      "author_url": "",
      "post_date": "09/05/2019 18:45:31",
      "content": "<p>You can stratify on attributes like number, size/area, and type of cloud classes. Does it help? Well it can’t hurt versus random. I usually average on pixel wise probs. Trust local CV not LB :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 619580,
      "author_name": "ybabakhin",
      "author_url": "",
      "post_date": "09/06/2019 10:09:25",
      "content": "<p>Pneumothorax competition was a 2-stage competition. So, the current comparison between Public and Private Leaderboards is not valid.</p>\n\n<ol>\n<li>I guess, it's always useful to use some kind of stratification instead of a random split. For this competition <a href=\"/robga\">@robga</a> mentioned a couple of possible options. Do not forget to place each image into a separate fold (as long as the train data contains 4 rows for each image).</li>\n<li>To get the local performance, the simplest way is to find the mean over all dice scores. However, a better approach is to look at the <code>mean(dice_scores) - std(dice_scores)</code>. It allows to take into account score deviation from one fold to another.</li>\n<li>Yes, for the test predictions you could find the average of mask probabilities over all folds. It could be a simple arithmetic mean, or geometric mean for each pixel. Also, you could combine the probabilities in some other ways: like, taking the maximum over all the folds.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 619690,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "09/06/2019 13:20:10",
          "content": "<p>Thanks! <code>mean(dice_scores) - std(dice_scores)</code> is a good criteria.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 647267,
      "author_name": "timetraveller98",
      "author_url": "",
      "post_date": "10/12/2019 10:08:40",
      "content": "<p>How is everyone dealing with k-fold cross-validation? My single model takes ~8 hours to train (30 epochs). If I were to take 5 folds it will take me ~40 hours. I have very limited GCP credits and am using a single P100 GPU.</p>",
      "votes": null,
      "replies": [
        {
          "id": 651669,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "10/17/2019 19:18:47",
          "content": "<p>What image size? And what model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 651687,
          "author_name": "timetraveller98",
          "author_url": "",
          "post_date": "10/17/2019 19:51:41",
          "content": "<p>From the beginning, I am working with images resized to 350 x 525 px (required output size for masks). Earlier I was using a resnet50 encoded UNet with ~30 epochs. I am using efficientnet-b2 now, it is much much faster. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 651692,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "10/17/2019 19:56:31",
          "content": "<p>Unfortunately image competitions require good GPUs. For now i'm using a similar size and efficientnetb3... taking around 4 hours/fold in a single V100</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "618878": "I am quite new to image segmentation and have just got started within past few days. I have been going through some recent image segmentation competition leaderboard like  [SIIM-ACR Pneumothorax Segmentation](https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/leaderboard) and notice that a lot of people that scored well in top 50 in public leaderboard perform poorly in private leaderboard. I presume this is mainly due to their model being too overfitted toward the public test set. If I am not wrong, this can be prevented if the model has a reliable local validation, such as K-fold validation, to assess their models performance. If that's the case, I have some questions on the implementing cross validation in image segmentation\n\n1. Do you stratify each fold with certain criteria? [In 1st Place TGS Competition Solution](https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/69291), additional information on salt depth level is given and the folds is stratified by the depth level. In this case, do we stratify by the type of cloud, amount of mask, both, or just randomly created folds is good enough?\n\n2. Is the model overall performance determine by the average validation accuracy in each fold?\n\n3. How to predict test set with folds, do we need to use all 5 folds model's prediction on the test set and average their pixel wise probability , etc?\n\n\nEdit: Fixed link for LB",
    "619020": "You can stratify on attributes like number, size/area, and type of cloud classes. Does it help? Well it can’t hurt versus random. I usually average on pixel wise probs. Trust local CV not LB :)",
    "619580": "Pneumothorax competition was a 2-stage competition. So, the current comparison between Public and Private Leaderboards is not valid.\n\n1. I guess, it's always useful to use some kind of stratification instead of a random split. For this competition @robga mentioned a couple of possible options. Do not forget to place each image into a separate fold (as long as the train data contains 4 rows for each image).\n2. To get the local performance, the simplest way is to find the mean over all dice scores. However, a better approach is to look at the `mean(dice_scores) - std(dice_scores)`. It allows to take into account score deviation from one fold to another.\n3. Yes, for the test predictions you could find the average of mask probabilities over all folds. It could be a simple arithmetic mean, or geometric mean for each pixel. Also, you could combine the probabilities in some other ways: like, taking the maximum over all the folds.",
    "619690": "Thanks! `mean(dice_scores) - std(dice_scores)` is a good criteria.",
    "647267": "How is everyone dealing with k-fold cross-validation? My single model takes ~8 hours to train (30 epochs). If I were to take 5 folds it will take me ~40 hours. I have very limited GCP credits and am using a single P100 GPU.",
    "651669": "What image size? And what model?",
    "651687": "From the beginning, I am working with images resized to 350 x 525 px (required output size for masks). Earlier I was using a resnet50 encoded UNet with ~30 epochs. I am using efficientnet-b2 now, it is much much faster.",
    "651692": "Unfortunately image competitions require good GPUs. For now i'm using a similar size and efficientnetb3... taking around 4 hours/fold in a single V100"
  },
  "source": "meta"
}