{
  "id": 118018,
  "title": "143rd place (bronze) solution with code",
  "url": "/competitions/understanding_cloud_organization/discussion/118018",
  "author_name": "Belinda Trotta",
  "post_date": "2019-11-19T08:06:06.997000",
  "votes": 26,
  "comment_count": 9,
  "views": 0,
  "content": "<h2>Code and more detail:  <a href=\"https://github.com/btrotta/kaggle-clouds\">https://github.com/btrotta/kaggle-clouds</a></h2>\n\n<h2>Pre-processing the images</h2>\n\n<p>I worked with grayscale images shrunken to 25% of original size.</p>\n\n<p>I got a large boost in model accuracy from filtering out the over-exposed areas in the images. Below is a sample image before and after correction (I also changed the missing area to grey).</p>\n\n<p><img src=\"https://raw.githubusercontent.com/btrotta/kaggle-clouds/master/img/before_after.png\" alt=\"Before and after correction\"></p>\n\n<h2>Model</h2>\n\n<p>I used a blend of efficientnetb4 effecientnetb5, both pre-trained from this library: <a href=\"https://github.com/qubvel/segmentation_models\">https://github.com/qubvel/segmentation_models</a>. I trained for 10 epochs with the encoder layers frozen, then fine-tuned the whole model for 10 epochs with a lower learning rate. I did horizontal and vertical flip augmentation; I tried others but found they didn't help.</p>\n\n<h2>Post-processing the model predictions</h2>\n\n<p>The key to post-processing is to observe that the dice metric is not continuous: if a class doesn't exist in an image, there is a huge difference in predicting 1 pixel (dice score 0) and predicting 0 pixels (dice score 1). So, to decide whether to make a non-zero prediction, we need to estimate 2 things: the probability that the class exists in the image, and the expected dice score given that the class does exist. Then we can calculate the expected dice score for a zero and a non-zero prediction, and choose between them accordingly. I built very simple models for these, all just using a single variable: the 95th percentile of the predicted class probabilities for each image. </p>\n\n<p>I didn't attempt to reshape the predicted areas into rectangles or polygons, as in some published kernels. I also didn't enforce a minimum predicted area. My hypothesis is that  this information is already built in to the neural network predictions, and that this is why augmentations that change the size or shape of the masks (e.g. skew, rotation, zoom) give poor results.</p>",
  "messages": [
    {
      "id": 676473,
      "postDate": "2019-11-19T08:06:06.997Z",
      "content": "<h2>Code and more detail:  <a href=\"https://github.com/btrotta/kaggle-clouds\">https://github.com/btrotta/kaggle-clouds</a></h2>\n\n<h2>Pre-processing the images</h2>\n\n<p>I worked with grayscale images shrunken to 25% of original size.</p>\n\n<p>I got a large boost in model accuracy from filtering out the over-exposed areas in the images. Below is a sample image before and after correction (I also changed the missing area to grey).</p>\n\n<p><img src=\"https://raw.githubusercontent.com/btrotta/kaggle-clouds/master/img/before_after.png\" alt=\"Before and after correction\"></p>\n\n<h2>Model</h2>\n\n<p>I used a blend of efficientnetb4 effecientnetb5, both pre-trained from this library: <a href=\"https://github.com/qubvel/segmentation_models\">https://github.com/qubvel/segmentation_models</a>. I trained for 10 epochs with the encoder layers frozen, then fine-tuned the whole model for 10 epochs with a lower learning rate. I did horizontal and vertical flip augmentation; I tried others but found they didn't help.</p>\n\n<h2>Post-processing the model predictions</h2>\n\n<p>The key to post-processing is to observe that the dice metric is not continuous: if a class doesn't exist in an image, there is a huge difference in predicting 1 pixel (dice score 0) and predicting 0 pixels (dice score 1). So, to decide whether to make a non-zero prediction, we need to estimate 2 things: the probability that the class exists in the image, and the expected dice score given that the class does exist. Then we can calculate the expected dice score for a zero and a non-zero prediction, and choose between them accordingly. I built very simple models for these, all just using a single variable: the 95th percentile of the predicted class probabilities for each image. </p>\n\n<p>I didn't attempt to reshape the predicted areas into rectangles or polygons, as in some published kernels. I also didn't enforce a minimum predicted area. My hypothesis is that  this information is already built in to the neural network predictions, and that this is why augmentations that change the size or shape of the masks (e.g. skew, rotation, zoom) give poor results.</p>",
      "rawMarkdown": "## Code and more detail:  https://github.com/btrotta/kaggle-clouds\n\n## Pre-processing the images\n\nI worked with grayscale images shrunken to 25% of original size.\n\nI got a large boost in model accuracy from filtering out the over-exposed areas in the images. Below is a sample image before and after correction (I also changed the missing area to grey).\n\n![Before and after correction](https://raw.githubusercontent.com/btrotta/kaggle-clouds/master/img/before_after.png)\n\n## Model\n\nI used a blend of efficientnetb4 effecientnetb5, both pre-trained from this library: https://github.com/qubvel/segmentation_models. I trained for 10 epochs with the encoder layers frozen, then fine-tuned the whole model for 10 epochs with a lower learning rate. I did horizontal and vertical flip augmentation; I tried others but found they didn't help.\n\n## Post-processing the model predictions\n\nThe key to post-processing is to observe that the dice metric is not continuous: if a class doesn't exist in an image, there is a huge difference in predicting 1 pixel (dice score 0) and predicting 0 pixels (dice score 1). So, to decide whether to make a non-zero prediction, we need to estimate 2 things: the probability that the class exists in the image, and the expected dice score given that the class does exist. Then we can calculate the expected dice score for a zero and a non-zero prediction, and choose between them accordingly. I built very simple models for these, all just using a single variable: the 95th percentile of the predicted class probabilities for each image. \n\nI didn't attempt to reshape the predicted areas into rectangles or polygons, as in some published kernels. I also didn't enforce a minimum predicted area. My hypothesis is that  this information is already built in to the neural network predictions, and that this is why augmentations that change the size or shape of the masks (e.g. skew, rotation, zoom) give poor results.\n",
      "votes": 26
    },
    {
      "id": 677429,
      "postDate": "2019-11-20T06:19:20.817Z",
      "content": "<p>Congrats Belinda. Your pre-processing is really cool.Its amazing that you got a medal without enforcing minimum predicted area. I as well as many other Kagglers got huge (really huge) boost by using this strategy.</p>",
      "rawMarkdown": "Congrats Belinda. Your pre-processing is really cool.Its amazing that you got a medal without enforcing minimum predicted area. I as well as many other Kagglers got huge (really huge) boost by using this strategy.",
      "replies": [
        {
          "id": 677463,
          "postDate": "2019-11-20T07:45:14.513Z",
          "content": "<p>I think the main reason that enforcing minimum predicted area gives a boost is because of the non-continuous dice metric. For images where the model identifies only a small area, there's a high chance the class doesn't exist in the image at all and it's better to predict an empty mask. This is the idea behind the minimum area bound. But in my approach the minimum bound is not needed, because I'm explicitly calculating the probability that the class exists in the image and can decide based on that whether to make an empty prediction.</p>",
          "rawMarkdown": "I think the main reason that enforcing minimum predicted area gives a boost is because of the non-continuous dice metric. For images where the model identifies only a small area, there's a high chance the class doesn't exist in the image at all and it's better to predict an empty mask. This is the idea behind the minimum area bound. But in my approach the minimum bound is not needed, because I'm explicitly calculating the probability that the class exists in the image and can decide based on that whether to make an empty prediction.",
          "votes": 2
        }
      ]
    },
    {
      "id": 676980,
      "postDate": "2019-11-19T17:18:24.500Z",
      "content": "<p>Congrats Belinda. Nice job, I saw that you jumped into the top 50 quickly in this comp.</p>\n\n<blockquote>\n  <p>So, to decide whether to make a non-zero prediction, we need to estimate 2 things: the probability that the class exists in the image, and the expected dice score given that the class does exist.</p>\n</blockquote>\n\n<p>I like this. I analyzed this formula too. An important observation is recognizing that the formula includes a <strong>conditional</strong> probability: \"the expected dice score <strong>given</strong> that the class does exist\". Therefore we can train segmentation models on all positive labels to increase this conditional probability even though it decreases prediction precision. Then we can use a second model (or second head on first model) to remove false positives. It seems that this a a common theme with Kaggle segmentation competitions since they use a discontinuous Dice metric.</p>\n\n<p>I like your preprocess. From the images you posted, your adjustment works very well. Using grayscale images was smart. That allowed bigger batch sizes with the large efficientnet b4 and b5.</p>",
      "rawMarkdown": "Congrats Belinda. Nice job, I saw that you jumped into the top 50 quickly in this comp.\n\n&gt; So, to decide whether to make a non-zero prediction, we need to estimate 2 things: the probability that the class exists in the image, and the expected dice score given that the class does exist.\n\nI like this. I analyzed this formula too. An important observation is recognizing that the formula includes a **conditional** probability: \"the expected dice score **given** that the class does exist\". Therefore we can train segmentation models on all positive labels to increase this conditional probability even though it decreases prediction precision. Then we can use a second model (or second head on first model) to remove false positives. It seems that this a a common theme with Kaggle segmentation competitions since they use a discontinuous Dice metric.\n\nI like your preprocess. From the images you posted, your adjustment works very well. Using grayscale images was smart. That allowed bigger batch sizes with the large efficientnet b4 and b5."
    },
    {
      "id": 676606,
      "postDate": "2019-11-19T10:46:38.277Z",
      "content": "<p>Thanks for sharing! And congratulations 🎉 </p>",
      "rawMarkdown": "Thanks for sharing! And congratulations 🎉 "
    },
    {
      "id": 676534,
      "postDate": "2019-11-19T09:16:08.993Z",
      "content": "<p>Like the pre-processing part, great job.</p>",
      "rawMarkdown": "Like the pre-processing part, great job.",
      "replies": [
        {
          "id": 676541,
          "postDate": "2019-11-19T09:27:58.923Z",
          "content": "<p>I'm sorry, if I understand right, the get backgound function is for searching over-exposed region, right? \n<code>(np.mean(block) &lt; 200) &amp; (np.mean(block) &gt; 2)</code>\nWith hard-coded threshold value 200 and 2, is this able to correct all the over-exposed images in training data? </p>",
          "rawMarkdown": "I'm sorry, if I understand right, the get backgound function is for searching over-exposed region, right? \n`(np.mean(block) &lt; 200) &amp; (np.mean(block) &gt; 2)`\nWith hard-coded threshold value 200 and 2, is this able to correct all the over-exposed images in training data? "
        },
        {
          "id": 676553,
          "postDate": "2019-11-19T09:47:40.513Z",
          "content": "<p>Yes, <code>get_background</code> finds the over-exposed areas. The lower bound of 2 is to exclude the black \"stripes\" (i.e. the missing parts in the images). The upper bound of 200 is to exclude areas which are genuinely white (e.g. the centers of \"flower\" areas). I tested on several different images, and it seems to work pretty well on all of them.</p>",
          "rawMarkdown": "Yes, `get_background` finds the over-exposed areas. The lower bound of 2 is to exclude the black \"stripes\" (i.e. the missing parts in the images). The upper bound of 200 is to exclude areas which are genuinely white (e.g. the centers of \"flower\" areas). I tested on several different images, and it seems to work pretty well on all of them.",
          "votes": 1
        }
      ]
    },
    {
      "id": 676524,
      "postDate": "2019-11-19T09:04:52.803Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 676886,
      "postDate": "2019-11-19T15:16:08.997Z",
      "content": "<p>Thanks for sharing:)</p>",
      "rawMarkdown": "Thanks for sharing:)"
    }
  ],
  "comments": [
    {
      "id": 677429,
      "author_name": "Raghawendra Singh",
      "author_url": "",
      "post_date": "2019-11-20T06:19:20.817000",
      "content": "<p>Congrats Belinda. Your pre-processing is really cool.Its amazing that you got a medal without enforcing minimum predicted area. I as well as many other Kagglers got huge (really huge) boost by using this strategy.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 677463,
          "author_name": "Belinda Trotta",
          "author_url": "",
          "post_date": "2019-11-20T07:45:14.513000",
          "content": "<p>I think the main reason that enforcing minimum predicted area gives a boost is because of the non-continuous dice metric. For images where the model identifies only a small area, there's a high chance the class doesn't exist in the image at all and it's better to predict an empty mask. This is the idea behind the minimum area bound. But in my approach the minimum bound is not needed, because I'm explicitly calculating the probability that the class exists in the image and can decide based on that whether to make an empty prediction.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 676980,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2019-11-19T17:18:24.500000",
      "content": "<p>Congrats Belinda. Nice job, I saw that you jumped into the top 50 quickly in this comp.</p>\n\n<blockquote>\n  <p>So, to decide whether to make a non-zero prediction, we need to estimate 2 things: the probability that the class exists in the image, and the expected dice score given that the class does exist.</p>\n</blockquote>\n\n<p>I like this. I analyzed this formula too. An important observation is recognizing that the formula includes a <strong>conditional</strong> probability: \"the expected dice score <strong>given</strong> that the class does exist\". Therefore we can train segmentation models on all positive labels to increase this conditional probability even though it decreases prediction precision. Then we can use a second model (or second head on first model) to remove false positives. It seems that this a a common theme with Kaggle segmentation competitions since they use a discontinuous Dice metric.</p>\n\n<p>I like your preprocess. From the images you posted, your adjustment works very well. Using grayscale images was smart. That allowed bigger batch sizes with the large efficientnet b4 and b5.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676606,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2019-11-19T10:46:38.277000",
      "content": "<p>Thanks for sharing! And congratulations 🎉 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676534,
      "author_name": "Tsai29",
      "author_url": "",
      "post_date": "2019-11-19T09:16:08.993000",
      "content": "<p>Like the pre-processing part, great job.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676541,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2019-11-19T09:27:58.923000",
          "content": "<p>I'm sorry, if I understand right, the get backgound function is for searching over-exposed region, right? \n<code>(np.mean(block) &lt; 200) &amp; (np.mean(block) &gt; 2)</code>\nWith hard-coded threshold value 200 and 2, is this able to correct all the over-exposed images in training data? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676553,
          "author_name": "Belinda Trotta",
          "author_url": "",
          "post_date": "2019-11-19T09:47:40.513000",
          "content": "<p>Yes, <code>get_background</code> finds the over-exposed areas. The lower bound of 2 is to exclude the black \"stripes\" (i.e. the missing parts in the images). The upper bound of 200 is to exclude areas which are genuinely white (e.g. the centers of \"flower\" areas). I tested on several different images, and it seems to work pretty well on all of them.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 676524,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-19T09:04:52.803000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676886,
      "author_name": "liuze",
      "author_url": "",
      "post_date": "2019-11-19T15:16:08.997000",
      "content": "<p>Thanks for sharing:)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "676473": "## Code and more detail:  https://github.com/btrotta/kaggle-clouds\n\n## Pre-processing the images\n\nI worked with grayscale images shrunken to 25% of original size.\n\nI got a large boost in model accuracy from filtering out the over-exposed areas in the images. Below is a sample image before and after correction (I also changed the missing area to grey).\n\n![Before and after correction](https://raw.githubusercontent.com/btrotta/kaggle-clouds/master/img/before_after.png)\n\n## Model\n\nI used a blend of efficientnetb4 effecientnetb5, both pre-trained from this library: https://github.com/qubvel/segmentation_models. I trained for 10 epochs with the encoder layers frozen, then fine-tuned the whole model for 10 epochs with a lower learning rate. I did horizontal and vertical flip augmentation; I tried others but found they didn't help.\n\n## Post-processing the model predictions\n\nThe key to post-processing is to observe that the dice metric is not continuous: if a class doesn't exist in an image, there is a huge difference in predicting 1 pixel (dice score 0) and predicting 0 pixels (dice score 1). So, to decide whether to make a non-zero prediction, we need to estimate 2 things: the probability that the class exists in the image, and the expected dice score given that the class does exist. Then we can calculate the expected dice score for a zero and a non-zero prediction, and choose between them accordingly. I built very simple models for these, all just using a single variable: the 95th percentile of the predicted class probabilities for each image. \n\nI didn't attempt to reshape the predicted areas into rectangles or polygons, as in some published kernels. I also didn't enforce a minimum predicted area. My hypothesis is that  this information is already built in to the neural network predictions, and that this is why augmentations that change the size or shape of the masks (e.g. skew, rotation, zoom) give poor results.\n",
    "677429": "Congrats Belinda. Your pre-processing is really cool.Its amazing that you got a medal without enforcing minimum predicted area. I as well as many other Kagglers got huge (really huge) boost by using this strategy.",
    "676980": "Congrats Belinda. Nice job, I saw that you jumped into the top 50 quickly in this comp.\n\n&gt; So, to decide whether to make a non-zero prediction, we need to estimate 2 things: the probability that the class exists in the image, and the expected dice score given that the class does exist.\n\nI like this. I analyzed this formula too. An important observation is recognizing that the formula includes a **conditional** probability: \"the expected dice score **given** that the class does exist\". Therefore we can train segmentation models on all positive labels to increase this conditional probability even though it decreases prediction precision. Then we can use a second model (or second head on first model) to remove false positives. It seems that this a a common theme with Kaggle segmentation competitions since they use a discontinuous Dice metric.\n\nI like your preprocess. From the images you posted, your adjustment works very well. Using grayscale images was smart. That allowed bigger batch sizes with the large efficientnet b4 and b5.",
    "676606": "Thanks for sharing! And congratulations 🎉 ",
    "676534": "Like the pre-processing part, great job.",
    "676524": "",
    "676886": "Thanks for sharing:)"
  }
}