{
  "id": 107711,
  "title": "Discussing post processing",
  "url": "/competitions/understanding_cloud_organization/discussion/107711",
  "author_name": "",
  "post_date": "2019-09-06T04:32:57.658729500Z",
  "votes": 30,
  "comment_count": 18,
  "views": 0,
  "content": "<p>In my code I use the following post processing:\n<code>\ndef post_process(probability, threshold, min_size):\n    \"\"\"\n    Post processing of each predicted mask, components with lesser number of pixels\n    than `min_size` are ignored\n    \"\"\"\n    mask = cv2.threshold(probability, threshold, 1, cv2.THRESH_BINARY)[1]\n    num_component, component = cv2.connectedComponents(mask.astype(np.uint8))\n    predictions = np.zeros((350, 525), np.float32)\n    num = 0\n    for c in range(1, num_component):\n        p = (component == c)\n        if p.sum() &gt; min_size:\n            predictions[p] = 1\n            num += 1\n    return predictions, num\n</code></p>\n\n<p>I have noticed that increasing <code>min_size</code> increases validation score, but decreases score on public leaderboard. What could be the reason: overfitting, wrong approach to post processing, different masks in test data or something else?</p>",
  "messages": [
    {
      "id": "619320",
      "postDate": "09/06/2019 04:32:57",
      "content": "<p>In my code I use the following post processing:\n<code>\ndef post_process(probability, threshold, min_size):\n    \"\"\"\n    Post processing of each predicted mask, components with lesser number of pixels\n    than `min_size` are ignored\n    \"\"\"\n    mask = cv2.threshold(probability, threshold, 1, cv2.THRESH_BINARY)[1]\n    num_component, component = cv2.connectedComponents(mask.astype(np.uint8))\n    predictions = np.zeros((350, 525), np.float32)\n    num = 0\n    for c in range(1, num_component):\n        p = (component == c)\n        if p.sum() &gt; min_size:\n            predictions[p] = 1\n            num += 1\n    return predictions, num\n</code></p>\n\n<p>I have noticed that increasing <code>min_size</code> increases validation score, but decreases score on public leaderboard. What could be the reason: overfitting, wrong approach to post processing, different masks in test data or something else?</p>",
      "rawMarkdown": "In my code I use the following post processing:\n```\ndef post_process(probability, threshold, min_size):\n    \"\"\"\n    Post processing of each predicted mask, components with lesser number of pixels\n    than `min_size` are ignored\n    \"\"\"\n    mask = cv2.threshold(probability, threshold, 1, cv2.THRESH_BINARY)[1]\n    num_component, component = cv2.connectedComponents(mask.astype(np.uint8))\n    predictions = np.zeros((350, 525), np.float32)\n    num = 0\n    for c in range(1, num_component):\n        p = (component == c)\n        if p.sum() &gt; min_size:\n            predictions[p] = 1\n            num += 1\n    return predictions, num\n```\n\nI have noticed that increasing `min_size` increases validation score, but decreases score on public leaderboard. What could be the reason: overfitting, wrong approach to post processing, different masks in test data or something else?",
      "votes": null
    },
    {
      "id": "619340",
      "postDate": "09/06/2019 05:12:57",
      "content": "<p>Maybe it is better to select threshold and min size separately for each class...</p>",
      "rawMarkdown": "Maybe it is better to select threshold and min size separately for each class...",
      "votes": null
    },
    {
      "id": "619343",
      "postDate": "09/06/2019 05:23:30",
      "content": "<p>I haven’t made many submissions, but I’ve not noticed this. My local and LB scores are very similar. Maybe I’ll notice a different effect with more submissions. </p>\n\n<p>The public leaderboard is only 25% of data so natural randomness may come into play. </p>\n\n<p>The ‘wrong scores 0’ metric is broadly satisfied with post processing similar to what you’ve performed. But it is easy to overfit the public LB by tuning this parameter for the LB. The Airbus challenge shakeup was partially due to this, and a significant difference between public and private Images, but that doesn’t appear to be the case here. </p>\n\n<p>On tasks like this I always trust local validation not public LB. especially since the organisers declared test is a random subset of the dataset. </p>",
      "rawMarkdown": "I haven’t made many submissions, but I’ve not noticed this. My local and LB scores are very similar. Maybe I’ll notice a different effect with more submissions. \n\nThe public leaderboard is only 25% of data so natural randomness may come into play. \n\nThe ‘wrong scores 0’ metric is broadly satisfied with post processing similar to what you’ve performed. But it is easy to overfit the public LB by tuning this parameter for the LB. The Airbus challenge shakeup was partially due to this, and a significant difference between public and private Images, but that doesn’t appear to be the case here. \n\nOn tasks like this I always trust local validation not public LB. especially since the organisers declared test is a random subset of the dataset.",
      "votes": null
    },
    {
      "id": "619344",
      "postDate": "09/06/2019 05:23:55",
      "content": "<blockquote>\n  <p><strong>Andrew Lukyanenko wrote:</strong></p>\n  \n  <p>Maybe it is better to select threshold and min size separately for each class...</p>\n</blockquote>\n\n<p>It is. </p>",
      "rawMarkdown": "&gt; **Andrew Lukyanenko wrote:**\n&gt; \n&gt; Maybe it is better to select threshold and min size separately for each class...\n\nIt is.",
      "votes": null
    },
    {
      "id": "619368",
      "postDate": "09/06/2019 06:11:01",
      "content": "<p>Also, on challenges with post processing, if you have enough data, I have found it useful to keep a separate holdout set to validate tuning of post-processing params. But then you're trsining with less images.</p>",
      "rawMarkdown": "Also, on challenges with post processing, if you have enough data, I have found it useful to keep a separate holdout set to validate tuning of post-processing params. But then you're trsining with less images.",
      "votes": null
    },
    {
      "id": "619481",
      "postDate": "09/06/2019 08:20:56",
      "content": "<p>Your function seems correct and I am also using it. This is my optimisation result and may be helpful. The optimal threshold and <code>min_size</code> pair are not the largest one in my case.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3177784%2F1eea287a7cc80a7fcdf16d33769c99a3%2F__results___18_0.png?generation=1567758053464121&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Your function seems correct and I am also using it. This is my optimisation result and may be helpful. The optimal threshold and `min_size` pair are not the largest one in my case.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3177784%2F1eea287a7cc80a7fcdf16d33769c99a3%2F__results___18_0.png?generation=1567758053464121&amp;alt=media)",
      "votes": null
    },
    {
      "id": "619487",
      "postDate": "09/06/2019 08:23:25",
      "content": "<p>I select these two together and it works fine.</p>",
      "rawMarkdown": "I select these two together and it works fine.",
      "votes": null
    },
    {
      "id": "619506",
      "postDate": "09/06/2019 08:42:53",
      "content": "<p>Hm, interesting. Maybe my model is simply too weak and I need to train it better.</p>",
      "rawMarkdown": "Hm, interesting. Maybe my model is simply too weak and I need to train it better.",
      "votes": null
    },
    {
      "id": "619509",
      "postDate": "09/06/2019 08:44:56",
      "content": "<p>Thanks for the function and it improves my LB for almost 0.05👍 </p>",
      "rawMarkdown": "Thanks for the function and it improves my LB for almost 0.05👍",
      "votes": null
    },
    {
      "id": "619522",
      "postDate": "09/06/2019 08:53:30",
      "content": "<p>I have found another wierd problem thought the optimisation process seems correct. That is, a strong model doesn't always work better than a weak model after optimisation. For example, if I have a 0.58+ model and a 0.60+ model, the former one may give better LB after optimisation.</p>",
      "rawMarkdown": "I have found another wierd problem thought the optimisation process seems correct. That is, a strong model doesn't always work better than a weak model after optimisation. For example, if I have a 0.58+ model and a 0.60+ model, the former one may give better LB after optimisation.",
      "votes": null
    },
    {
      "id": "620090",
      "postDate": "09/07/2019 02:44:28",
      "content": "<p>I have used a very similar bit of code a couple of times but never tried to optimize it.  In my use I simply wanted to prevent small noise pixels from occurring.  I always based the value I used on the distribution of sizes of segments in the training data.  In those cases a min_size that was the same as the smallest actual area worked.  As I used lower threshold values more of these small noise segments occurred and the min_size helped keep the metric in good shape by preventing noise.</p>\n\n<p>Using your kernel with a variety of encoders the optimized threshold is always .8 or higher.  I hope to get a plot that looks more like Yirun's in future as I believe that high thresholds indicate a poor model.  When the model is poor I don't think you are always able to make sense of the interactions of other features on the LB score.  </p>\n\n<p>The value's I am getting for min_size are much larger than what I think is the minimum size area in the training data. - I have that EDA on my todo list.  So my assumption is that in addition to eliminating noise we might be also eliminating legitimate small areas.  I am pretty sure that once I run the EDA on segment area in training that I will use that distribution to pick my min_size rather than your method of finding it.</p>",
      "rawMarkdown": "I have used a very similar bit of code a couple of times but never tried to optimize it.  In my use I simply wanted to prevent small noise pixels from occurring.  I always based the value I used on the distribution of sizes of segments in the training data.  In those cases a min_size that was the same as the smallest actual area worked.  As I used lower threshold values more of these small noise segments occurred and the min_size helped keep the metric in good shape by preventing noise.\n\nUsing your kernel with a variety of encoders the optimized threshold is always .8 or higher.  I hope to get a plot that looks more like Yirun's in future as I believe that high thresholds indicate a poor model.  When the model is poor I don't think you are always able to make sense of the interactions of other features on the LB score.  \n\nThe value's I am getting for min_size are much larger than what I think is the minimum size area in the training data. - I have that EDA on my todo list.  So my assumption is that in addition to eliminating noise we might be also eliminating legitimate small areas.  I am pretty sure that once I run the EDA on segment area in training that I will use that distribution to pick my min_size rather than your method of finding it.",
      "votes": null
    },
    {
      "id": "620198",
      "postDate": "09/07/2019 06:18:08",
      "content": "<p>Perhaps the training data we are given does not correlate as well we hope (enough) for the threshold-optimization to guarantee it's improvement. Almost all competitions I joined seemed to follow the trend of optimized threshold improving LB score only up to extent (e.g. In TGS my 0.75 model saw benefit of CRF (not same as technique discussed happening here but just as an example) until it stopped when model on its own reached 0.8, in APTOS optimized kappa gave no improvement at all due to train/test mismatch).</p>",
      "rawMarkdown": "Perhaps the training data we are given does not correlate as well we hope (enough) for the threshold-optimization to guarantee it's improvement. Almost all competitions I joined seemed to follow the trend of optimized threshold improving LB score only up to extent (e.g. In TGS my 0.75 model saw benefit of CRF (not same as technique discussed happening here but just as an example) until it stopped when model on its own reached 0.8, in APTOS optimized kappa gave no improvement at all due to train/test mismatch).",
      "votes": null
    },
    {
      "id": "622412",
      "postDate": "09/09/2019 16:09:28",
      "content": "<p>Hi Andrew,</p>\n\n<p>first, thank you for the great work you share, it's very much appreciated!</p>\n\n<p>Regarding your question, it seems to me that a possible reason might be a tiny bug in your validation score estimation. In the <a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">notebook</a> you compute mean Dice with this peace of code:\n```\nd = []\nfor i, j in zip(masks, valid_masks):\n    if i.sum() != 0:\n        d.append(dice(i, j))</p>\n\n<p>np.mean(d)\n```</p>\n\n<p>This way, false negative pictures are ignored. And if you increase the <code>min_size</code> parameter, then you have more FNs, which are ignored during local validation, yet contribute with zeros during LB estimation on Kaggle.\nI believe mean Dice should be estimated like this:\n```\ndef estimate_dice(masks, valid_masks):\n    d = []\n    for i, j in zip(masks, valid_masks):\n        if i.sum() != 0:\n            d.append(dice(i, j))\n        else:\n            d.append(int(j.sum() == 0))              </p>\n\n<pre><code>return np.mean(d)\n</code></pre>\n\n<p>```</p>\n\n<p>Hopefully, I haven't missed something :) Please, let me know if this observation seems valid to you and if it helps.</p>",
      "rawMarkdown": "Hi Andrew,\n\nfirst, thank you for the great work you share, it's very much appreciated!\n\nRegarding your question, it seems to me that a possible reason might be a tiny bug in your validation score estimation. In the [notebook](https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools) you compute mean Dice with this peace of code:\n```\nd = []\nfor i, j in zip(masks, valid_masks):\n    if i.sum() != 0:\n        d.append(dice(i, j))\n\nnp.mean(d)\n```\n\nThis way, false negative pictures are ignored. And if you increase the `min_size` parameter, then you have more FNs, which are ignored during local validation, yet contribute with zeros during LB estimation on Kaggle.\nI believe mean Dice should be estimated like this:\n```\ndef estimate_dice(masks, valid_masks):\n    d = []\n    for i, j in zip(masks, valid_masks):\n        if i.sum() != 0:\n            d.append(dice(i, j))\n        else:\n            d.append(int(j.sum() == 0))              \n\n    return np.mean(d)\n```\n\nHopefully, I haven't missed something :) Please, let me know if this observation seems valid to you and if it helps.",
      "votes": null
    },
    {
      "id": "622429",
      "postDate": "09/09/2019 16:40:32",
      "content": "<p>Nice work! I am using a similar one.\n<code>\nif (i.sum() == 0) &amp; (j.sum() == 0):\n    d.append(1)\nelse:\n    d.append(dice(i, j))\n</code></p>",
      "rawMarkdown": "Nice work! I am using a similar one.\n```\nif (i.sum() == 0) &amp; (j.sum() == 0):\n    d.append(1)\nelse:\n    d.append(dice(i, j))\n```",
      "votes": null
    },
    {
      "id": "622791",
      "postDate": "09/10/2019 05:19:05",
      "content": "<p>Oh, thank you! I suppose that this was the reason.</p>",
      "rawMarkdown": "Oh, thank you! I suppose that this was the reason.",
      "votes": null
    },
    {
      "id": "627552",
      "postDate": "09/16/2019 05:53:19",
      "content": "<p>Hi, are you using K-fold validation? It is quite surprising that you could get validation to be similar to LB</p>",
      "rawMarkdown": "Hi, are you using K-fold validation? It is quite surprising that you could get validation to be similar to LB",
      "votes": null
    },
    {
      "id": "631558",
      "postDate": "09/22/2019 08:20:41",
      "content": "<p>I share another way of post-processing by converting masks to rectangle shape which can be found here : \n<a href=\"https://www.kaggle.com/ratthachat/cloud-rectangle-mask-postprocessing-no-gpu/\">https://www.kaggle.com/ratthachat/cloud-rectangle-mask-postprocessing-no-gpu/</a></p>\n\n<p><strong>EDIT : Update to have more and better choices : convex-shape or approximate polygon shape</strong></p>\n\n<p><img src=\"https://i.ibb.co/w45jCdW/convex-mask.jpg\" alt=\"rectangle masks\"></p>",
      "rawMarkdown": "I share another way of post-processing by converting masks to rectangle shape which can be found here : \nhttps://www.kaggle.com/ratthachat/cloud-rectangle-mask-postprocessing-no-gpu/\n\n**EDIT : Update to have more and better choices : convex-shape or approximate polygon shape**\n\n![rectangle masks](https://i.ibb.co/w45jCdW/convex-mask.jpg)",
      "votes": null
    },
    {
      "id": "631999",
      "postDate": "09/23/2019 04:58:52",
      "content": "<p>I did a look at this and found that there are some masks even below 2000 and about 10% falls under 20000 so by setting the value that high it makes it physically impossible to get about 10% of the answers right</p>",
      "rawMarkdown": "I did a look at this and found that there are some masks even below 2000 and about 10% falls under 20000 so by setting the value that high it makes it physically impossible to get about 10% of the answers right",
      "votes": null
    },
    {
      "id": "1120827",
      "postDate": "12/21/2020 06:24:43",
      "content": "<p>Thanks it will be useful for my predictions </p>",
      "rawMarkdown": "Thanks it will be useful for my predictions",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1120827,
      "author_name": "gobikrish27",
      "author_url": "",
      "post_date": "12/21/2020 06:24:43",
      "content": "<p>Thanks it will be useful for my predictions </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 619340,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "09/06/2019 05:12:57",
      "content": "<p>Maybe it is better to select threshold and min size separately for each class...</p>",
      "votes": null,
      "replies": [
        {
          "id": 619344,
          "author_name": "robga",
          "author_url": "",
          "post_date": "09/06/2019 05:23:55",
          "content": "<blockquote>\n  <p><strong>Andrew Lukyanenko wrote:</strong></p>\n  \n  <p>Maybe it is better to select threshold and min size separately for each class...</p>\n</blockquote>\n\n<p>It is. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 619487,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "09/06/2019 08:23:25",
          "content": "<p>I select these two together and it works fine.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 619343,
      "author_name": "robga",
      "author_url": "",
      "post_date": "09/06/2019 05:23:30",
      "content": "<p>I haven’t made many submissions, but I’ve not noticed this. My local and LB scores are very similar. Maybe I’ll notice a different effect with more submissions. </p>\n\n<p>The public leaderboard is only 25% of data so natural randomness may come into play. </p>\n\n<p>The ‘wrong scores 0’ metric is broadly satisfied with post processing similar to what you’ve performed. But it is easy to overfit the public LB by tuning this parameter for the LB. The Airbus challenge shakeup was partially due to this, and a significant difference between public and private Images, but that doesn’t appear to be the case here. </p>\n\n<p>On tasks like this I always trust local validation not public LB. especially since the organisers declared test is a random subset of the dataset. </p>",
      "votes": null,
      "replies": [
        {
          "id": 627552,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "09/16/2019 05:53:19",
          "content": "<p>Hi, are you using K-fold validation? It is quite surprising that you could get validation to be similar to LB</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 619368,
      "author_name": "robga",
      "author_url": "",
      "post_date": "09/06/2019 06:11:01",
      "content": "<p>Also, on challenges with post processing, if you have enough data, I have found it useful to keep a separate holdout set to validate tuning of post-processing params. But then you're trsining with less images.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 619481,
      "author_name": "gogo827jz",
      "author_url": "",
      "post_date": "09/06/2019 08:20:56",
      "content": "<p>Your function seems correct and I am also using it. This is my optimisation result and may be helpful. The optimal threshold and <code>min_size</code> pair are not the largest one in my case.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3177784%2F1eea287a7cc80a7fcdf16d33769c99a3%2F__results___18_0.png?generation=1567758053464121&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 619506,
          "author_name": "artgor",
          "author_url": "",
          "post_date": "09/06/2019 08:42:53",
          "content": "<p>Hm, interesting. Maybe my model is simply too weak and I need to train it better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 619509,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "09/06/2019 08:44:56",
          "content": "<p>Thanks for the function and it improves my LB for almost 0.05👍 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 619522,
      "author_name": "gogo827jz",
      "author_url": "",
      "post_date": "09/06/2019 08:53:30",
      "content": "<p>I have found another wierd problem thought the optimisation process seems correct. That is, a strong model doesn't always work better than a weak model after optimisation. For example, if I have a 0.58+ model and a 0.60+ model, the former one may give better LB after optimisation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 620090,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "09/07/2019 02:44:28",
      "content": "<p>I have used a very similar bit of code a couple of times but never tried to optimize it.  In my use I simply wanted to prevent small noise pixels from occurring.  I always based the value I used on the distribution of sizes of segments in the training data.  In those cases a min_size that was the same as the smallest actual area worked.  As I used lower threshold values more of these small noise segments occurred and the min_size helped keep the metric in good shape by preventing noise.</p>\n\n<p>Using your kernel with a variety of encoders the optimized threshold is always .8 or higher.  I hope to get a plot that looks more like Yirun's in future as I believe that high thresholds indicate a poor model.  When the model is poor I don't think you are always able to make sense of the interactions of other features on the LB score.  </p>\n\n<p>The value's I am getting for min_size are much larger than what I think is the minimum size area in the training data. - I have that EDA on my todo list.  So my assumption is that in addition to eliminating noise we might be also eliminating legitimate small areas.  I am pretty sure that once I run the EDA on segment area in training that I will use that distribution to pick my min_size rather than your method of finding it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 631999,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "09/23/2019 04:58:52",
          "content": "<p>I did a look at this and found that there are some masks even below 2000 and about 10% falls under 20000 so by setting the value that high it makes it physically impossible to get about 10% of the answers right</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 620198,
      "author_name": "joonl04",
      "author_url": "",
      "post_date": "09/07/2019 06:18:08",
      "content": "<p>Perhaps the training data we are given does not correlate as well we hope (enough) for the threshold-optimization to guarantee it's improvement. Almost all competitions I joined seemed to follow the trend of optimized threshold improving LB score only up to extent (e.g. In TGS my 0.75 model saw benefit of CRF (not same as technique discussed happening here but just as an example) until it stopped when model on its own reached 0.8, in APTOS optimized kappa gave no improvement at all due to train/test mismatch).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 622412,
      "author_name": "samusram",
      "author_url": "",
      "post_date": "09/09/2019 16:09:28",
      "content": "<p>Hi Andrew,</p>\n\n<p>first, thank you for the great work you share, it's very much appreciated!</p>\n\n<p>Regarding your question, it seems to me that a possible reason might be a tiny bug in your validation score estimation. In the <a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">notebook</a> you compute mean Dice with this peace of code:\n```\nd = []\nfor i, j in zip(masks, valid_masks):\n    if i.sum() != 0:\n        d.append(dice(i, j))</p>\n\n<p>np.mean(d)\n```</p>\n\n<p>This way, false negative pictures are ignored. And if you increase the <code>min_size</code> parameter, then you have more FNs, which are ignored during local validation, yet contribute with zeros during LB estimation on Kaggle.\nI believe mean Dice should be estimated like this:\n```\ndef estimate_dice(masks, valid_masks):\n    d = []\n    for i, j in zip(masks, valid_masks):\n        if i.sum() != 0:\n            d.append(dice(i, j))\n        else:\n            d.append(int(j.sum() == 0))              </p>\n\n<pre><code>return np.mean(d)\n</code></pre>\n\n<p>```</p>\n\n<p>Hopefully, I haven't missed something :) Please, let me know if this observation seems valid to you and if it helps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 622429,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "09/09/2019 16:40:32",
          "content": "<p>Nice work! I am using a similar one.\n<code>\nif (i.sum() == 0) &amp; (j.sum() == 0):\n    d.append(1)\nelse:\n    d.append(dice(i, j))\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 622791,
          "author_name": "artgor",
          "author_url": "",
          "post_date": "09/10/2019 05:19:05",
          "content": "<p>Oh, thank you! I suppose that this was the reason.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 631558,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "09/22/2019 08:20:41",
      "content": "<p>I share another way of post-processing by converting masks to rectangle shape which can be found here : \n<a href=\"https://www.kaggle.com/ratthachat/cloud-rectangle-mask-postprocessing-no-gpu/\">https://www.kaggle.com/ratthachat/cloud-rectangle-mask-postprocessing-no-gpu/</a></p>\n\n<p><strong>EDIT : Update to have more and better choices : convex-shape or approximate polygon shape</strong></p>\n\n<p><img src=\"https://i.ibb.co/w45jCdW/convex-mask.jpg\" alt=\"rectangle masks\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "619320": "In my code I use the following post processing:\n```\ndef post_process(probability, threshold, min_size):\n    \"\"\"\n    Post processing of each predicted mask, components with lesser number of pixels\n    than `min_size` are ignored\n    \"\"\"\n    mask = cv2.threshold(probability, threshold, 1, cv2.THRESH_BINARY)[1]\n    num_component, component = cv2.connectedComponents(mask.astype(np.uint8))\n    predictions = np.zeros((350, 525), np.float32)\n    num = 0\n    for c in range(1, num_component):\n        p = (component == c)\n        if p.sum() &gt; min_size:\n            predictions[p] = 1\n            num += 1\n    return predictions, num\n```\n\nI have noticed that increasing `min_size` increases validation score, but decreases score on public leaderboard. What could be the reason: overfitting, wrong approach to post processing, different masks in test data or something else?",
    "619340": "Maybe it is better to select threshold and min size separately for each class...",
    "619343": "I haven’t made many submissions, but I’ve not noticed this. My local and LB scores are very similar. Maybe I’ll notice a different effect with more submissions. \n\nThe public leaderboard is only 25% of data so natural randomness may come into play. \n\nThe ‘wrong scores 0’ metric is broadly satisfied with post processing similar to what you’ve performed. But it is easy to overfit the public LB by tuning this parameter for the LB. The Airbus challenge shakeup was partially due to this, and a significant difference between public and private Images, but that doesn’t appear to be the case here. \n\nOn tasks like this I always trust local validation not public LB. especially since the organisers declared test is a random subset of the dataset.",
    "619344": "&gt; **Andrew Lukyanenko wrote:**\n&gt; \n&gt; Maybe it is better to select threshold and min size separately for each class...\n\nIt is.",
    "619368": "Also, on challenges with post processing, if you have enough data, I have found it useful to keep a separate holdout set to validate tuning of post-processing params. But then you're trsining with less images.",
    "619481": "Your function seems correct and I am also using it. This is my optimisation result and may be helpful. The optimal threshold and `min_size` pair are not the largest one in my case.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3177784%2F1eea287a7cc80a7fcdf16d33769c99a3%2F__results___18_0.png?generation=1567758053464121&amp;alt=media)",
    "619487": "I select these two together and it works fine.",
    "619506": "Hm, interesting. Maybe my model is simply too weak and I need to train it better.",
    "619509": "Thanks for the function and it improves my LB for almost 0.05👍",
    "619522": "I have found another wierd problem thought the optimisation process seems correct. That is, a strong model doesn't always work better than a weak model after optimisation. For example, if I have a 0.58+ model and a 0.60+ model, the former one may give better LB after optimisation.",
    "620090": "I have used a very similar bit of code a couple of times but never tried to optimize it.  In my use I simply wanted to prevent small noise pixels from occurring.  I always based the value I used on the distribution of sizes of segments in the training data.  In those cases a min_size that was the same as the smallest actual area worked.  As I used lower threshold values more of these small noise segments occurred and the min_size helped keep the metric in good shape by preventing noise.\n\nUsing your kernel with a variety of encoders the optimized threshold is always .8 or higher.  I hope to get a plot that looks more like Yirun's in future as I believe that high thresholds indicate a poor model.  When the model is poor I don't think you are always able to make sense of the interactions of other features on the LB score.  \n\nThe value's I am getting for min_size are much larger than what I think is the minimum size area in the training data. - I have that EDA on my todo list.  So my assumption is that in addition to eliminating noise we might be also eliminating legitimate small areas.  I am pretty sure that once I run the EDA on segment area in training that I will use that distribution to pick my min_size rather than your method of finding it.",
    "620198": "Perhaps the training data we are given does not correlate as well we hope (enough) for the threshold-optimization to guarantee it's improvement. Almost all competitions I joined seemed to follow the trend of optimized threshold improving LB score only up to extent (e.g. In TGS my 0.75 model saw benefit of CRF (not same as technique discussed happening here but just as an example) until it stopped when model on its own reached 0.8, in APTOS optimized kappa gave no improvement at all due to train/test mismatch).",
    "622412": "Hi Andrew,\n\nfirst, thank you for the great work you share, it's very much appreciated!\n\nRegarding your question, it seems to me that a possible reason might be a tiny bug in your validation score estimation. In the [notebook](https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools) you compute mean Dice with this peace of code:\n```\nd = []\nfor i, j in zip(masks, valid_masks):\n    if i.sum() != 0:\n        d.append(dice(i, j))\n\nnp.mean(d)\n```\n\nThis way, false negative pictures are ignored. And if you increase the `min_size` parameter, then you have more FNs, which are ignored during local validation, yet contribute with zeros during LB estimation on Kaggle.\nI believe mean Dice should be estimated like this:\n```\ndef estimate_dice(masks, valid_masks):\n    d = []\n    for i, j in zip(masks, valid_masks):\n        if i.sum() != 0:\n            d.append(dice(i, j))\n        else:\n            d.append(int(j.sum() == 0))              \n\n    return np.mean(d)\n```\n\nHopefully, I haven't missed something :) Please, let me know if this observation seems valid to you and if it helps.",
    "622429": "Nice work! I am using a similar one.\n```\nif (i.sum() == 0) &amp; (j.sum() == 0):\n    d.append(1)\nelse:\n    d.append(dice(i, j))\n```",
    "622791": "Oh, thank you! I suppose that this was the reason.",
    "627552": "Hi, are you using K-fold validation? It is quite surprising that you could get validation to be similar to LB",
    "631558": "I share another way of post-processing by converting masks to rectangle shape which can be found here : \nhttps://www.kaggle.com/ratthachat/cloud-rectangle-mask-postprocessing-no-gpu/\n\n**EDIT : Update to have more and better choices : convex-shape or approximate polygon shape**\n\n![rectangle masks](https://i.ibb.co/w45jCdW/convex-mask.jpg)",
    "631999": "I did a look at this and found that there are some masks even below 2000 and about 10% falls under 20000 so by setting the value that high it makes it physically impossible to get about 10% of the answers right",
    "1120827": "Thanks it will be useful for my predictions"
  },
  "source": "meta"
}