{
  "id": 117310,
  "title": "How many masks do you have in your best submission?",
  "url": "/competitions/understanding_cloud_organization/discussion/117310",
  "author_name": "Yirun Zhang",
  "post_date": "2019-11-14T15:27:59.293000",
  "votes": 10,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I have noticed that there are around 53% masks in the training set. However, in our best submission we only have approximate 40% masks. Assuming the distribution of the test set is similar to that of the training set, I could imagine that this missing 13% masks may lead to a possible LB shakeup.</p>",
  "messages": [
    {
      "id": 673149,
      "postDate": "2019-11-14T15:27:59.293Z",
      "content": "<p>I have noticed that there are around 53% masks in the training set. However, in our best submission we only have approximate 40% masks. Assuming the distribution of the test set is similar to that of the training set, I could imagine that this missing 13% masks may lead to a possible LB shakeup.</p>",
      "rawMarkdown": "I have noticed that there are around 53% masks in the training set. However, in our best submission we only have approximate 40% masks. Assuming the distribution of the test set is similar to that of the training set, I could imagine that this missing 13% masks may lead to a possible LB shakeup.",
      "votes": 10
    },
    {
      "id": 673209,
      "postDate": "2019-11-14T17:06:54.190Z",
      "content": "<p>I agree with <a href=\"/hengck23\">@hengck23</a>. In my best submission, I have around 38.5% masks. It looks like it always better to predict less than wrong predictions.  </p>\n\n<p>However, in the submission between 0.670 to 0.672 I found that for different mask percentage I am getting similar LB for different thresholds. But be generating the same amount of wrong and right predictions. </p>",
      "rawMarkdown": "I agree with @hengck23. In my best submission, I have around 38.5% masks. It looks like it always better to predict less than wrong predictions.  \n\nHowever, in the submission between 0.670 to 0.672 I found that for different mask percentage I am getting similar LB for different thresholds. But be generating the same amount of wrong and right predictions. ",
      "votes": 2,
      "replies": [
        {
          "id": 674540,
          "postDate": "2019-11-16T17:13:20.580Z",
          "content": "<p>Many people complained that you shouldn't share a magic. But what you wrote was actually magic for me :) I was using a really low label threshold 0.55, but today I started to adjust the threshold. Now our main ensemble predicts around 38.9% masks and LB is increased around 0.003. </p>",
          "rawMarkdown": "Many people complained that you shouldn't share a magic. But what you wrote was actually magic for me :) I was using a really low label threshold 0.55, but today I started to adjust the threshold. Now our main ensemble predicts around 38.9% masks and LB is increased around 0.003. ",
          "votes": 1
        },
        {
          "id": 674560,
          "postDate": "2019-11-16T17:46:03.427Z",
          "content": "<p>45/3698/4 = 0.003042184964845863</p>\n\n<p>about 45 + ? false positive instance were removed. ? Accounts for wrongly removing positive mask</p>\n\n<p>```\nnegative precision  in range (0.55, new_threshold) \n  = (num_correct_neg_label) /num_of_neg_label\n  = (num_of_neg_label-45-?) /num_of_neg_label</p>\n\n<p>see if this is the same as your validation set ... then you can estimate shakeup \n```</p>\n\n<p>38% don't work for me. my classifier is not as precise as yours. I think each of us will have different magic numbers</p>",
          "rawMarkdown": "45/3698/4 = 0.003042184964845863\n\nabout 45 + ? false positive instance were removed. ? Accounts for wrongly removing positive mask\n\n```\nnegative precision  in range (0.55, new_threshold) \n  = (num_correct_neg_label) /num_of_neg_label\n  = (num_of_neg_label-45-?) /num_of_neg_label\n\n\nsee if this is the same as your validation set ... then you can estimate shakeup \n```\n\n38% don't work for me. my classifier is not as precise as yours. I think each of us will have different magic numbers",
          "votes": 1
        },
        {
          "id": 674590,
          "postDate": "2019-11-16T18:39:21.670Z",
          "content": "<p>I don't use classifier now, the best submission is only with the segmentation.</p>",
          "rawMarkdown": "I don't use classifier now, the best submission is only with the segmentation.",
          "votes": 1
        },
        {
          "id": 674706,
          "postDate": "2019-11-16T23:46:39.960Z",
          "content": "<p>Use classifier. Classifier is working good really. Just try the one already available in the notebooks. </p>",
          "rawMarkdown": "Use classifier. Classifier is working good really. Just try the one already available in the notebooks. ",
          "votes": 1
        },
        {
          "id": 674707,
          "postDate": "2019-11-16T23:49:35.197Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> I found literary good outcome from the available classifier in notebooks. Still, I believe the classifier will going to make the difference in this competition.</p>",
          "rawMarkdown": "@hengck23 I found literary good outcome from the available classifier in notebooks. Still, I believe the classifier will going to make the difference in this competition.",
          "votes": 1
        },
        {
          "id": 674964,
          "postDate": "2019-11-17T10:51:20.480Z",
          "content": "<p>@Md Yasin Kabir</p>\n\n<p>thanks for your suggestion. I always though segmentation is enough because i was my segmentation label prediction is better than my previous classifier.</p>\n\n<p>after your suggestion, i take more efforts into designing better classifier. And now i can can get classifier better or as good as  label prediction from segmentation. I realized that different input size, depth, etc is required and that is why i failed previously.</p>",
          "rawMarkdown": "@Md Yasin Kabir\n\nthanks for your suggestion. I always though segmentation is enough because i was my segmentation label prediction is better than my previous classifier.\n\nafter your suggestion, i take more efforts into designing better classifier. And now i can can get classifier better or as good as  label prediction from segmentation. I realized that different input size, depth, etc is required and that is why i failed previously."
        }
      ]
    },
    {
      "id": 673559,
      "postDate": "2019-11-15T06:26:57.027Z",
      "content": "<p>Related metric, my files are typically about 22MB</p>",
      "rawMarkdown": "Related metric, my files are typically about 22MB"
    },
    {
      "id": 673168,
      "postDate": "2019-11-14T15:55:26.213Z",
      "content": "<p>\"missing 13% masks\"\na null prediction is better than wrong prediction.</p>\n\n<p>in the steel competition, high thresholds that reject false positive works the best. In summary, predict mask only when you are confident. (to see the effect, check if your LB score is better for image label threshold 0.5, 0.6 , 0.7, etc)</p>",
      "rawMarkdown": "\"missing 13% masks\"\na null prediction is better than wrong prediction.\n\nin the steel competition, high thresholds that reject false positive works the best. In summary, predict mask only when you are confident. (to see the effect, check if your LB score is better for image label threshold 0.5, 0.6 , 0.7, etc)",
      "votes": 1,
      "replies": [
        {
          "id": 673173,
          "postDate": "2019-11-14T16:00:38.930Z",
          "content": "<p>Thanks, you are right. Because the dice will be 1 if both are null.</p>",
          "rawMarkdown": "Thanks, you are right. Because the dice will be 1 if both are null.",
          "votes": 2
        },
        {
          "id": 673175,
          "postDate": "2019-11-14T16:04:21.857Z",
          "content": "<p>in my experiments, e.g. i can get best local cv score using 0.6 image label threshold.\nbut for several submissions i made, the thresholds needs to be higher for the optimum public LB score after submitting values from 0.6 to 0.8</p>",
          "rawMarkdown": "in my experiments, e.g. i can get best local cv score using 0.6 image label threshold.\nbut for several submissions i made, the thresholds needs to be higher for the optimum public LB score after submitting values from 0.6 to 0.8",
          "votes": 2
        },
        {
          "id": 673549,
          "postDate": "2019-11-15T06:21:29.553Z",
          "content": "<p>Maybe minsize matters? What's your minsize?\nI used to use 20000 for all class but recently i find 10000 is much better.\nMaybe large minsize eliminate to much tps</p>",
          "rawMarkdown": "Maybe minsize matters? What's your minsize?\nI used to use 20000 for all class but recently i find 10000 is much better.\nMaybe large minsize eliminate to much tps"
        },
        {
          "id": 673613,
          "postDate": "2019-11-15T08:11:36.963Z",
          "content": "<p>i don't use min size.</p>\n\n<p>i find that if max_pixel_probabilty threshold is high,  small size predicted mask disappear too.</p>",
          "rawMarkdown": "i don't use min size.\n\ni find that if max\\_pixel\\_probabilty threshold is high,  small size predicted mask disappear too."
        },
        {
          "id": 673647,
          "postDate": "2019-11-15T09:12:49.887Z",
          "content": "<p>wow! it's the thing that i haven't noticed.Thanks for your sharing.\nBut i still think small min_size is necessary such as 100 in case of outliers</p>",
          "rawMarkdown": "wow! it's the thing that i haven't noticed.Thanks for your sharing.\nBut i still think small min_size is necessary such as 100 in case of outliers"
        },
        {
          "id": 673648,
          "postDate": "2019-11-15T09:25:03.523Z",
          "content": "<p>my results looks like these. low max pixel probability will gives negative image label.\nthese that pass through seems to have large area:</p>\n\n<p>number in white = max pixel probability = image label\ncontour = pixel probability 0.3  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fae126899bba8ffa2d385a069899cfb96%2F0a245ba.png?generation=1573809774763824&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa8011a3a2d4eac6ed853a29d0c9c0d1d%2F0a8b542.png?generation=1573809831919112&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "my results looks like these. low max pixel probability will gives negative image label.\nthese that pass through seems to have large area:\n\nnumber in white = max pixel probability = image label\ncontour = pixel probability 0.3  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fae126899bba8ffa2d385a069899cfb96%2F0a245ba.png?generation=1573809774763824&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa8011a3a2d4eac6ed853a29d0c9c0d1d%2F0a8b542.png?generation=1573809831919112&amp;alt=media)\n"
        }
      ]
    },
    {
      "id": 674691,
      "postDate": "2019-11-16T23:15:11.650Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 673209,
      "author_name": "Md Yasin Kabir",
      "author_url": "",
      "post_date": "2019-11-14T17:06:54.190000",
      "content": "<p>I agree with <a href=\"/hengck23\">@hengck23</a>. In my best submission, I have around 38.5% masks. It looks like it always better to predict less than wrong predictions.  </p>\n\n<p>However, in the submission between 0.670 to 0.672 I found that for different mask percentage I am getting similar LB for different thresholds. But be generating the same amount of wrong and right predictions. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 674540,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2019-11-16T17:13:20.580000",
          "content": "<p>Many people complained that you shouldn't share a magic. But what you wrote was actually magic for me :) I was using a really low label threshold 0.55, but today I started to adjust the threshold. Now our main ensemble predicts around 38.9% masks and LB is increased around 0.003. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 674560,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-16T17:46:03.427000",
          "content": "<p>45/3698/4 = 0.003042184964845863</p>\n\n<p>about 45 + ? false positive instance were removed. ? Accounts for wrongly removing positive mask</p>\n\n<p>```\nnegative precision  in range (0.55, new_threshold) \n  = (num_correct_neg_label) /num_of_neg_label\n  = (num_of_neg_label-45-?) /num_of_neg_label</p>\n\n<p>see if this is the same as your validation set ... then you can estimate shakeup \n```</p>\n\n<p>38% don't work for me. my classifier is not as precise as yours. I think each of us will have different magic numbers</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 674590,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2019-11-16T18:39:21.670000",
          "content": "<p>I don't use classifier now, the best submission is only with the segmentation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 674706,
          "author_name": "Md Yasin Kabir",
          "author_url": "",
          "post_date": "2019-11-16T23:46:39.960000",
          "content": "<p>Use classifier. Classifier is working good really. Just try the one already available in the notebooks. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 674707,
          "author_name": "Md Yasin Kabir",
          "author_url": "",
          "post_date": "2019-11-16T23:49:35.197000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> I found literary good outcome from the available classifier in notebooks. Still, I believe the classifier will going to make the difference in this competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 674964,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-17T10:51:20.480000",
          "content": "<p>@Md Yasin Kabir</p>\n\n<p>thanks for your suggestion. I always though segmentation is enough because i was my segmentation label prediction is better than my previous classifier.</p>\n\n<p>after your suggestion, i take more efforts into designing better classifier. And now i can can get classifier better or as good as  label prediction from segmentation. I realized that different input size, depth, etc is required and that is why i failed previously.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 673559,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2019-11-15T06:26:57.027000",
      "content": "<p>Related metric, my files are typically about 22MB</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 673168,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-14T15:55:26.213000",
      "content": "<p>\"missing 13% masks\"\na null prediction is better than wrong prediction.</p>\n\n<p>in the steel competition, high thresholds that reject false positive works the best. In summary, predict mask only when you are confident. (to see the effect, check if your LB score is better for image label threshold 0.5, 0.6 , 0.7, etc)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 673173,
          "author_name": "Yirun Zhang",
          "author_url": "",
          "post_date": "2019-11-14T16:00:38.930000",
          "content": "<p>Thanks, you are right. Because the dice will be 1 if both are null.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 673175,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-14T16:04:21.857000",
          "content": "<p>in my experiments, e.g. i can get best local cv score using 0.6 image label threshold.\nbut for several submissions i made, the thresholds needs to be higher for the optimum public LB score after submitting values from 0.6 to 0.8</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 673549,
          "author_name": "llh1818",
          "author_url": "",
          "post_date": "2019-11-15T06:21:29.553000",
          "content": "<p>Maybe minsize matters? What's your minsize?\nI used to use 20000 for all class but recently i find 10000 is much better.\nMaybe large minsize eliminate to much tps</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673613,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-15T08:11:36.963000",
          "content": "<p>i don't use min size.</p>\n\n<p>i find that if max_pixel_probabilty threshold is high,  small size predicted mask disappear too.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673647,
          "author_name": "llh1818",
          "author_url": "",
          "post_date": "2019-11-15T09:12:49.887000",
          "content": "<p>wow! it's the thing that i haven't noticed.Thanks for your sharing.\nBut i still think small min_size is necessary such as 100 in case of outliers</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673648,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-15T09:25:03.523000",
          "content": "<p>my results looks like these. low max pixel probability will gives negative image label.\nthese that pass through seems to have large area:</p>\n\n<p>number in white = max pixel probability = image label\ncontour = pixel probability 0.3  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fae126899bba8ffa2d385a069899cfb96%2F0a245ba.png?generation=1573809774763824&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa8011a3a2d4eac6ed853a29d0c9c0d1d%2F0a8b542.png?generation=1573809831919112&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 674691,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-16T23:15:11.650000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "673149": "I have noticed that there are around 53% masks in the training set. However, in our best submission we only have approximate 40% masks. Assuming the distribution of the test set is similar to that of the training set, I could imagine that this missing 13% masks may lead to a possible LB shakeup.",
    "673209": "I agree with @hengck23. In my best submission, I have around 38.5% masks. It looks like it always better to predict less than wrong predictions.  \n\nHowever, in the submission between 0.670 to 0.672 I found that for different mask percentage I am getting similar LB for different thresholds. But be generating the same amount of wrong and right predictions. ",
    "673559": "Related metric, my files are typically about 22MB",
    "673168": "\"missing 13% masks\"\na null prediction is better than wrong prediction.\n\nin the steel competition, high thresholds that reject false positive works the best. In summary, predict mask only when you are confident. (to see the effect, check if your LB score is better for image label threshold 0.5, 0.6 , 0.7, etc)",
    "674691": ""
  }
}