{
  "id": 118065,
  "title": "3rd silver place key points",
  "url": "/competitions/understanding_cloud_organization/discussion/118065",
  "author_name": "Kha Vo",
  "post_date": "2019-11-19T12:42:07.211000",
  "votes": 27,
  "comment_count": 12,
  "views": 0,
  "content": "<p>First I would like to thank Max Planck Institute and Kaggle for hosting this interesting competition.</p>\n\n<p>I would like to share some of the key points of my 3rd place (silver) solution :-) It sounds cool right?  (well I love to make that top silver medal the most out of it, forgive me :P)</p>\n\n<h3>1) Cutmix augmentation</h3>\n\n<p>Naturally thinking, cutmix is the best way to deal with this competition. We can cut a part of this image and paste to another image. This idea came off from my mind without knowing its existence academically, which I later found an official paper about it.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1829450%2Fb8053a93d2452d3417e3bfe0100ea953%2Fcutmix.png?generation=1574165517927124&amp;alt=media\" alt=\"\"></p>\n\n<p>How to do that in code?\nI search for some augmentation package, but find it hard to flexibly code it my way. So I decided to do it manually.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1829450%2Fbc80b1cb83404da2281afe27ec2f7f72%2Fcutmixcode.png?generation=1574165662592266&amp;alt=media\" alt=\"\">\nThe different of doing cutmix or not is just an extra section of  <code>__getitem__</code> in data generator. Here, <code>indexes_augment</code> is the random indexes pick from the training data, <code>w_cutmix</code> and <code>h_cutmix</code> are width and height of the crop. So I just get the random starting position of width and height in the original drawn image (<code>X</code>) and insert part of other image (<code>Xc</code>) into it. </p>\n\n<p><strong>Cutmix boosted both LB and CV by 0.004.</strong></p>\n\n<h3>2) Pseudo-label</h3>\n\n<p>Pseudo-label only works if we correctly select good samples, as well as the correct number of samples. I did this by assessing the <code>quality</code> of each predicted validation image by calculating:\n<code>quality = (number of pixels with probability &amp;gt; top) + w*(number of pixels with probability &amp;lt; bot)</code>. Here, <code>top</code> can take values from [.7, .75, .8, .85, .9], <code>bot</code> can take values from [.1, .15, .2, .25, .3], and <code>w</code> is the weight of low-value pixels as compared to high-value pixels, which can be taken from, say, [.1, .5, 1, 2, 10]. </p>\n\n<p>I get the <code>quality</code> of all validation data, rank it, and select <code>nb_samples</code> most confident samples from it, and see the score. I search through a full set of validation data and had a result something like this\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1829450%2Fbdb7de2de3781797e891e7eb8ec93508%2Fpseudo.png?generation=1574166277986726&amp;alt=media\" alt=\"\"></p>\n\n<p>So, I can manually decide <code>bot</code>, <code>top</code>, <code>w</code>, and <code>nb_samples</code> as long as <code>nb_samples</code> are reasonable with the corresponding score. For example, <code>bot=.1</code>, <code>top=.7</code>, <code>w=1</code>, and <code>nb_samples</code>=1000 (with corresponding <code>dice=0.77xx</code>), which means the most 1000 confident predictions out of 5546 train images can have that good dice. Then I can pick up the same ratio of images from test predictions, which is (1000/5546*3698).</p>\n\n<p><strong>Pseudo labelling boosted around 0.003 on both CV and LB.</strong></p>\n\n<h3>3) Estimating private LB distribution and decide to trust CV</h3>\n\n<p>First, I did a test on private LB, based on <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/109793#latest-631950\">this topic</a>. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F43771b2c3fd35e39047f55b552604852%2FLB_probing.png?generation=1569149948996763&amp;alt=media\" alt=\"\"></p>\n\n<p>From that probing result, we need to make 1 assumption:\n\"Train set and full test set should have the same distribution of classes.\"</p>\n\n<p>Then, by submitting each class as empty (others as 1-pixel masks), we can know the percentage of each class in the public LB. Then the assumption I make will allow us to know the percentage of each class in the private LB. The result is:</p>\n\n<p>Train data: Fish 49.85%, Flower 57.35%, Gravel 47.00%, Sugar 32.36%.\nPrivate test data: Fish 49.94%, Flower 56.76%, Gravel 47.20%, Sugar 31.30%.</p>\n\n<p>As you can see, the distributions of private test and train very similar, allowing me to completely trust CV. Therefore during the whole competition, I never probed LB by submissions, but only stick with full k-fold to search for post-processing parameters. <strong>This is important, as it guides the way we do everyday in the competition</strong>. And you can see that I jumped on private LB, and I also selected my possibly best submission.</p>\n\n<p>Finally, I still would like to emphasize again that late sharing should not be encouraged. I have a bad thought that whenever I see the excessive sharers around in future competitions, I would be very disappointed, and discouraged from competing. In other words I am somehow \"scared\" of their existence. </p>\n\n<p>Thanks for reading!</p>",
  "messages": [
    {
      "id": 676697,
      "postDate": "2019-11-19T12:42:07.210Z",
      "content": "<p>First I would like to thank Max Planck Institute and Kaggle for hosting this interesting competition.</p>\n\n<p>I would like to share some of the key points of my 3rd place (silver) solution :-) It sounds cool right?  (well I love to make that top silver medal the most out of it, forgive me :P)</p>\n\n<h3>1) Cutmix augmentation</h3>\n\n<p>Naturally thinking, cutmix is the best way to deal with this competition. We can cut a part of this image and paste to another image. This idea came off from my mind without knowing its existence academically, which I later found an official paper about it.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1829450%2Fb8053a93d2452d3417e3bfe0100ea953%2Fcutmix.png?generation=1574165517927124&amp;alt=media\" alt=\"\"></p>\n\n<p>How to do that in code?\nI search for some augmentation package, but find it hard to flexibly code it my way. So I decided to do it manually.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1829450%2Fbc80b1cb83404da2281afe27ec2f7f72%2Fcutmixcode.png?generation=1574165662592266&amp;alt=media\" alt=\"\">\nThe different of doing cutmix or not is just an extra section of  <code>__getitem__</code> in data generator. Here, <code>indexes_augment</code> is the random indexes pick from the training data, <code>w_cutmix</code> and <code>h_cutmix</code> are width and height of the crop. So I just get the random starting position of width and height in the original drawn image (<code>X</code>) and insert part of other image (<code>Xc</code>) into it. </p>\n\n<p><strong>Cutmix boosted both LB and CV by 0.004.</strong></p>\n\n<h3>2) Pseudo-label</h3>\n\n<p>Pseudo-label only works if we correctly select good samples, as well as the correct number of samples. I did this by assessing the <code>quality</code> of each predicted validation image by calculating:\n<code>quality = (number of pixels with probability &amp;gt; top) + w*(number of pixels with probability &amp;lt; bot)</code>. Here, <code>top</code> can take values from [.7, .75, .8, .85, .9], <code>bot</code> can take values from [.1, .15, .2, .25, .3], and <code>w</code> is the weight of low-value pixels as compared to high-value pixels, which can be taken from, say, [.1, .5, 1, 2, 10]. </p>\n\n<p>I get the <code>quality</code> of all validation data, rank it, and select <code>nb_samples</code> most confident samples from it, and see the score. I search through a full set of validation data and had a result something like this\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1829450%2Fbdb7de2de3781797e891e7eb8ec93508%2Fpseudo.png?generation=1574166277986726&amp;alt=media\" alt=\"\"></p>\n\n<p>So, I can manually decide <code>bot</code>, <code>top</code>, <code>w</code>, and <code>nb_samples</code> as long as <code>nb_samples</code> are reasonable with the corresponding score. For example, <code>bot=.1</code>, <code>top=.7</code>, <code>w=1</code>, and <code>nb_samples</code>=1000 (with corresponding <code>dice=0.77xx</code>), which means the most 1000 confident predictions out of 5546 train images can have that good dice. Then I can pick up the same ratio of images from test predictions, which is (1000/5546*3698).</p>\n\n<p><strong>Pseudo labelling boosted around 0.003 on both CV and LB.</strong></p>\n\n<h3>3) Estimating private LB distribution and decide to trust CV</h3>\n\n<p>First, I did a test on private LB, based on <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/109793#latest-631950\">this topic</a>. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F43771b2c3fd35e39047f55b552604852%2FLB_probing.png?generation=1569149948996763&amp;alt=media\" alt=\"\"></p>\n\n<p>From that probing result, we need to make 1 assumption:\n\"Train set and full test set should have the same distribution of classes.\"</p>\n\n<p>Then, by submitting each class as empty (others as 1-pixel masks), we can know the percentage of each class in the public LB. Then the assumption I make will allow us to know the percentage of each class in the private LB. The result is:</p>\n\n<p>Train data: Fish 49.85%, Flower 57.35%, Gravel 47.00%, Sugar 32.36%.\nPrivate test data: Fish 49.94%, Flower 56.76%, Gravel 47.20%, Sugar 31.30%.</p>\n\n<p>As you can see, the distributions of private test and train very similar, allowing me to completely trust CV. Therefore during the whole competition, I never probed LB by submissions, but only stick with full k-fold to search for post-processing parameters. <strong>This is important, as it guides the way we do everyday in the competition</strong>. And you can see that I jumped on private LB, and I also selected my possibly best submission.</p>\n\n<p>Finally, I still would like to emphasize again that late sharing should not be encouraged. I have a bad thought that whenever I see the excessive sharers around in future competitions, I would be very disappointed, and discouraged from competing. In other words I am somehow \"scared\" of their existence. </p>\n\n<p>Thanks for reading!</p>",
      "rawMarkdown": "First I would like to thank Max Planck Institute and Kaggle for hosting this interesting competition.\n\nI would like to share some of the key points of my 3rd place (silver) solution :-) It sounds cool right?  (well I love to make that top silver medal the most out of it, forgive me :P)\n\n### 1) Cutmix augmentation\nNaturally thinking, cutmix is the best way to deal with this competition. We can cut a part of this image and paste to another image. This idea came off from my mind without knowing its existence academically, which I later found an official paper about it.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1829450%2Fb8053a93d2452d3417e3bfe0100ea953%2Fcutmix.png?generation=1574165517927124&amp;alt=media)\n\nHow to do that in code?\nI search for some augmentation package, but find it hard to flexibly code it my way. So I decided to do it manually.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1829450%2Fbc80b1cb83404da2281afe27ec2f7f72%2Fcutmixcode.png?generation=1574165662592266&amp;alt=media)\nThe different of doing cutmix or not is just an extra section of  `__getitem__` in data generator. Here, `indexes_augment` is the random indexes pick from the training data, `w_cutmix` and `h_cutmix` are width and height of the crop. So I just get the random starting position of width and height in the original drawn image (`X`) and insert part of other image (`Xc`) into it. \n\n**Cutmix boosted both LB and CV by 0.004.**\n\n### 2) Pseudo-label\nPseudo-label only works if we correctly select good samples, as well as the correct number of samples. I did this by assessing the `quality` of each predicted validation image by calculating:\n`quality = (number of pixels with probability &gt; top) + w*(number of pixels with probability &lt; bot)`. Here, `top` can take values from [.7, .75, .8, .85, .9], `bot` can take values from [.1, .15, .2, .25, .3], and `w` is the weight of low-value pixels as compared to high-value pixels, which can be taken from, say, [.1, .5, 1, 2, 10]. \n\nI get the `quality` of all validation data, rank it, and select `nb_samples` most confident samples from it, and see the score. I search through a full set of validation data and had a result something like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1829450%2Fbdb7de2de3781797e891e7eb8ec93508%2Fpseudo.png?generation=1574166277986726&amp;alt=media)\n\nSo, I can manually decide `bot`, `top`, `w`, and `nb_samples` as long as `nb_samples` are reasonable with the corresponding score. For example, `bot=.1`, `top=.7`, `w=1`, and `nb_samples`=1000 (with corresponding `dice=0.77xx`), which means the most 1000 confident predictions out of 5546 train images can have that good dice. Then I can pick up the same ratio of images from test predictions, which is (1000/5546*3698).\n\n**Pseudo labelling boosted around 0.003 on both CV and LB.**\n\n### 3) Estimating private LB distribution and decide to trust CV\nFirst, I did a test on private LB, based on [this topic](https://www.kaggle.com/c/understanding_cloud_organization/discussion/109793#latest-631950). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F43771b2c3fd35e39047f55b552604852%2FLB_probing.png?generation=1569149948996763&amp;alt=media)\n\nFrom that probing result, we need to make 1 assumption:\n\"Train set and full test set should have the same distribution of classes.\"\n\nThen, by submitting each class as empty (others as 1-pixel masks), we can know the percentage of each class in the public LB. Then the assumption I make will allow us to know the percentage of each class in the private LB. The result is:\n\n\nTrain data: Fish 49.85%, Flower 57.35%, Gravel 47.00%, Sugar 32.36%.\nPrivate test data: Fish 49.94%, Flower 56.76%, Gravel 47.20%, Sugar 31.30%.\n\nAs you can see, the distributions of private test and train very similar, allowing me to completely trust CV. Therefore during the whole competition, I never probed LB by submissions, but only stick with full k-fold to search for post-processing parameters. **This is important, as it guides the way we do everyday in the competition**. And you can see that I jumped on private LB, and I also selected my possibly best submission.\n\nFinally, I still would like to emphasize again that late sharing should not be encouraged. I have a bad thought that whenever I see the excessive sharers around in future competitions, I would be very disappointed, and discouraged from competing. In other words I am somehow \"scared\" of their existence. \n\nThanks for reading!",
      "votes": 26
    },
    {
      "id": 678494,
      "postDate": "2019-11-21T13:19:39.840Z",
      "content": "<p>congrats !!</p>",
      "rawMarkdown": "congrats !!",
      "votes": 3
    },
    {
      "id": 677270,
      "postDate": "2019-11-20T01:33:58.817Z",
      "content": "<p>Congrats and thank you for sharing the cutmix augmentation technique. I will try it in the next competition. </p>",
      "rawMarkdown": "Congrats and thank you for sharing the cutmix augmentation technique. I will try it in the next competition. ",
      "votes": 1
    },
    {
      "id": 677016,
      "postDate": "2019-11-19T17:59:45.810Z",
      "content": "<p>Congrats KhaVo. Wow, CutMix augmentation is cool. And your method for selecting Pseudo labeling is ingenious. I plan to use both in the future. Thanks for sharing.</p>\n\n<p>I saw that you jumped 20 spots on the last day. Was it one of these two tricks? </p>",
      "rawMarkdown": "Congrats KhaVo. Wow, CutMix augmentation is cool. And your method for selecting Pseudo labeling is ingenious. I plan to use both in the future. Thanks for sharing.\n\nI saw that you jumped 20 spots on the last day. Was it one of these two tricks? ",
      "votes": 1,
      "replies": [
        {
          "id": 677297,
          "postDate": "2019-11-20T02:31:23.430Z",
          "content": "<p>Thanks Chris. The jump was just from ensembling, which I did only in the last 2 days. Before jumping, my single models scored around 0.669 public LB, after ensemble public 0.674 LB.</p>",
          "rawMarkdown": "Thanks Chris. The jump was just from ensembling, which I did only in the last 2 days. Before jumping, my single models scored around 0.669 public LB, after ensemble public 0.674 LB."
        }
      ]
    },
    {
      "id": 676852,
      "postDate": "2019-11-19T14:47:34.663Z",
      "content": "<p>Congratulations\nGreat Write-Up\nThank you for Sharing your Insights &amp; Approach! <a href=\"/khahuras\">@khahuras</a> </p>",
      "rawMarkdown": "Congratulations\nGreat Write-Up\nThank you for Sharing your Insights &amp; Approach! @khahuras ",
      "votes": 1
    },
    {
      "id": 676707,
      "postDate": "2019-11-19T12:52:10.423Z",
      "content": "<p>Thank you for your writeup. Cutmix always looked promising and I am glad it worked for you. \nI'm curious about your models and loss functions</p>",
      "rawMarkdown": "Thank you for your writeup. Cutmix always looked promising and I am glad it worked for you. \nI'm curious about your models and loss functions",
      "votes": 1,
      "replies": [
        {
          "id": 676710,
          "postDate": "2019-11-19T12:55:24.043Z",
          "content": "<p>My models are not special, since I am completely new to computer vision and segmentation in particular. What worked is that I trained a bunch of different encoders, pseudo-label thresholds, then combine around 10 models, search for best postprocessing parameters through CV. That's simply it.</p>",
          "rawMarkdown": "My models are not special, since I am completely new to computer vision and segmentation in particular. What worked is that I trained a bunch of different encoders, pseudo-label thresholds, then combine around 10 models, search for best postprocessing parameters through CV. That's simply it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 677081,
      "postDate": "2019-11-19T19:14:23.660Z",
      "content": "<p>Hey glad to see cutmix helped you even we tried this type of augmentation in our initial training but model took a lot of time to converge so we dropped it did it happened with you too?</p>",
      "rawMarkdown": "Hey glad to see cutmix helped you even we tried this type of augmentation in our initial training but model took a lot of time to converge so we dropped it did it happened with you too?",
      "votes": 2,
      "replies": [
        {
          "id": 677294,
          "postDate": "2019-11-20T02:30:25.733Z",
          "content": "<p>It needs around 20 epochs I think</p>",
          "rawMarkdown": "It needs around 20 epochs I think",
          "votes": 1
        },
        {
          "id": 677303,
          "postDate": "2019-11-20T02:43:54.943Z",
          "content": "<p>I think then we dropped the idea too early</p>",
          "rawMarkdown": "I think then we dropped the idea too early"
        }
      ]
    },
    {
      "id": 680911,
      "postDate": "2019-11-25T11:45:16.067Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 678492,
      "postDate": "2019-11-21T13:19:13.990Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 678494,
      "author_name": "ravi tanwar",
      "author_url": "",
      "post_date": "2019-11-21T13:19:39.840000",
      "content": "<p>congrats !!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 677270,
      "author_name": "sabari nathan",
      "author_url": "",
      "post_date": "2019-11-20T01:33:58.817000",
      "content": "<p>Congrats and thank you for sharing the cutmix augmentation technique. I will try it in the next competition. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 677016,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2019-11-19T17:59:45.810000",
      "content": "<p>Congrats KhaVo. Wow, CutMix augmentation is cool. And your method for selecting Pseudo labeling is ingenious. I plan to use both in the future. Thanks for sharing.</p>\n\n<p>I saw that you jumped 20 spots on the last day. Was it one of these two tricks? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 677297,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-11-20T02:31:23.430000",
          "content": "<p>Thanks Chris. The jump was just from ensembling, which I did only in the last 2 days. Before jumping, my single models scored around 0.669 public LB, after ensemble public 0.674 LB.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676852,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2019-11-19T14:47:34.663000",
      "content": "<p>Congratulations\nGreat Write-Up\nThank you for Sharing your Insights &amp; Approach! <a href=\"/khahuras\">@khahuras</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 676707,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2019-11-19T12:52:10.423000",
      "content": "<p>Thank you for your writeup. Cutmix always looked promising and I am glad it worked for you. \nI'm curious about your models and loss functions</p>",
      "votes": 1,
      "replies": [
        {
          "id": 676710,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-11-19T12:55:24.043000",
          "content": "<p>My models are not special, since I am completely new to computer vision and segmentation in particular. What worked is that I trained a bunch of different encoders, pseudo-label thresholds, then combine around 10 models, search for best postprocessing parameters through CV. That's simply it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 677081,
      "author_name": "Udbhav Bamba",
      "author_url": "",
      "post_date": "2019-11-19T19:14:23.660000",
      "content": "<p>Hey glad to see cutmix helped you even we tried this type of augmentation in our initial training but model took a lot of time to converge so we dropped it did it happened with you too?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 677294,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-11-20T02:30:25.733000",
          "content": "<p>It needs around 20 epochs I think</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 677303,
          "author_name": "Udbhav Bamba",
          "author_url": "",
          "post_date": "2019-11-20T02:43:54.943000",
          "content": "<p>I think then we dropped the idea too early</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 680911,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-25T11:45:16.067000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 678492,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-21T13:19:13.990000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "676697": "First I would like to thank Max Planck Institute and Kaggle for hosting this interesting competition.\n\nI would like to share some of the key points of my 3rd place (silver) solution :-) It sounds cool right?  (well I love to make that top silver medal the most out of it, forgive me :P)\n\n### 1) Cutmix augmentation\nNaturally thinking, cutmix is the best way to deal with this competition. We can cut a part of this image and paste to another image. This idea came off from my mind without knowing its existence academically, which I later found an official paper about it.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1829450%2Fb8053a93d2452d3417e3bfe0100ea953%2Fcutmix.png?generation=1574165517927124&amp;alt=media)\n\nHow to do that in code?\nI search for some augmentation package, but find it hard to flexibly code it my way. So I decided to do it manually.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1829450%2Fbc80b1cb83404da2281afe27ec2f7f72%2Fcutmixcode.png?generation=1574165662592266&amp;alt=media)\nThe different of doing cutmix or not is just an extra section of  `__getitem__` in data generator. Here, `indexes_augment` is the random indexes pick from the training data, `w_cutmix` and `h_cutmix` are width and height of the crop. So I just get the random starting position of width and height in the original drawn image (`X`) and insert part of other image (`Xc`) into it. \n\n**Cutmix boosted both LB and CV by 0.004.**\n\n### 2) Pseudo-label\nPseudo-label only works if we correctly select good samples, as well as the correct number of samples. I did this by assessing the `quality` of each predicted validation image by calculating:\n`quality = (number of pixels with probability &gt; top) + w*(number of pixels with probability &lt; bot)`. Here, `top` can take values from [.7, .75, .8, .85, .9], `bot` can take values from [.1, .15, .2, .25, .3], and `w` is the weight of low-value pixels as compared to high-value pixels, which can be taken from, say, [.1, .5, 1, 2, 10]. \n\nI get the `quality` of all validation data, rank it, and select `nb_samples` most confident samples from it, and see the score. I search through a full set of validation data and had a result something like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1829450%2Fbdb7de2de3781797e891e7eb8ec93508%2Fpseudo.png?generation=1574166277986726&amp;alt=media)\n\nSo, I can manually decide `bot`, `top`, `w`, and `nb_samples` as long as `nb_samples` are reasonable with the corresponding score. For example, `bot=.1`, `top=.7`, `w=1`, and `nb_samples`=1000 (with corresponding `dice=0.77xx`), which means the most 1000 confident predictions out of 5546 train images can have that good dice. Then I can pick up the same ratio of images from test predictions, which is (1000/5546*3698).\n\n**Pseudo labelling boosted around 0.003 on both CV and LB.**\n\n### 3) Estimating private LB distribution and decide to trust CV\nFirst, I did a test on private LB, based on [this topic](https://www.kaggle.com/c/understanding_cloud_organization/discussion/109793#latest-631950). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F43771b2c3fd35e39047f55b552604852%2FLB_probing.png?generation=1569149948996763&amp;alt=media)\n\nFrom that probing result, we need to make 1 assumption:\n\"Train set and full test set should have the same distribution of classes.\"\n\nThen, by submitting each class as empty (others as 1-pixel masks), we can know the percentage of each class in the public LB. Then the assumption I make will allow us to know the percentage of each class in the private LB. The result is:\n\n\nTrain data: Fish 49.85%, Flower 57.35%, Gravel 47.00%, Sugar 32.36%.\nPrivate test data: Fish 49.94%, Flower 56.76%, Gravel 47.20%, Sugar 31.30%.\n\nAs you can see, the distributions of private test and train very similar, allowing me to completely trust CV. Therefore during the whole competition, I never probed LB by submissions, but only stick with full k-fold to search for post-processing parameters. **This is important, as it guides the way we do everyday in the competition**. And you can see that I jumped on private LB, and I also selected my possibly best submission.\n\nFinally, I still would like to emphasize again that late sharing should not be encouraged. I have a bad thought that whenever I see the excessive sharers around in future competitions, I would be very disappointed, and discouraged from competing. In other words I am somehow \"scared\" of their existence. \n\nThanks for reading!",
    "678494": "congrats !!",
    "677270": "Congrats and thank you for sharing the cutmix augmentation technique. I will try it in the next competition. ",
    "677016": "Congrats KhaVo. Wow, CutMix augmentation is cool. And your method for selecting Pseudo labeling is ingenious. I plan to use both in the future. Thanks for sharing.\n\nI saw that you jumped 20 spots on the last day. Was it one of these two tricks? ",
    "676852": "Congratulations\nGreat Write-Up\nThank you for Sharing your Insights &amp; Approach! @khahuras ",
    "676707": "Thank you for your writeup. Cutmix always looked promising and I am glad it worked for you. \nI'm curious about your models and loss functions",
    "677081": "Hey glad to see cutmix helped you even we tried this type of augmentation in our initial training but model took a lot of time to converge so we dropped it did it happened with you too?",
    "680911": "",
    "678492": ""
  }
}