{
  "id": 118017,
  "title": "6th simple solution, pre-training, single model private 0.66927",
  "url": "/competitions/understanding_cloud_organization/discussion/118017",
  "author_name": "Limerobot",
  "post_date": "2019-11-19T07:59:21.451000",
  "votes": 55,
  "comment_count": 45,
  "views": 0,
  "content": "<p>First of all, I would like to thank the hosting organization that hosted this competition and Kaggle. Like any competition, this competition was also hot until the end.  So, I want to congratulate Kagglers who struggled until the end of this competition.</p>\n\n<p>I will summarize and write down the part of my solution that you will be interested in. It's <code>pre-training</code></p>\n\n<h1>pre-training</h1>\n\n<p>The challenge of this competition is to segment according to the shape of the cloud.  Therefore, I tried to pre-train the model to learn the shape of the cloud.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F7e5d1fb4cbda36a5bcad97879a424176%2F1st-training.png?generation=1574148986382141&amp;alt=media\" alt=\"\"></p>\n\n<p>Because the clouds are white, I generated <code>cloud_mask</code> with the threshold of \"pixel &gt; 115\". Then, I used it as a label. (Since the total number of image files is 9244, the cloud_mask also generates 9244.)</p>\n\n<p>After pre-training, I tried a 2nd-stage training.\nPre-trained(1st stage-training) model are used as the initial value of 2nd-stage model weights.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F6ed4418967a6461295eed180c047feda%2F2nd-training.png?generation=1574149060524535&amp;alt=media\" alt=\"\"></p>\n\n<p>This training process boosted my CV 0.005~0.01. So, my single model score is as follows.</p>\n\n<p>| model | private | public |\n| --- | --- | --- |\n| efficientnet-b4, unet | 0.66927 | 0.67437 |\n| efficientnet-b4, fpn | 0.66827 | 0.67508 |</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F95009ca5bb429732cb523d971978e0fd%2F.png?generation=1574149018151413&amp;alt=media\" alt=\"\"></p>\n\n<p>The rest is not special, so I'll skip the description. 😁 \nThanks for your reading!</p>",
  "messages": [
    {
      "id": 676467,
      "postDate": "2019-11-19T07:59:21.450Z",
      "content": "<p>First of all, I would like to thank the hosting organization that hosted this competition and Kaggle. Like any competition, this competition was also hot until the end.  So, I want to congratulate Kagglers who struggled until the end of this competition.</p>\n\n<p>I will summarize and write down the part of my solution that you will be interested in. It's <code>pre-training</code></p>\n\n<h1>pre-training</h1>\n\n<p>The challenge of this competition is to segment according to the shape of the cloud.  Therefore, I tried to pre-train the model to learn the shape of the cloud.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F7e5d1fb4cbda36a5bcad97879a424176%2F1st-training.png?generation=1574148986382141&amp;alt=media\" alt=\"\"></p>\n\n<p>Because the clouds are white, I generated <code>cloud_mask</code> with the threshold of \"pixel &gt; 115\". Then, I used it as a label. (Since the total number of image files is 9244, the cloud_mask also generates 9244.)</p>\n\n<p>After pre-training, I tried a 2nd-stage training.\nPre-trained(1st stage-training) model are used as the initial value of 2nd-stage model weights.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F6ed4418967a6461295eed180c047feda%2F2nd-training.png?generation=1574149060524535&amp;alt=media\" alt=\"\"></p>\n\n<p>This training process boosted my CV 0.005~0.01. So, my single model score is as follows.</p>\n\n<p>| model | private | public |\n| --- | --- | --- |\n| efficientnet-b4, unet | 0.66927 | 0.67437 |\n| efficientnet-b4, fpn | 0.66827 | 0.67508 |</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F95009ca5bb429732cb523d971978e0fd%2F.png?generation=1574149018151413&amp;alt=media\" alt=\"\"></p>\n\n<p>The rest is not special, so I'll skip the description. 😁 \nThanks for your reading!</p>",
      "rawMarkdown": "First of all, I would like to thank the hosting organization that hosted this competition and Kaggle. Like any competition, this competition was also hot until the end.  So, I want to congratulate Kagglers who struggled until the end of this competition.\n\nI will summarize and write down the part of my solution that you will be interested in. It's `pre-training`\n\n# pre-training\nThe challenge of this competition is to segment according to the shape of the cloud.  Therefore, I tried to pre-train the model to learn the shape of the cloud.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F7e5d1fb4cbda36a5bcad97879a424176%2F1st-training.png?generation=1574148986382141&amp;alt=media)\n\nBecause the clouds are white, I generated `cloud_mask` with the threshold of \"pixel &gt; 115\". Then, I used it as a label. (Since the total number of image files is 9244, the cloud_mask also generates 9244.)\n\nAfter pre-training, I tried a 2nd-stage training.\nPre-trained(1st stage-training) model are used as the initial value of 2nd-stage model weights.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F6ed4418967a6461295eed180c047feda%2F2nd-training.png?generation=1574149060524535&amp;alt=media)\n\nThis training process boosted my CV 0.005~0.01. So, my single model score is as follows.\n\n| model | private | public |\n| --- | --- | --- |\n| efficientnet-b4, unet | 0.66927 | 0.67437 |\n| efficientnet-b4, fpn | 0.66827 | 0.67508 |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F95009ca5bb429732cb523d971978e0fd%2F.png?generation=1574149018151413&amp;alt=media)\n\nThe rest is not special, so I'll skip the description. 😁 \nThanks for your reading!\n",
      "votes": 54
    },
    {
      "id": 676680,
      "postDate": "2019-11-19T12:18:02.260Z",
      "content": "<p>Congratulations, the pre-training strategy is very smart.</p>",
      "rawMarkdown": "Congratulations, the pre-training strategy is very smart.",
      "votes": 1,
      "replies": [
        {
          "id": 676894,
          "postDate": "2019-11-19T15:28:52.453Z",
          "content": "<p>Thanks for your congratulations!</p>",
          "rawMarkdown": "Thanks for your congratulations!"
        }
      ]
    },
    {
      "id": 681494,
      "postDate": "2019-11-26T07:01:17.917Z",
      "content": "<p>Congratulation!</p>\n\n<p>How do you find out this way?   I think pixel &gt; 115 is very useful for reducing the noise(illumination）， that increase the quanlity of learning.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1965832%2F01d8ed5d1b73087170a4c97ee634c216%2Forg.png?generation=1574751538876480&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1965832%2Fc80ada396099f2f2b866535d1e5aa48a%2Fnew.png?generation=1574751613493251&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Congratulation!\n\nHow do you find out this way?   I think pixel &gt; 115 is very useful for reducing the noise(illumination）， that increase the quanlity of learning.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1965832%2F01d8ed5d1b73087170a4c97ee634c216%2Forg.png?generation=1574751538876480&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1965832%2Fc80ada396099f2f2b866535d1e5aa48a%2Fnew.png?generation=1574751613493251&amp;alt=media)\n"
    },
    {
      "id": 677586,
      "postDate": "2019-11-20T10:58:37.227Z",
      "content": "<p>Very clever!</p>",
      "rawMarkdown": "Very clever!",
      "replies": [
        {
          "id": 678055,
          "postDate": "2019-11-21T00:28:37.367Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 677485,
      "postDate": "2019-11-20T08:22:25.087Z",
      "content": "<p><a href=\"/limerobot\">@limerobot</a> Congratulation on your solid solo gold, and thanks for sharing!!</p>\n\n<p>Allow me to ask a bit more details : </p>\n\n<p>(1) what are your optimizers and learning rates on the 1st and 2nd stages?\nIn order to be the most effective, do we have to carefully adjust small learning rates in the 2nd-stage? (or do we have to freeze some layers first?)</p>\n\n<p><strong>(UPDATED)</strong>\n(2) in the 1st-stage, what is the label of the classification head?</p>\n\n<p>(3) I am re-thinking about the 1st-stage, isn't it should be easy for neural network to learn the cloud mask rule? (pixel &gt;115) . (Even shallow networks should be able to learn this? )</p>\n\n<p>PS. i notice many dog guys now change their avatars to cats ... I have no idea why this happen since I am a dog fan.</p>",
      "rawMarkdown": "@limerobot Congratulation on your solid solo gold, and thanks for sharing!!\n\nAllow me to ask a bit more details : \n\n(1) what are your optimizers and learning rates on the 1st and 2nd stages?\nIn order to be the most effective, do we have to carefully adjust small learning rates in the 2nd-stage? (or do we have to freeze some layers first?)\n\n**(UPDATED)**\n(2) in the 1st-stage, what is the label of the classification head?\n\n(3) I am re-thinking about the 1st-stage, isn't it should be easy for neural network to learn the cloud mask rule? (pixel &gt;115) . (Even shallow networks should be able to learn this? )\n\nPS. i notice many dog guys now change their avatars to cats ... I have no idea why this happen since I am a dog fan.",
      "replies": [
        {
          "id": 677491,
          "postDate": "2019-11-20T08:26:59.460Z",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> it is funny, maybe next time, you can see the dogs come back.</p>",
          "rawMarkdown": "@ratthachat it is funny, maybe next time, you can see the dogs come back.",
          "votes": 2
        },
        {
          "id": 677818,
          "postDate": "2019-11-20T16:08:38.230Z",
          "content": "<blockquote>\n  <p>(3) I am re-thinking about the 1st-stage, isn't it should be easy for neural network to learn the cloud mask rule? (pixel &gt;115) . (Even shallow networks should be able to learn this? )</p>\n</blockquote>\n\n<p>The main challenge with neural networks is training. For example, a fully connected neural network can be an exact copy of a convolutional neural network with the correct weights, but a fully connected neural network will never find those weights with gradient descent learning. Similarily pretraining a network is a trick to encourage a neural network to find weights it would not find otherwise. Because you are beginning your gradient descent learning from a different starting point. </p>",
          "rawMarkdown": "&gt; (3) I am re-thinking about the 1st-stage, isn't it should be easy for neural network to learn the cloud mask rule? (pixel &gt;115) . (Even shallow networks should be able to learn this? )\n\nThe main challenge with neural networks is training. For example, a fully connected neural network can be an exact copy of a convolutional neural network with the correct weights, but a fully connected neural network will never find those weights with gradient descent learning. Similarily pretraining a network is a trick to encourage a neural network to find weights it would not find otherwise. Because you are beginning your gradient descent learning from a different starting point. ",
          "votes": 1
        },
        {
          "id": 678053,
          "postDate": "2019-11-21T00:27:05.917Z",
          "content": "<p>Thanks for your congratulations!</p>\n\n<p><code>(1) what are your optimizers and learning rates on the 1st and 2nd stages?\nIn order to be the most effective, do we have to carefully adjust small learning rates in the 2nd-stage? (or do we have to freeze some layers first?)</code></p>\n\n<p><strong>1st-stage</strong>\noptimizer: AdamW\nscheduler: WarmupLinearSchedule\nlearning_rate: 5e-04</p>\n\n<p><strong>2nd-stage</strong>\noptimizer: AdamW\nscheduler: MultiStepLR (milestones=[5, 10], gamma=0.1)\nlearning_rate: 5e-04</p>\n\n<p>In my experiments, the small initial learning rate(5e-05) was not good.\n(Like the BERT encoder, I think the fine-tuning stage(2nd stage) is better to learn the entire weight of the model.)</p>\n\n<p><code>(2) in the 1st-stage, what is the label of the classification head?</code></p>\n\n<p>Both the 1st and 2nd stages used the same model architecture; therefore, the classification head is not learned at 1st-stage. (no need label)</p>\n\n<p><code>(3) I am re-thinking about the 1st-stage, isn't it should be easy for neural network to learn the cloud mask rule? (pixel &amp;gt;115) . (Even shallow networks should be able to learn this? )</code></p>\n\n<p>Yeah, I think it's easy. So I used only 5 epochs(with WarmupLinearSchedule) to prevent overfitting.</p>\n\n<p>PS. My wife loves cats more than dogs. 😂 </p>",
          "rawMarkdown": "Thanks for your congratulations!\n\n`(1) what are your optimizers and learning rates on the 1st and 2nd stages?\nIn order to be the most effective, do we have to carefully adjust small learning rates in the 2nd-stage? (or do we have to freeze some layers first?)`\n\n**1st-stage**\noptimizer: AdamW\nscheduler: WarmupLinearSchedule\nlearning_rate: 5e-04\n\n**2nd-stage**\noptimizer: AdamW\nscheduler: MultiStepLR (milestones=[5, 10], gamma=0.1)\nlearning_rate: 5e-04\n\nIn my experiments, the small initial learning rate(5e-05) was not good.\n(Like the BERT encoder, I think the fine-tuning stage(2nd stage) is better to learn the entire weight of the model.)\n\n`(2) in the 1st-stage, what is the label of the classification head?`\n\nBoth the 1st and 2nd stages used the same model architecture; therefore, the classification head is not learned at 1st-stage. (no need label)\n\n`(3) I am re-thinking about the 1st-stage, isn't it should be easy for neural network to learn the cloud mask rule? (pixel &gt;115) . (Even shallow networks should be able to learn this? )`\n\nYeah, I think it's easy. So I used only 5 epochs(with WarmupLinearSchedule) to prevent overfitting.\n\nPS. My wife loves cats more than dogs. 😂 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 677283,
      "postDate": "2019-11-20T01:49:29.057Z",
      "content": "<p>Congratulations, your method to get initial weights is impressive! learned.</p>",
      "rawMarkdown": "Congratulations, your method to get initial weights is impressive! learned.",
      "replies": [
        {
          "id": 678054,
          "postDate": "2019-11-21T00:28:27.363Z",
          "content": "<p>Thanks for your congratulations!</p>",
          "rawMarkdown": "Thanks for your congratulations!"
        }
      ]
    },
    {
      "id": 677211,
      "postDate": "2019-11-19T23:23:50.120Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "replies": [
        {
          "id": 677247,
          "postDate": "2019-11-20T00:35:07.120Z",
          "content": "<p>Thanks for your congrats! </p>",
          "rawMarkdown": "Thanks for your congrats! "
        }
      ]
    },
    {
      "id": 676957,
      "postDate": "2019-11-19T16:47:34.050Z",
      "content": "<p>Congrats for the solo gold. And appreciate for the well-written post.  Creating mast is very interesting. \nCould you explain why did you choose \"115\" as the threshold ?</p>",
      "rawMarkdown": "Congrats for the solo gold. And appreciate for the well-written post.  Creating mast is very interesting. \nCould you explain why did you choose \"115\" as the threshold ?",
      "replies": [
        {
          "id": 677246,
          "postDate": "2019-11-20T00:35:00.370Z",
          "content": "<p>Thanks for your congrats! \nI just chose roughly. 😁  (A rule of thumb, 눈대중)</p>",
          "rawMarkdown": "Thanks for your congrats! \nI just chose roughly. 😁  (A rule of thumb, 눈대중)",
          "votes": 1
        }
      ]
    },
    {
      "id": 676948,
      "postDate": "2019-11-19T16:36:46.203Z",
      "content": "<p>Congratulations on solo Gold. Pretraining on cloud shape is brilliant !! You discovered a way to use the test data during training. I also like how your model has a classification and segmentation head. And you balanced all the losses nicely by using weights 0.2, 0.2, and 0.8</p>",
      "rawMarkdown": "Congratulations on solo Gold. Pretraining on cloud shape is brilliant !! You discovered a way to use the test data during training. I also like how your model has a classification and segmentation head. And you balanced all the losses nicely by using weights 0.2, 0.2, and 0.8",
      "replies": [
        {
          "id": 677245,
          "postDate": "2019-11-20T00:33:17.363Z",
          "content": "<p>Thanks for your congratulations! 😊 </p>",
          "rawMarkdown": "Thanks for your congratulations! 😊 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 676655,
      "postDate": "2019-11-19T11:49:50.850Z",
      "content": "<p>congratulations, limerobot, good method</p>\n\n<p>for these part, 'label' is cloud in each box, and then generated cloud mask, am I right?</p>\n\n<p>for example, for pic1, it has two box, flower and sugar, and you would generate two cloud mask flower and sugar in each box?</p>\n\n<blockquote>\n  <p>Because the clouds are white, I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label. (Since the total number of image files is 9244, the cloud_mask also generates 9244.)</p>\n</blockquote>",
      "rawMarkdown": "congratulations, limerobot, good method\n\nfor these part, 'label' is cloud in each box, and then generated cloud mask, am I right?\n\nfor example, for pic1, it has two box, flower and sugar, and you would generate two cloud mask flower and sugar in each box?\n\n&gt; Because the clouds are white, I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label. (Since the total number of image files is 9244, the cloud_mask also generates 9244.)",
      "replies": [
        {
          "id": 676890,
          "postDate": "2019-11-19T15:24:23.997Z",
          "content": "<p>No, I just used this code to generate a cloud_mask.\n<code>cloud_mask = (image &gt; 115).astype(int)</code></p>\n\n<ul>\n<li><p>image\n![image](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2Fbfb36270c4b0ef8ff920d3d633519fea%2F009e2f3.jpg?generation=1574176746316088&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2Fbfb36270c4b0ef8ff920d3d633519fea%2F009e2f3.jpg?generation=1574176746316088&amp;alt=media</a> =200x300)</p></li>\n<li><p>cloud_mask = (image &gt; 115).astype(int)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F1482ca9ded9080c3200e0607852587ea%2F009e2f3_true%20(1\" alt=\"image\">.jpg?generation=1574152348586165&amp;alt=media =200x300)</p></li>\n</ul>",
          "rawMarkdown": "No, I just used this code to generate a cloud_mask.\n`cloud_mask = (image &gt; 115).astype(int)`\n\n- image\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2Fbfb36270c4b0ef8ff920d3d633519fea%2F009e2f3.jpg?generation=1574176746316088&amp;alt=media =200x300)\n\n- cloud_mask = (image &gt; 115).astype(int)\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F1482ca9ded9080c3200e0607852587ea%2F009e2f3_true%20(1).jpg?generation=1574152348586165&amp;alt=media =200x300)\n"
        }
      ]
    },
    {
      "id": 676622,
      "postDate": "2019-11-19T11:12:04.413Z",
      "content": "<p>Congratulation on the gold medal!</p>",
      "rawMarkdown": "Congratulation on the gold medal!",
      "replies": [
        {
          "id": 676878,
          "postDate": "2019-11-19T15:10:11.270Z",
          "content": "<p>wow~ thanks!</p>",
          "rawMarkdown": "wow~ thanks!"
        }
      ]
    },
    {
      "id": 676611,
      "postDate": "2019-11-19T10:53:10.397Z",
      "content": "<p>Congratulations！！！\nPre-training stage is impressive.</p>",
      "rawMarkdown": "Congratulations！！！\nPre-training stage is impressive.",
      "replies": [
        {
          "id": 676876,
          "postDate": "2019-11-19T15:09:53.877Z",
          "content": "<p>Thanks. I congratulate you too :)</p>",
          "rawMarkdown": "Thanks. I congratulate you too :)"
        }
      ]
    },
    {
      "id": 676610,
      "postDate": "2019-11-19T10:52:02.167Z",
      "content": "<p>Very interesting, <a href=\"/limerobot\">@limerobot</a> . Have you tried to put your <code>cloud_mask</code> as fourth channel to original images? </p>",
      "rawMarkdown": "Very interesting, @limerobot . Have you tried to put your `cloud_mask` as fourth channel to original images? ",
      "replies": [
        {
          "id": 676874,
          "postDate": "2019-11-19T15:07:37.220Z",
          "content": "<p>Yep, you're right.</p>",
          "rawMarkdown": "Yep, you're right."
        }
      ]
    },
    {
      "id": 676607,
      "postDate": "2019-11-19T10:49:45.390Z",
      "content": "<p>very good training skills!\nNice work!</p>",
      "rawMarkdown": "very good training skills!\nNice work!",
      "replies": [
        {
          "id": 676615,
          "postDate": "2019-11-19T11:02:42.810Z",
          "content": "<p>this is just an idea, and i am not sure if it would work.</p>\n\n<p>\"I generated cloud_mask with the threshold of \"pixel &gt; 115\"</p>\n\n<p>you can try multi label for pretraining: e.g. \"x&gt;115\", \"80&gt;x&gt;155\"... it can capture some fine cloud? </p>",
          "rawMarkdown": "this is just an idea, and i am not sure if it would work.\n\n\"I generated cloud\\_mask with the threshold of \"pixel &gt; 115\"\n\nyou can try multi label for pretraining: e.g. \"x&gt;115\", \"80&gt;x&gt;155\"... it can capture some fine cloud? ",
          "votes": 1
        },
        {
          "id": 676633,
          "postDate": "2019-11-19T11:27:50.277Z",
          "content": "<p>&gt; Because the clouds are white, I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label.</p>\n\n<p><a href=\"/hengck23\">@hengck23</a> , Could you help me understand his pipeline as I am finding it hard to understand his Stage1 training pipeline?</p>",
          "rawMarkdown": "&gt; Because the clouds are white, I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label.\n\n\n\n@hengck23 , Could you help me understand his pipeline as I am finding it hard to understand his Stage1 training pipeline?"
        },
        {
          "id": 676666,
          "postDate": "2019-11-19T12:06:29.270Z",
          "content": "<p>He used the original images and created masks without any neural networks. He just took each image and said <code>new_mask = (original_image&gt;115).astype(int)</code> where the original image has pixel values between 0 and 255. (By doing this he captured where the clouds are located but not what type of cloud). Then he trained his first epochs on these masks then switched to real masks.</p>",
          "rawMarkdown": "He used the original images and created masks without any neural networks. He just took each image and said `new_mask = (original_image&gt;115).astype(int)` where the original image has pixel values between 0 and 255. (By doing this he captured where the clouds are located but not what type of cloud). Then he trained his first epochs on these masks then switched to real masks.",
          "votes": 2
        },
        {
          "id": 676677,
          "postDate": "2019-11-19T12:16:11.347Z",
          "content": "<p>Thank you Chris for making this clear. Now I understand what he meant when he says\n&gt;  I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label.</p>\n\n<p>In Stage1, he trained using <code>train+test</code> images and their corresponding labels obtained by doing <code>(image&amp;gt;115).astype(int)</code> \nFor stage2, he used Stage1 model to initialize training for <code>train images</code> and their corresponding <code>masks</code>. \nI hope this helps others like me who were confused about his Stage1 pipeline</p>",
          "rawMarkdown": "Thank you Chris for making this clear. Now I understand what he meant when he says\n&gt;  I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label.\n\nIn Stage1, he trained using `train+test` images and their corresponding labels obtained by doing `(image&gt;115).astype(int)` \nFor stage2, he used Stage1 model to initialize training for `train images` and their corresponding `masks`. \nI hope this helps others like me who were confused about his Stage1 pipeline"
        },
        {
          "id": 676693,
          "postDate": "2019-11-19T12:39:35.883Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I was trying to play with the pixel values this way . I didn't think of pretraining , more of post processing . However , Sugar always threw me off .. it had pixels values starting 60+ .. so I had dropped that idea . Does &gt; 115 capture all clouds ? </p>",
          "rawMarkdown": "@cdeotte I was trying to play with the pixel values this way . I didn't think of pretraining , more of post processing . However , Sugar always threw me off .. it had pixels values starting 60+ .. so I had dropped that idea . Does &gt; 115 capture all clouds ? "
        },
        {
          "id": 676875,
          "postDate": "2019-11-19T15:09:11.710Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thank you for your clarification! 👍 </p>",
          "rawMarkdown": "@cdeotte Thank you for your clarification! 👍 "
        },
        {
          "id": 676893,
          "postDate": "2019-11-19T15:27:17.170Z",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a>  No, &gt; 115 doesn't seem to capture all the clouds.</p>",
          "rawMarkdown": "@phoenix9032  No, &gt; 115 doesn't seem to capture all the clouds."
        },
        {
          "id": 676900,
          "postDate": "2019-11-19T15:31:01.330Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> Thanks for your congratulations, and I didn’t think like that. Interesting!</p>",
          "rawMarkdown": "@hengck23 Thanks for your congratulations, and I didn’t think like that. Interesting!"
        }
      ]
    },
    {
      "id": 676594,
      "postDate": "2019-11-19T10:34:17.520Z",
      "content": "<p>Congratulations 🎉  and thanks for sharing your solution!</p>",
      "rawMarkdown": "Congratulations 🎉  and thanks for sharing your solution!",
      "replies": [
        {
          "id": 676904,
          "postDate": "2019-11-19T15:36:30.897Z",
          "content": "<p>Thanks for your congratulations!</p>",
          "rawMarkdown": "Thanks for your congratulations!"
        }
      ]
    },
    {
      "id": 676487,
      "postDate": "2019-11-19T08:17:29.057Z",
      "content": "<p>Congratulations.\nPretrained on white cloud is interesting.</p>",
      "rawMarkdown": "Congratulations.\nPretrained on white cloud is interesting.",
      "replies": [
        {
          "id": 676511,
          "postDate": "2019-11-19T08:45:37.770Z",
          "content": "<p>I congratulate you too :)</p>",
          "rawMarkdown": "I congratulate you too :)"
        }
      ]
    },
    {
      "id": 676472,
      "postDate": "2019-11-19T08:05:54.457Z",
      "content": "<p>Thank you for sharing and congrats for your gold finish.\nI have few things to ask about your approach.\n&gt; Therefore, I tried to pre-train the model to learn the shape of the cloud.</p>\n\n<p>In stage1, the label is the mask? or the image itself?</p>\n\n<p>&gt; Because the clouds are white, I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label.</p>\n\n<p>Could you elaborate on this part? this increased the number of training samples for 2nd stage?</p>",
      "rawMarkdown": "Thank you for sharing and congrats for your gold finish.\nI have few things to ask about your approach.\n&gt; Therefore, I tried to pre-train the model to learn the shape of the cloud.\n\nIn stage1, the label is the mask? or the image itself?\n\n&gt; Because the clouds are white, I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label.\n\nCould you elaborate on this part? this increased the number of training samples for 2nd stage?",
      "replies": [
        {
          "id": 676497,
          "postDate": "2019-11-19T08:33:05.127Z",
          "content": "<p>I'm sorry I didn't explain it clearly. 😂 </p>\n\n<blockquote>\n  <p>In stage1, the label is the mask? or the image itself?</p>\n</blockquote>\n\n<p>The label is the mask. <br>\nOne example of mask\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F1482ca9ded9080c3200e0607852587ea%2F009e2f3_true%20(1\" alt=\"\">.jpg?generation=1574152348586165&amp;alt=media)</p>\n\n<blockquote>\n  <p>Could you elaborate on this part? this increased the number of training samples for 2nd stage?</p>\n</blockquote>\n\n<p>Pre-trained(1st stage-training) model are used as the initial value of 2nd-stage model weights.</p>",
          "rawMarkdown": "I'm sorry I didn't explain it clearly. 😂 \n\n&gt; In stage1, the label is the mask? or the image itself?\n\nThe label is the mask.  \nOne example of mask\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F1482ca9ded9080c3200e0607852587ea%2F009e2f3_true%20(1).jpg?generation=1574152348586165&amp;alt=media)\n\n\n&gt; Could you elaborate on this part? this increased the number of training samples for 2nd stage?\n\nPre-trained(1st stage-training) model are used as the initial value of 2nd-stage model weights.",
          "votes": 1
        },
        {
          "id": 676519,
          "postDate": "2019-11-19T08:56:25.777Z",
          "content": "<p>Ok!! I drew a sketch of your pipeline(skipping classification part of stage2) based on my understanding of your writeup....could you please confirm?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2Fc34eb0c5ea50c5f789a85b72c5520a85%2Fmask_limerobot.jpg?generation=1574153776348792&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Ok!! I drew a sketch of your pipeline(skipping classification part of stage2) based on my understanding of your writeup....could you please confirm?![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2Fc34eb0c5ea50c5f789a85b72c5520a85%2Fmask_limerobot.jpg?generation=1574153776348792&amp;alt=media)\n"
        },
        {
          "id": 676555,
          "postDate": "2019-11-19T09:49:03.410Z",
          "content": "<p>Yeah, that's right.</p>\n\n<ul>\n<li>1st-stage training: all images, generatedmask </li>\n<li>2nd-stage training: train images, realmask</li>\n</ul>",
          "rawMarkdown": "Yeah, that's right.\n\n- 1st-stage training: all images, generatedmask \n- 2nd-stage training: train images, realmask",
          "votes": 1
        },
        {
          "id": 676563,
          "postDate": "2019-11-19T10:01:13.873Z",
          "content": "<p>Thank you for your clarification. If what I understood is correct(as you confirmed it), where is this part?</p>\n\n<blockquote>\n  <p>Then, I used it as a label.</p>\n</blockquote>",
          "rawMarkdown": "Thank you for your clarification. If what I understood is correct(as you confirmed it), where is this part?\n&gt; Then, I used it as a label."
        }
      ]
    },
    {
      "id": 676525,
      "postDate": "2019-11-19T09:07:35.340Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 676905,
          "postDate": "2019-11-19T15:36:37.640Z",
          "content": "<p>Thanks for your congratulations!</p>",
          "rawMarkdown": "Thanks for your congratulations!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 676680,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-11-19T12:18:02.260000",
      "content": "<p>Congratulations, the pre-training strategy is very smart.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 676894,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T15:28:52.453000",
          "content": "<p>Thanks for your congratulations!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 681494,
      "author_name": "imagination",
      "author_url": "",
      "post_date": "2019-11-26T07:01:17.917000",
      "content": "<p>Congratulation!</p>\n\n<p>How do you find out this way?   I think pixel &gt; 115 is very useful for reducing the noise(illumination）， that increase the quanlity of learning.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1965832%2F01d8ed5d1b73087170a4c97ee634c216%2Forg.png?generation=1574751538876480&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1965832%2Fc80ada396099f2f2b866535d1e5aa48a%2Fnew.png?generation=1574751613493251&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 677586,
      "author_name": "Karl Hornlund",
      "author_url": "",
      "post_date": "2019-11-20T10:58:37.227000",
      "content": "<p>Very clever!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 678055,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-21T00:28:37.367000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 677485,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-11-20T08:22:25.087000",
      "content": "<p><a href=\"/limerobot\">@limerobot</a> Congratulation on your solid solo gold, and thanks for sharing!!</p>\n\n<p>Allow me to ask a bit more details : </p>\n\n<p>(1) what are your optimizers and learning rates on the 1st and 2nd stages?\nIn order to be the most effective, do we have to carefully adjust small learning rates in the 2nd-stage? (or do we have to freeze some layers first?)</p>\n\n<p><strong>(UPDATED)</strong>\n(2) in the 1st-stage, what is the label of the classification head?</p>\n\n<p>(3) I am re-thinking about the 1st-stage, isn't it should be easy for neural network to learn the cloud mask rule? (pixel &gt;115) . (Even shallow networks should be able to learn this? )</p>\n\n<p>PS. i notice many dog guys now change their avatars to cats ... I have no idea why this happen since I am a dog fan.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 677491,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-11-20T08:26:59.460000",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> it is funny, maybe next time, you can see the dogs come back.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 677818,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-11-20T16:08:38.230000",
          "content": "<blockquote>\n  <p>(3) I am re-thinking about the 1st-stage, isn't it should be easy for neural network to learn the cloud mask rule? (pixel &gt;115) . (Even shallow networks should be able to learn this? )</p>\n</blockquote>\n\n<p>The main challenge with neural networks is training. For example, a fully connected neural network can be an exact copy of a convolutional neural network with the correct weights, but a fully connected neural network will never find those weights with gradient descent learning. Similarily pretraining a network is a trick to encourage a neural network to find weights it would not find otherwise. Because you are beginning your gradient descent learning from a different starting point. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 678053,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-21T00:27:05.917000",
          "content": "<p>Thanks for your congratulations!</p>\n\n<p><code>(1) what are your optimizers and learning rates on the 1st and 2nd stages?\nIn order to be the most effective, do we have to carefully adjust small learning rates in the 2nd-stage? (or do we have to freeze some layers first?)</code></p>\n\n<p><strong>1st-stage</strong>\noptimizer: AdamW\nscheduler: WarmupLinearSchedule\nlearning_rate: 5e-04</p>\n\n<p><strong>2nd-stage</strong>\noptimizer: AdamW\nscheduler: MultiStepLR (milestones=[5, 10], gamma=0.1)\nlearning_rate: 5e-04</p>\n\n<p>In my experiments, the small initial learning rate(5e-05) was not good.\n(Like the BERT encoder, I think the fine-tuning stage(2nd stage) is better to learn the entire weight of the model.)</p>\n\n<p><code>(2) in the 1st-stage, what is the label of the classification head?</code></p>\n\n<p>Both the 1st and 2nd stages used the same model architecture; therefore, the classification head is not learned at 1st-stage. (no need label)</p>\n\n<p><code>(3) I am re-thinking about the 1st-stage, isn't it should be easy for neural network to learn the cloud mask rule? (pixel &amp;gt;115) . (Even shallow networks should be able to learn this? )</code></p>\n\n<p>Yeah, I think it's easy. So I used only 5 epochs(with WarmupLinearSchedule) to prevent overfitting.</p>\n\n<p>PS. My wife loves cats more than dogs. 😂 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 677283,
      "author_name": "langzi",
      "author_url": "",
      "post_date": "2019-11-20T01:49:29.057000",
      "content": "<p>Congratulations, your method to get initial weights is impressive! learned.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 678054,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-21T00:28:27.363000",
          "content": "<p>Thanks for your congratulations!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 677211,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2019-11-19T23:23:50.120000",
      "content": "<p>Congrats!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 677247,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-20T00:35:07.120000",
          "content": "<p>Thanks for your congrats! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676957,
      "author_name": "giba.kim",
      "author_url": "",
      "post_date": "2019-11-19T16:47:34.050000",
      "content": "<p>Congrats for the solo gold. And appreciate for the well-written post.  Creating mast is very interesting. \nCould you explain why did you choose \"115\" as the threshold ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 677246,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-20T00:35:00.370000",
          "content": "<p>Thanks for your congrats! \nI just chose roughly. 😁  (A rule of thumb, 눈대중)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 676948,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2019-11-19T16:36:46.203000",
      "content": "<p>Congratulations on solo Gold. Pretraining on cloud shape is brilliant !! You discovered a way to use the test data during training. I also like how your model has a classification and segmentation head. And you balanced all the losses nicely by using weights 0.2, 0.2, and 0.8</p>",
      "votes": 0,
      "replies": [
        {
          "id": 677245,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-20T00:33:17.363000",
          "content": "<p>Thanks for your congratulations! 😊 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 676655,
      "author_name": "liuze",
      "author_url": "",
      "post_date": "2019-11-19T11:49:50.850000",
      "content": "<p>congratulations, limerobot, good method</p>\n\n<p>for these part, 'label' is cloud in each box, and then generated cloud mask, am I right?</p>\n\n<p>for example, for pic1, it has two box, flower and sugar, and you would generate two cloud mask flower and sugar in each box?</p>\n\n<blockquote>\n  <p>Because the clouds are white, I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label. (Since the total number of image files is 9244, the cloud_mask also generates 9244.)</p>\n</blockquote>",
      "votes": 0,
      "replies": [
        {
          "id": 676890,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T15:24:23.997000",
          "content": "<p>No, I just used this code to generate a cloud_mask.\n<code>cloud_mask = (image &gt; 115).astype(int)</code></p>\n\n<ul>\n<li><p>image\n![image](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2Fbfb36270c4b0ef8ff920d3d633519fea%2F009e2f3.jpg?generation=1574176746316088&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2Fbfb36270c4b0ef8ff920d3d633519fea%2F009e2f3.jpg?generation=1574176746316088&amp;alt=media</a> =200x300)</p></li>\n<li><p>cloud_mask = (image &gt; 115).astype(int)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F1482ca9ded9080c3200e0607852587ea%2F009e2f3_true%20(1\" alt=\"image\">.jpg?generation=1574152348586165&amp;alt=media =200x300)</p></li>\n</ul>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676622,
      "author_name": "Taemyung Heo",
      "author_url": "",
      "post_date": "2019-11-19T11:12:04.413000",
      "content": "<p>Congratulation on the gold medal!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676878,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T15:10:11.270000",
          "content": "<p>wow~ thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676611,
      "author_name": "He",
      "author_url": "",
      "post_date": "2019-11-19T10:53:10.397000",
      "content": "<p>Congratulations！！！\nPre-training stage is impressive.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676876,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T15:09:53.877000",
          "content": "<p>Thanks. I congratulate you too :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676610,
      "author_name": "Nikita Detkov",
      "author_url": "",
      "post_date": "2019-11-19T10:52:02.167000",
      "content": "<p>Very interesting, <a href=\"/limerobot\">@limerobot</a> . Have you tried to put your <code>cloud_mask</code> as fourth channel to original images? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 676874,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T15:07:37.220000",
          "content": "<p>Yep, you're right.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676607,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-19T10:49:45.390000",
      "content": "<p>very good training skills!\nNice work!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676615,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-19T11:02:42.810000",
          "content": "<p>this is just an idea, and i am not sure if it would work.</p>\n\n<p>\"I generated cloud_mask with the threshold of \"pixel &gt; 115\"</p>\n\n<p>you can try multi label for pretraining: e.g. \"x&gt;115\", \"80&gt;x&gt;155\"... it can capture some fine cloud? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 676633,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2019-11-19T11:27:50.277000",
          "content": "<p>&gt; Because the clouds are white, I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label.</p>\n\n<p><a href=\"/hengck23\">@hengck23</a> , Could you help me understand his pipeline as I am finding it hard to understand his Stage1 training pipeline?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676666,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-11-19T12:06:29.270000",
          "content": "<p>He used the original images and created masks without any neural networks. He just took each image and said <code>new_mask = (original_image&gt;115).astype(int)</code> where the original image has pixel values between 0 and 255. (By doing this he captured where the clouds are located but not what type of cloud). Then he trained his first epochs on these masks then switched to real masks.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 676677,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2019-11-19T12:16:11.347000",
          "content": "<p>Thank you Chris for making this clear. Now I understand what he meant when he says\n&gt;  I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label.</p>\n\n<p>In Stage1, he trained using <code>train+test</code> images and their corresponding labels obtained by doing <code>(image&amp;gt;115).astype(int)</code> \nFor stage2, he used Stage1 model to initialize training for <code>train images</code> and their corresponding <code>masks</code>. \nI hope this helps others like me who were confused about his Stage1 pipeline</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676693,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2019-11-19T12:39:35.883000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I was trying to play with the pixel values this way . I didn't think of pretraining , more of post processing . However , Sugar always threw me off .. it had pixels values starting 60+ .. so I had dropped that idea . Does &gt; 115 capture all clouds ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676875,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T15:09:11.710000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thank you for your clarification! 👍 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676893,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T15:27:17.170000",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a>  No, &gt; 115 doesn't seem to capture all the clouds.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676900,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T15:31:01.330000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> Thanks for your congratulations, and I didn’t think like that. Interesting!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676594,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2019-11-19T10:34:17.520000",
      "content": "<p>Congratulations 🎉  and thanks for sharing your solution!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676904,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T15:36:30.897000",
          "content": "<p>Thanks for your congratulations!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676487,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2019-11-19T08:17:29.057000",
      "content": "<p>Congratulations.\nPretrained on white cloud is interesting.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676511,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T08:45:37.770000",
          "content": "<p>I congratulate you too :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676472,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2019-11-19T08:05:54.457000",
      "content": "<p>Thank you for sharing and congrats for your gold finish.\nI have few things to ask about your approach.\n&gt; Therefore, I tried to pre-train the model to learn the shape of the cloud.</p>\n\n<p>In stage1, the label is the mask? or the image itself?</p>\n\n<p>&gt; Because the clouds are white, I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label.</p>\n\n<p>Could you elaborate on this part? this increased the number of training samples for 2nd stage?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676497,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T08:33:05.127000",
          "content": "<p>I'm sorry I didn't explain it clearly. 😂 </p>\n\n<blockquote>\n  <p>In stage1, the label is the mask? or the image itself?</p>\n</blockquote>\n\n<p>The label is the mask. <br>\nOne example of mask\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F1482ca9ded9080c3200e0607852587ea%2F009e2f3_true%20(1\" alt=\"\">.jpg?generation=1574152348586165&amp;alt=media)</p>\n\n<blockquote>\n  <p>Could you elaborate on this part? this increased the number of training samples for 2nd stage?</p>\n</blockquote>\n\n<p>Pre-trained(1st stage-training) model are used as the initial value of 2nd-stage model weights.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 676519,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2019-11-19T08:56:25.777000",
          "content": "<p>Ok!! I drew a sketch of your pipeline(skipping classification part of stage2) based on my understanding of your writeup....could you please confirm?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2Fc34eb0c5ea50c5f789a85b72c5520a85%2Fmask_limerobot.jpg?generation=1574153776348792&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676555,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T09:49:03.410000",
          "content": "<p>Yeah, that's right.</p>\n\n<ul>\n<li>1st-stage training: all images, generatedmask </li>\n<li>2nd-stage training: train images, realmask</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 676563,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2019-11-19T10:01:13.873000",
          "content": "<p>Thank you for your clarification. If what I understood is correct(as you confirmed it), where is this part?</p>\n\n<blockquote>\n  <p>Then, I used it as a label.</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676525,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-19T09:07:35.340000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 676905,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2019-11-19T15:36:37.640000",
          "content": "<p>Thanks for your congratulations!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "676467": "First of all, I would like to thank the hosting organization that hosted this competition and Kaggle. Like any competition, this competition was also hot until the end.  So, I want to congratulate Kagglers who struggled until the end of this competition.\n\nI will summarize and write down the part of my solution that you will be interested in. It's `pre-training`\n\n# pre-training\nThe challenge of this competition is to segment according to the shape of the cloud.  Therefore, I tried to pre-train the model to learn the shape of the cloud.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F7e5d1fb4cbda36a5bcad97879a424176%2F1st-training.png?generation=1574148986382141&amp;alt=media)\n\nBecause the clouds are white, I generated `cloud_mask` with the threshold of \"pixel &gt; 115\". Then, I used it as a label. (Since the total number of image files is 9244, the cloud_mask also generates 9244.)\n\nAfter pre-training, I tried a 2nd-stage training.\nPre-trained(1st stage-training) model are used as the initial value of 2nd-stage model weights.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F6ed4418967a6461295eed180c047feda%2F2nd-training.png?generation=1574149060524535&amp;alt=media)\n\nThis training process boosted my CV 0.005~0.01. So, my single model score is as follows.\n\n| model | private | public |\n| --- | --- | --- |\n| efficientnet-b4, unet | 0.66927 | 0.67437 |\n| efficientnet-b4, fpn | 0.66827 | 0.67508 |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1658108%2F95009ca5bb429732cb523d971978e0fd%2F.png?generation=1574149018151413&amp;alt=media)\n\nThe rest is not special, so I'll skip the description. 😁 \nThanks for your reading!\n",
    "676680": "Congratulations, the pre-training strategy is very smart.",
    "681494": "Congratulation!\n\nHow do you find out this way?   I think pixel &gt; 115 is very useful for reducing the noise(illumination）， that increase the quanlity of learning.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1965832%2F01d8ed5d1b73087170a4c97ee634c216%2Forg.png?generation=1574751538876480&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1965832%2Fc80ada396099f2f2b866535d1e5aa48a%2Fnew.png?generation=1574751613493251&amp;alt=media)\n",
    "677586": "Very clever!",
    "677485": "@limerobot Congratulation on your solid solo gold, and thanks for sharing!!\n\nAllow me to ask a bit more details : \n\n(1) what are your optimizers and learning rates on the 1st and 2nd stages?\nIn order to be the most effective, do we have to carefully adjust small learning rates in the 2nd-stage? (or do we have to freeze some layers first?)\n\n**(UPDATED)**\n(2) in the 1st-stage, what is the label of the classification head?\n\n(3) I am re-thinking about the 1st-stage, isn't it should be easy for neural network to learn the cloud mask rule? (pixel &gt;115) . (Even shallow networks should be able to learn this? )\n\nPS. i notice many dog guys now change their avatars to cats ... I have no idea why this happen since I am a dog fan.",
    "677283": "Congratulations, your method to get initial weights is impressive! learned.",
    "677211": "Congrats!",
    "676957": "Congrats for the solo gold. And appreciate for the well-written post.  Creating mast is very interesting. \nCould you explain why did you choose \"115\" as the threshold ?",
    "676948": "Congratulations on solo Gold. Pretraining on cloud shape is brilliant !! You discovered a way to use the test data during training. I also like how your model has a classification and segmentation head. And you balanced all the losses nicely by using weights 0.2, 0.2, and 0.8",
    "676655": "congratulations, limerobot, good method\n\nfor these part, 'label' is cloud in each box, and then generated cloud mask, am I right?\n\nfor example, for pic1, it has two box, flower and sugar, and you would generate two cloud mask flower and sugar in each box?\n\n&gt; Because the clouds are white, I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label. (Since the total number of image files is 9244, the cloud_mask also generates 9244.)",
    "676622": "Congratulation on the gold medal!",
    "676611": "Congratulations！！！\nPre-training stage is impressive.",
    "676610": "Very interesting, @limerobot . Have you tried to put your `cloud_mask` as fourth channel to original images? ",
    "676607": "very good training skills!\nNice work!",
    "676594": "Congratulations 🎉  and thanks for sharing your solution!",
    "676487": "Congratulations.\nPretrained on white cloud is interesting.",
    "676472": "Thank you for sharing and congrats for your gold finish.\nI have few things to ask about your approach.\n&gt; Therefore, I tried to pre-train the model to learn the shape of the cloud.\n\nIn stage1, the label is the mask? or the image itself?\n\n&gt; Because the clouds are white, I generated cloud_mask with the threshold of \"pixel &gt; 115\". Then, I used it as a label.\n\nCould you elaborate on this part? this increased the number of training samples for 2nd stage?",
    "676525": ""
  }
}