{
  "id": 40199,
  "title": "3rd place solution (U-Net + Dilated Conv)",
  "url": "/competitions/carvana-image-masking-challenge/writeups/lyakaap-3rd-place-solution-u-net-dilated-conv",
  "author_name": "",
  "post_date": "2017-10-01T13:49:50.510Z",
  "votes": 96,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Thanks for hosting such an exciting competition! I was really enthusiastic for spending my time for this competition! And thanks for helpful code by @Peter and excellent ideas by @HengCher Keng. I learned a lot from them.</p>\n\n<p>The competition repository is here. I put two scripts (My network script &amp; loss functions script) in it.\n<a href=\"https://github.com/lyakaap/Kaggle-Carvana-3rd-place-solution\">https://github.com/lyakaap/Kaggle-Carvana-3rd-place-solution</a></p>\n\n<h2>My solution overview</h2>\n\n<ul>\n<li><p>I used 1536x1024 &amp; 1920x1280 resolution.</p></li>\n<li><p>I used modified U-Net. It has several dilated convolution layers in bottleneck block. (i.e. where the resolution of feature maps are lowest)</p></li>\n</ul>\n\n<p>Detailed figure of my network architecture is here.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/225523/7428/network.png\" alt=\"my_network\" title=\"\"></p>\n\n<p>The best score of this model is <strong>0.997193</strong> only around 8.5 million parameters. (trained one of 6 folds, no TTA &amp; no ensemble, input resolution: 1920x1280)\nAveraging two predictions(TTA, original image &amp; flipped image) by 0.997193 model reached <strong>0.997222</strong>. They are ranked 6th place and 5th place on LB respectively!</p>\n\n<p>I tried normal convolution layers instead of dilated convolution layers in bottleneck block, and its score is significantly lower than using dilated convolution. (normal: 0.9905, using dilated conv: 0.9918 @256x256)</p>\n\n<p>I also tried parallelized dilated convolution layers instead of stacking them, but it gave me lower score than stacked architecture.</p>\n\n<ul>\n<li><p>Optimizer: RMSprop lr = 0.0002, reducing learning rate by using ReduceLROnPlateau() that is Keras callback function. Reducing factor is 0.2 &amp; 0.5</p></li>\n<li><p>Data Augmentation: only horizontal flip. Scaling, Shifting, and Shifting HSV were results of overfitting for me.</p></li>\n<li><p>Batchsize: 1, and no BN.</p></li>\n<li><p>Training whole time on single model takes around 2 days.</p></li>\n<li><p>Pseudo Labeling: learning simultaneously or only using pretraining phase.</p></li>\n<li><p>Loss function: bce + dice loss (I also tried weighing boundary pixel loss, it gave similar result. Fear of overfitting, I finally decided not to use it.)</p></li>\n<li><p>Ensemble: 5 fold ensemble @1536x1024 + 6 fold ensemble @1920x1280, weighted average. I weighted by LB ranking in my submissions.</p></li>\n<li><p>TTA: only horizontal flip.</p></li>\n<li><p>Adjusting threshold: I decided threshold which gives best score on validation set. I set the threshold to 0.508. In LB, it makes score improving only 0.000001.</p></li>\n</ul>\n\n<h2>Other</h2>\n\n<p>One of the best contributer of improving score is training on pseudo labeling data. I think why pseudo labeling contribute so much is the amount of test data, and we can get predictions close ground truth.</p>\n\n<p>I tried post processing by using pydensecrf for only difficult to mask car images. But it gave me no improvement.\nAs for how to choose \"difficult images\", I calculated multi class version of dice coeficient (I'm afraid that I shouldn't say so) of predictions by several models.</p>",
  "messages": [
    {
      "id": "225523",
      "postDate": "09/29/2017 11:23:41",
      "content": "<p>Thanks for hosting such an exciting competition! I was really enthusiastic for spending my time for this competition! And thanks for helpful code by @Peter and excellent ideas by @HengCher Keng. I learned a lot from them.</p>\n\n<p>The competition repository is here. I put two scripts (My network script &amp; loss functions script) in it.\n<a href=\"https://github.com/lyakaap/Kaggle-Carvana-3rd-place-solution\">https://github.com/lyakaap/Kaggle-Carvana-3rd-place-solution</a></p>\n\n<h2>My solution overview</h2>\n\n<ul>\n<li><p>I used 1536x1024 &amp; 1920x1280 resolution.</p></li>\n<li><p>I used modified U-Net. It has several dilated convolution layers in bottleneck block. (i.e. where the resolution of feature maps are lowest)</p></li>\n</ul>\n\n<p>Detailed figure of my network architecture is here.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/225523/7428/network.png\" alt=\"my_network\" title=\"\"></p>\n\n<p>The best score of this model is <strong>0.997193</strong> only around 8.5 million parameters. (trained one of 6 folds, no TTA &amp; no ensemble, input resolution: 1920x1280)\nAveraging two predictions(TTA, original image &amp; flipped image) by 0.997193 model reached <strong>0.997222</strong>. They are ranked 6th place and 5th place on LB respectively!</p>\n\n<p>I tried normal convolution layers instead of dilated convolution layers in bottleneck block, and its score is significantly lower than using dilated convolution. (normal: 0.9905, using dilated conv: 0.9918 @256x256)</p>\n\n<p>I also tried parallelized dilated convolution layers instead of stacking them, but it gave me lower score than stacked architecture.</p>\n\n<ul>\n<li><p>Optimizer: RMSprop lr = 0.0002, reducing learning rate by using ReduceLROnPlateau() that is Keras callback function. Reducing factor is 0.2 &amp; 0.5</p></li>\n<li><p>Data Augmentation: only horizontal flip. Scaling, Shifting, and Shifting HSV were results of overfitting for me.</p></li>\n<li><p>Batchsize: 1, and no BN.</p></li>\n<li><p>Training whole time on single model takes around 2 days.</p></li>\n<li><p>Pseudo Labeling: learning simultaneously or only using pretraining phase.</p></li>\n<li><p>Loss function: bce + dice loss (I also tried weighing boundary pixel loss, it gave similar result. Fear of overfitting, I finally decided not to use it.)</p></li>\n<li><p>Ensemble: 5 fold ensemble @1536x1024 + 6 fold ensemble @1920x1280, weighted average. I weighted by LB ranking in my submissions.</p></li>\n<li><p>TTA: only horizontal flip.</p></li>\n<li><p>Adjusting threshold: I decided threshold which gives best score on validation set. I set the threshold to 0.508. In LB, it makes score improving only 0.000001.</p></li>\n</ul>\n\n<h2>Other</h2>\n\n<p>One of the best contributer of improving score is training on pseudo labeling data. I think why pseudo labeling contribute so much is the amount of test data, and we can get predictions close ground truth.</p>\n\n<p>I tried post processing by using pydensecrf for only difficult to mask car images. But it gave me no improvement.\nAs for how to choose \"difficult images\", I calculated multi class version of dice coeficient (I'm afraid that I shouldn't say so) of predictions by several models.</p>",
      "rawMarkdown": "Thanks for hosting such an exciting competition! I was really enthusiastic for spending my time for this competition! And thanks for helpful code by @Peter and excellent ideas by @HengCher Keng. I learned a lot from them.\n\nThe competition repository is here. I put two scripts (My network script &amp; loss functions script) in it.\nhttps://github.com/lyakaap/Kaggle-Carvana-3rd-place-solution\n\n## My solution overview\n\n* I used 1536x1024 &amp; 1920x1280 resolution.\n\n* I used modified U-Net. It has several dilated convolution layers in bottleneck block. (i.e. where the resolution of feature maps are lowest)\n\nDetailed figure of my network architecture is here.\n\n![my_network][1]\n\nThe best score of this model is **0.997193** only around 8.5 million parameters. (trained one of 6 folds, no TTA &amp; no ensemble, input resolution: 1920x1280)\nAveraging two predictions(TTA, original image &amp; flipped image) by 0.997193 model reached **0.997222**. They are ranked 6th place and 5th place on LB respectively!\n\nI tried normal convolution layers instead of dilated convolution layers in bottleneck block, and its score is significantly lower than using dilated convolution. (normal: 0.9905, using dilated conv: 0.9918 @256x256)\n\nI also tried parallelized dilated convolution layers instead of stacking them, but it gave me lower score than stacked architecture.\n\n* Optimizer: RMSprop lr = 0.0002, reducing learning rate by using ReduceLROnPlateau() that is Keras callback function. Reducing factor is 0.2 &amp; 0.5\n\n* Data Augmentation: only horizontal flip. Scaling, Shifting, and Shifting HSV were results of overfitting for me.\n\n* Batchsize: 1, and no BN.\n\n* Training whole time on single model takes around 2 days.\n\n* Pseudo Labeling: learning simultaneously or only using pretraining phase.\n\n* Loss function: bce + dice loss (I also tried weighing boundary pixel loss, it gave similar result. Fear of overfitting, I finally decided not to use it.)\n\n* Ensemble: 5 fold ensemble @1536x1024 + 6 fold ensemble @1920x1280, weighted average. I weighted by LB ranking in my submissions.\n\n* TTA: only horizontal flip.\n\n* Adjusting threshold: I decided threshold which gives best score on validation set. I set the threshold to 0.508. In LB, it makes score improving only 0.000001.\n\n## Other\n\nOne of the best contributer of improving score is training on pseudo labeling data. I think why pseudo labeling contribute so much is the amount of test data, and we can get predictions close ground truth.\n\nI tried post processing by using pydensecrf for only difficult to mask car images. But it gave me no improvement.\nAs for how to choose \"difficult images\", I calculated multi class version of dice coeficient (I'm afraid that I shouldn't say so) of predictions by several models.\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225523/7428/network.png",
      "votes": null
    },
    {
      "id": "225528",
      "postDate": "09/29/2017 11:36:10",
      "content": "<p>Principal Components of feature maps that is output of Dilated Convolution layers each has a different dilation_rate. I think it's interesting.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/225528/7429/dilation_rate=1.png\" alt=\"1\" title=\"\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/225528/7430/dilation_rate=2.png\" alt=\"2\" title=\"\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/225528/7434/dilation_rate=4.png\" alt=\"4\" title=\"\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/225528/7431/dilation_rate=8.png\" alt=\"8\" title=\"\"></p>",
      "rawMarkdown": "Principal Components of feature maps that is output of Dilated Convolution layers each has a different dilation_rate. I think it's interesting.\n![1][1]\n![2][2]\n![4][3]\n![8][4]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225528/7429/dilation_rate=1.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225528/7430/dilation_rate=2.png\n  [3]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225528/7434/dilation_rate=4.png\n  [4]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225528/7431/dilation_rate=8.png\n  [5]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225528/7433/dilation_rate=16.png\n  [6]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225528/7432/dilation_rate=32.png",
      "votes": null
    },
    {
      "id": "225540",
      "postDate": "09/29/2017 12:24:14",
      "content": "<p>Really nice work!</p>",
      "rawMarkdown": "Really nice work!",
      "votes": null
    },
    {
      "id": "225597",
      "postDate": "09/29/2017 15:30:07",
      "content": "<p>Thanks, Did you split by car on cross-validation?</p>",
      "rawMarkdown": "Thanks, Did you split by car on cross-validation?",
      "votes": null
    },
    {
      "id": "225604",
      "postDate": "09/29/2017 15:52:45",
      "content": "<p>Amazing! You are the winner in simplicity category, man!</p>",
      "rawMarkdown": "Amazing! You are the winner in simplicity category, man!",
      "votes": null
    },
    {
      "id": "225616",
      "postDate": "09/29/2017 16:30:45",
      "content": "<p>Yes. I used normal K-fold CV.</p>",
      "rawMarkdown": "Yes. I used normal K-fold CV.",
      "votes": null
    },
    {
      "id": "225620",
      "postDate": "09/29/2017 16:41:35",
      "content": "<p>Thanks! I tried many things make my model further complicated(post processing, various augmentation method, weighted loss, etc.), but no one helps me improving score.\nI realized that simple method is good for simple task like this competition :)</p>",
      "rawMarkdown": "Thanks! I tried many things make my model further complicated(post processing, various augmentation method, weighted loss, etc.), but no one helps me improving score.\nI realized that simple method is good for simple task like this competition :)",
      "votes": null
    },
    {
      "id": "225635",
      "postDate": "09/29/2017 17:21:05",
      "content": "<p>Thanks @Iyakaap for your detailed overview of your solution and sharing the scripts.</p>",
      "rawMarkdown": "Thanks @Iyakaap for your detailed overview of your solution and sharing the scripts.",
      "votes": null
    },
    {
      "id": "225661",
      "postDate": "09/29/2017 18:31:17",
      "content": "<p>Thanks for sharing! what do you exactly mean by saying pseudo labeling data?</p>",
      "rawMarkdown": "Thanks for sharing! what do you exactly mean by saying pseudo labeling data?",
      "votes": null
    },
    {
      "id": "225775",
      "postDate": "09/30/2017 01:55:34",
      "content": "<p>Very nice work. By the way i tries residual connections at the center layers. It also improve my results. It seems that the center layers are important but I don't know why. It will be fun to investigate. </p>\n\n<p>Being the smallest, the center layers are the easiest to work will. I am going to try other structure like densenet or senet at the center layers later</p>",
      "rawMarkdown": "Very nice work. By the way i tries residual connections at the center layers. It also improve my results. It seems that the center layers are important but I don't know why. It will be fun to investigate. \n\nBeing the smallest, the center layers are the easiest to work will. I am going to try other structure like densenet or senet at the center layers later",
      "votes": null
    },
    {
      "id": "225807",
      "postDate": "09/30/2017 02:05:55",
      "content": "<p>Congratulation and thanks for sharing! By the way, could you share the setting of ReduceLROnPlateau, e.g. the monitor is val_loss or val_metrics?</p>",
      "rawMarkdown": "Congratulation and thanks for sharing! By the way, could you share the setting of ReduceLROnPlateau, e.g. the monitor is val_loss or val_metrics?",
      "votes": null
    },
    {
      "id": "225824",
      "postDate": "09/30/2017 03:21:46",
      "content": "<p>Thank you! I'm looking forward to your experiment results. I'm very curious.</p>",
      "rawMarkdown": "Thank you! I'm looking forward to your experiment results. I'm very curious.",
      "votes": null
    },
    {
      "id": "225826",
      "postDate": "09/30/2017 03:25:51",
      "content": "<p>Using test data predictions by other model as if them were ground truth. In the training phase, you can use (X_train, y_train) and (X_test, y_test) for training.</p>",
      "rawMarkdown": "Using test data predictions by other model as if them were ground truth. In the training phase, you can use (X_train, y_train) and (X_test, y_test) for training.",
      "votes": null
    },
    {
      "id": "225828",
      "postDate": "09/30/2017 03:31:25",
      "content": "<pre><code>ReduceLROnPlateau(monitor='val_dice_coef',\n                           factor=0.2,\n                           patience=3,\n                           verbose=1,\n                           epsilon=1e-4,\n                           mode='max')\n</code></pre>\n\n<p>I used above settings, but I think it is better to use validation loss as metrics.</p>",
      "rawMarkdown": "ReduceLROnPlateau(monitor='val_dice_coef',\n                               factor=0.2,\n                               patience=3,\n                               verbose=1,\n                               epsilon=1e-4,\n                               mode='max')\nI used above settings, but I think it is better to use validation loss as metrics.",
      "votes": null
    },
    {
      "id": "225976",
      "postDate": "09/30/2017 15:13:09",
      "content": "<p>Awesome solution, I love the simplicity of it. I too tried dilated convs /Unet but I guess I did something wrong, b/c it was a bit worse; although I did not add up (or stack) the feature layers at the bottom botteneck. </p>\n\n<p>What was the intuition behind adding them up? </p>",
      "rawMarkdown": "Awesome solution, I love the simplicity of it. I too tried dilated convs /Unet but I guess I did something wrong, b/c it was a bit worse; although I did not add up (or stack) the feature layers at the bottom botteneck. \n\nWhat was the intuition behind adding them up?",
      "votes": null
    },
    {
      "id": "226162",
      "postDate": "10/01/2017 08:12:38",
      "content": "<p>Thank you for sharing your solution. Does it help to remove BN after each convolution layer?</p>",
      "rawMarkdown": "Thank you for sharing your solution. Does it help to remove BN after each convolution layer?",
      "votes": null
    },
    {
      "id": "254247",
      "postDate": "12/06/2017 15:07:13",
      "content": "<p>I tried it (with input 512x512). The dice_coef was  .91 by the first training epoch itself. This outperforms some other deeper U-Nets I also tried. </p>",
      "rawMarkdown": "I tried it (with input 512x512). The dice_coef was  .91 by the first training epoch itself. This outperforms some other deeper U-Nets I also tried.",
      "votes": null
    },
    {
      "id": "330147",
      "postDate": "05/18/2018 06:27:56",
      "content": "<p>Amazing solution ever!</p>",
      "rawMarkdown": "Amazing solution ever!",
      "votes": null
    },
    {
      "id": "344866",
      "postDate": "06/18/2018 19:59:56",
      "content": "<p>I tried executing your solution but it is unable to fetch the .png files from the train_masks folder as it contains .gif files only. Am I missing something? Can someone please help me in executing the same?</p>",
      "rawMarkdown": "I tried executing your solution but it is unable to fetch the .png files from the train_masks folder as it contains .gif files only. Am I missing something? Can someone please help me in executing the same?",
      "votes": null
    },
    {
      "id": "383705",
      "postDate": "09/09/2018 12:06:35",
      "content": "<p>Hi, <a href=\"/lyakaap\">@lyakaap</a>, thank you for your sharing.\nwhat tool did you use to draw the picture?</p>",
      "rawMarkdown": "Hi, @lyakaap, thank you for your sharing.\nwhat tool did you use to draw the picture?",
      "votes": null
    },
    {
      "id": "383755",
      "postDate": "09/09/2018 14:38:03",
      "content": "<p>I believe he used draw.io. Very easy to use. I did a U-Net diagram and it came out just like this.</p>",
      "rawMarkdown": "I believe he used draw.io. Very easy to use. I did a U-Net diagram and it came out just like this.",
      "votes": null
    },
    {
      "id": "385133",
      "postDate": "09/10/2018 11:37:39",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    },
    {
      "id": "431294",
      "postDate": "12/02/2018 02:17:18",
      "content": "<p>Hi, I am still wondering about the intuition for adding the dilation layers. I couldnt find research papers suggesting this but it does make sense. I would love to read related papers if someone can point me to them.  I could only find the U-Net paper and Multi-scale context aggregation by dilated convolutions paper but did not find any research paper using dilated convolutions like used here.</p>",
      "rawMarkdown": "Hi, I am still wondering about the intuition for adding the dilation layers. I couldnt find research papers suggesting this but it does make sense. I would love to read related papers if someone can point me to them.  I could only find the U-Net paper and Multi-scale context aggregation by dilated convolutions paper but did not find any research paper using dilated convolutions like used here.",
      "votes": null
    },
    {
      "id": "461123",
      "postDate": "01/25/2019 09:49:56",
      "content": "<p>Hi, I recently looked up into this model, what could be intuition for adding binary cross entropy loss with dice loss, is there any specific reference to that, could anyone help me out </p>",
      "rawMarkdown": "Hi, I recently looked up into this model, what could be intuition for adding binary cross entropy loss with dice loss, is there any specific reference to that, could anyone help me out",
      "votes": null
    },
    {
      "id": "501212",
      "postDate": "03/27/2019 03:25:04",
      "content": "<p>very nice work！</p>\n\n<p>what does the meaning of weight w1 and w0 in loss function of “weighted_bce_dice_loss”~</p>\n\n<p>thx！</p>",
      "rawMarkdown": "very nice work！\n\nwhat does the meaning of weight w1 and w0 in loss function of “weighted_bce_dice_loss”~\n\n\nthx！",
      "votes": null
    },
    {
      "id": "539573",
      "postDate": "05/30/2019 08:20:23",
      "content": "<p>Hi, Thanks for the solution. Did you convert  '.gif' files into a format such as '.png' ?. The reason behind, I wasn't able to fetch gif files directly. </p>",
      "rawMarkdown": "Hi, Thanks for the solution. Did you convert  '.gif' files into a format such as '.png' ?. The reason behind, I wasn't able to fetch gif files directly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 225528,
      "author_name": "lyakaap",
      "author_url": "",
      "post_date": "09/29/2017 11:36:10",
      "content": "<p>Principal Components of feature maps that is output of Dilated Convolution layers each has a different dilation_rate. I think it's interesting.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/225528/7429/dilation_rate=1.png\" alt=\"1\" title=\"\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/225528/7430/dilation_rate=2.png\" alt=\"2\" title=\"\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/225528/7434/dilation_rate=4.png\" alt=\"4\" title=\"\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/225528/7431/dilation_rate=8.png\" alt=\"8\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 225540,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "09/29/2017 12:24:14",
      "content": "<p>Really nice work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 225597,
      "author_name": "ironbar",
      "author_url": "",
      "post_date": "09/29/2017 15:30:07",
      "content": "<p>Thanks, Did you split by car on cross-validation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 225616,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "09/29/2017 16:30:45",
          "content": "<p>Yes. I used normal K-fold CV.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225604,
      "author_name": "ceperaang",
      "author_url": "",
      "post_date": "09/29/2017 15:52:45",
      "content": "<p>Amazing! You are the winner in simplicity category, man!</p>",
      "votes": null,
      "replies": [
        {
          "id": 225620,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "09/29/2017 16:41:35",
          "content": "<p>Thanks! I tried many things make my model further complicated(post processing, various augmentation method, weighted loss, etc.), but no one helps me improving score.\nI realized that simple method is good for simple task like this competition :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225635,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "09/29/2017 17:21:05",
      "content": "<p>Thanks @Iyakaap for your detailed overview of your solution and sharing the scripts.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 225661,
      "author_name": "amirbarwin",
      "author_url": "",
      "post_date": "09/29/2017 18:31:17",
      "content": "<p>Thanks for sharing! what do you exactly mean by saying pseudo labeling data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 225826,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "09/30/2017 03:25:51",
          "content": "<p>Using test data predictions by other model as if them were ground truth. In the training phase, you can use (X_train, y_train) and (X_test, y_test) for training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225775,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/30/2017 01:55:34",
      "content": "<p>Very nice work. By the way i tries residual connections at the center layers. It also improve my results. It seems that the center layers are important but I don't know why. It will be fun to investigate. </p>\n\n<p>Being the smallest, the center layers are the easiest to work will. I am going to try other structure like densenet or senet at the center layers later</p>",
      "votes": null,
      "replies": [
        {
          "id": 225824,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "09/30/2017 03:21:46",
          "content": "<p>Thank you! I'm looking forward to your experiment results. I'm very curious.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225807,
      "author_name": "tangzy",
      "author_url": "",
      "post_date": "09/30/2017 02:05:55",
      "content": "<p>Congratulation and thanks for sharing! By the way, could you share the setting of ReduceLROnPlateau, e.g. the monitor is val_loss or val_metrics?</p>",
      "votes": null,
      "replies": [
        {
          "id": 225828,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "09/30/2017 03:31:25",
          "content": "<pre><code>ReduceLROnPlateau(monitor='val_dice_coef',\n                           factor=0.2,\n                           patience=3,\n                           verbose=1,\n                           epsilon=1e-4,\n                           mode='max')\n</code></pre>\n\n<p>I used above settings, but I think it is better to use validation loss as metrics.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225976,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "09/30/2017 15:13:09",
      "content": "<p>Awesome solution, I love the simplicity of it. I too tried dilated convs /Unet but I guess I did something wrong, b/c it was a bit worse; although I did not add up (or stack) the feature layers at the bottom botteneck. </p>\n\n<p>What was the intuition behind adding them up? </p>",
      "votes": null,
      "replies": [
        {
          "id": 431294,
          "author_name": "nirvedh",
          "author_url": "",
          "post_date": "12/02/2018 02:17:18",
          "content": "<p>Hi, I am still wondering about the intuition for adding the dilation layers. I couldnt find research papers suggesting this but it does make sense. I would love to read related papers if someone can point me to them.  I could only find the U-Net paper and Multi-scale context aggregation by dilated convolutions paper but did not find any research paper using dilated convolutions like used here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 226162,
      "author_name": "unixnme",
      "author_url": "",
      "post_date": "10/01/2017 08:12:38",
      "content": "<p>Thank you for sharing your solution. Does it help to remove BN after each convolution layer?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 254247,
      "author_name": "shir0mani",
      "author_url": "",
      "post_date": "12/06/2017 15:07:13",
      "content": "<p>I tried it (with input 512x512). The dice_coef was  .91 by the first training epoch itself. This outperforms some other deeper U-Nets I also tried. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 330147,
      "author_name": "cafeal",
      "author_url": "",
      "post_date": "05/18/2018 06:27:56",
      "content": "<p>Amazing solution ever!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 344866,
      "author_name": "nikhilnik11",
      "author_url": "",
      "post_date": "06/18/2018 19:59:56",
      "content": "<p>I tried executing your solution but it is unable to fetch the .png files from the train_masks folder as it contains .gif files only. Am I missing something? Can someone please help me in executing the same?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 383705,
      "author_name": "tommao",
      "author_url": "",
      "post_date": "09/09/2018 12:06:35",
      "content": "<p>Hi, <a href=\"/lyakaap\">@lyakaap</a>, thank you for your sharing.\nwhat tool did you use to draw the picture?</p>",
      "votes": null,
      "replies": [
        {
          "id": 383755,
          "author_name": "shir0mani",
          "author_url": "",
          "post_date": "09/09/2018 14:38:03",
          "content": "<p>I believe he used draw.io. Very easy to use. I did a U-Net diagram and it came out just like this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 385133,
          "author_name": "tommao",
          "author_url": "",
          "post_date": "09/10/2018 11:37:39",
          "content": "<p>Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 461123,
      "author_name": "goutham4deepu",
      "author_url": "",
      "post_date": "01/25/2019 09:49:56",
      "content": "<p>Hi, I recently looked up into this model, what could be intuition for adding binary cross entropy loss with dice loss, is there any specific reference to that, could anyone help me out </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 501212,
      "author_name": "wanglilin",
      "author_url": "",
      "post_date": "03/27/2019 03:25:04",
      "content": "<p>very nice work！</p>\n\n<p>what does the meaning of weight w1 and w0 in loss function of “weighted_bce_dice_loss”~</p>\n\n<p>thx！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 539573,
      "author_name": "s1782662",
      "author_url": "",
      "post_date": "05/30/2019 08:20:23",
      "content": "<p>Hi, Thanks for the solution. Did you convert  '.gif' files into a format such as '.png' ?. The reason behind, I wasn't able to fetch gif files directly. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "225523": "Thanks for hosting such an exciting competition! I was really enthusiastic for spending my time for this competition! And thanks for helpful code by @Peter and excellent ideas by @HengCher Keng. I learned a lot from them.\n\nThe competition repository is here. I put two scripts (My network script &amp; loss functions script) in it.\nhttps://github.com/lyakaap/Kaggle-Carvana-3rd-place-solution\n\n## My solution overview\n\n* I used 1536x1024 &amp; 1920x1280 resolution.\n\n* I used modified U-Net. It has several dilated convolution layers in bottleneck block. (i.e. where the resolution of feature maps are lowest)\n\nDetailed figure of my network architecture is here.\n\n![my_network][1]\n\nThe best score of this model is **0.997193** only around 8.5 million parameters. (trained one of 6 folds, no TTA &amp; no ensemble, input resolution: 1920x1280)\nAveraging two predictions(TTA, original image &amp; flipped image) by 0.997193 model reached **0.997222**. They are ranked 6th place and 5th place on LB respectively!\n\nI tried normal convolution layers instead of dilated convolution layers in bottleneck block, and its score is significantly lower than using dilated convolution. (normal: 0.9905, using dilated conv: 0.9918 @256x256)\n\nI also tried parallelized dilated convolution layers instead of stacking them, but it gave me lower score than stacked architecture.\n\n* Optimizer: RMSprop lr = 0.0002, reducing learning rate by using ReduceLROnPlateau() that is Keras callback function. Reducing factor is 0.2 &amp; 0.5\n\n* Data Augmentation: only horizontal flip. Scaling, Shifting, and Shifting HSV were results of overfitting for me.\n\n* Batchsize: 1, and no BN.\n\n* Training whole time on single model takes around 2 days.\n\n* Pseudo Labeling: learning simultaneously or only using pretraining phase.\n\n* Loss function: bce + dice loss (I also tried weighing boundary pixel loss, it gave similar result. Fear of overfitting, I finally decided not to use it.)\n\n* Ensemble: 5 fold ensemble @1536x1024 + 6 fold ensemble @1920x1280, weighted average. I weighted by LB ranking in my submissions.\n\n* TTA: only horizontal flip.\n\n* Adjusting threshold: I decided threshold which gives best score on validation set. I set the threshold to 0.508. In LB, it makes score improving only 0.000001.\n\n## Other\n\nOne of the best contributer of improving score is training on pseudo labeling data. I think why pseudo labeling contribute so much is the amount of test data, and we can get predictions close ground truth.\n\nI tried post processing by using pydensecrf for only difficult to mask car images. But it gave me no improvement.\nAs for how to choose \"difficult images\", I calculated multi class version of dice coeficient (I'm afraid that I shouldn't say so) of predictions by several models.\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225523/7428/network.png",
    "225528": "Principal Components of feature maps that is output of Dilated Convolution layers each has a different dilation_rate. I think it's interesting.\n![1][1]\n![2][2]\n![4][3]\n![8][4]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225528/7429/dilation_rate=1.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225528/7430/dilation_rate=2.png\n  [3]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225528/7434/dilation_rate=4.png\n  [4]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225528/7431/dilation_rate=8.png\n  [5]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225528/7433/dilation_rate=16.png\n  [6]: https://kaggle2.blob.core.windows.net/forum-message-attachments/225528/7432/dilation_rate=32.png",
    "225540": "Really nice work!",
    "225597": "Thanks, Did you split by car on cross-validation?",
    "225604": "Amazing! You are the winner in simplicity category, man!",
    "225616": "Yes. I used normal K-fold CV.",
    "225620": "Thanks! I tried many things make my model further complicated(post processing, various augmentation method, weighted loss, etc.), but no one helps me improving score.\nI realized that simple method is good for simple task like this competition :)",
    "225635": "Thanks @Iyakaap for your detailed overview of your solution and sharing the scripts.",
    "225661": "Thanks for sharing! what do you exactly mean by saying pseudo labeling data?",
    "225775": "Very nice work. By the way i tries residual connections at the center layers. It also improve my results. It seems that the center layers are important but I don't know why. It will be fun to investigate. \n\nBeing the smallest, the center layers are the easiest to work will. I am going to try other structure like densenet or senet at the center layers later",
    "225807": "Congratulation and thanks for sharing! By the way, could you share the setting of ReduceLROnPlateau, e.g. the monitor is val_loss or val_metrics?",
    "225824": "Thank you! I'm looking forward to your experiment results. I'm very curious.",
    "225826": "Using test data predictions by other model as if them were ground truth. In the training phase, you can use (X_train, y_train) and (X_test, y_test) for training.",
    "225828": "ReduceLROnPlateau(monitor='val_dice_coef',\n                               factor=0.2,\n                               patience=3,\n                               verbose=1,\n                               epsilon=1e-4,\n                               mode='max')\nI used above settings, but I think it is better to use validation loss as metrics.",
    "225976": "Awesome solution, I love the simplicity of it. I too tried dilated convs /Unet but I guess I did something wrong, b/c it was a bit worse; although I did not add up (or stack) the feature layers at the bottom botteneck. \n\nWhat was the intuition behind adding them up?",
    "226162": "Thank you for sharing your solution. Does it help to remove BN after each convolution layer?",
    "254247": "I tried it (with input 512x512). The dice_coef was  .91 by the first training epoch itself. This outperforms some other deeper U-Nets I also tried.",
    "330147": "Amazing solution ever!",
    "344866": "I tried executing your solution but it is unable to fetch the .png files from the train_masks folder as it contains .gif files only. Am I missing something? Can someone please help me in executing the same?",
    "383705": "Hi, @lyakaap, thank you for your sharing.\nwhat tool did you use to draw the picture?",
    "383755": "I believe he used draw.io. Very easy to use. I did a U-Net diagram and it came out just like this.",
    "385133": "Thanks a lot!",
    "431294": "Hi, I am still wondering about the intuition for adding the dilation layers. I couldnt find research papers suggesting this but it does make sense. I would love to read related papers if someone can point me to them.  I could only find the U-Net paper and Multi-scale context aggregation by dilated convolutions paper but did not find any research paper using dilated convolutions like used here.",
    "461123": "Hi, I recently looked up into this model, what could be intuition for adding binary cross entropy loss with dice loss, is there any specific reference to that, could anyone help me out",
    "501212": "very nice work！\n\nwhat does the meaning of weight w1 and w0 in loss function of “weighted_bce_dice_loss”~\n\n\nthx！",
    "539573": "Hi, Thanks for the solution. Did you convert  '.gif' files into a format such as '.png' ?. The reason behind, I wasn't able to fetch gif files directly."
  },
  "source": "meta"
}