{
  "id": 35422,
  "title": "Quick overview my solution: UNet + regression on masks",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/writeups/konstantin-lopuhin-quick-overview-my-solution-unet",
  "author_name": "",
  "post_date": "2017-06-28T10:13:15.330Z",
  "votes": 65,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Hi, congrats to @outrunner for an amazing performance, and congrats to everyone who participated! Special thanks to @threeplusone for the lion coordinates!</p>\n\n<p>Quick overview of my solution:</p>\n\n<p><strong>UNet</strong> with batch normalization and upsampling instead of deconvolution, trained on 256x256 patches with scale (about 0.8 -- 1.6) and rotation augmentation. With probability 0.2 patches were sampled close to some random sea lion. The task was to predict fixed sized squares centered at lion coordinates, softmax output with log loss was used. Here is what it looked like:\n<img src=\"http://imgur.com/I9JlU9g.png\" alt=\"UNet training\" title=\"\"></p>\n\n<p>And here is what predictions (here all classes are combined) looked like:</p>\n\n<p><img src=\"http://i.imgur.com/YST2jU1.png\" alt=\"UNet predictions (all classes)\" title=\"\"></p>\n\n<p>In order to <strong>predict the lion count</strong>, patches of about the same size (240x240) were extracted from predicted masks, and several features were extracted: mostly sum of outputs (with several thresholds) and also blobs were detected and counted. This features were extracted from predictions of each class separately, and for each class a separate regressor was trained (an ensemble of two extra trees regressors from sklearn and one xgboost regressor).</p>\n\n<p>The only <strong>trick for Test</strong> was to predict at 0.5 scale (so I downscaled test 2x) - I determined this value by just looking at the images.</p>\n\n<p>All in all I was very unsure of the leaderboard score as it often did not correlate with my local validation, so I didn't do any per-class LB tuning. For the final submission I averaged several models that were better on the LB (mid-13 public LB / mid-12 private LB), and some models that were better on my validation (they gave about 14 on public LB / 13 on private LB). </p>\n\n<p>Training took about 12-24 hours on a single 1070, and prediction took another 12 hours. Plus about 3 hours for training and running count prediction.</p>",
  "messages": [
    {
      "id": "196850",
      "postDate": "06/28/2017 08:41:53",
      "content": "<p>Hi, congrats to @outrunner for an amazing performance, and congrats to everyone who participated! Special thanks to @threeplusone for the lion coordinates!</p>\n\n<p>Quick overview of my solution:</p>\n\n<p><strong>UNet</strong> with batch normalization and upsampling instead of deconvolution, trained on 256x256 patches with scale (about 0.8 -- 1.6) and rotation augmentation. With probability 0.2 patches were sampled close to some random sea lion. The task was to predict fixed sized squares centered at lion coordinates, softmax output with log loss was used. Here is what it looked like:\n<img src=\"http://imgur.com/I9JlU9g.png\" alt=\"UNet training\" title=\"\"></p>\n\n<p>And here is what predictions (here all classes are combined) looked like:</p>\n\n<p><img src=\"http://i.imgur.com/YST2jU1.png\" alt=\"UNet predictions (all classes)\" title=\"\"></p>\n\n<p>In order to <strong>predict the lion count</strong>, patches of about the same size (240x240) were extracted from predicted masks, and several features were extracted: mostly sum of outputs (with several thresholds) and also blobs were detected and counted. This features were extracted from predictions of each class separately, and for each class a separate regressor was trained (an ensemble of two extra trees regressors from sklearn and one xgboost regressor).</p>\n\n<p>The only <strong>trick for Test</strong> was to predict at 0.5 scale (so I downscaled test 2x) - I determined this value by just looking at the images.</p>\n\n<p>All in all I was very unsure of the leaderboard score as it often did not correlate with my local validation, so I didn't do any per-class LB tuning. For the final submission I averaged several models that were better on the LB (mid-13 public LB / mid-12 private LB), and some models that were better on my validation (they gave about 14 on public LB / 13 on private LB). </p>\n\n<p>Training took about 12-24 hours on a single 1070, and prediction took another 12 hours. Plus about 3 hours for training and running count prediction.</p>",
      "rawMarkdown": "Hi, congrats to @outrunner for an amazing performance, and congrats to everyone who participated! Special thanks to @threeplusone for the lion coordinates!\n\nQuick overview of my solution:\n\n**UNet** with batch normalization and upsampling instead of deconvolution, trained on 256x256 patches with scale (about 0.8 -- 1.6) and rotation augmentation. With probability 0.2 patches were sampled close to some random sea lion. The task was to predict fixed sized squares centered at lion coordinates, softmax output with log loss was used. Here is what it looked like:\n![UNet training][1]\n\nAnd here is what predictions (here all classes are combined) looked like:\n\n![UNet predictions (all classes)][2]\n\nIn order to **predict the lion count**, patches of about the same size (240x240) were extracted from predicted masks, and several features were extracted: mostly sum of outputs (with several thresholds) and also blobs were detected and counted. This features were extracted from predictions of each class separately, and for each class a separate regressor was trained (an ensemble of two extra trees regressors from sklearn and one xgboost regressor).\n\nThe only **trick for Test** was to predict at 0.5 scale (so I downscaled test 2x) - I determined this value by just looking at the images.\n\nAll in all I was very unsure of the leaderboard score as it often did not correlate with my local validation, so I didn't do any per-class LB tuning. For the final submission I averaged several models that were better on the LB (mid-13 public LB / mid-12 private LB), and some models that were better on my validation (they gave about 14 on public LB / 13 on private LB). \n\nTraining took about 12-24 hours on a single 1070, and prediction took another 12 hours. Plus about 3 hours for training and running count prediction.\n\n  [1]: http://imgur.com/I9JlU9g.png\n  [2]: http://i.imgur.com/YST2jU1.png",
      "votes": null
    },
    {
      "id": "196857",
      "postDate": "06/28/2017 09:05:53",
      "content": "<p>Congratulations and thanks.</p>",
      "rawMarkdown": "Congratulations and thanks.",
      "votes": null
    },
    {
      "id": "196863",
      "postDate": "06/28/2017 09:16:03",
      "content": "<p>Congrats.. well done!</p>\n\n<p>Can we see the code? our team was trying something very similar but we were stuck somewhere in the implementation.</p>",
      "rawMarkdown": "Congrats.. well done!\n\nCan we see the code? our team was trying something very similar but we were stuck somewhere in the implementation.",
      "votes": null
    },
    {
      "id": "196875",
      "postDate": "06/28/2017 09:32:15",
      "content": "<p>Yes, sure - I will post the code in 1-2 weeks or earlier, will post the link here. </p>",
      "rawMarkdown": "Yes, sure - I will post the code in 1-2 weeks or earlier, will post the link here.",
      "votes": null
    },
    {
      "id": "196882",
      "postDate": "06/28/2017 09:51:36",
      "content": "<p>Nice work. Congrats on getting UNET to work so successfully. I've tried it out on a couple of projects and have found it to be a pain in the arse to train :)</p>",
      "rawMarkdown": "Nice work. Congrats on getting UNET to work so successfully. I've tried it out on a couple of projects and have found it to be a pain in the arse to train :)",
      "votes": null
    },
    {
      "id": "196886",
      "postDate": "06/28/2017 10:16:59",
      "content": "<p>Thanks! Yeah, training UNet can be painful!</p>\n\n<p>Forgot to mention one more difference from vanilla UNet: I'm using upsampling instead of deconvolutions, on other problems it was much more robust and produced better looking masks.</p>\n\n<p>I also could not make UNet to work properly with dice loss (which should work better in case of class imbalance), and also UNet that predicted 4x smaller output also didn't work so well (which is a pity because it was much faster and more convenient).</p>",
      "rawMarkdown": "Thanks! Yeah, training UNet can be painful!\n\nForgot to mention one more difference from vanilla UNet: I'm using upsampling instead of deconvolutions, on other problems it was much more robust and produced better looking masks.\n\nI also could not make UNet to work properly with dice loss (which should work better in case of class imbalance), and also UNet that predicted 4x smaller output also didn't work so well (which is a pity because it was much faster and more convenient).",
      "votes": null
    },
    {
      "id": "196888",
      "postDate": "06/28/2017 10:22:16",
      "content": "<p>Congrats for getting segmentation to work, I have struggled to get decent results with Tiramisu. Looking forward to look at your code :P</p>",
      "rawMarkdown": "Congrats for getting segmentation to work, I have struggled to get decent results with Tiramisu. Looking forward to look at your code :P",
      "votes": null
    },
    {
      "id": "196891",
      "postDate": "06/28/2017 10:34:20",
      "content": "<p>Excellent work Konstantin, congratulations and thanks also for sharing a lot on forum throughout the competition.    Wondering if you might be willing to try outrunner's simple and elegant postprocessing trick:  add 50% juveniles, subtract adult_females with the same amount, and add 20% pups.   Does this improve your public/private scores to his amazing levels?   For us it improved public by 2.3 and private by 1.8.   You can likely tune the percentages even better given the powerful properties of your nice UNet model.</p>",
      "rawMarkdown": "Excellent work Konstantin, congratulations and thanks also for sharing a lot on forum throughout the competition.    Wondering if you might be willing to try outrunner's simple and elegant postprocessing trick:  add 50% juveniles, subtract adult_females with the same amount, and add 20% pups.   Does this improve your public/private scores to his amazing levels?   For us it improved public by 2.3 and private by 1.8.   You can likely tune the percentages even better given the powerful properties of your nice UNet model.",
      "votes": null
    },
    {
      "id": "196926",
      "postDate": "06/28/2017 12:14:06",
      "content": "<p>Congratulations. Thanks for sharing the details both before and after the competition deadline :). It helped us a lot. I was working a similar approach but was not getting good accuracy. Would you mind answering a few questions - </p>\n\n<ol>\n<li>How was the performance with deconvolution layers instead of upsampling? </li>\n<li>I am guessing, you used 6 layer masks as Unet output with cross entropy without weighs for your Unet.</li>\n<li>Did you encounter problems with misclassification/ crowded sea lions ? ( I was using Unet and most of the sea-lions were being misclassified as adult-females).</li>\n<li>How was the training done ? I mean, how many epochs were trained to get these results ? Also, what kind of optimizer and lr was used? In my case, after 1-2 epochs, the log loss stops decreasing and remains constant. The output masks at this point remain inaccurate. The output masks had sea-lion shaped predictions but they were being misclassified and some other structures were also classified as sea-lions.</li>\n<li>Is the purple background in the output image coming from the background label?</li>\n</ol>\n\n<p>Thanks again and congrats.</p>",
      "rawMarkdown": "Congratulations. Thanks for sharing the details both before and after the competition deadline :). It helped us a lot. I was working a similar approach but was not getting good accuracy. Would you mind answering a few questions - \n\n1. How was the performance with deconvolution layers instead of upsampling? \n2. I am guessing, you used 6 layer masks as Unet output with cross entropy without weighs for your Unet.\n3.  Did you encounter problems with misclassification/ crowded sea lions ? ( I was using Unet and most of the sea-lions were being misclassified as adult-females).\n4. How was the training done ? I mean, how many epochs were trained to get these results ? Also, what kind of optimizer and lr was used? In my case, after 1-2 epochs, the log loss stops decreasing and remains constant. The output masks at this point remain inaccurate. The output masks had sea-lion shaped predictions but they were being misclassified and some other structures were also classified as sea-lions.\n5. Is the purple background in the output image coming from the background label?\n\nThanks again and congrats.",
      "votes": null
    },
    {
      "id": "196927",
      "postDate": "06/28/2017 12:17:09",
      "content": "<p>Thanks Konstantin and great work. Quick question, did you learn the categories separately? In other words your log loss was over 5 categories?</p>",
      "rawMarkdown": "Thanks Konstantin and great work. Quick question, did you learn the categories separately? In other words your log loss was over 5 categories?",
      "votes": null
    },
    {
      "id": "196942",
      "postDate": "06/28/2017 13:01:39",
      "content": "<p>Thanks Kapil!</p>\n\n<blockquote>\n  <p>How was the performance with deconvolution layers instead of upsampling?</p>\n</blockquote>\n\n<p>I didn't check on this task, but on DSTL Satellite image segmentation I just could not get good results with deconvolutions at all - maybe it was a problem with my implementation (because others used them just fine), or bad init - I was using PyTorch. So here I just went with what I knew worked.</p>\n\n<blockquote>\n  <p>I am guessing, you used 6 layer masks as Unet output with cross entropy without weighs for your Unet.</p>\n</blockquote>\n\n<p>Right. I tried applying a different weight to the background, but it resulted in more false positives and lower quality masks. On the other hand, oversampling of images with lions helped - which makes sense in retrospect.</p>\n\n<blockquote>\n  <p>Did you encounter problems with misclassification/ crowded sea lions ? ( I was using Unet and most of the sea-lions were being misclassified as adult-females).</p>\n</blockquote>\n\n<p>Yes to both, and to be honest, I didn't manage to implement a good solution to any of those. For crowded areas maybe xgboost regressor + custom features helped a little. For misclassificatoin, I added a VGG-based classifier that tried to re-classify blobs, but I implemented it only two days before the deadline and didn't manage to get any gain from it - still curious if this would help.</p>\n\n<blockquote>\n  <p>How was the training done ? I mean, how many epochs were trained to get these results ? Also, what kind of optimizer and lr was used? In my case, after 1-2 epochs, the log loss stops decreasing and remains constant. The output masks at this point remain inaccurate. </p>\n</blockquote>\n\n<p>For me one epoch covered entire training set (but I sampled randomly and did all augmentations on the fly), and I got best results with training for 10-20 epochs. I used Adam with lr 0.0001 and batch size 32 (but I multiplied loss by batch size) - it was much better than SGD. I divided lr by 5 when validation loss stopped decreasing for more than two epochs (just once usually). Both training and validation losses continued to go down, but it turned out it's better to stop training even before the best validation loss is reached - probably due to difference between train and test.</p>\n\n<blockquote>\n  <p>The output masks had sea-lion shaped predictions but they were being misclassified and some other structures were also classified as sea-lions.</p>\n</blockquote>\n\n<p>I've seen such shapes early in training, maybe after the first few epochs. I also saw some misclassifications and some false positives (especially on the test set), didn't get to solve it.</p>\n\n<blockquote>\n  <p>Is the purple background in the output image coming from the background label?</p>\n</blockquote>\n\n<p>Yes - it's just some default way of visualising masks with matplotlib, definitely not the best way to show it.</p>",
      "rawMarkdown": "Thanks Kapil!\n\n&gt; How was the performance with deconvolution layers instead of upsampling?\n\nI didn't check on this task, but on DSTL Satellite image segmentation I just could not get good results with deconvolutions at all - maybe it was a problem with my implementation (because others used them just fine), or bad init - I was using PyTorch. So here I just went with what I knew worked.\n\n&gt; I am guessing, you used 6 layer masks as Unet output with cross entropy without weighs for your Unet.\n\nRight. I tried applying a different weight to the background, but it resulted in more false positives and lower quality masks. On the other hand, oversampling of images with lions helped - which makes sense in retrospect.\n\n&gt; Did you encounter problems with misclassification/ crowded sea lions ? ( I was using Unet and most of the sea-lions were being misclassified as adult-females).\n\nYes to both, and to be honest, I didn't manage to implement a good solution to any of those. For crowded areas maybe xgboost regressor + custom features helped a little. For misclassificatoin, I added a VGG-based classifier that tried to re-classify blobs, but I implemented it only two days before the deadline and didn't manage to get any gain from it - still curious if this would help.\n\n&gt; How was the training done ? I mean, how many epochs were trained to get these results ? Also, what kind of optimizer and lr was used? In my case, after 1-2 epochs, the log loss stops decreasing and remains constant. The output masks at this point remain inaccurate. \n\nFor me one epoch covered entire training set (but I sampled randomly and did all augmentations on the fly), and I got best results with training for 10-20 epochs. I used Adam with lr 0.0001 and batch size 32 (but I multiplied loss by batch size) - it was much better than SGD. I divided lr by 5 when validation loss stopped decreasing for more than two epochs (just once usually). Both training and validation losses continued to go down, but it turned out it's better to stop training even before the best validation loss is reached - probably due to difference between train and test.\n\n&gt; The output masks had sea-lion shaped predictions but they were being misclassified and some other structures were also classified as sea-lions.\n\nI've seen such shapes early in training, maybe after the first few epochs. I also saw some misclassifications and some false positives (especially on the test set), didn't get to solve it.\n\n&gt; Is the purple background in the output image coming from the background label?\n\nYes - it's just some default way of visualising masks with matplotlib, definitely not the best way to show it.",
      "votes": null
    },
    {
      "id": "196944",
      "postDate": "06/28/2017 13:04:35",
      "content": "<p>Thanks @gbhalla! I used a log loss with 6 classes, one for background and the other for different lions - so it's the same loss you normally use for classification problems, but applied per-pixel on the output mask.</p>",
      "rawMarkdown": "Thanks @gbhalla! I used a log loss with 6 classes, one for background and the other for different lions - so it's the same loss you normally use for classification problems, but applied per-pixel on the output mask.",
      "votes": null
    },
    {
      "id": "196962",
      "postDate": "06/28/2017 13:44:47",
      "content": "<p>Thanks for the detailed insights of your approach.</p>\n\n<p>Just one more question and then I will stop bugging you :)  - </p>\n\n<p>By \"With probability 0.2 patches were sampled close to some random sea lion\", do you mean that a sea-lion in train images was selected randomly and then several patches were cropped within some specified distance from that sea-lion ? But then where does the 0.2 probability come into play ? </p>",
      "rawMarkdown": "Thanks for the detailed insights of your approach.\n\nJust one more question and then I will stop bugging you :)  - \n\nBy \"With probability 0.2 patches were sampled close to some random sea lion\", do you mean that a sea-lion in train images was selected randomly and then several patches were cropped within some specified distance from that sea-lion ? But then where does the 0.2 probability come into play ?",
      "votes": null
    },
    {
      "id": "196969",
      "postDate": "06/28/2017 13:50:01",
      "content": "<p>I'm happy to answer the questions :)</p>\n\n<p>I sampled patches that formed the batch completely independently. With probability 0.8, I sampled a random image, a random location in that image, a random scale and random angle, and created a patch from it. With probability 0.2, I sampled a random lion (uniformly across all lions, so crowded images got more samples), and then sampled a random location within a pre-defined distance (~patch size) from that lion, and then again the angle and scale.</p>",
      "rawMarkdown": "I'm happy to answer the questions :)\n\nI sampled patches that formed the batch completely independently. With probability 0.8, I sampled a random image, a random location in that image, a random scale and random angle, and created a patch from it. With probability 0.2, I sampled a random lion (uniformly across all lions, so crowded images got more samples), and then sampled a random location within a pre-defined distance (~patch size) from that lion, and then again the angle and scale.",
      "votes": null
    },
    {
      "id": "197016",
      "postDate": "06/28/2017 16:12:09",
      "content": "<p>Kostia,\nWhich UNet specifically did you use? Based on VGG ?</p>",
      "rawMarkdown": "Kostia,\nWhich UNet specifically did you use? Based on VGG ?",
      "votes": null
    },
    {
      "id": "197018",
      "postDate": "06/28/2017 16:13:04",
      "content": "<p>Could you please elaborate a little bit on blob detection and counting?</p>",
      "rawMarkdown": "Could you please elaborate a little bit on blob detection and counting?",
      "votes": null
    },
    {
      "id": "197024",
      "postDate": "06/28/2017 16:21:05",
      "content": "<p>Yes, I would say it's a classical UNet, and it's similar to bn-VGG, just conv3s everywhere. It's almost the same as this <a href=\"https://github.com/lopuhin/kaggle-dstl/blob/master/models.py#L188\">https://github.com/lopuhin/kaggle-dstl/blob/master/models.py#L188</a> with a small twist: it has 4x pooling instead of 2x in the \"deepest\" part - this save a bit of memory and in theory gives a larger receptive field.</p>",
      "rawMarkdown": "Yes, I would say it's a classical UNet, and it's similar to bn-VGG, just conv3s everywhere. It's almost the same as this https://github.com/lopuhin/kaggle-dstl/blob/master/models.py#L188 with a small twist: it has 4x pooling instead of 2x in the \"deepest\" part - this save a bit of memory and in theory gives a larger receptive field.",
      "votes": null
    },
    {
      "id": "197025",
      "postDate": "06/28/2017 16:24:07",
      "content": "<p>Sure - I used blob_log from skimage to detect blobs <a href=\"http://scikit-image.org/docs/dev/api/skimage.feature.html#skimage.feature.blob_log\">http://scikit-image.org/docs/dev/api/skimage.feature.html#skimage.feature.blob_log</a> and then generated two kinds of features from them: \"blob-sum\" is the sum of probabilities at blob centers (so if we have two blobs with max probability 0.3 and 0.05 the value would be 0.35), and \"blob-count\" is just the number of blobs (so here the value of the feature would be 2). I also did this for several blob thresholds (minimal blob intensities). In order to speed things up and to save space, I saved 4x downscaled prediction, and ran blob detection at this scale too, so it was not so slow.</p>",
      "rawMarkdown": "Sure - I used blob_log from skimage to detect blobs http://scikit-image.org/docs/dev/api/skimage.feature.html#skimage.feature.blob_log and then generated two kinds of features from them: \"blob-sum\" is the sum of probabilities at blob centers (so if we have two blobs with max probability 0.3 and 0.05 the value would be 0.35), and \"blob-count\" is just the number of blobs (so here the value of the feature would be 2). I also did this for several blob thresholds (minimal blob intensities). In order to speed things up and to save space, I saved 4x downscaled prediction, and ran blob detection at this scale too, so it was not so slow.",
      "votes": null
    },
    {
      "id": "197076",
      "postDate": "06/28/2017 18:32:01",
      "content": "<p>Congratulation!!  and thanks for sharing!</p>\n\n<p>So you resized the test images by half. Did it not make all the sea lions smaller compare to training? (in pixels). Or did you resize the train images too?</p>\n\n<p>How did you deal with half cut sea lions on the patches (like the tail in on patch and the other half with the head on another patch)?</p>",
      "rawMarkdown": "Congratulation!!  and thanks for sharing!\n\nSo you resized the test images by half. Did it not make all the sea lions smaller compare to training? (in pixels). Or did you resize the train images too?\n\nHow did you deal with half cut sea lions on the patches (like the tail in on patch and the other half with the head on another patch)?",
      "votes": null
    },
    {
      "id": "197077",
      "postDate": "06/28/2017 18:32:33",
      "content": "<p>that's very cool of you! thanks!</p>",
      "rawMarkdown": "that's very cool of you! thanks!",
      "votes": null
    },
    {
      "id": "197084",
      "postDate": "06/28/2017 18:58:12",
      "content": "<p>Thanks!</p>\n\n<blockquote>\n  <p>So you resized the test images by half. Did it not make all the sea lions smaller compare to training? (in pixels). Or did you resize the train images too?</p>\n</blockquote>\n\n<p>I think that images in test are about 2x bigger than what we have in the training set. Probably they used a different camera, or didn't resize them, or flew lower, or something like this :)</p>\n\n<p>Still the images in test had different scales, so I applied scale augmentation during training to make the model more robust. Also I was not sure if 2x was the precise difference, and didn't really want to fit it using the public LB.</p>\n\n<blockquote>\n  <p>How did you deal with half cut sea lions on the patches (like the tail in on patch and the other half with the head on another patch)?</p>\n</blockquote>\n\n<p>I averaged overlapping predictions at test time. At train time I didn't do anything special.</p>",
      "rawMarkdown": "Thanks!\n\n&gt; So you resized the test images by half. Did it not make all the sea lions smaller compare to training? (in pixels). Or did you resize the train images too?\n\nI think that images in test are about 2x bigger than what we have in the training set. Probably they used a different camera, or didn't resize them, or flew lower, or something like this :)\n\nStill the images in test had different scales, so I applied scale augmentation during training to make the model more robust. Also I was not sure if 2x was the precise difference, and didn't really want to fit it using the public LB.\n\n&gt; How did you deal with half cut sea lions on the patches (like the tail in on patch and the other half with the head on another patch)?\n\nI averaged overlapping predictions at test time. At train time I didn't do anything special.",
      "votes": null
    },
    {
      "id": "197110",
      "postDate": "06/28/2017 20:03:08",
      "content": "<p>Wonderful and congratulations</p>",
      "rawMarkdown": "Wonderful and congratulations",
      "votes": null
    },
    {
      "id": "197120",
      "postDate": "06/28/2017 20:44:17",
      "content": "<p>@Konstantin, congrats and thanks a lot for sharing hints on the forum during the competition: there were really helpful and I learnt a lot despite the fact I never take off from the bottom of the leaderboard.</p>\n\n<p>@KapilYadav: if I may despite my worst result ever in a kaggle competition, I had the same problem to initiate unet and many versions of those (using sigmoid in particular) didn't give me any results. I found that initiate unet with a small portion of the pictures (1/4 to 1/10 of the most crowed pictures) and using elu instead of relu until it reaches a plateau allows the unet to converge pretty quickly usually in a quarter of the time of relu and then saving the weights and restarting the unet with relu activation makes the trick.</p>",
      "rawMarkdown": "Konstantin, congrats and thanks a lot for sharing hints on the forum during the competition: there were really helpful and I learnt a lot despite the fact I never take off from the bottom of the leaderboard.\n\n@KapilYadav: if I may despite my worst result ever in a kaggle competition, I had the same problem to initiate unet and many versions of those (using sigmoid in particular) didn't give me any results. I found that initiate unet with a small portion of the pictures (1/4 to 1/10 of the most crowed pictures) and using elu instead of relu until it reaches a plateau allows the unet to converge pretty quickly usually in a quarter of the time of relu and then saving the weights and restarting the unet with relu activation makes the trick.",
      "votes": null
    },
    {
      "id": "197126",
      "postDate": "06/28/2017 21:38:18",
      "content": "<p>Great work Konstantin! Thanks for sharing</p>",
      "rawMarkdown": "Great work Konstantin! Thanks for sharing",
      "votes": null
    },
    {
      "id": "197153",
      "postDate": "06/28/2017 22:38:45",
      "content": "<p>@Charles Jansen, You can check our solution. <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35442\">https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35442</a> <br>\nTo mitigate the effect of cutting a lion on half I placed gaussians on top of each lion and regressed the sum over all Gaussians in the current tile. </p>",
      "rawMarkdown": "Charles Jansen, You can check our solution. https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35442  \nTo mitigate the effect of cutting a lion on half I placed gaussians on top of each lion and regressed the sum over all Gaussians in the current tile.",
      "votes": null
    },
    {
      "id": "197157",
      "postDate": "06/28/2017 22:43:40",
      "content": "<p>It's very  high level overview, so maybe it's better to look at the code after, but I have one question: I haven't tried semantic segmentation losses yet, but I was able to train density map regression with U-net only for all classes collapsed to one class and only for 'positive' tiles, so I wonder is it so hard to train U-net or maybe it's due to class imbalance? do you balance your data?</p>\n\n<p>P.S. thanks for your comments in Discusssion threads.</p>",
      "rawMarkdown": "It's very  high level overview, so maybe it's better to look at the code after, but I have one question: I haven't tried semantic segmentation losses yet, but I was able to train density map regression with U-net only for all classes collapsed to one class and only for 'positive' tiles, so I wonder is it so hard to train U-net or maybe it's due to class imbalance? do you balance your data?\n\nP.S. thanks for your comments in Discusssion threads.",
      "votes": null
    },
    {
      "id": "197193",
      "postDate": "06/29/2017 01:20:29",
      "content": "<p>I tried the density map regression as well and had classes collapsing too. I think this is just a limitation of the density map regression. It doesn't work with more than one class.</p>",
      "rawMarkdown": "I tried the density map regression as well and had classes collapsing too. I think this is just a limitation of the density map regression. It doesn't work with more than one class.",
      "votes": null
    },
    {
      "id": "197198",
      "postDate": "06/29/2017 01:32:47",
      "content": "<p>It's very impressive and clever to extract features from the first stage prediction results, and regress to the final counts. I also used a method based on masks regression, but I used the peak detection results as my final submission directly, which introduced lots of uncontrollable errors.</p>",
      "rawMarkdown": "It's very impressive and clever to extract features from the first stage prediction results, and regress to the final counts. I also used a method based on masks regression, but I used the peak detection results as my final submission directly, which introduced lots of uncontrollable errors.",
      "votes": null
    },
    {
      "id": "203361",
      "postDate": "07/14/2017 17:54:27",
      "content": "<p>Thanks Russ, I really liked your solution, I think it's the most principled and robust of the ones shared due to explicit accounting for scale.</p>\n\n<p>Finally got to check this clever hack, a milder version of it (just increasing pups by 20%) also gives a moderate boost: 12.5 -&gt; 12.0 on private (and 13.2 -&gt; 12.9 on public), but the optimal settings are probably different  for my model.</p>",
      "rawMarkdown": "Thanks Russ, I really liked your solution, I think it's the most principled and robust of the ones shared due to explicit accounting for scale.\n\nFinally got to check this clever hack, a milder version of it (just increasing pups by 20%) also gives a moderate boost: 12.5 -&gt; 12.0 on private (and 13.2 -&gt; 12.9 on public), but the optimal settings are probably different  for my model.",
      "votes": null
    },
    {
      "id": "217600",
      "postDate": "08/31/2017 09:18:07",
      "content": "<p>Here is the code: <a href=\"https://github.com/lopuhin/kaggle-lions-2017\">https://github.com/lopuhin/kaggle-lions-2017</a> (including all failed experiments, so in some places the code is less clean then I'd like)</p>",
      "rawMarkdown": "Here is the code: https://github.com/lopuhin/kaggle-lions-2017 (including all failed experiments, so in some places the code is less clean then I'd like)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 196857,
      "author_name": "outrunner",
      "author_url": "",
      "post_date": "06/28/2017 09:05:53",
      "content": "<p>Congratulations and thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 196863,
      "author_name": "tour1st",
      "author_url": "",
      "post_date": "06/28/2017 09:16:03",
      "content": "<p>Congrats.. well done!</p>\n\n<p>Can we see the code? our team was trying something very similar but we were stuck somewhere in the implementation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 196875,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/28/2017 09:32:15",
          "content": "<p>Yes, sure - I will post the code in 1-2 weeks or earlier, will post the link here. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 197077,
          "author_name": "cjansen",
          "author_url": "",
          "post_date": "06/28/2017 18:32:33",
          "content": "<p>that's very cool of you! thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 196882,
      "author_name": "garethjns",
      "author_url": "",
      "post_date": "06/28/2017 09:51:36",
      "content": "<p>Nice work. Congrats on getting UNET to work so successfully. I've tried it out on a couple of projects and have found it to be a pain in the arse to train :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 196886,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/28/2017 10:16:59",
          "content": "<p>Thanks! Yeah, training UNet can be painful!</p>\n\n<p>Forgot to mention one more difference from vanilla UNet: I'm using upsampling instead of deconvolutions, on other problems it was much more robust and produced better looking masks.</p>\n\n<p>I also could not make UNet to work properly with dice loss (which should work better in case of class imbalance), and also UNet that predicted 4x smaller output also didn't work so well (which is a pity because it was much faster and more convenient).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 196888,
      "author_name": "syeddanish",
      "author_url": "",
      "post_date": "06/28/2017 10:22:16",
      "content": "<p>Congrats for getting segmentation to work, I have struggled to get decent results with Tiramisu. Looking forward to look at your code :P</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 196891,
      "author_name": "sasrdw",
      "author_url": "",
      "post_date": "06/28/2017 10:34:20",
      "content": "<p>Excellent work Konstantin, congratulations and thanks also for sharing a lot on forum throughout the competition.    Wondering if you might be willing to try outrunner's simple and elegant postprocessing trick:  add 50% juveniles, subtract adult_females with the same amount, and add 20% pups.   Does this improve your public/private scores to his amazing levels?   For us it improved public by 2.3 and private by 1.8.   You can likely tune the percentages even better given the powerful properties of your nice UNet model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 203361,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "07/14/2017 17:54:27",
          "content": "<p>Thanks Russ, I really liked your solution, I think it's the most principled and robust of the ones shared due to explicit accounting for scale.</p>\n\n<p>Finally got to check this clever hack, a milder version of it (just increasing pups by 20%) also gives a moderate boost: 12.5 -&gt; 12.0 on private (and 13.2 -&gt; 12.9 on public), but the optimal settings are probably different  for my model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 196926,
      "author_name": "kapilyadav",
      "author_url": "",
      "post_date": "06/28/2017 12:14:06",
      "content": "<p>Congratulations. Thanks for sharing the details both before and after the competition deadline :). It helped us a lot. I was working a similar approach but was not getting good accuracy. Would you mind answering a few questions - </p>\n\n<ol>\n<li>How was the performance with deconvolution layers instead of upsampling? </li>\n<li>I am guessing, you used 6 layer masks as Unet output with cross entropy without weighs for your Unet.</li>\n<li>Did you encounter problems with misclassification/ crowded sea lions ? ( I was using Unet and most of the sea-lions were being misclassified as adult-females).</li>\n<li>How was the training done ? I mean, how many epochs were trained to get these results ? Also, what kind of optimizer and lr was used? In my case, after 1-2 epochs, the log loss stops decreasing and remains constant. The output masks at this point remain inaccurate. The output masks had sea-lion shaped predictions but they were being misclassified and some other structures were also classified as sea-lions.</li>\n<li>Is the purple background in the output image coming from the background label?</li>\n</ol>\n\n<p>Thanks again and congrats.</p>",
      "votes": null,
      "replies": [
        {
          "id": 196942,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/28/2017 13:01:39",
          "content": "<p>Thanks Kapil!</p>\n\n<blockquote>\n  <p>How was the performance with deconvolution layers instead of upsampling?</p>\n</blockquote>\n\n<p>I didn't check on this task, but on DSTL Satellite image segmentation I just could not get good results with deconvolutions at all - maybe it was a problem with my implementation (because others used them just fine), or bad init - I was using PyTorch. So here I just went with what I knew worked.</p>\n\n<blockquote>\n  <p>I am guessing, you used 6 layer masks as Unet output with cross entropy without weighs for your Unet.</p>\n</blockquote>\n\n<p>Right. I tried applying a different weight to the background, but it resulted in more false positives and lower quality masks. On the other hand, oversampling of images with lions helped - which makes sense in retrospect.</p>\n\n<blockquote>\n  <p>Did you encounter problems with misclassification/ crowded sea lions ? ( I was using Unet and most of the sea-lions were being misclassified as adult-females).</p>\n</blockquote>\n\n<p>Yes to both, and to be honest, I didn't manage to implement a good solution to any of those. For crowded areas maybe xgboost regressor + custom features helped a little. For misclassificatoin, I added a VGG-based classifier that tried to re-classify blobs, but I implemented it only two days before the deadline and didn't manage to get any gain from it - still curious if this would help.</p>\n\n<blockquote>\n  <p>How was the training done ? I mean, how many epochs were trained to get these results ? Also, what kind of optimizer and lr was used? In my case, after 1-2 epochs, the log loss stops decreasing and remains constant. The output masks at this point remain inaccurate. </p>\n</blockquote>\n\n<p>For me one epoch covered entire training set (but I sampled randomly and did all augmentations on the fly), and I got best results with training for 10-20 epochs. I used Adam with lr 0.0001 and batch size 32 (but I multiplied loss by batch size) - it was much better than SGD. I divided lr by 5 when validation loss stopped decreasing for more than two epochs (just once usually). Both training and validation losses continued to go down, but it turned out it's better to stop training even before the best validation loss is reached - probably due to difference between train and test.</p>\n\n<blockquote>\n  <p>The output masks had sea-lion shaped predictions but they were being misclassified and some other structures were also classified as sea-lions.</p>\n</blockquote>\n\n<p>I've seen such shapes early in training, maybe after the first few epochs. I also saw some misclassifications and some false positives (especially on the test set), didn't get to solve it.</p>\n\n<blockquote>\n  <p>Is the purple background in the output image coming from the background label?</p>\n</blockquote>\n\n<p>Yes - it's just some default way of visualising masks with matplotlib, definitely not the best way to show it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 196962,
          "author_name": "kapilyadav",
          "author_url": "",
          "post_date": "06/28/2017 13:44:47",
          "content": "<p>Thanks for the detailed insights of your approach.</p>\n\n<p>Just one more question and then I will stop bugging you :)  - </p>\n\n<p>By \"With probability 0.2 patches were sampled close to some random sea lion\", do you mean that a sea-lion in train images was selected randomly and then several patches were cropped within some specified distance from that sea-lion ? But then where does the 0.2 probability come into play ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 196969,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/28/2017 13:50:01",
          "content": "<p>I'm happy to answer the questions :)</p>\n\n<p>I sampled patches that formed the batch completely independently. With probability 0.8, I sampled a random image, a random location in that image, a random scale and random angle, and created a patch from it. With probability 0.2, I sampled a random lion (uniformly across all lions, so crowded images got more samples), and then sampled a random location within a pre-defined distance (~patch size) from that lion, and then again the angle and scale.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 196927,
      "author_name": "gbhalla",
      "author_url": "",
      "post_date": "06/28/2017 12:17:09",
      "content": "<p>Thanks Konstantin and great work. Quick question, did you learn the categories separately? In other words your log loss was over 5 categories?</p>",
      "votes": null,
      "replies": [
        {
          "id": 196944,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/28/2017 13:04:35",
          "content": "<p>Thanks @gbhalla! I used a log loss with 6 classes, one for background and the other for different lions - so it's the same loss you normally use for classification problems, but applied per-pixel on the output mask.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 197016,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "06/28/2017 16:12:09",
      "content": "<p>Kostia,\nWhich UNet specifically did you use? Based on VGG ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 197024,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/28/2017 16:21:05",
          "content": "<p>Yes, I would say it's a classical UNet, and it's similar to bn-VGG, just conv3s everywhere. It's almost the same as this <a href=\"https://github.com/lopuhin/kaggle-dstl/blob/master/models.py#L188\">https://github.com/lopuhin/kaggle-dstl/blob/master/models.py#L188</a> with a small twist: it has 4x pooling instead of 2x in the \"deepest\" part - this save a bit of memory and in theory gives a larger receptive field.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 197018,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "06/28/2017 16:13:04",
      "content": "<p>Could you please elaborate a little bit on blob detection and counting?</p>",
      "votes": null,
      "replies": [
        {
          "id": 197025,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/28/2017 16:24:07",
          "content": "<p>Sure - I used blob_log from skimage to detect blobs <a href=\"http://scikit-image.org/docs/dev/api/skimage.feature.html#skimage.feature.blob_log\">http://scikit-image.org/docs/dev/api/skimage.feature.html#skimage.feature.blob_log</a> and then generated two kinds of features from them: \"blob-sum\" is the sum of probabilities at blob centers (so if we have two blobs with max probability 0.3 and 0.05 the value would be 0.35), and \"blob-count\" is just the number of blobs (so here the value of the feature would be 2). I also did this for several blob thresholds (minimal blob intensities). In order to speed things up and to save space, I saved 4x downscaled prediction, and ran blob detection at this scale too, so it was not so slow.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 197076,
      "author_name": "cjansen",
      "author_url": "",
      "post_date": "06/28/2017 18:32:01",
      "content": "<p>Congratulation!!  and thanks for sharing!</p>\n\n<p>So you resized the test images by half. Did it not make all the sea lions smaller compare to training? (in pixels). Or did you resize the train images too?</p>\n\n<p>How did you deal with half cut sea lions on the patches (like the tail in on patch and the other half with the head on another patch)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 197084,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/28/2017 18:58:12",
          "content": "<p>Thanks!</p>\n\n<blockquote>\n  <p>So you resized the test images by half. Did it not make all the sea lions smaller compare to training? (in pixels). Or did you resize the train images too?</p>\n</blockquote>\n\n<p>I think that images in test are about 2x bigger than what we have in the training set. Probably they used a different camera, or didn't resize them, or flew lower, or something like this :)</p>\n\n<p>Still the images in test had different scales, so I applied scale augmentation during training to make the model more robust. Also I was not sure if 2x was the precise difference, and didn't really want to fit it using the public LB.</p>\n\n<blockquote>\n  <p>How did you deal with half cut sea lions on the patches (like the tail in on patch and the other half with the head on another patch)?</p>\n</blockquote>\n\n<p>I averaged overlapping predictions at test time. At train time I didn't do anything special.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 197153,
          "author_name": "asanakoev",
          "author_url": "",
          "post_date": "06/28/2017 22:38:45",
          "content": "<p>@Charles Jansen, You can check our solution. <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35442\">https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35442</a> <br>\nTo mitigate the effect of cutting a lion on half I placed gaussians on top of each lion and regressed the sum over all Gaussians in the current tile. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 197110,
      "author_name": "abriosi",
      "author_url": "",
      "post_date": "06/28/2017 20:03:08",
      "content": "<p>Wonderful and congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 197120,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "06/28/2017 20:44:17",
      "content": "<p>@Konstantin, congrats and thanks a lot for sharing hints on the forum during the competition: there were really helpful and I learnt a lot despite the fact I never take off from the bottom of the leaderboard.</p>\n\n<p>@KapilYadav: if I may despite my worst result ever in a kaggle competition, I had the same problem to initiate unet and many versions of those (using sigmoid in particular) didn't give me any results. I found that initiate unet with a small portion of the pictures (1/4 to 1/10 of the most crowed pictures) and using elu instead of relu until it reaches a plateau allows the unet to converge pretty quickly usually in a quarter of the time of relu and then saving the weights and restarting the unet with relu activation makes the trick.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 197126,
      "author_name": "darraghdog",
      "author_url": "",
      "post_date": "06/28/2017 21:38:18",
      "content": "<p>Great work Konstantin! Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 197157,
      "author_name": "mrgloom",
      "author_url": "",
      "post_date": "06/28/2017 22:43:40",
      "content": "<p>It's very  high level overview, so maybe it's better to look at the code after, but I have one question: I haven't tried semantic segmentation losses yet, but I was able to train density map regression with U-net only for all classes collapsed to one class and only for 'positive' tiles, so I wonder is it so hard to train U-net or maybe it's due to class imbalance? do you balance your data?</p>\n\n<p>P.S. thanks for your comments in Discusssion threads.</p>",
      "votes": null,
      "replies": [
        {
          "id": 197193,
          "author_name": "asanakoev",
          "author_url": "",
          "post_date": "06/29/2017 01:20:29",
          "content": "<p>I tried the density map regression as well and had classes collapsing too. I think this is just a limitation of the density map regression. It doesn't work with more than one class.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 197198,
      "author_name": "lzhang57",
      "author_url": "",
      "post_date": "06/29/2017 01:32:47",
      "content": "<p>It's very impressive and clever to extract features from the first stage prediction results, and regress to the final counts. I also used a method based on masks regression, but I used the peak detection results as my final submission directly, which introduced lots of uncontrollable errors.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 217600,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "08/31/2017 09:18:07",
      "content": "<p>Here is the code: <a href=\"https://github.com/lopuhin/kaggle-lions-2017\">https://github.com/lopuhin/kaggle-lions-2017</a> (including all failed experiments, so in some places the code is less clean then I'd like)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "196850": "Hi, congrats to @outrunner for an amazing performance, and congrats to everyone who participated! Special thanks to @threeplusone for the lion coordinates!\n\nQuick overview of my solution:\n\n**UNet** with batch normalization and upsampling instead of deconvolution, trained on 256x256 patches with scale (about 0.8 -- 1.6) and rotation augmentation. With probability 0.2 patches were sampled close to some random sea lion. The task was to predict fixed sized squares centered at lion coordinates, softmax output with log loss was used. Here is what it looked like:\n![UNet training][1]\n\nAnd here is what predictions (here all classes are combined) looked like:\n\n![UNet predictions (all classes)][2]\n\nIn order to **predict the lion count**, patches of about the same size (240x240) were extracted from predicted masks, and several features were extracted: mostly sum of outputs (with several thresholds) and also blobs were detected and counted. This features were extracted from predictions of each class separately, and for each class a separate regressor was trained (an ensemble of two extra trees regressors from sklearn and one xgboost regressor).\n\nThe only **trick for Test** was to predict at 0.5 scale (so I downscaled test 2x) - I determined this value by just looking at the images.\n\nAll in all I was very unsure of the leaderboard score as it often did not correlate with my local validation, so I didn't do any per-class LB tuning. For the final submission I averaged several models that were better on the LB (mid-13 public LB / mid-12 private LB), and some models that were better on my validation (they gave about 14 on public LB / 13 on private LB). \n\nTraining took about 12-24 hours on a single 1070, and prediction took another 12 hours. Plus about 3 hours for training and running count prediction.\n\n  [1]: http://imgur.com/I9JlU9g.png\n  [2]: http://i.imgur.com/YST2jU1.png",
    "196857": "Congratulations and thanks.",
    "196863": "Congrats.. well done!\n\nCan we see the code? our team was trying something very similar but we were stuck somewhere in the implementation.",
    "196875": "Yes, sure - I will post the code in 1-2 weeks or earlier, will post the link here.",
    "196882": "Nice work. Congrats on getting UNET to work so successfully. I've tried it out on a couple of projects and have found it to be a pain in the arse to train :)",
    "196886": "Thanks! Yeah, training UNet can be painful!\n\nForgot to mention one more difference from vanilla UNet: I'm using upsampling instead of deconvolutions, on other problems it was much more robust and produced better looking masks.\n\nI also could not make UNet to work properly with dice loss (which should work better in case of class imbalance), and also UNet that predicted 4x smaller output also didn't work so well (which is a pity because it was much faster and more convenient).",
    "196888": "Congrats for getting segmentation to work, I have struggled to get decent results with Tiramisu. Looking forward to look at your code :P",
    "196891": "Excellent work Konstantin, congratulations and thanks also for sharing a lot on forum throughout the competition.    Wondering if you might be willing to try outrunner's simple and elegant postprocessing trick:  add 50% juveniles, subtract adult_females with the same amount, and add 20% pups.   Does this improve your public/private scores to his amazing levels?   For us it improved public by 2.3 and private by 1.8.   You can likely tune the percentages even better given the powerful properties of your nice UNet model.",
    "196926": "Congratulations. Thanks for sharing the details both before and after the competition deadline :). It helped us a lot. I was working a similar approach but was not getting good accuracy. Would you mind answering a few questions - \n\n1. How was the performance with deconvolution layers instead of upsampling? \n2. I am guessing, you used 6 layer masks as Unet output with cross entropy without weighs for your Unet.\n3.  Did you encounter problems with misclassification/ crowded sea lions ? ( I was using Unet and most of the sea-lions were being misclassified as adult-females).\n4. How was the training done ? I mean, how many epochs were trained to get these results ? Also, what kind of optimizer and lr was used? In my case, after 1-2 epochs, the log loss stops decreasing and remains constant. The output masks at this point remain inaccurate. The output masks had sea-lion shaped predictions but they were being misclassified and some other structures were also classified as sea-lions.\n5. Is the purple background in the output image coming from the background label?\n\nThanks again and congrats.",
    "196927": "Thanks Konstantin and great work. Quick question, did you learn the categories separately? In other words your log loss was over 5 categories?",
    "196942": "Thanks Kapil!\n\n&gt; How was the performance with deconvolution layers instead of upsampling?\n\nI didn't check on this task, but on DSTL Satellite image segmentation I just could not get good results with deconvolutions at all - maybe it was a problem with my implementation (because others used them just fine), or bad init - I was using PyTorch. So here I just went with what I knew worked.\n\n&gt; I am guessing, you used 6 layer masks as Unet output with cross entropy without weighs for your Unet.\n\nRight. I tried applying a different weight to the background, but it resulted in more false positives and lower quality masks. On the other hand, oversampling of images with lions helped - which makes sense in retrospect.\n\n&gt; Did you encounter problems with misclassification/ crowded sea lions ? ( I was using Unet and most of the sea-lions were being misclassified as adult-females).\n\nYes to both, and to be honest, I didn't manage to implement a good solution to any of those. For crowded areas maybe xgboost regressor + custom features helped a little. For misclassificatoin, I added a VGG-based classifier that tried to re-classify blobs, but I implemented it only two days before the deadline and didn't manage to get any gain from it - still curious if this would help.\n\n&gt; How was the training done ? I mean, how many epochs were trained to get these results ? Also, what kind of optimizer and lr was used? In my case, after 1-2 epochs, the log loss stops decreasing and remains constant. The output masks at this point remain inaccurate. \n\nFor me one epoch covered entire training set (but I sampled randomly and did all augmentations on the fly), and I got best results with training for 10-20 epochs. I used Adam with lr 0.0001 and batch size 32 (but I multiplied loss by batch size) - it was much better than SGD. I divided lr by 5 when validation loss stopped decreasing for more than two epochs (just once usually). Both training and validation losses continued to go down, but it turned out it's better to stop training even before the best validation loss is reached - probably due to difference between train and test.\n\n&gt; The output masks had sea-lion shaped predictions but they were being misclassified and some other structures were also classified as sea-lions.\n\nI've seen such shapes early in training, maybe after the first few epochs. I also saw some misclassifications and some false positives (especially on the test set), didn't get to solve it.\n\n&gt; Is the purple background in the output image coming from the background label?\n\nYes - it's just some default way of visualising masks with matplotlib, definitely not the best way to show it.",
    "196944": "Thanks @gbhalla! I used a log loss with 6 classes, one for background and the other for different lions - so it's the same loss you normally use for classification problems, but applied per-pixel on the output mask.",
    "196962": "Thanks for the detailed insights of your approach.\n\nJust one more question and then I will stop bugging you :)  - \n\nBy \"With probability 0.2 patches were sampled close to some random sea lion\", do you mean that a sea-lion in train images was selected randomly and then several patches were cropped within some specified distance from that sea-lion ? But then where does the 0.2 probability come into play ?",
    "196969": "I'm happy to answer the questions :)\n\nI sampled patches that formed the batch completely independently. With probability 0.8, I sampled a random image, a random location in that image, a random scale and random angle, and created a patch from it. With probability 0.2, I sampled a random lion (uniformly across all lions, so crowded images got more samples), and then sampled a random location within a pre-defined distance (~patch size) from that lion, and then again the angle and scale.",
    "197016": "Kostia,\nWhich UNet specifically did you use? Based on VGG ?",
    "197018": "Could you please elaborate a little bit on blob detection and counting?",
    "197024": "Yes, I would say it's a classical UNet, and it's similar to bn-VGG, just conv3s everywhere. It's almost the same as this https://github.com/lopuhin/kaggle-dstl/blob/master/models.py#L188 with a small twist: it has 4x pooling instead of 2x in the \"deepest\" part - this save a bit of memory and in theory gives a larger receptive field.",
    "197025": "Sure - I used blob_log from skimage to detect blobs http://scikit-image.org/docs/dev/api/skimage.feature.html#skimage.feature.blob_log and then generated two kinds of features from them: \"blob-sum\" is the sum of probabilities at blob centers (so if we have two blobs with max probability 0.3 and 0.05 the value would be 0.35), and \"blob-count\" is just the number of blobs (so here the value of the feature would be 2). I also did this for several blob thresholds (minimal blob intensities). In order to speed things up and to save space, I saved 4x downscaled prediction, and ran blob detection at this scale too, so it was not so slow.",
    "197076": "Congratulation!!  and thanks for sharing!\n\nSo you resized the test images by half. Did it not make all the sea lions smaller compare to training? (in pixels). Or did you resize the train images too?\n\nHow did you deal with half cut sea lions on the patches (like the tail in on patch and the other half with the head on another patch)?",
    "197077": "that's very cool of you! thanks!",
    "197084": "Thanks!\n\n&gt; So you resized the test images by half. Did it not make all the sea lions smaller compare to training? (in pixels). Or did you resize the train images too?\n\nI think that images in test are about 2x bigger than what we have in the training set. Probably they used a different camera, or didn't resize them, or flew lower, or something like this :)\n\nStill the images in test had different scales, so I applied scale augmentation during training to make the model more robust. Also I was not sure if 2x was the precise difference, and didn't really want to fit it using the public LB.\n\n&gt; How did you deal with half cut sea lions on the patches (like the tail in on patch and the other half with the head on another patch)?\n\nI averaged overlapping predictions at test time. At train time I didn't do anything special.",
    "197110": "Wonderful and congratulations",
    "197120": "Konstantin, congrats and thanks a lot for sharing hints on the forum during the competition: there were really helpful and I learnt a lot despite the fact I never take off from the bottom of the leaderboard.\n\n@KapilYadav: if I may despite my worst result ever in a kaggle competition, I had the same problem to initiate unet and many versions of those (using sigmoid in particular) didn't give me any results. I found that initiate unet with a small portion of the pictures (1/4 to 1/10 of the most crowed pictures) and using elu instead of relu until it reaches a plateau allows the unet to converge pretty quickly usually in a quarter of the time of relu and then saving the weights and restarting the unet with relu activation makes the trick.",
    "197126": "Great work Konstantin! Thanks for sharing",
    "197153": "Charles Jansen, You can check our solution. https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35442  \nTo mitigate the effect of cutting a lion on half I placed gaussians on top of each lion and regressed the sum over all Gaussians in the current tile.",
    "197157": "It's very  high level overview, so maybe it's better to look at the code after, but I have one question: I haven't tried semantic segmentation losses yet, but I was able to train density map regression with U-net only for all classes collapsed to one class and only for 'positive' tiles, so I wonder is it so hard to train U-net or maybe it's due to class imbalance? do you balance your data?\n\nP.S. thanks for your comments in Discusssion threads.",
    "197193": "I tried the density map regression as well and had classes collapsing too. I think this is just a limitation of the density map regression. It doesn't work with more than one class.",
    "197198": "It's very impressive and clever to extract features from the first stage prediction results, and regress to the final counts. I also used a method based on masks regression, but I used the peak detection results as my final submission directly, which introduced lots of uncontrollable errors.",
    "203361": "Thanks Russ, I really liked your solution, I think it's the most principled and robust of the ones shared due to explicit accounting for scale.\n\nFinally got to check this clever hack, a milder version of it (just increasing pups by 20%) also gives a moderate boost: 12.5 -&gt; 12.0 on private (and 13.2 -&gt; 12.9 on public), but the optimal settings are probably different  for my model.",
    "217600": "Here is the code: https://github.com/lopuhin/kaggle-lions-2017 (including all failed experiments, so in some places the code is less clean then I'd like)"
  },
  "source": "meta"
}