{
  "id": 35408,
  "title": "Finally, we have 1,178,071 sea lions",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/35408",
  "author_name": "outrunner",
  "post_date": "2017-06-28T02:28:12.197000",
  "votes": 67,
  "comment_count": 66,
  "views": 0,
  "content": "<p>Thanks to everyone. I love this competition.</p>\n\n<p><a href=\"https://www.kaggle.com/outrunner/use-keras-to-count-sea-lions/notebook\">Here is my solution.</a> </p>\n\n<p>I use VGG16 without top, add FC-1024, and FC-5 with linear output.</p>\n\n<p>SGD optimization and mean_squared_error loss.</p>\n\n<p>And post processing for better regression.</p>\n\n<p>I think it is a simple way to count. Looking forward to hearing your thoughts, thanks.</p>",
  "messages": [
    {
      "id": 196735,
      "postDate": "2017-06-28T02:28:12.197Z",
      "content": "<p>Thanks to everyone. I love this competition.</p>\n\n<p><a href=\"https://www.kaggle.com/outrunner/use-keras-to-count-sea-lions/notebook\">Here is my solution.</a> </p>\n\n<p>I use VGG16 without top, add FC-1024, and FC-5 with linear output.</p>\n\n<p>SGD optimization and mean_squared_error loss.</p>\n\n<p>And post processing for better regression.</p>\n\n<p>I think it is a simple way to count. Looking forward to hearing your thoughts, thanks.</p>",
      "rawMarkdown": "Thanks to everyone. I love this competition.\n\n[Here is my solution.][1] \n\nI use VGG16 without top, add FC-1024, and FC-5 with linear output.\n\nSGD optimization and mean_squared_error loss.\n\nAnd post processing for better regression.\n\nI think it is a simple way to count. Looking forward to hearing your thoughts, thanks.\n\n  [1]: https://www.kaggle.com/outrunner/use-keras-to-count-sea-lions/notebook",
      "votes": 67
    },
    {
      "id": 196751,
      "postDate": "2017-06-28T03:02:21.327Z",
      "content": "<p>I'm proud to see my blob detection in the winning solution.</p>\n\n<p>Good job!</p>",
      "rawMarkdown": "I'm proud to see my blob detection in the winning solution.\n\nGood job!",
      "votes": 8,
      "replies": [
        {
          "id": 196783,
          "postDate": "2017-06-28T05:13:23.207Z",
          "content": "<p>Thanks again.</p>",
          "rawMarkdown": "Thanks again.",
          "votes": 4
        }
      ]
    },
    {
      "id": 197128,
      "postDate": "2017-06-28T21:43:51.967Z",
      "content": "<p>Outstanding work outrunner! Thanks for sharing; may I ask... did you have any intuition, before starting out, that this network would work ? And what do you think made the big difference in this architecture - was it the 1024 FC or the SGD or slow training or sthing else ... i am sure you tried lots of changes in this area. \nI gather you started with pretrained weights. </p>",
      "rawMarkdown": "Outstanding work outrunner! Thanks for sharing; may I ask... did you have any intuition, before starting out, that this network would work ? And what do you think made the big difference in this architecture - was it the 1024 FC or the SGD or slow training or sthing else ... i am sure you tried lots of changes in this area. \nI gather you started with pretrained weights. ",
      "votes": 1,
      "replies": [
        {
          "id": 197219,
          "postDate": "2017-06-29T02:39:36.913Z",
          "content": "<p>Just intuition, and I think it would work undoubtedly even when I got trouble at the beginning. VGG is a basic CNN, I took the 1024 FC as a buffer between FCN and output. The size is according to RAM limit and model size, and intuitively I think it is enough. When I got scale problem, I tried more neurons and nothing happened. SGD is more stable than other aggressive optimizers in my work, maybe it is due to the large variation in output labels.</p>",
          "rawMarkdown": "Just intuition, and I think it would work undoubtedly even when I got trouble at the beginning. VGG is a basic CNN, I took the 1024 FC as a buffer between FCN and output. The size is according to RAM limit and model size, and intuitively I think it is enough. When I got scale problem, I tried more neurons and nothing happened. SGD is more stable than other aggressive optimizers in my work, maybe it is due to the large variation in output labels.",
          "votes": 3
        }
      ]
    },
    {
      "id": 196752,
      "postDate": "2017-06-28T03:02:34.203Z",
      "content": "<p>Congratulations and Excellent work. Your kaggle ranking  will be top 40.</p>",
      "rawMarkdown": "Congratulations and Excellent work. Your kaggle ranking  will be top 40.",
      "votes": 1
    },
    {
      "id": 197692,
      "postDate": "2017-06-30T03:08:24.533Z",
      "content": "<p>Something I could share about why misclassification on females and juveniles is worse more than I expected.</p>\n\n<p><img src=\"http://i.imgur.com/rJBOLir.gif\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>I generated random dot for explanation. I wish the threshold between females and juveniles in the model could be the green line, but it is too naive. Then I think at least it should be the red line by statistically, but in fact the threshold is dotted red line. (It is about size feature, there are other features like lion distribution etc. in the model let the prediction better.)</p>\n\n<p>I guess it is because in the small scale patch, lion count is larger, and the squared error even more. I tried add large scale training patch with more juveniles and it helps a little, but overall score is not significant.</p>",
      "rawMarkdown": "Something I could share about why misclassification on females and juveniles is worse more than I expected.\n\n![enter image description here][1]\n\nI generated random dot for explanation. I wish the threshold between females and juveniles in the model could be the green line, but it is too naive. Then I think at least it should be the red line by statistically, but in fact the threshold is dotted red line. (It is about size feature, there are other features like lion distribution etc. in the model let the prediction better.)\n\nI guess it is because in the small scale patch, lion count is larger, and the squared error even more. I tried add large scale training patch with more juveniles and it helps a little, but overall score is not significant.\n\n  [1]: http://i.imgur.com/rJBOLir.gif",
      "votes": 2
    },
    {
      "id": 198656,
      "postDate": "2017-07-03T09:16:21.487Z",
      "content": "<p>Would you mind creating a visualization of the predictions that the network produces similar to the ones by Konstantin Lopuhin, e.g. using <code>for i, p in enumerate(batch_of_predictions): scipy.misc.imsave('preds_%i.png' % i, normalize(np.concatenate(np.moveaxis(p, -1, 0))))</code>? I think it would be interesting to see how well the network manages to tell the classes apart in various cases and what the predicted dots look like.</p>",
      "rawMarkdown": "Would you mind creating a visualization of the predictions that the network produces similar to the ones by Konstantin Lopuhin, e.g. using `for i, p in enumerate(batch_of_predictions): scipy.misc.imsave('preds_%i.png' % i, normalize(np.concatenate(np.moveaxis(p, -1, 0))))`? I think it would be interesting to see how well the network manages to tell the classes apart in various cases and what the predicted dots look like.",
      "replies": [
        {
          "id": 198664,
          "postDate": "2017-07-03T09:54:53.283Z",
          "content": "<p>I don't quite understand what is batch_of_predictions?</p>",
          "rawMarkdown": "I don't quite understand what is batch_of_predictions?"
        },
        {
          "id": 198671,
          "postDate": "2017-07-03T10:40:24.903Z",
          "content": "<p>The variable <code>batch_of_predictions</code> is be the output of your network. Usually, all operations in deep learning frameworks operate on batches of training examples. When you run the model to test it on a couple of images, you need to feed a batch of n≥1 images and you get n predictions back. This is how you can save all images from the batch. I basically just copied &amp; modified a line of code from my solution.</p>",
          "rawMarkdown": "The variable `batch_of_predictions` is be the output of your network. Usually, all operations in deep learning frameworks operate on batches of training examples. When you run the model to test it on a couple of images, you need to feed a batch of n≥1 images and you get n predictions back. This is how you can save all images from the batch. I basically just copied &amp; modified a line of code from my solution."
        },
        {
          "id": 198681,
          "postDate": "2017-07-03T11:22:55.327Z",
          "content": "<p>I mean, in my network, predictions are only numbers. Or you want to see some testing patches with its output?</p>",
          "rawMarkdown": "I mean, in my network, predictions are only numbers. Or you want to see some testing patches with its output?"
        },
        {
          "id": 198816,
          "postDate": "2017-07-03T19:17:09.003Z",
          "content": "<p>Oops, I misunderstood the way you create the <code>trainY</code>s. I thought your <code>trainY</code>s consisted of dots at the animal coordinates. Never mind then. I could imagine your solution works well because only predicting the total counts (not the locations) in a large window like that reduces noise from the label positions (sometimes the labels are even outside of the animals). For your method, imprecise label positions only cause noise at the borders of the patch.</p>",
          "rawMarkdown": "Oops, I misunderstood the way you create the `trainY`s. I thought your `trainY`s consisted of dots at the animal coordinates. Never mind then. I could imagine your solution works well because only predicting the total counts (not the locations) in a large window like that reduces noise from the label positions (sometimes the labels are even outside of the animals). For your method, imprecise label positions only cause noise at the borders of the patch."
        },
        {
          "id": 198888,
          "postDate": "2017-07-04T02:07:23.307Z",
          "content": "<p>By the way, this is how I verify the prediction: （more count more bright)\n<img src=\"http://i.imgur.com/FKQdbpX.jpg\" alt=\"enter image description here\" title=\"\"></p>",
          "rawMarkdown": "By the way, this is how I verify the prediction: （more count more bright)\n![enter image description here][1]\n\n  [1]: http://i.imgur.com/FKQdbpX.jpg",
          "votes": 4
        },
        {
          "id": 198899,
          "postDate": "2017-07-04T03:22:00.263Z",
          "content": "<p>It is cool! How can you show the picture in this way? Could you please show some code on this?</p>",
          "rawMarkdown": "It is cool! How can you show the picture in this way? Could you please show some code on this?"
        },
        {
          "id": 198901,
          "postDate": "2017-07-04T03:50:06.773Z",
          "content": "<p>for p in test_patches: p = p*f(count(p)) #make sure 0 &lt; f &lt; 1 or do some normalized processing</p>\n\n<p>then stitch them side by side</p>",
          "rawMarkdown": "for p in test_patches: p = p*f(count(p)) #make sure 0 &lt; f &lt; 1 or do some normalized processing\n\nthen stitch them side by side"
        }
      ]
    },
    {
      "id": 198600,
      "postDate": "2017-07-03T05:14:55.613Z",
      "content": "<p>Congrats on winning. Great solution!! \nI noticed that the test was also run on the patches that were cut from the original test images and the overall count for an image is the sum of the counts from all the patches cut from the image. Since the patch cut is random, would a sea lion at the boundary of two patches be double-counted? How would the double counting be avoided in the solution?\nThanks!!</p>",
      "rawMarkdown": "Congrats on winning. Great solution!! \nI noticed that the test was also run on the patches that were cut from the original test images and the overall count for an image is the sum of the counts from all the patches cut from the image. Since the patch cut is random, would a sea lion at the boundary of two patches be double-counted? How would the double counting be avoided in the solution?\nThanks!!",
      "replies": [
        {
          "id": 198653,
          "postDate": "2017-07-03T08:54:14.403Z",
          "content": "<p>Ideally, output should be 0.5 if the lion has been cut in half. Because in the same situation at training stage, one patch is 1 and the other one is 0.</p>",
          "rawMarkdown": "Ideally, output should be 0.5 if the lion has been cut in half. Because in the same situation at training stage, one patch is 1 and the other one is 0."
        }
      ]
    },
    {
      "id": 197253,
      "postDate": "2017-06-29T06:33:34.793Z",
      "rawMarkdown": ""
    },
    {
      "id": 197252,
      "postDate": "2017-06-29T06:33:22.160Z",
      "content": "<p>Hi Outrunner,</p>\n\n<p>Congratulations!</p>\n\n<p>I'm not very familiar with Keras code. Could you explain the last layers of your neural network? What happens after the convolutional layers? Did you just use linear activations to generate a 5-layer heatmap and compared that to the GT heatmap using RMSE?</p>",
      "rawMarkdown": "Hi Outrunner,\n\nCongratulations!\n\nI'm not very familiar with Keras code. Could you explain the last layers of your neural network? What happens after the convolutional layers? Did you just use linear activations to generate a 5-layer heatmap and compared that to the GT heatmap using RMSE?",
      "replies": [
        {
          "id": 197267,
          "postDate": "2017-06-29T07:00:52.650Z",
          "content": "<p>The \"Patches looks like\" section in notebook, image is training patch, [1 0 3 0 1] is label (lion count per type in the patch), the last is 5 neurons fully-connected layer with linear output.</p>",
          "rawMarkdown": "The \"Patches looks like\" section in notebook, image is training patch, [1 0 3 0 1] is label (lion count per type in the patch), the last is 5 neurons fully-connected layer with linear output."
        },
        {
          "id": 197570,
          "postDate": "2017-06-29T20:52:45.550Z",
          "content": "<p>Aha! So it's a kind of regression?</p>",
          "rawMarkdown": "Aha! So it's a kind of regression?"
        },
        {
          "id": 197595,
          "postDate": "2017-06-29T23:11:52.107Z",
          "content": "<p>I think so.</p>",
          "rawMarkdown": " I think so."
        }
      ]
    },
    {
      "id": 197238,
      "postDate": "2017-06-29T04:45:26.363Z",
      "content": "<p>Great job here!</p>",
      "rawMarkdown": "Great job here!"
    },
    {
      "id": 197143,
      "postDate": "2017-06-28T22:17:25.833Z",
      "content": "<ol>\n<li>In your kernel you use only 1 image, how training was setup in practice? it was something like Keras ImageDataGenerator or you load all data to RAM?</li>\n<li>Does LB score depend on tile size? How you handle images at prediction time, i.e. with tile overlapping? In my opinion in case of pretrained network it should be about 'original' size, i.e. 224x224 for VGG.</li>\n<li>Why mean_squared_error was used that is MSE and in competition we have RMSE?</li>\n<li>What LB score you get without tricks described in Experience  section?</li>\n</ol>",
      "rawMarkdown": "1. In your kernel you use only 1 image, how training was setup in practice? it was something like Keras ImageDataGenerator or you load all data to RAM?\n2.  Does LB score depend on tile size? How you handle images at prediction time, i.e. with tile overlapping? In my opinion in case of pretrained network it should be about 'original' size, i.e. 224x224 for VGG.\n3. Why mean_squared_error was used that is MSE and in competition we have RMSE?\n4. What LB score you get without tricks described in Experience  section?",
      "replies": [
        {
          "id": 197225,
          "postDate": "2017-06-29T03:10:59.207Z",
          "content": "<ol>\n<li><p>I use custom data generator.</p></li>\n<li><p>I trained a 200x200 and it performed worse. Maybe it is because 300 contains more relationship information between pups and adult_females. Flipping(rotating) 4 times at prediction time improved 0.04 RMSE. Using only convolutional layer in VGG is size free.</p></li>\n<li><p>Just a normal loss function.</p></li>\n<li><p>You can reference the table in the section.</p></li>\n</ol>",
          "rawMarkdown": "1. I use custom data generator.\n\n2. I trained a 200x200 and it performed worse. Maybe it is because 300 contains more relationship information between pups and adult_females. Flipping(rotating) 4 times at prediction time improved 0.04 RMSE. Using only convolutional layer in VGG is size free.\n\n3. Just a normal loss function.\n\n4. You can reference the table in the section."
        }
      ]
    },
    {
      "id": 197139,
      "postDate": "2017-06-28T22:07:43.777Z",
      "content": "<p>Respect! Never underestimate the basics.</p>",
      "rawMarkdown": "Respect! Never underestimate the basics."
    },
    {
      "id": 197074,
      "postDate": "2017-06-28T18:22:34.977Z",
      "content": "<p>@outrunner\nCongrats and thanks for the code and the explanations.</p>\n\n<p>As you requested, please find attached the bounding boxes. the first batch are the 5000 first bboxes manually labeled and the other 75 to 80,000 are the ones generated using SSD (no guarantee on the quality and some subadult male disappeared in the mix).</p>\n\n<p>Keep us informed on your progress with them please.</p>\n\n<p>I generated them as I started this competition as standard detection problem but later switched to a segmentation where they were not the best idea and dragged me to the bottom of the leader-board: I was more concerned by crowed area where bboxes are an handicap rather than image scale...</p>",
      "rawMarkdown": "@outrunner\nCongrats and thanks for the code and the explanations.\n\nAs you requested, please find attached the bounding boxes. the first batch are the 5000 first bboxes manually labeled and the other 75 to 80,000 are the ones generated using SSD (no guarantee on the quality and some subadult male disappeared in the mix).\n\nKeep us informed on your progress with them please.\n\nI generated them as I started this competition as standard detection problem but later switched to a segmentation where they were not the best idea and dragged me to the bottom of the leader-board: I was more concerned by crowed area where bboxes are an handicap rather than image scale...",
      "replies": [
        {
          "id": 197214,
          "postDate": "2017-06-29T02:14:31.817Z",
          "content": "<p>Thanks. But I may try it when I have time or need to do farther research on this topic.</p>",
          "rawMarkdown": "Thanks. But I may try it when I have time or need to do farther research on this topic."
        }
      ]
    },
    {
      "id": 196987,
      "postDate": "2017-06-28T14:38:50.380Z",
      "content": "<p>It is interesting that the model trained at 1.0x scale had a much better performance being evaluated on the test set at 0.67x scale than on 1.0x. What might be the reason for this? Perhaps the receptive field of VGG16 is just not big enough to actually recognize the animals at 1.0x scale. VGG16 has an effective receptive field size of ~220px at the output neurons, but the largest adult males are up to ~260-270px in length. Also the largest distance of the pubs to other animals is about 250px.</p>",
      "rawMarkdown": "It is interesting that the model trained at 1.0x scale had a much better performance being evaluated on the test set at 0.67x scale than on 1.0x. What might be the reason for this? Perhaps the receptive field of VGG16 is just not big enough to actually recognize the animals at 1.0x scale. VGG16 has an effective receptive field size of ~220px at the output neurons, but the largest adult males are up to ~260-270px in length. Also the largest distance of the pubs to other animals is about 250px.",
      "replies": [
        {
          "id": 197003,
          "postDate": "2017-06-28T15:27:55.667Z",
          "content": "<p>I think it is typically because the testing image's sea lion is larger than training in average. In my opinion, the model don't know what sea lion is and where it is.</p>",
          "rawMarkdown": "I think it is typically because the testing image's sea lion is larger than training in average. In my opinion, the model don't know what sea lion is and where it is."
        },
        {
          "id": 197088,
          "postDate": "2017-06-28T19:22:08.303Z",
          "content": "<p>&gt; In my opinion, the model don't know what sea lion is and where it is.</p>\n\n<p>I agree. :) 'Not recognizing' was just shorthand for 'it might be harder to find distinguishing features if the animals are often only partially in sight for the network'.  My info was incomplete/wrong anyway though: VGG's receptive field is 224 × 224 with zero-padding, but 404 × 404 if the net is applied convolutionally^<a href=\"https://arxiv.org/pdf/1412.7062.pdf\">1</a>.</p>",
          "rawMarkdown": "&gt; In my opinion, the model don't know what sea lion is and where it is.\n\nI agree. :) 'Not recognizing' was just shorthand for 'it might be harder to find distinguishing features if the animals are often only partially in sight for the network'.  My info was incomplete/wrong anyway though: VGG's receptive field is 224 × 224 with zero-padding, but 404 × 404 if the net is applied convolutionally^[1].\n\n  [1]: https://arxiv.org/pdf/1412.7062.pdf"
        },
        {
          "id": 197204,
          "postDate": "2017-06-29T01:53:17.547Z",
          "content": "<p>I haven't studied receptive field in detail yet. So, thanks for sharing.</p>",
          "rawMarkdown": "I haven't studied receptive field in detail yet. So, thanks for sharing."
        }
      ]
    },
    {
      "id": 196976,
      "postDate": "2017-06-28T14:12:33.797Z",
      "content": "<p>Fantastic result, congrats!</p>\n\n<p>Am I understanding correctly that the multi-scale model is essentially heavy data augmentation with different scales?</p>",
      "rawMarkdown": "Fantastic result, congrats!\n\nAm I understanding correctly that the multi-scale model is essentially heavy data augmentation with different scales?"
    },
    {
      "id": 196967,
      "postDate": "2017-06-28T13:48:30.440Z",
      "content": "<p>Such a simple solution. It really goes to show how far machine learning, convolutional neural networks, and a perfect implementation, application, and execution can take you. <strong>Congrats @outrunner on 1st place</strong>, and much thanks for the informative write-up. </p>",
      "rawMarkdown": "Such a simple solution. It really goes to show how far machine learning, convolutional neural networks, and a perfect implementation, application, and execution can take you. **Congrats @outrunner on 1st place**, and much thanks for the informative write-up. ",
      "replies": [
        {
          "id": 196970,
          "postDate": "2017-06-28T13:52:27.350Z",
          "content": "<p>I also had a question about the values on your table, I don't quite understand why a higher value on the table like the values in your final submission column yielded better results. What is the significance of <strong>average of juveniles# / (adult_females# + juveniles#)</strong>?</p>",
          "rawMarkdown": "I also had a question about the values on your table, I don't quite understand why a higher value on the table like the values in your final submission column yielded better results. What is the significance of **average of juveniles# / (adult_females# + juveniles#)**?",
          "votes": 1
        },
        {
          "id": 196982,
          "postDate": "2017-06-28T14:27:33.503Z",
          "content": "<p>I think the numbers just show that on the test set outrunner's models predicted much fewer juveniles vs adult_females than during training, thus he/she corrected roughly for that ratio (with great success). E.g. the the patches with 46-50 juveniles had a number of 0.93 during training with a single zoom level (1x), but on the test set only 0.68, so it was corrected to 0.76. In that particular case it would have been interesting to see the performance of the r0.59 model without post-processing on the public LB.</p>\n\n<p>This might imply that the test set had different juvenile/adult_female ratios than the training set.</p>",
          "rawMarkdown": "I think the numbers just show that on the test set outrunner's models predicted much fewer juveniles vs adult_females than during training, thus he/she corrected roughly for that ratio (with great success). E.g. the the patches with 46-50 juveniles had a number of 0.93 during training with a single zoom level (1x), but on the test set only 0.68, so it was corrected to 0.76. In that particular case it would have been interesting to see the performance of the r0.59 model without post-processing on the public LB.\n\nThis might imply that the test set had different juvenile/adult_female ratios than the training set."
        },
        {
          "id": 196997,
          "postDate": "2017-06-28T15:14:00.110Z",
          "content": "<p>Sorry for being unclear and thanks to erde. It is the ratio between juveniles and juveniles + adult_females.</p>\n\n<p>I think the ratio problem comes from scale problem. The CNN basically is a statistical model, simply distinguish juveniles and adult_females by size. Scale down image can get more juveniles, but the problem is if out of range, total lions count will go down. For example:</p>\n\n<ul>\n<li>scale   predict</li>\n<li>0.59    [ 13   5 180 120   2]</li>\n<li>0.53    [ 14   8 141 149   0]</li>\n<li>0.48    [ 10   5 132 165   2]</li>\n<li>0.43    [  9   7 122 182   0]</li>\n<li>0.39    [  6   7 101 185   0]</li>\n<li>0.35    [  7   7  74 176   0]</li>\n<li>0.31    [  4   6  62 168   1]</li>\n<li>0.28    [  5   7  48 137   0]</li>\n</ul>\n\n<p>This is the predict result of test/12690 on multi-scale model with different test scale. The juveniles + adult_females keeps around 300 for 0.59~0.43, and began to decrease from 0.39.</p>\n\n<p>So, all you are interesting in without post processing's performance. It is sure if I scale down more. But I don't want to take the risk if sea lions in private set are smaller then public.</p>\n\n<blockquote>\n  <p>the multi-scale model is essentially heavy data augmentation with different scales?</p>\n</blockquote>\n\n<p>Yes, and single scale model's keeping number scale range is narrower.</p>",
          "rawMarkdown": " Sorry for being unclear and thanks to erde. It is the ratio between juveniles and juveniles + adult_females.\n\nI think the ratio problem comes from scale problem. The CNN basically is a statistical model, simply distinguish juveniles and adult_females by size. Scale down image can get more juveniles, but the problem is if out of range, total lions count will go down. For example:\n\n - scale   predict\n - 0.59    [ 13   5 180 120   2]\n - 0.53    [ 14   8 141 149   0]\n - 0.48    [ 10   5 132 165   2]\n - 0.43    [  9   7 122 182   0]\n - 0.39    [  6   7 101 185   0]\n - 0.35    [  7   7  74 176   0]\n - 0.31    [  4   6  62 168   1]\n - 0.28    [  5   7  48 137   0]\n\nThis is the predict result of test/12690 on multi-scale model with different test scale. The juveniles + adult_females keeps around 300 for 0.59~0.43, and began to decrease from 0.39.\n\nSo, all you are interesting in without post processing's performance. It is sure if I scale down more. But I don't want to take the risk if sea lions in private set are smaller then public.\n\n&gt; the multi-scale model is essentially heavy data augmentation with different scales?\n\nYes, and single scale model's keeping number scale range is narrower.\n\n",
          "votes": 1
        },
        {
          "id": 197004,
          "postDate": "2017-06-28T15:29:00.683Z",
          "content": "<p>Ah I see so scaled down images tend to have more juveniles than adult females versus not scaled down. If it were possible to teach the method how to compare the sizes of sea lions within a single image's scale, then the model could learn to compare the juvenile's relative size to female's relative size and your post processing trick might not have been needed? Maybe if first the image is processed for females then those features are fed to another network that searches for juveniles, the features can be reused. I think your post processing basically gave the model domain knowledge about the problem that could be replaced with a different architecture, would be interesting to test this!</p>",
          "rawMarkdown": "Ah I see so scaled down images tend to have more juveniles than adult females versus not scaled down. If it were possible to teach the method how to compare the sizes of sea lions within a single image's scale, then the model could learn to compare the juvenile's relative size to female's relative size and your post processing trick might not have been needed? Maybe if first the image is processed for females then those features are fed to another network that searches for juveniles, the features can be reused. I think your post processing basically gave the model domain knowledge about the problem that could be replaced with a different architecture, would be interesting to test this!"
        },
        {
          "id": 197007,
          "postDate": "2017-06-28T15:39:33.807Z",
          "content": "<p>Sure, and it is not easy work. By the way, if they have grounding truth bounding box, I want to train a object detection for image scale detect by sea lion's size, just a thought.</p>",
          "rawMarkdown": "Sure, and it is not easy work. By the way, if they have grounding truth bounding box, I want to train a object detection for image scale detect by sea lion's size, just a thought."
        },
        {
          "id": 197012,
          "postDate": "2017-06-28T15:59:07.277Z",
          "content": "<p>Ideally, there should be preprocessing that makes all images the same scale. Much like satellite images are orthorectified to make objects on the ground to be the same scale and angle (so square on ground is square on image no matter how camera looks at Earth -- at right or some different angle). The same should be done here -- knowing flight path, UAV elevation, camera characteristics, capture angle and so on, it is possible to make all images the same scale. </p>",
          "rawMarkdown": "Ideally, there should be preprocessing that makes all images the same scale. Much like satellite images are orthorectified to make objects on the ground to be the same scale and angle (so square on ground is square on image no matter how camera looks at Earth -- at right or some different angle). The same should be done here -- knowing flight path, UAV elevation, camera characteristics, capture angle and so on, it is possible to make all images the same scale. \n",
          "votes": 5
        },
        {
          "id": 197023,
          "postDate": "2017-06-28T16:20:39.583Z",
          "content": "<p>We labeled some images  from train set according to scale and tried to train a CNN to regress a scale of the image. But it didn't work out. I recon this is due to high variation in terrain and not being able to estimate scale if you look at tiles with a small spatial context (even with my own eyes).</p>",
          "rawMarkdown": "We labeled some images  from train set according to scale and tried to train a CNN to regress a scale of the image. But it didn't work out. I recon this is due to high variation in terrain and not being able to estimate scale if you look at tiles with a small spatial context (even with my own eyes)."
        },
        {
          "id": 197028,
          "postDate": "2017-06-28T16:32:43.247Z",
          "content": "<p>I tried a little different, and failed as expected.</p>",
          "rawMarkdown": "I tried a little different, and failed as expected."
        },
        {
          "id": 197037,
          "postDate": "2017-06-28T16:51:15.843Z",
          "content": "<p>agreed. including such data would make for a more robust approach.</p>",
          "rawMarkdown": "agreed. including such data would make for a more robust approach."
        }
      ]
    },
    {
      "id": 196839,
      "postDate": "2017-06-28T08:16:13.257Z",
      "content": "<p>Well done &amp; congratz @outrunner.</p>",
      "rawMarkdown": "Well done &amp; congratz @outrunner."
    },
    {
      "id": 196830,
      "postDate": "2017-06-28T08:02:23.047Z",
      "content": "<p>Сongratulations!</p>",
      "rawMarkdown": "Сongratulations!"
    },
    {
      "id": 196769,
      "postDate": "2017-06-28T03:44:00.367Z",
      "content": "<p>Congrats! Great solution!</p>\n\n<p>A couple of questions: What was the difference in public/private score between r0.59 and r0.59*1.3  ?</p>\n\n<p>Did you try other optimizers/did the learning rate schedule make a large difference in score?</p>\n\n<p>For the multi-scale model did you take larger images and resize them down to 300x300 or did you train seperate networks on different scaling of the 300x300 patches?</p>\n\n<p>Thanks, and congratulations on the win again.</p>",
      "rawMarkdown": "Congrats! Great solution!\n\nA couple of questions: What was the difference in public/private score between r0.59 and r0.59*1.3  ?\n\nDid you try other optimizers/did the learning rate schedule make a large difference in score?\n\nFor the multi-scale model did you take larger images and resize them down to 300x300 or did you train seperate networks on different scaling of the 300x300 patches?\n\nThanks, and congratulations on the win again.\n",
      "replies": [
        {
          "id": 196780,
          "postDate": "2017-06-28T05:05:25.763Z",
          "content": "<p>Sorry, do you mean different score or what make the score difference? I did not submit r0.59, I think it may be 14~15. About post processing: if a image predict in r0.59 = [10,10,100,100,20], then submit r0.59*1.3 = [10,10,70,130,20]</p>\n\n<p>I tried other optimizers/learning rate schedule, just for observe the learning loss convergence.</p>\n\n<p>I scaled down the original image (*0.9, ..., *0.6561), cut into 300x300, and use those patches together to train a network.</p>",
          "rawMarkdown": "Sorry, do you mean different score or what make the score difference? I did not submit r0.59, I think it may be 14~15. About post processing: if a image predict in r0.59 = [10,10,100,100,20], then submit r0.59*1.3 = [10,10,70,130,20]\n\nI tried other optimizers/learning rate schedule, just for observe the learning loss convergence.\n\nI scaled down the original image (*0.9, ..., *0.6561), cut into 300x300, and use those patches together to train a network.",
          "votes": 1
        },
        {
          "id": 196788,
          "postDate": "2017-06-28T05:21:11.980Z",
          "content": "<p>Thanks for the info. You are correct I was wondering what the leaderboard score of r0.59 was. From 14~15 to ~11 is a big boost in score.</p>",
          "rawMarkdown": "Thanks for the info. You are correct I was wondering what the leaderboard score of r0.59 was. From 14~15 to ~11 is a big boost in score."
        },
        {
          "id": 196791,
          "postDate": "2017-06-28T05:39:33.490Z",
          "content": "<p>Additionally, post processing also add pups, in this case, it should be [10,10,70,130,24]. But the main contribution comes from juveniles/adult_females adjustment.</p>",
          "rawMarkdown": "Additionally, post processing also add pups, in this case, it should be [10,10,70,130,24]. But the main contribution comes from juveniles/adult_females adjustment.",
          "votes": 1
        },
        {
          "id": 196822,
          "postDate": "2017-06-28T07:34:11.220Z",
          "content": "<p>Great result, congratulation with winning the competition!</p>\n\n<p>Interesting, what was your score before post processing? It's amazing a relatively simple cnn managed to predict number of lion so well.</p>",
          "rawMarkdown": "Great result, congratulation with winning the competition!\n\nInteresting, what was your score before post processing? It's amazing a relatively simple cnn managed to predict number of lion so well."
        },
        {
          "id": 196828,
          "postDate": "2017-06-28T07:56:48.270Z",
          "content": "<p>I found a submission got 13.46393 private LB score is made by:</p>\n\n<ul>\n<li>training scale: *1.23~0.66</li>\n<li>testing scale: *0.48</li>\n</ul>\n\n<p>only cut pups more than 1.8x adult_females after predict, and this affect a little in my experience.</p>\n\n<p>So, without post processing on other combination, maybe it could be around 13.</p>",
          "rawMarkdown": "I found a submission got 13.46393 private LB score is made by:\n\n - training scale: *1.23~0.66\n - testing scale: *0.48\n\nonly cut pups more than 1.8x adult_females after predict, and this affect a little in my experience.\n\nSo, without post processing on other combination, maybe it could be around 13.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 196766,
      "postDate": "2017-06-28T03:35:23.173Z",
      "content": "<p>congratulations and amazing that you achieved  #1 position with in four months .. </p>",
      "rawMarkdown": "congratulations and amazing that you achieved  #1 position with in four months .. "
    },
    {
      "id": 196761,
      "postDate": "2017-06-28T03:27:50.767Z",
      "content": "<p>Congrats for the 1st position. Thanks for sharing the details. Its a nice approach.</p>\n\n<p>I have a few questions - </p>\n\n<ol>\n<li>Did you make all the VGG16 layers trainable ? ( I tried VGG16, but I couldnt make it stop overfit on train data. :(  )</li>\n<li>How were the final output of linear layers compared ? ( I guess, a sigmoid layer was used ... )</li>\n<li>Just a clarification - In cases when a patch has multiple sea-lions, the network will output high values for their respective nodes in the last layer, right? </li>\n<li>Did you face any difficulty for patches where sea-lions were all crowded in one place?</li>\n</ol>",
      "rawMarkdown": "Congrats for the 1st position. Thanks for sharing the details. Its a nice approach.\n\nI have a few questions - \n\n 1. Did you make all the VGG16 layers trainable ? ( I tried VGG16, but I couldnt make it stop overfit on train data. :(  )\n 2. How were the final output of linear layers compared ? ( I guess, a sigmoid layer was used ... )\n 3. Just a clarification - In cases when a patch has multiple sea-lions, the network will output high values for their respective nodes in the last layer, right? \n 4. Did you face any difficulty for patches where sea-lions were all crowded in one place?",
      "replies": [
        {
          "id": 196772,
          "postDate": "2017-06-28T03:52:21.703Z",
          "content": "<ol>\n<li><p>Yes, train the new FC layers first, then all layers.</p></li>\n<li><p>I use linear activation at output layer. (because 3.)</p></li>\n<li><p>Yes.</p></li>\n<li><p>Training got exploded sometimes, but it is not problem. When testing, distinguish juveniles and adult_females is difficult. I think the model don't know what is \"relative\" size. Actually, in the training set, we also got a lot of human label noise at juveniles and adult_females.</p></li>\n</ol>",
          "rawMarkdown": "1. Yes, train the new FC layers first, then all layers.\n\n2. I use linear activation at output layer. (because 3.)\n\n3. Yes.\n\n4. Training got exploded sometimes, but it is not problem. When testing, distinguish juveniles and adult_females is difficult. I think the model don't know what is \"relative\" size. Actually, in the training set, we also got a lot of human label noise at juveniles and adult_females.",
          "votes": 2
        },
        {
          "id": 196966,
          "postDate": "2017-06-28T13:47:34.633Z",
          "content": "<p>Thanks.</p>",
          "rawMarkdown": "Thanks."
        },
        {
          "id": 197134,
          "postDate": "2017-06-28T21:52:43.907Z",
          "content": "<p>Why linear activation? is that mean that it can predict negative values?</p>",
          "rawMarkdown": "Why linear activation? is that mean that it can predict negative values?"
        },
        {
          "id": 197206,
          "postDate": "2017-06-29T01:55:00.797Z",
          "content": "<p>Why not?</p>",
          "rawMarkdown": "Why not?"
        },
        {
          "id": 197257,
          "postDate": "2017-06-29T06:37:05.960Z",
          "content": "<p>I mean it will work, but it can produce negative values as side affect? are you perform some sort of clipping?</p>\n\n<p>Will it work with Relu activation?</p>",
          "rawMarkdown": "I mean it will work, but it can produce negative values as side affect? are you perform some sort of clipping?\n\nWill it work with Relu activation?"
        },
        {
          "id": 197279,
          "postDate": "2017-06-29T07:36:51.033Z",
          "content": "<p>Relu got problem at 0, and also need to do clip.</p>",
          "rawMarkdown": "Relu got problem at 0, and also need to do clip."
        },
        {
          "id": 197284,
          "postDate": "2017-06-29T07:44:24.763Z",
          "content": "<p>Can you elaborate on this? I mean that if final prediction will be [-10, 0, 0, 4, 5] it should be clipped to  [0, 0, 0, 4, 5] (to make things cleaner I don't talking about gradient clipping).</p>",
          "rawMarkdown": "Can you elaborate on this? I mean that if final prediction will be [-10, 0, 0, 4, 5] it should be clipped to  [0, 0, 0, 4, 5] (to make things cleaner I don't talking about gradient clipping)."
        },
        {
          "id": 197288,
          "postDate": "2017-06-29T07:49:59.657Z",
          "content": "<p>Yes, and if trained well, it is really hard to predict -10, since training labels never less then 0. In statistically, it is no need to clip even use linear, but we can simply clip the noise around 0 and get better performance.</p>",
          "rawMarkdown": "Yes, and if trained well, it is really hard to predict -10, since training labels never less then 0. In statistically, it is no need to clip even use linear, but we can simply clip the noise around 0 and get better performance."
        },
        {
          "id": 199823,
          "postDate": "2017-07-06T11:51:31.807Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 199843,
          "postDate": "2017-07-06T12:38:03.150Z",
          "content": "<p>58 training epochs is a example in the kernel. I train the new FC several epochs, observe the loss, then train more layers until the train loss close to 0.3. The total epochs before training the whole network is about 60(I think it is too slow due to my poor skill). And the whole network takes 27 epochs to got 0.126/0.247 train/validation loss with private LB 11.7. </p>",
          "rawMarkdown": "58 training epochs is a example in the kernel. I train the new FC several epochs, observe the loss, then train more layers until the train loss close to 0.3. The total epochs before training the whole network is about 60(I think it is too slow due to my poor skill). And the whole network takes 27 epochs to got 0.126/0.247 train/validation loss with private LB 11.7. "
        }
      ]
    },
    {
      "id": 196749,
      "postDate": "2017-06-28T02:53:58.883Z",
      "content": "<p>Congratulations, thanks for sharing your solution.</p>\n\n<p>A little question is, can you share what computer did you use and how long did training 30 epochs took?</p>",
      "rawMarkdown": "Congratulations, thanks for sharing your solution.\n\nA little question is, can you share what computer did you use and how long did training 30 epochs took?",
      "replies": [
        {
          "id": 196758,
          "postDate": "2017-06-28T03:17:14.697Z",
          "content": "<p>i7 7770K, 16GB RAM, 1 GTX 1080.</p>\n\n<p>For single scale model, positive training patches = 14237, so total training patches = 54100 (positive*4*0.95). If one epoch means total training patches, it took about 30 mins a epoch. (Multi-scale takes 4 times) It took several days for training.</p>",
          "rawMarkdown": "i7 7770K, 16GB RAM, 1 GTX 1080.\n\nFor single scale model, positive training patches = 14237, so total training patches = 54100 (positive*4*0.95). If one epoch means total training patches, it took about 30 mins a epoch. (Multi-scale takes 4 times) It took several days for training.",
          "votes": 2
        }
      ]
    },
    {
      "id": 197468,
      "postDate": "2017-06-29T16:37:56.817Z",
      "content": "<p>Great work! Thanks for sharing!</p>",
      "rawMarkdown": "Great work! Thanks for sharing!"
    },
    {
      "id": 196760,
      "postDate": "2017-06-28T03:20:34.857Z",
      "content": "<p>congrats and thanks for sharing.</p>",
      "rawMarkdown": "congrats and thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 196751,
      "author_name": "Radu Stoicescu",
      "author_url": "",
      "post_date": "2017-06-28T03:02:21.327000",
      "content": "<p>I'm proud to see my blob detection in the winning solution.</p>\n\n<p>Good job!</p>",
      "votes": 8,
      "replies": [
        {
          "id": 196783,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-28T05:13:23.207000",
          "content": "<p>Thanks again.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 197128,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2017-06-28T21:43:51.967000",
      "content": "<p>Outstanding work outrunner! Thanks for sharing; may I ask... did you have any intuition, before starting out, that this network would work ? And what do you think made the big difference in this architecture - was it the 1024 FC or the SGD or slow training or sthing else ... i am sure you tried lots of changes in this area. \nI gather you started with pretrained weights. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 197219,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-29T02:39:36.913000",
          "content": "<p>Just intuition, and I think it would work undoubtedly even when I got trouble at the beginning. VGG is a basic CNN, I took the 1024 FC as a buffer between FCN and output. The size is according to RAM limit and model size, and intuitively I think it is enough. When I got scale problem, I tried more neurons and nothing happened. SGD is more stable than other aggressive optimizers in my work, maybe it is due to the large variation in output labels.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 196752,
      "author_name": "sfx7rail",
      "author_url": "",
      "post_date": "2017-06-28T03:02:34.203000",
      "content": "<p>Congratulations and Excellent work. Your kaggle ranking  will be top 40.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 197692,
      "author_name": "outrunner",
      "author_url": "",
      "post_date": "2017-06-30T03:08:24.533000",
      "content": "<p>Something I could share about why misclassification on females and juveniles is worse more than I expected.</p>\n\n<p><img src=\"http://i.imgur.com/rJBOLir.gif\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>I generated random dot for explanation. I wish the threshold between females and juveniles in the model could be the green line, but it is too naive. Then I think at least it should be the red line by statistically, but in fact the threshold is dotted red line. (It is about size feature, there are other features like lion distribution etc. in the model let the prediction better.)</p>\n\n<p>I guess it is because in the small scale patch, lion count is larger, and the squared error even more. I tried add large scale training patch with more juveniles and it helps a little, but overall score is not significant.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 198656,
      "author_name": "erde",
      "author_url": "",
      "post_date": "2017-07-03T09:16:21.487000",
      "content": "<p>Would you mind creating a visualization of the predictions that the network produces similar to the ones by Konstantin Lopuhin, e.g. using <code>for i, p in enumerate(batch_of_predictions): scipy.misc.imsave('preds_%i.png' % i, normalize(np.concatenate(np.moveaxis(p, -1, 0))))</code>? I think it would be interesting to see how well the network manages to tell the classes apart in various cases and what the predicted dots look like.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 198664,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-07-03T09:54:53.283000",
          "content": "<p>I don't quite understand what is batch_of_predictions?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 198671,
          "author_name": "erde",
          "author_url": "",
          "post_date": "2017-07-03T10:40:24.903000",
          "content": "<p>The variable <code>batch_of_predictions</code> is be the output of your network. Usually, all operations in deep learning frameworks operate on batches of training examples. When you run the model to test it on a couple of images, you need to feed a batch of n≥1 images and you get n predictions back. This is how you can save all images from the batch. I basically just copied &amp; modified a line of code from my solution.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 198681,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-07-03T11:22:55.327000",
          "content": "<p>I mean, in my network, predictions are only numbers. Or you want to see some testing patches with its output?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 198816,
          "author_name": "erde",
          "author_url": "",
          "post_date": "2017-07-03T19:17:09.003000",
          "content": "<p>Oops, I misunderstood the way you create the <code>trainY</code>s. I thought your <code>trainY</code>s consisted of dots at the animal coordinates. Never mind then. I could imagine your solution works well because only predicting the total counts (not the locations) in a large window like that reduces noise from the label positions (sometimes the labels are even outside of the animals). For your method, imprecise label positions only cause noise at the borders of the patch.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 198888,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-07-04T02:07:23.307000",
          "content": "<p>By the way, this is how I verify the prediction: （more count more bright)\n<img src=\"http://i.imgur.com/FKQdbpX.jpg\" alt=\"enter image description here\" title=\"\"></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 198899,
          "author_name": "JamesGoGo",
          "author_url": "",
          "post_date": "2017-07-04T03:22:00.263000",
          "content": "<p>It is cool! How can you show the picture in this way? Could you please show some code on this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 198901,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-07-04T03:50:06.773000",
          "content": "<p>for p in test_patches: p = p*f(count(p)) #make sure 0 &lt; f &lt; 1 or do some normalized processing</p>\n\n<p>then stitch them side by side</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 198600,
      "author_name": "MingT",
      "author_url": "",
      "post_date": "2017-07-03T05:14:55.613000",
      "content": "<p>Congrats on winning. Great solution!! \nI noticed that the test was also run on the patches that were cut from the original test images and the overall count for an image is the sum of the counts from all the patches cut from the image. Since the patch cut is random, would a sea lion at the boundary of two patches be double-counted? How would the double counting be avoided in the solution?\nThanks!!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 198653,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-07-03T08:54:14.403000",
          "content": "<p>Ideally, output should be 0.5 if the lion has been cut in half. Because in the same situation at training stage, one patch is 1 and the other one is 0.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 197253,
      "author_name": "Sander",
      "author_url": "",
      "post_date": "2017-06-29T06:33:34.793000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 197252,
      "author_name": "Sander",
      "author_url": "",
      "post_date": "2017-06-29T06:33:22.160000",
      "content": "<p>Hi Outrunner,</p>\n\n<p>Congratulations!</p>\n\n<p>I'm not very familiar with Keras code. Could you explain the last layers of your neural network? What happens after the convolutional layers? Did you just use linear activations to generate a 5-layer heatmap and compared that to the GT heatmap using RMSE?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 197267,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-29T07:00:52.650000",
          "content": "<p>The \"Patches looks like\" section in notebook, image is training patch, [1 0 3 0 1] is label (lion count per type in the patch), the last is 5 neurons fully-connected layer with linear output.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197570,
          "author_name": "Sander",
          "author_url": "",
          "post_date": "2017-06-29T20:52:45.550000",
          "content": "<p>Aha! So it's a kind of regression?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197595,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-29T23:11:52.107000",
          "content": "<p>I think so.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 197238,
      "author_name": "Jerome Zhou",
      "author_url": "",
      "post_date": "2017-06-29T04:45:26.363000",
      "content": "<p>Great job here!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 197143,
      "author_name": "mrgloom",
      "author_url": "",
      "post_date": "2017-06-28T22:17:25.833000",
      "content": "<ol>\n<li>In your kernel you use only 1 image, how training was setup in practice? it was something like Keras ImageDataGenerator or you load all data to RAM?</li>\n<li>Does LB score depend on tile size? How you handle images at prediction time, i.e. with tile overlapping? In my opinion in case of pretrained network it should be about 'original' size, i.e. 224x224 for VGG.</li>\n<li>Why mean_squared_error was used that is MSE and in competition we have RMSE?</li>\n<li>What LB score you get without tricks described in Experience  section?</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 197225,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-29T03:10:59.207000",
          "content": "<ol>\n<li><p>I use custom data generator.</p></li>\n<li><p>I trained a 200x200 and it performed worse. Maybe it is because 300 contains more relationship information between pups and adult_females. Flipping(rotating) 4 times at prediction time improved 0.04 RMSE. Using only convolutional layer in VGG is size free.</p></li>\n<li><p>Just a normal loss function.</p></li>\n<li><p>You can reference the table in the section.</p></li>\n</ol>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 197139,
      "author_name": "harshml",
      "author_url": "",
      "post_date": "2017-06-28T22:07:43.777000",
      "content": "<p>Respect! Never underestimate the basics.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 197074,
      "author_name": "eagle4",
      "author_url": "",
      "post_date": "2017-06-28T18:22:34.977000",
      "content": "<p>@outrunner\nCongrats and thanks for the code and the explanations.</p>\n\n<p>As you requested, please find attached the bounding boxes. the first batch are the 5000 first bboxes manually labeled and the other 75 to 80,000 are the ones generated using SSD (no guarantee on the quality and some subadult male disappeared in the mix).</p>\n\n<p>Keep us informed on your progress with them please.</p>\n\n<p>I generated them as I started this competition as standard detection problem but later switched to a segmentation where they were not the best idea and dragged me to the bottom of the leader-board: I was more concerned by crowed area where bboxes are an handicap rather than image scale...</p>",
      "votes": 0,
      "replies": [
        {
          "id": 197214,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-29T02:14:31.817000",
          "content": "<p>Thanks. But I may try it when I have time or need to do farther research on this topic.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 196987,
      "author_name": "erde",
      "author_url": "",
      "post_date": "2017-06-28T14:38:50.380000",
      "content": "<p>It is interesting that the model trained at 1.0x scale had a much better performance being evaluated on the test set at 0.67x scale than on 1.0x. What might be the reason for this? Perhaps the receptive field of VGG16 is just not big enough to actually recognize the animals at 1.0x scale. VGG16 has an effective receptive field size of ~220px at the output neurons, but the largest adult males are up to ~260-270px in length. Also the largest distance of the pubs to other animals is about 250px.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 197003,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-28T15:27:55.667000",
          "content": "<p>I think it is typically because the testing image's sea lion is larger than training in average. In my opinion, the model don't know what sea lion is and where it is.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197088,
          "author_name": "erde",
          "author_url": "",
          "post_date": "2017-06-28T19:22:08.303000",
          "content": "<p>&gt; In my opinion, the model don't know what sea lion is and where it is.</p>\n\n<p>I agree. :) 'Not recognizing' was just shorthand for 'it might be harder to find distinguishing features if the animals are often only partially in sight for the network'.  My info was incomplete/wrong anyway though: VGG's receptive field is 224 × 224 with zero-padding, but 404 × 404 if the net is applied convolutionally^<a href=\"https://arxiv.org/pdf/1412.7062.pdf\">1</a>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197204,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-29T01:53:17.547000",
          "content": "<p>I haven't studied receptive field in detail yet. So, thanks for sharing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 196976,
      "author_name": "erde",
      "author_url": "",
      "post_date": "2017-06-28T14:12:33.797000",
      "content": "<p>Fantastic result, congrats!</p>\n\n<p>Am I understanding correctly that the multi-scale model is essentially heavy data augmentation with different scales?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 196967,
      "author_name": "LivingProgram",
      "author_url": "",
      "post_date": "2017-06-28T13:48:30.440000",
      "content": "<p>Such a simple solution. It really goes to show how far machine learning, convolutional neural networks, and a perfect implementation, application, and execution can take you. <strong>Congrats @outrunner on 1st place</strong>, and much thanks for the informative write-up. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 196970,
          "author_name": "LivingProgram",
          "author_url": "",
          "post_date": "2017-06-28T13:52:27.350000",
          "content": "<p>I also had a question about the values on your table, I don't quite understand why a higher value on the table like the values in your final submission column yielded better results. What is the significance of <strong>average of juveniles# / (adult_females# + juveniles#)</strong>?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 196982,
          "author_name": "erde",
          "author_url": "",
          "post_date": "2017-06-28T14:27:33.503000",
          "content": "<p>I think the numbers just show that on the test set outrunner's models predicted much fewer juveniles vs adult_females than during training, thus he/she corrected roughly for that ratio (with great success). E.g. the the patches with 46-50 juveniles had a number of 0.93 during training with a single zoom level (1x), but on the test set only 0.68, so it was corrected to 0.76. In that particular case it would have been interesting to see the performance of the r0.59 model without post-processing on the public LB.</p>\n\n<p>This might imply that the test set had different juvenile/adult_female ratios than the training set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 196997,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-28T15:14:00.110000",
          "content": "<p>Sorry for being unclear and thanks to erde. It is the ratio between juveniles and juveniles + adult_females.</p>\n\n<p>I think the ratio problem comes from scale problem. The CNN basically is a statistical model, simply distinguish juveniles and adult_females by size. Scale down image can get more juveniles, but the problem is if out of range, total lions count will go down. For example:</p>\n\n<ul>\n<li>scale   predict</li>\n<li>0.59    [ 13   5 180 120   2]</li>\n<li>0.53    [ 14   8 141 149   0]</li>\n<li>0.48    [ 10   5 132 165   2]</li>\n<li>0.43    [  9   7 122 182   0]</li>\n<li>0.39    [  6   7 101 185   0]</li>\n<li>0.35    [  7   7  74 176   0]</li>\n<li>0.31    [  4   6  62 168   1]</li>\n<li>0.28    [  5   7  48 137   0]</li>\n</ul>\n\n<p>This is the predict result of test/12690 on multi-scale model with different test scale. The juveniles + adult_females keeps around 300 for 0.59~0.43, and began to decrease from 0.39.</p>\n\n<p>So, all you are interesting in without post processing's performance. It is sure if I scale down more. But I don't want to take the risk if sea lions in private set are smaller then public.</p>\n\n<blockquote>\n  <p>the multi-scale model is essentially heavy data augmentation with different scales?</p>\n</blockquote>\n\n<p>Yes, and single scale model's keeping number scale range is narrower.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 197004,
          "author_name": "LivingProgram",
          "author_url": "",
          "post_date": "2017-06-28T15:29:00.683000",
          "content": "<p>Ah I see so scaled down images tend to have more juveniles than adult females versus not scaled down. If it were possible to teach the method how to compare the sizes of sea lions within a single image's scale, then the model could learn to compare the juvenile's relative size to female's relative size and your post processing trick might not have been needed? Maybe if first the image is processed for females then those features are fed to another network that searches for juveniles, the features can be reused. I think your post processing basically gave the model domain knowledge about the problem that could be replaced with a different architecture, would be interesting to test this!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197007,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-28T15:39:33.807000",
          "content": "<p>Sure, and it is not easy work. By the way, if they have grounding truth bounding box, I want to train a object detection for image scale detect by sea lion's size, just a thought.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197012,
          "author_name": "Cut Onion",
          "author_url": "",
          "post_date": "2017-06-28T15:59:07.277000",
          "content": "<p>Ideally, there should be preprocessing that makes all images the same scale. Much like satellite images are orthorectified to make objects on the ground to be the same scale and angle (so square on ground is square on image no matter how camera looks at Earth -- at right or some different angle). The same should be done here -- knowing flight path, UAV elevation, camera characteristics, capture angle and so on, it is possible to make all images the same scale. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 197023,
          "author_name": "Artem.Sanakoev",
          "author_url": "",
          "post_date": "2017-06-28T16:20:39.583000",
          "content": "<p>We labeled some images  from train set according to scale and tried to train a CNN to regress a scale of the image. But it didn't work out. I recon this is due to high variation in terrain and not being able to estimate scale if you look at tiles with a small spatial context (even with my own eyes).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197028,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-28T16:32:43.247000",
          "content": "<p>I tried a little different, and failed as expected.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197037,
          "author_name": "zero zero",
          "author_url": "",
          "post_date": "2017-06-28T16:51:15.843000",
          "content": "<p>agreed. including such data would make for a more robust approach.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 196839,
      "author_name": "Faron",
      "author_url": "",
      "post_date": "2017-06-28T08:16:13.257000",
      "content": "<p>Well done &amp; congratz @outrunner.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 196830,
      "author_name": "Vladimir Savinov",
      "author_url": "",
      "post_date": "2017-06-28T08:02:23.047000",
      "content": "<p>Сongratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 196769,
      "author_name": "Devin Anzelmo",
      "author_url": "",
      "post_date": "2017-06-28T03:44:00.367000",
      "content": "<p>Congrats! Great solution!</p>\n\n<p>A couple of questions: What was the difference in public/private score between r0.59 and r0.59*1.3  ?</p>\n\n<p>Did you try other optimizers/did the learning rate schedule make a large difference in score?</p>\n\n<p>For the multi-scale model did you take larger images and resize them down to 300x300 or did you train seperate networks on different scaling of the 300x300 patches?</p>\n\n<p>Thanks, and congratulations on the win again.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 196780,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-28T05:05:25.763000",
          "content": "<p>Sorry, do you mean different score or what make the score difference? I did not submit r0.59, I think it may be 14~15. About post processing: if a image predict in r0.59 = [10,10,100,100,20], then submit r0.59*1.3 = [10,10,70,130,20]</p>\n\n<p>I tried other optimizers/learning rate schedule, just for observe the learning loss convergence.</p>\n\n<p>I scaled down the original image (*0.9, ..., *0.6561), cut into 300x300, and use those patches together to train a network.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 196788,
          "author_name": "Devin Anzelmo",
          "author_url": "",
          "post_date": "2017-06-28T05:21:11.980000",
          "content": "<p>Thanks for the info. You are correct I was wondering what the leaderboard score of r0.59 was. From 14~15 to ~11 is a big boost in score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 196791,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-28T05:39:33.490000",
          "content": "<p>Additionally, post processing also add pups, in this case, it should be [10,10,70,130,24]. But the main contribution comes from juveniles/adult_females adjustment.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 196822,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2017-06-28T07:34:11.220000",
          "content": "<p>Great result, congratulation with winning the competition!</p>\n\n<p>Interesting, what was your score before post processing? It's amazing a relatively simple cnn managed to predict number of lion so well.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 196828,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-28T07:56:48.270000",
          "content": "<p>I found a submission got 13.46393 private LB score is made by:</p>\n\n<ul>\n<li>training scale: *1.23~0.66</li>\n<li>testing scale: *0.48</li>\n</ul>\n\n<p>only cut pups more than 1.8x adult_females after predict, and this affect a little in my experience.</p>\n\n<p>So, without post processing on other combination, maybe it could be around 13.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 196766,
      "author_name": "Kishore M",
      "author_url": "",
      "post_date": "2017-06-28T03:35:23.173000",
      "content": "<p>congratulations and amazing that you achieved  #1 position with in four months .. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 196761,
      "author_name": "KapilYadav",
      "author_url": "",
      "post_date": "2017-06-28T03:27:50.767000",
      "content": "<p>Congrats for the 1st position. Thanks for sharing the details. Its a nice approach.</p>\n\n<p>I have a few questions - </p>\n\n<ol>\n<li>Did you make all the VGG16 layers trainable ? ( I tried VGG16, but I couldnt make it stop overfit on train data. :(  )</li>\n<li>How were the final output of linear layers compared ? ( I guess, a sigmoid layer was used ... )</li>\n<li>Just a clarification - In cases when a patch has multiple sea-lions, the network will output high values for their respective nodes in the last layer, right? </li>\n<li>Did you face any difficulty for patches where sea-lions were all crowded in one place?</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 196772,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-28T03:52:21.703000",
          "content": "<ol>\n<li><p>Yes, train the new FC layers first, then all layers.</p></li>\n<li><p>I use linear activation at output layer. (because 3.)</p></li>\n<li><p>Yes.</p></li>\n<li><p>Training got exploded sometimes, but it is not problem. When testing, distinguish juveniles and adult_females is difficult. I think the model don't know what is \"relative\" size. Actually, in the training set, we also got a lot of human label noise at juveniles and adult_females.</p></li>\n</ol>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 196966,
          "author_name": "KapilYadav",
          "author_url": "",
          "post_date": "2017-06-28T13:47:34.633000",
          "content": "<p>Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197134,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "2017-06-28T21:52:43.907000",
          "content": "<p>Why linear activation? is that mean that it can predict negative values?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197206,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-29T01:55:00.797000",
          "content": "<p>Why not?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197257,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "2017-06-29T06:37:05.960000",
          "content": "<p>I mean it will work, but it can produce negative values as side affect? are you perform some sort of clipping?</p>\n\n<p>Will it work with Relu activation?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197279,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-29T07:36:51.033000",
          "content": "<p>Relu got problem at 0, and also need to do clip.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197284,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "2017-06-29T07:44:24.763000",
          "content": "<p>Can you elaborate on this? I mean that if final prediction will be [-10, 0, 0, 4, 5] it should be clipped to  [0, 0, 0, 4, 5] (to make things cleaner I don't talking about gradient clipping).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 197288,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-29T07:49:59.657000",
          "content": "<p>Yes, and if trained well, it is really hard to predict -10, since training labels never less then 0. In statistically, it is no need to clip even use linear, but we can simply clip the noise around 0 and get better performance.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 199823,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-07-06T11:51:31.807000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 199843,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-07-06T12:38:03.150000",
          "content": "<p>58 training epochs is a example in the kernel. I train the new FC several epochs, observe the loss, then train more layers until the train loss close to 0.3. The total epochs before training the whole network is about 60(I think it is too slow due to my poor skill). And the whole network takes 27 epochs to got 0.126/0.247 train/validation loss with private LB 11.7. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 196749,
      "author_name": "InfiniteWing",
      "author_url": "",
      "post_date": "2017-06-28T02:53:58.883000",
      "content": "<p>Congratulations, thanks for sharing your solution.</p>\n\n<p>A little question is, can you share what computer did you use and how long did training 30 epochs took?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 196758,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2017-06-28T03:17:14.697000",
          "content": "<p>i7 7770K, 16GB RAM, 1 GTX 1080.</p>\n\n<p>For single scale model, positive training patches = 14237, so total training patches = 54100 (positive*4*0.95). If one epoch means total training patches, it took about 30 mins a epoch. (Multi-scale takes 4 times) It took several days for training.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 197468,
      "author_name": "Josue De Santiago",
      "author_url": "",
      "post_date": "2017-06-29T16:37:56.817000",
      "content": "<p>Great work! Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 196760,
      "author_name": "zero zero",
      "author_url": "",
      "post_date": "2017-06-28T03:20:34.857000",
      "content": "<p>congrats and thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "196735": "Thanks to everyone. I love this competition.\n\n[Here is my solution.][1] \n\nI use VGG16 without top, add FC-1024, and FC-5 with linear output.\n\nSGD optimization and mean_squared_error loss.\n\nAnd post processing for better regression.\n\nI think it is a simple way to count. Looking forward to hearing your thoughts, thanks.\n\n  [1]: https://www.kaggle.com/outrunner/use-keras-to-count-sea-lions/notebook",
    "196751": "I'm proud to see my blob detection in the winning solution.\n\nGood job!",
    "197128": "Outstanding work outrunner! Thanks for sharing; may I ask... did you have any intuition, before starting out, that this network would work ? And what do you think made the big difference in this architecture - was it the 1024 FC or the SGD or slow training or sthing else ... i am sure you tried lots of changes in this area. \nI gather you started with pretrained weights. ",
    "196752": "Congratulations and Excellent work. Your kaggle ranking  will be top 40.",
    "197692": "Something I could share about why misclassification on females and juveniles is worse more than I expected.\n\n![enter image description here][1]\n\nI generated random dot for explanation. I wish the threshold between females and juveniles in the model could be the green line, but it is too naive. Then I think at least it should be the red line by statistically, but in fact the threshold is dotted red line. (It is about size feature, there are other features like lion distribution etc. in the model let the prediction better.)\n\nI guess it is because in the small scale patch, lion count is larger, and the squared error even more. I tried add large scale training patch with more juveniles and it helps a little, but overall score is not significant.\n\n  [1]: http://i.imgur.com/rJBOLir.gif",
    "198656": "Would you mind creating a visualization of the predictions that the network produces similar to the ones by Konstantin Lopuhin, e.g. using `for i, p in enumerate(batch_of_predictions): scipy.misc.imsave('preds_%i.png' % i, normalize(np.concatenate(np.moveaxis(p, -1, 0))))`? I think it would be interesting to see how well the network manages to tell the classes apart in various cases and what the predicted dots look like.",
    "198600": "Congrats on winning. Great solution!! \nI noticed that the test was also run on the patches that were cut from the original test images and the overall count for an image is the sum of the counts from all the patches cut from the image. Since the patch cut is random, would a sea lion at the boundary of two patches be double-counted? How would the double counting be avoided in the solution?\nThanks!!",
    "197253": "",
    "197252": "Hi Outrunner,\n\nCongratulations!\n\nI'm not very familiar with Keras code. Could you explain the last layers of your neural network? What happens after the convolutional layers? Did you just use linear activations to generate a 5-layer heatmap and compared that to the GT heatmap using RMSE?",
    "197238": "Great job here!",
    "197143": "1. In your kernel you use only 1 image, how training was setup in practice? it was something like Keras ImageDataGenerator or you load all data to RAM?\n2.  Does LB score depend on tile size? How you handle images at prediction time, i.e. with tile overlapping? In my opinion in case of pretrained network it should be about 'original' size, i.e. 224x224 for VGG.\n3. Why mean_squared_error was used that is MSE and in competition we have RMSE?\n4. What LB score you get without tricks described in Experience  section?",
    "197139": "Respect! Never underestimate the basics.",
    "197074": "@outrunner\nCongrats and thanks for the code and the explanations.\n\nAs you requested, please find attached the bounding boxes. the first batch are the 5000 first bboxes manually labeled and the other 75 to 80,000 are the ones generated using SSD (no guarantee on the quality and some subadult male disappeared in the mix).\n\nKeep us informed on your progress with them please.\n\nI generated them as I started this competition as standard detection problem but later switched to a segmentation where they were not the best idea and dragged me to the bottom of the leader-board: I was more concerned by crowed area where bboxes are an handicap rather than image scale...",
    "196987": "It is interesting that the model trained at 1.0x scale had a much better performance being evaluated on the test set at 0.67x scale than on 1.0x. What might be the reason for this? Perhaps the receptive field of VGG16 is just not big enough to actually recognize the animals at 1.0x scale. VGG16 has an effective receptive field size of ~220px at the output neurons, but the largest adult males are up to ~260-270px in length. Also the largest distance of the pubs to other animals is about 250px.",
    "196976": "Fantastic result, congrats!\n\nAm I understanding correctly that the multi-scale model is essentially heavy data augmentation with different scales?",
    "196967": "Such a simple solution. It really goes to show how far machine learning, convolutional neural networks, and a perfect implementation, application, and execution can take you. **Congrats @outrunner on 1st place**, and much thanks for the informative write-up. ",
    "196839": "Well done &amp; congratz @outrunner.",
    "196830": "Сongratulations!",
    "196769": "Congrats! Great solution!\n\nA couple of questions: What was the difference in public/private score between r0.59 and r0.59*1.3  ?\n\nDid you try other optimizers/did the learning rate schedule make a large difference in score?\n\nFor the multi-scale model did you take larger images and resize them down to 300x300 or did you train seperate networks on different scaling of the 300x300 patches?\n\nThanks, and congratulations on the win again.\n",
    "196766": "congratulations and amazing that you achieved  #1 position with in four months .. ",
    "196761": "Congrats for the 1st position. Thanks for sharing the details. Its a nice approach.\n\nI have a few questions - \n\n 1. Did you make all the VGG16 layers trainable ? ( I tried VGG16, but I couldnt make it stop overfit on train data. :(  )\n 2. How were the final output of linear layers compared ? ( I guess, a sigmoid layer was used ... )\n 3. Just a clarification - In cases when a patch has multiple sea-lions, the network will output high values for their respective nodes in the last layer, right? \n 4. Did you face any difficulty for patches where sea-lions were all crowded in one place?",
    "196749": "Congratulations, thanks for sharing your solution.\n\nA little question is, can you share what computer did you use and how long did training 30 epochs took?",
    "197468": "Great work! Thanks for sharing!",
    "196760": "congrats and thanks for sharing."
  }
}