{
  "id": 35442,
  "title": "Solution in a nutshell. 2th Public LB / 4th Private LB",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/writeups/dl-sucks-solution-in-a-nutshell-2th-public-lb-4th-",
  "author_name": "",
  "post_date": "2017-09-10T22:30:40.410Z",
  "votes": 36,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Here I will give a short description of the approach of our team (me and <a href=\"https://www.kaggle.com/chelovekparohod\">DmitryKotovenko</a>). <br>\nThe source code is available on GitHub <a href=\"https://github.com/asanakoy/kaggle_sea_lions_counting\">https://github.com/asanakoy/kaggle_sea_lions_counting</a>.</p>\n\n<h2>Preprocessing</h2>\n\n<p>Thanks to @Radu Stoicescu for his blob detection to get corrected counts from dotted images.</p>\n\n<p>GT count for each tile was generated as a sum over heatmap (to overcome cases with lions on the border of the tile).   On top of each lion we put a Gaussian with a standard deviation heuristically estimated by calculating the smallest distance between lion on the image. <br>\nWe set the standard deviation to be 50 at least for each Gaussian and adjust it according to size of the animals from different classes (\nmultiplied by 2 for adult males and by 0.5 for pups).</p>\n\n<p><a href=\"https://ibb.co/hhAda5\">GT Gaussians for tile</a>\n<img src=\"https://preview.ibb.co/bJ5Bv5/Screenshot_from_2017_06_28_16_28_57.png\" alt=\"GT Gaussians for tile\" title=\"\"></p>\n\n<h2>Model and Training</h2>\n\n<p>Our model incarnates <strong>regression for 5 classes on tiles of the images</strong>. (In similar spirit as the approach of @outrunner)</p>\n\n<p><strong>Inception Resnet v2</strong> pretrainedon Imagenet. <br>\nWe substituted the last layer with 256-way FC layer + dropout + 5-way FC layer on top. + RMSE loss.\nThen we fine-tuned the model on 299x299 image tiles with Adam optimizer. </p>\n\n<p><strong>Augmentations:</strong>  random rotation on 90/180/270 grads, random flip left-right, bottom-up.</p>\n\n<p><strong>Scale augmentations:</strong> one model without them, one model with 0.83-1.25 random scaling, one model with 0.66 - 1.5 random scaling.</p>\n\n<p><a href=\"https://ibb.co/kVhwTQ\">RMSE on val for 3 best models</a>\n<img src=\"https://preview.ibb.co/jx6O8Q/Screenshot_from_2017_06_28_17_02_17.png\" alt=\"RMSE on val for 3 best models\" title=\"\"></p>\n\n<h2>Testing</h2>\n\n<p><strong>During test</strong> we made predictions up to 5 times for each model using different shifts of the tiles in the image. \nTest images were downscaled in 0.4-0.5 times.</p>\n\n<p>The final ensemble was made by averaging all the predictions. <br>\n<strong>Private LB RMSE:</strong> <em>13.18968</em> <br>\n<strong>Public LB RMSE:</strong> <em>13.29065</em></p>\n\n<p>Applying further postprocessing as suggested by @outrunner, could improve results. <br>\nJust increasing the number of pups by 20% gives a huge improvement: <br>\n<strong>Private LB RMSE:</strong>  <em>12.58131</em> <br>\n<strong>Public LB RMSE:</strong> <em>12.75510</em></p>\n\n<h2>Some negative experiments</h2>\n\n<p>We labeled some images from train set according to scale and trained a CNN to regress a scale of the image.   This could unify all the images to have the same approximate size of the lions of corresponding classes and simplify the CNN training to count animals.</p>\n\n<p>But it didn't work out.   I reckon, the reason is the high variation in terrain and inability to estimate scale of objects if you look at them within a small spatial context (even with my own eyes).</p>",
  "messages": [
    {
      "id": "197005",
      "postDate": "06/28/2017 15:31:26",
      "content": "<p>Here I will give a short description of the approach of our team (me and <a href=\"https://www.kaggle.com/chelovekparohod\">DmitryKotovenko</a>). <br>\nThe source code is available on GitHub <a href=\"https://github.com/asanakoy/kaggle_sea_lions_counting\">https://github.com/asanakoy/kaggle_sea_lions_counting</a>.</p>\n\n<h2>Preprocessing</h2>\n\n<p>Thanks to @Radu Stoicescu for his blob detection to get corrected counts from dotted images.</p>\n\n<p>GT count for each tile was generated as a sum over heatmap (to overcome cases with lions on the border of the tile).   On top of each lion we put a Gaussian with a standard deviation heuristically estimated by calculating the smallest distance between lion on the image. <br>\nWe set the standard deviation to be 50 at least for each Gaussian and adjust it according to size of the animals from different classes (\nmultiplied by 2 for adult males and by 0.5 for pups).</p>\n\n<p><a href=\"https://ibb.co/hhAda5\">GT Gaussians for tile</a>\n<img src=\"https://preview.ibb.co/bJ5Bv5/Screenshot_from_2017_06_28_16_28_57.png\" alt=\"GT Gaussians for tile\" title=\"\"></p>\n\n<h2>Model and Training</h2>\n\n<p>Our model incarnates <strong>regression for 5 classes on tiles of the images</strong>. (In similar spirit as the approach of @outrunner)</p>\n\n<p><strong>Inception Resnet v2</strong> pretrainedon Imagenet. <br>\nWe substituted the last layer with 256-way FC layer + dropout + 5-way FC layer on top. + RMSE loss.\nThen we fine-tuned the model on 299x299 image tiles with Adam optimizer. </p>\n\n<p><strong>Augmentations:</strong>  random rotation on 90/180/270 grads, random flip left-right, bottom-up.</p>\n\n<p><strong>Scale augmentations:</strong> one model without them, one model with 0.83-1.25 random scaling, one model with 0.66 - 1.5 random scaling.</p>\n\n<p><a href=\"https://ibb.co/kVhwTQ\">RMSE on val for 3 best models</a>\n<img src=\"https://preview.ibb.co/jx6O8Q/Screenshot_from_2017_06_28_17_02_17.png\" alt=\"RMSE on val for 3 best models\" title=\"\"></p>\n\n<h2>Testing</h2>\n\n<p><strong>During test</strong> we made predictions up to 5 times for each model using different shifts of the tiles in the image. \nTest images were downscaled in 0.4-0.5 times.</p>\n\n<p>The final ensemble was made by averaging all the predictions. <br>\n<strong>Private LB RMSE:</strong> <em>13.18968</em> <br>\n<strong>Public LB RMSE:</strong> <em>13.29065</em></p>\n\n<p>Applying further postprocessing as suggested by @outrunner, could improve results. <br>\nJust increasing the number of pups by 20% gives a huge improvement: <br>\n<strong>Private LB RMSE:</strong>  <em>12.58131</em> <br>\n<strong>Public LB RMSE:</strong> <em>12.75510</em></p>\n\n<h2>Some negative experiments</h2>\n\n<p>We labeled some images from train set according to scale and trained a CNN to regress a scale of the image.   This could unify all the images to have the same approximate size of the lions of corresponding classes and simplify the CNN training to count animals.</p>\n\n<p>But it didn't work out.   I reckon, the reason is the high variation in terrain and inability to estimate scale of objects if you look at them within a small spatial context (even with my own eyes).</p>",
      "rawMarkdown": "Here I will give a short description of the approach of our team (me and [DmitryKotovenko][1]).  \nThe source code is available on GitHub https://github.com/asanakoy/kaggle_sea_lions_counting.\n\n\nPreprocessing\n--------------\n\nThanks to @Radu Stoicescu for his blob detection to get corrected counts from dotted images.\n\nGT count for each tile was generated as a sum over heatmap (to overcome cases with lions on the border of the tile).   On top of each lion we put a Gaussian with a standard deviation heuristically estimated by calculating the smallest distance between lion on the image.   \nWe set the standard deviation to be 50 at least for each Gaussian and adjust it according to size of the animals from different classes (\nmultiplied by 2 for adult males and by 0.5 for pups).\n\n<a href=\"https://ibb.co/hhAda5\">GT Gaussians for tile</a>\n![GT Gaussians for tile][2]\n\n\nModel and Training\n------\nOur model incarnates **regression for 5 classes on tiles of the images**. (In similar spirit as the approach of @outrunner)\n\n**Inception Resnet v2** pretrainedon Imagenet.   \nWe substituted the last layer with 256-way FC layer + dropout + 5-way FC layer on top. + RMSE loss.\nThen we fine-tuned the model on 299x299 image tiles with Adam optimizer. \n\n**Augmentations:**  random rotation on 90/180/270 grads, random flip left-right, bottom-up.\n\n**Scale augmentations:** one model without them, one model with 0.83-1.25 random scaling, one model with 0.66 - 1.5 random scaling.\n\n<a href=\"https://ibb.co/kVhwTQ\">RMSE on val for 3 best models</a>\n![RMSE on val for 3 best models][3]\n    \n\nTesting\n-------\n\n**During test** we made predictions up to 5 times for each model using different shifts of the tiles in the image. \nTest images were downscaled in 0.4-0.5 times.\n\nThe final ensemble was made by averaging all the predictions.   \n**Private LB RMSE:** *13.18968*  \n**Public LB RMSE:** *13.29065*\n\nApplying further postprocessing as suggested by @outrunner, could improve results.  \nJust increasing the number of pups by 20% gives a huge improvement:  \n**Private LB RMSE:**  *12.58131*  \n**Public LB RMSE:** *12.75510*\n\n\nSome negative experiments\n-------\n\nWe labeled some images from train set according to scale and trained a CNN to regress a scale of the image.   This could unify all the images to have the same approximate size of the lions of corresponding classes and simplify the CNN training to count animals.\n\nBut it didn't work out.   I reckon, the reason is the high variation in terrain and inability to estimate scale of objects if you look at them within a small spatial context (even with my own eyes).\n\n  [1]: https://www.kaggle.com/chelovekparohod\n  [2]: https://preview.ibb.co/bJ5Bv5/Screenshot_from_2017_06_28_16_28_57.png\n  [3]: https://preview.ibb.co/jx6O8Q/Screenshot_from_2017_06_28_17_02_17.png",
      "votes": null
    },
    {
      "id": "197021",
      "postDate": "06/28/2017 16:18:27",
      "content": "<p>Thanks for your sharing~</p>",
      "rawMarkdown": "Thanks for your sharing~",
      "votes": null
    },
    {
      "id": "200697",
      "postDate": "07/08/2017 23:07:14",
      "content": "<p>UPD: I have added a link to the source code on GitHub.</p>",
      "rawMarkdown": "UPD: I have added a link to the source code on GitHub.",
      "votes": null
    },
    {
      "id": "203552",
      "postDate": "07/15/2017 13:27:26",
      "content": "<p>Thanks a lot for sharing knowledge! </p>",
      "rawMarkdown": "Thanks a lot for sharing knowledge!",
      "votes": null
    },
    {
      "id": "204543",
      "postDate": "07/18/2017 22:32:24",
      "content": "<p>Thanks for sharing! Will enjoy delving into the code.</p>",
      "rawMarkdown": "Thanks for sharing! Will enjoy delving into the code.",
      "votes": null
    },
    {
      "id": "212189",
      "postDate": "08/10/2017 22:01:20",
      "content": "<p>@ Artem,\nThanks a lot for sharing the code. I am still intrigued by your solution which seems to combine a segmentation and a regression in the same model.\nI have a question since I am not tf fluent:</p>\n\n<p>As far as I understand, you take the train images, apply the gaussian on the train pictures (are these gaussian pictures a secondary output or the input in your model?) and you do a regression using a resnet like neural network on them.</p>\n\n<p>Can you re-explain please the part with the Gaussian kernels and how these gaussian kernels works at prediction time?</p>\n\n<p>Thanks in advance for your help.</p>",
      "rawMarkdown": "Artem,\nThanks a lot for sharing the code. I am still intrigued by your solution which seems to combine a segmentation and a regression in the same model.\nI have a question since I am not tf fluent:\n\nAs far as I understand, you take the train images, apply the gaussian on the train pictures (are these gaussian pictures a secondary output or the input in your model?) and you do a regression using a resnet like neural network on them.\n\nCan you re-explain please the part with the Gaussian kernels and how these gaussian kernels works at prediction time?\n\nThanks in advance for your help.",
      "votes": null
    },
    {
      "id": "212202",
      "postDate": "08/10/2017 22:55:40",
      "content": "<p>Hey @eagle4. \nGaussians are used to estimate the ground truth number of lions on the tile. Gaussians are not used during inference. Putting a gaussian on top of each sea lion allows us to cut this lion with the tile border and still have a fraction of the lion  in each adjacent tile. Otherwise, we would have to assign this lion to the only one tile.</p>",
      "rawMarkdown": "Hey @eagle4. \nGaussians are used to estimate the ground truth number of lions on the tile. Gaussians are not used during inference. Putting a gaussian on top of each sea lion allows us to cut this lion with the tile border and still have a fraction of the lion  in each adjacent tile. Otherwise, we would have to assign this lion to the only one tile.",
      "votes": null
    },
    {
      "id": "212347",
      "postDate": "08/11/2017 12:28:34",
      "content": "<p>Arthem,\nThanks a lot. It makes much more sense now. Congrats again on the solution.</p>",
      "rawMarkdown": "Arthem,\nThanks a lot. It makes much more sense now. Congrats again on the solution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 197021,
      "author_name": "infinitewing",
      "author_url": "",
      "post_date": "06/28/2017 16:18:27",
      "content": "<p>Thanks for your sharing~</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 200697,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "07/08/2017 23:07:14",
      "content": "<p>UPD: I have added a link to the source code on GitHub.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 203552,
      "author_name": "craymond",
      "author_url": "",
      "post_date": "07/15/2017 13:27:26",
      "content": "<p>Thanks a lot for sharing knowledge! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 204543,
      "author_name": "saiguysci",
      "author_url": "",
      "post_date": "07/18/2017 22:32:24",
      "content": "<p>Thanks for sharing! Will enjoy delving into the code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 212189,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "08/10/2017 22:01:20",
      "content": "<p>@ Artem,\nThanks a lot for sharing the code. I am still intrigued by your solution which seems to combine a segmentation and a regression in the same model.\nI have a question since I am not tf fluent:</p>\n\n<p>As far as I understand, you take the train images, apply the gaussian on the train pictures (are these gaussian pictures a secondary output or the input in your model?) and you do a regression using a resnet like neural network on them.</p>\n\n<p>Can you re-explain please the part with the Gaussian kernels and how these gaussian kernels works at prediction time?</p>\n\n<p>Thanks in advance for your help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 212202,
          "author_name": "asanakoev",
          "author_url": "",
          "post_date": "08/10/2017 22:55:40",
          "content": "<p>Hey @eagle4. \nGaussians are used to estimate the ground truth number of lions on the tile. Gaussians are not used during inference. Putting a gaussian on top of each sea lion allows us to cut this lion with the tile border and still have a fraction of the lion  in each adjacent tile. Otherwise, we would have to assign this lion to the only one tile.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 212347,
          "author_name": "chabir",
          "author_url": "",
          "post_date": "08/11/2017 12:28:34",
          "content": "<p>Arthem,\nThanks a lot. It makes much more sense now. Congrats again on the solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "197005": "Here I will give a short description of the approach of our team (me and [DmitryKotovenko][1]).  \nThe source code is available on GitHub https://github.com/asanakoy/kaggle_sea_lions_counting.\n\n\nPreprocessing\n--------------\n\nThanks to @Radu Stoicescu for his blob detection to get corrected counts from dotted images.\n\nGT count for each tile was generated as a sum over heatmap (to overcome cases with lions on the border of the tile).   On top of each lion we put a Gaussian with a standard deviation heuristically estimated by calculating the smallest distance between lion on the image.   \nWe set the standard deviation to be 50 at least for each Gaussian and adjust it according to size of the animals from different classes (\nmultiplied by 2 for adult males and by 0.5 for pups).\n\n<a href=\"https://ibb.co/hhAda5\">GT Gaussians for tile</a>\n![GT Gaussians for tile][2]\n\n\nModel and Training\n------\nOur model incarnates **regression for 5 classes on tiles of the images**. (In similar spirit as the approach of @outrunner)\n\n**Inception Resnet v2** pretrainedon Imagenet.   \nWe substituted the last layer with 256-way FC layer + dropout + 5-way FC layer on top. + RMSE loss.\nThen we fine-tuned the model on 299x299 image tiles with Adam optimizer. \n\n**Augmentations:**  random rotation on 90/180/270 grads, random flip left-right, bottom-up.\n\n**Scale augmentations:** one model without them, one model with 0.83-1.25 random scaling, one model with 0.66 - 1.5 random scaling.\n\n<a href=\"https://ibb.co/kVhwTQ\">RMSE on val for 3 best models</a>\n![RMSE on val for 3 best models][3]\n    \n\nTesting\n-------\n\n**During test** we made predictions up to 5 times for each model using different shifts of the tiles in the image. \nTest images were downscaled in 0.4-0.5 times.\n\nThe final ensemble was made by averaging all the predictions.   \n**Private LB RMSE:** *13.18968*  \n**Public LB RMSE:** *13.29065*\n\nApplying further postprocessing as suggested by @outrunner, could improve results.  \nJust increasing the number of pups by 20% gives a huge improvement:  \n**Private LB RMSE:**  *12.58131*  \n**Public LB RMSE:** *12.75510*\n\n\nSome negative experiments\n-------\n\nWe labeled some images from train set according to scale and trained a CNN to regress a scale of the image.   This could unify all the images to have the same approximate size of the lions of corresponding classes and simplify the CNN training to count animals.\n\nBut it didn't work out.   I reckon, the reason is the high variation in terrain and inability to estimate scale of objects if you look at them within a small spatial context (even with my own eyes).\n\n  [1]: https://www.kaggle.com/chelovekparohod\n  [2]: https://preview.ibb.co/bJ5Bv5/Screenshot_from_2017_06_28_16_28_57.png\n  [3]: https://preview.ibb.co/jx6O8Q/Screenshot_from_2017_06_28_17_02_17.png",
    "197021": "Thanks for your sharing~",
    "200697": "UPD: I have added a link to the source code on GitHub.",
    "203552": "Thanks a lot for sharing knowledge!",
    "204543": "Thanks for sharing! Will enjoy delving into the code.",
    "212189": "Artem,\nThanks a lot for sharing the code. I am still intrigued by your solution which seems to combine a segmentation and a regression in the same model.\nI have a question since I am not tf fluent:\n\nAs far as I understand, you take the train images, apply the gaussian on the train pictures (are these gaussian pictures a secondary output or the input in your model?) and you do a regression using a resnet like neural network on them.\n\nCan you re-explain please the part with the Gaussian kernels and how these gaussian kernels works at prediction time?\n\nThanks in advance for your help.",
    "212202": "Hey @eagle4. \nGaussians are used to estimate the ground truth number of lions on the tile. Gaussians are not used during inference. Putting a gaussian on top of each sea lion allows us to cut this lion with the tile border and still have a fraction of the lion  in each adjacent tile. Otherwise, we would have to assign this lion to the only one tile.",
    "212347": "Arthem,\nThanks a lot. It makes much more sense now. Congrats again on the solution."
  },
  "source": "meta"
}