{
  "id": 35491,
  "title": "R&D&G team solution, 5th place",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/writeups/r-d-g-r-d-g-team-solution-5th-place",
  "author_name": "",
  "post_date": "2017-06-29T13:40:41.037409400Z",
  "votes": 19,
  "comment_count": 7,
  "views": 1,
  "content": "<p>Congrats to outrunner, Konstantin and bestfitting!</p>\n\n<p>A short overview of our team solution:</p>\n\n<p>We estimated image size, implemented u-net line network and with sum of heatmaps as resulting count.</p>\n\n<p>Our image scale estimation solution:</p>\n\n<p>Firstly I have labelled the scale of the training dataset images.\nTo do this, I have done a simple application to display all sea lions from the image next to lions from the reference image, and adjusted scale of lions until it visually matches for all classes:\n<a href=\"https://ibb.co/mWg5oQ\">https://ibb.co/mWg5oQ</a></p>\n\n<p>Instead I trained the SSD network to detect sea lion rects using the average sea lion size for known  image scale and used predicted sea lion rects as input to xgboost network, predicting actual image scale.</p>\n\n<p>So training pipeline looks like:</p>\n\n<ol>\n<li><p>Label image scale for train dataset</p></li>\n<li><p>Extra scale augmentation, prepare set of sea lion squares</p></li>\n<li><p>Train SSD, predict sea lion squares</p></li>\n<li><p>Combine results of SSD, like mean of predicted crops mean size, size of different percentiles of rects sorted by detection confidence, the same per class, etc and trained XGBoost to predict image scale</p></li>\n</ol>\n\n<p>Prediction pipeline:</p>\n\n<ol>\n<li><p>Downscale image 2x as sea lions at some images are over 3x larger than on the train dataset</p></li>\n<li><p>Predict sea lion rects with SSD and image scale with XGBoost</p></li>\n<li><p>Scale test to scale predicted with step 2, it should be the better match than original 0.5x</p></li>\n<li><p>Predict scale with SSD and XGBoost again</p></li>\n</ol>\n\n<p>Applying estimated scale allowed to improve score from ~23 to 15.9.</p>\n\n<p>Distribution of image scale of the train and test datasets (on this plot - how much image has to be rescaled to match train/0.jpg):</p>\n\n<p><a href=\"https://ibb.co/i3X38Q\">https://ibb.co/i3X38Q</a></p>\n\n<p>For some reason the test dataset is significantly different from the train dataset, image scale is approx 2x larger.</p>\n\n<p>To count sea lions I combined the imagenet pretrained networks like Resnet and VGG with decoder part of U-net.</p>\n\n<p>I used pretrained and fine tuned on sea lion crops Resnet50 network,\nattached 3 stages of U-net like decoder to network output and internal layers.</p>\n\n<p>The resulting resolution was 4x lower than resolution of original image.\nI tried to use original U-Net or added more stages to increase resolution, but results were worse. Adding of residual blocks to U-Net decoder allowed to improve the results.</p>\n\n<p>I marked each lion as 3x3 square regardless of class and predicted heatmap with softmax loss for 5 sea lion classes + 1 background class. </p>\n\n<p>My idea was - if logloss/softmax has minimum when output is equal to probability of particular  class, the count of sea lions should be equal to sum of per class heatmap / area of marker used for training. It worked reasonably well, on images with wrong scale the count shifted between classes as network learned lion sizes but the total count matched.</p>\n\n<p>examples of heatmaps, number is predicted count, as sum(heatmap)/(3*3)</p>\n\n<p><a href=\"https://ibb.co/ndUmv5\">https://ibb.co/ndUmv5</a></p>\n\n<p><a href=\"https://ibb.co/jg6ANk\">https://ibb.co/jg6ANk</a></p>\n\n<p>For the final submission, Gilberto has trained a few 2nd level models we averaged together with predictions of Resnet50 and VGG16 based models.</p>",
  "messages": [
    {
      "id": "197382",
      "postDate": "06/29/2017 13:40:41",
      "content": "<p>Congrats to outrunner, Konstantin and bestfitting!</p>\n\n<p>A short overview of our team solution:</p>\n\n<p>We estimated image size, implemented u-net line network and with sum of heatmaps as resulting count.</p>\n\n<p>Our image scale estimation solution:</p>\n\n<p>Firstly I have labelled the scale of the training dataset images.\nTo do this, I have done a simple application to display all sea lions from the image next to lions from the reference image, and adjusted scale of lions until it visually matches for all classes:\n<a href=\"https://ibb.co/mWg5oQ\">https://ibb.co/mWg5oQ</a></p>\n\n<p>Instead I trained the SSD network to detect sea lion rects using the average sea lion size for known  image scale and used predicted sea lion rects as input to xgboost network, predicting actual image scale.</p>\n\n<p>So training pipeline looks like:</p>\n\n<ol>\n<li><p>Label image scale for train dataset</p></li>\n<li><p>Extra scale augmentation, prepare set of sea lion squares</p></li>\n<li><p>Train SSD, predict sea lion squares</p></li>\n<li><p>Combine results of SSD, like mean of predicted crops mean size, size of different percentiles of rects sorted by detection confidence, the same per class, etc and trained XGBoost to predict image scale</p></li>\n</ol>\n\n<p>Prediction pipeline:</p>\n\n<ol>\n<li><p>Downscale image 2x as sea lions at some images are over 3x larger than on the train dataset</p></li>\n<li><p>Predict sea lion rects with SSD and image scale with XGBoost</p></li>\n<li><p>Scale test to scale predicted with step 2, it should be the better match than original 0.5x</p></li>\n<li><p>Predict scale with SSD and XGBoost again</p></li>\n</ol>\n\n<p>Applying estimated scale allowed to improve score from ~23 to 15.9.</p>\n\n<p>Distribution of image scale of the train and test datasets (on this plot - how much image has to be rescaled to match train/0.jpg):</p>\n\n<p><a href=\"https://ibb.co/i3X38Q\">https://ibb.co/i3X38Q</a></p>\n\n<p>For some reason the test dataset is significantly different from the train dataset, image scale is approx 2x larger.</p>\n\n<p>To count sea lions I combined the imagenet pretrained networks like Resnet and VGG with decoder part of U-net.</p>\n\n<p>I used pretrained and fine tuned on sea lion crops Resnet50 network,\nattached 3 stages of U-net like decoder to network output and internal layers.</p>\n\n<p>The resulting resolution was 4x lower than resolution of original image.\nI tried to use original U-Net or added more stages to increase resolution, but results were worse. Adding of residual blocks to U-Net decoder allowed to improve the results.</p>\n\n<p>I marked each lion as 3x3 square regardless of class and predicted heatmap with softmax loss for 5 sea lion classes + 1 background class. </p>\n\n<p>My idea was - if logloss/softmax has minimum when output is equal to probability of particular  class, the count of sea lions should be equal to sum of per class heatmap / area of marker used for training. It worked reasonably well, on images with wrong scale the count shifted between classes as network learned lion sizes but the total count matched.</p>\n\n<p>examples of heatmaps, number is predicted count, as sum(heatmap)/(3*3)</p>\n\n<p><a href=\"https://ibb.co/ndUmv5\">https://ibb.co/ndUmv5</a></p>\n\n<p><a href=\"https://ibb.co/jg6ANk\">https://ibb.co/jg6ANk</a></p>\n\n<p>For the final submission, Gilberto has trained a few 2nd level models we averaged together with predictions of Resnet50 and VGG16 based models.</p>",
      "rawMarkdown": "Congrats to outrunner, Konstantin and bestfitting!\n\nA short overview of our team solution:\n\nWe estimated image size, implemented u-net line network and with sum of heatmaps as resulting count.\n\nOur image scale estimation solution:\n\nFirstly I have labelled the scale of the training dataset images.\nTo do this, I have done a simple application to display all sea lions from the image next to lions from the reference image, and adjusted scale of lions until it visually matches for all classes:\n<img>https://ibb.co/mWg5oQ\n\nInstead I trained the SSD network to detect sea lion rects using the average sea lion size for known  image scale and used predicted sea lion rects as input to xgboost network, predicting actual image scale.\n\nSo training pipeline looks like:\n\n1. Label image scale for train dataset\n\n2. Extra scale augmentation, prepare set of sea lion squares\n\n3. Train SSD, predict sea lion squares\n\n4. Combine results of SSD, like mean of predicted crops mean size, size of different percentiles of rects sorted by detection confidence, the same per class, etc and trained XGBoost to predict image scale\n\nPrediction pipeline:\n\n1. Downscale image 2x as sea lions at some images are over 3x larger than on the train dataset\n\n2. Predict sea lion rects with SSD and image scale with XGBoost\n\n3. Scale test to scale predicted with step 2, it should be the better match than original 0.5x\n\n4. Predict scale with SSD and XGBoost again\n\nApplying estimated scale allowed to improve score from ~23 to 15.9.\n\nDistribution of image scale of the train and test datasets (on this plot - how much image has to be rescaled to match train/0.jpg):\n\nhttps://ibb.co/i3X38Q\n\nFor some reason the test dataset is significantly different from the train dataset, image scale is approx 2x larger.\n\n\n\nTo count sea lions I combined the imagenet pretrained networks like Resnet and VGG with decoder part of U-net.\n\nI used pretrained and fine tuned on sea lion crops Resnet50 network,\nattached 3 stages of U-net like decoder to network output and internal layers.\n\nThe resulting resolution was 4x lower than resolution of original image.\nI tried to use original U-Net or added more stages to increase resolution, but results were worse. Adding of residual blocks to U-Net decoder allowed to improve the results.\n\nI marked each lion as 3x3 square regardless of class and predicted heatmap with softmax loss for 5 sea lion classes + 1 background class. \n\nMy idea was - if logloss/softmax has minimum when output is equal to probability of particular  class, the count of sea lions should be equal to sum of per class heatmap / area of marker used for training. It worked reasonably well, on images with wrong scale the count shifted between classes as network learned lion sizes but the total count matched.\n\nexamples of heatmaps, number is predicted count, as sum(heatmap)/(3*3)\n\nhttps://ibb.co/ndUmv5\n\nhttps://ibb.co/jg6ANk\n\nFor the final submission, Gilberto has trained a few 2nd level models we averaged together with predictions of Resnet50 and VGG16 based models.",
      "votes": null
    },
    {
      "id": "197384",
      "postDate": "06/29/2017 13:48:23",
      "content": "<p>Amazing and impressive work !</p>",
      "rawMarkdown": "Amazing and impressive work !",
      "votes": null
    },
    {
      "id": "197390",
      "postDate": "06/29/2017 13:58:21",
      "content": "<p>What I tried and did not work:</p>\n\n<p>I tried to build a conv net to predict scale from patches, downscaled image or even image after cosine transform, but nothing worked reliably, partially due to different terrain on the test images.</p>\n\n<p>I tried to build the second level model to predict count from heatmaps, the results looked reasonable and improved the local cross validation score, but the leaderboard result was always worse. One possible explanation would be difference between training and test datasets.</p>\n\n<p>Applying the same postprocessing as used by outrunner has improved the public/private score to 13.1/12.2</p>",
      "rawMarkdown": "What I tried and did not work:\n\nI tried to build a conv net to predict scale from patches, downscaled image or even image after cosine transform, but nothing worked reliably, partially due to different terrain on the test images.\n\nI tried to build the second level model to predict count from heatmaps, the results looked reasonable and improved the local cross validation score, but the leaderboard result was always worse. One possible explanation would be difference between training and test datasets.\n\nApplying the same postprocessing as used by outrunner has improved the public/private score to 13.1/12.2",
      "votes": null
    },
    {
      "id": "197533",
      "postDate": "06/29/2017 19:00:38",
      "content": "<p>Nice approach for scale estimation, thanks!</p>",
      "rawMarkdown": "Nice approach for scale estimation, thanks!",
      "votes": null
    },
    {
      "id": "197572",
      "postDate": "06/29/2017 20:56:23",
      "content": "<p>Nice work Dmytro and team</p>",
      "rawMarkdown": "Nice work Dmytro and team",
      "votes": null
    },
    {
      "id": "197598",
      "postDate": "06/29/2017 23:27:01",
      "content": "<p>Nice work on scale estimation, thanks.</p>",
      "rawMarkdown": "Nice work on scale estimation, thanks.",
      "votes": null
    },
    {
      "id": "197635",
      "postDate": "06/30/2017 00:59:32",
      "content": "<p>Great job!I also tried to estimate the scale of the test image but I did not finish it until last minute.<br>I did not notice the scale variance of the train set images and build a stable model from train-set.Instead,I tried to get average area of the sealions and compare them with predicted test set images.<br>I also think,perhaps I can predict scale of the test-set images  if I do scale argument on train set images to simuate the test-set images' scale and build a model  predict the scale of test-set images since I can get the area of each sealions.</p>",
      "rawMarkdown": "Great job!I also tried to estimate the scale of the test image but I did not finish it until last minute.<br>I did not notice the scale variance of the train set images and build a stable model from train-set.Instead,I tried to get average area of the sealions and compare them with predicted test set images.<br>I also think,perhaps I can predict scale of the test-set images  if I do scale argument on train set images to simuate the test-set images' scale and build a model  predict the scale of test-set images since I can get the area of each sealions.",
      "votes": null
    },
    {
      "id": "198059",
      "postDate": "06/30/2017 21:26:56",
      "content": "<p>I think for scale estimation your sea lions area would be even more reliable data source than area of rects predicted by SSD, especially when you predict sea lions class as well in addition to area and do a few iterations of initial scale -&gt; geneate segmentation -&gt; estimate scale -&gt; scale images by estimated scale -&gt; generate heatmaps -&gt; improved scale estimation.</p>\n\n<p>Multiple iterations help to improve quality of classification as after each iteration the size of sea lions is closer to size used during training.</p>",
      "rawMarkdown": "I think for scale estimation your sea lions area would be even more reliable data source than area of rects predicted by SSD, especially when you predict sea lions class as well in addition to area and do a few iterations of initial scale -&gt; geneate segmentation -&gt; estimate scale -&gt; scale images by estimated scale -&gt; generate heatmaps -&gt; improved scale estimation.\n\nMultiple iterations help to improve quality of classification as after each iteration the size of sea lions is closer to size used during training.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 197384,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "06/29/2017 13:48:23",
      "content": "<p>Amazing and impressive work !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 197390,
      "author_name": "dmytropoplavskiy",
      "author_url": "",
      "post_date": "06/29/2017 13:58:21",
      "content": "<p>What I tried and did not work:</p>\n\n<p>I tried to build a conv net to predict scale from patches, downscaled image or even image after cosine transform, but nothing worked reliably, partially due to different terrain on the test images.</p>\n\n<p>I tried to build the second level model to predict count from heatmaps, the results looked reasonable and improved the local cross validation score, but the leaderboard result was always worse. One possible explanation would be difference between training and test datasets.</p>\n\n<p>Applying the same postprocessing as used by outrunner has improved the public/private score to 13.1/12.2</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 197533,
      "author_name": "harshml",
      "author_url": "",
      "post_date": "06/29/2017 19:00:38",
      "content": "<p>Nice approach for scale estimation, thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 197572,
      "author_name": "darraghdog",
      "author_url": "",
      "post_date": "06/29/2017 20:56:23",
      "content": "<p>Nice work Dmytro and team</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 197598,
      "author_name": "outrunner",
      "author_url": "",
      "post_date": "06/29/2017 23:27:01",
      "content": "<p>Nice work on scale estimation, thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 197635,
      "author_name": "bestfitting",
      "author_url": "",
      "post_date": "06/30/2017 00:59:32",
      "content": "<p>Great job!I also tried to estimate the scale of the test image but I did not finish it until last minute.<br>I did not notice the scale variance of the train set images and build a stable model from train-set.Instead,I tried to get average area of the sealions and compare them with predicted test set images.<br>I also think,perhaps I can predict scale of the test-set images  if I do scale argument on train set images to simuate the test-set images' scale and build a model  predict the scale of test-set images since I can get the area of each sealions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 198059,
          "author_name": "dmytropoplavskiy",
          "author_url": "",
          "post_date": "06/30/2017 21:26:56",
          "content": "<p>I think for scale estimation your sea lions area would be even more reliable data source than area of rects predicted by SSD, especially when you predict sea lions class as well in addition to area and do a few iterations of initial scale -&gt; geneate segmentation -&gt; estimate scale -&gt; scale images by estimated scale -&gt; generate heatmaps -&gt; improved scale estimation.</p>\n\n<p>Multiple iterations help to improve quality of classification as after each iteration the size of sea lions is closer to size used during training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "197382": "Congrats to outrunner, Konstantin and bestfitting!\n\nA short overview of our team solution:\n\nWe estimated image size, implemented u-net line network and with sum of heatmaps as resulting count.\n\nOur image scale estimation solution:\n\nFirstly I have labelled the scale of the training dataset images.\nTo do this, I have done a simple application to display all sea lions from the image next to lions from the reference image, and adjusted scale of lions until it visually matches for all classes:\n<img>https://ibb.co/mWg5oQ\n\nInstead I trained the SSD network to detect sea lion rects using the average sea lion size for known  image scale and used predicted sea lion rects as input to xgboost network, predicting actual image scale.\n\nSo training pipeline looks like:\n\n1. Label image scale for train dataset\n\n2. Extra scale augmentation, prepare set of sea lion squares\n\n3. Train SSD, predict sea lion squares\n\n4. Combine results of SSD, like mean of predicted crops mean size, size of different percentiles of rects sorted by detection confidence, the same per class, etc and trained XGBoost to predict image scale\n\nPrediction pipeline:\n\n1. Downscale image 2x as sea lions at some images are over 3x larger than on the train dataset\n\n2. Predict sea lion rects with SSD and image scale with XGBoost\n\n3. Scale test to scale predicted with step 2, it should be the better match than original 0.5x\n\n4. Predict scale with SSD and XGBoost again\n\nApplying estimated scale allowed to improve score from ~23 to 15.9.\n\nDistribution of image scale of the train and test datasets (on this plot - how much image has to be rescaled to match train/0.jpg):\n\nhttps://ibb.co/i3X38Q\n\nFor some reason the test dataset is significantly different from the train dataset, image scale is approx 2x larger.\n\n\n\nTo count sea lions I combined the imagenet pretrained networks like Resnet and VGG with decoder part of U-net.\n\nI used pretrained and fine tuned on sea lion crops Resnet50 network,\nattached 3 stages of U-net like decoder to network output and internal layers.\n\nThe resulting resolution was 4x lower than resolution of original image.\nI tried to use original U-Net or added more stages to increase resolution, but results were worse. Adding of residual blocks to U-Net decoder allowed to improve the results.\n\nI marked each lion as 3x3 square regardless of class and predicted heatmap with softmax loss for 5 sea lion classes + 1 background class. \n\nMy idea was - if logloss/softmax has minimum when output is equal to probability of particular  class, the count of sea lions should be equal to sum of per class heatmap / area of marker used for training. It worked reasonably well, on images with wrong scale the count shifted between classes as network learned lion sizes but the total count matched.\n\nexamples of heatmaps, number is predicted count, as sum(heatmap)/(3*3)\n\nhttps://ibb.co/ndUmv5\n\nhttps://ibb.co/jg6ANk\n\nFor the final submission, Gilberto has trained a few 2nd level models we averaged together with predictions of Resnet50 and VGG16 based models.",
    "197384": "Amazing and impressive work !",
    "197390": "What I tried and did not work:\n\nI tried to build a conv net to predict scale from patches, downscaled image or even image after cosine transform, but nothing worked reliably, partially due to different terrain on the test images.\n\nI tried to build the second level model to predict count from heatmaps, the results looked reasonable and improved the local cross validation score, but the leaderboard result was always worse. One possible explanation would be difference between training and test datasets.\n\nApplying the same postprocessing as used by outrunner has improved the public/private score to 13.1/12.2",
    "197533": "Nice approach for scale estimation, thanks!",
    "197572": "Nice work Dmytro and team",
    "197598": "Nice work on scale estimation, thanks.",
    "197635": "Great job!I also tried to estimate the scale of the test image but I did not finish it until last minute.<br>I did not notice the scale variance of the train set images and build a stable model from train-set.Instead,I tried to get average area of the sealions and compare them with predicted test set images.<br>I also think,perhaps I can predict scale of the test-set images  if I do scale argument on train set images to simuate the test-set images' scale and build a model  predict the scale of test-set images since I can get the area of each sealions.",
    "198059": "I think for scale estimation your sea lions area would be even more reliable data source than area of rects predicted by SSD, especially when you predict sea lions class as well in addition to area and do a few iterations of initial scale -&gt; geneate segmentation -&gt; estimate scale -&gt; scale images by estimated scale -&gt; generate heatmaps -&gt; improved scale estimation.\n\nMultiple iterations help to improve quality of classification as after each iteration the size of sea lions is closer to size used during training."
  },
  "source": "meta"
}