{
  "id": 35465,
  "title": "My heatmap regression+segmentation solution and possible ways to normalize scale changes",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/writeups/lionheart-my-heatmap-regression-segmentation-solut",
  "author_name": "",
  "post_date": "2017-06-29T00:29:50.517Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Congratulations to the winners! \nI really enjoy the time I spend on  that competition, though for a while I almost <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/33788#190898\">gave up finishing it</a> .\nMy approach is kind of similar to some already released approaches, I really learned a lot by reading those posts, so I decide to share my approach here. In general it is a combination of gaussian heatmap regression and semantic segmentation. At the beginning I tried to regress 5 heatmaps for each individual class with Unet, later I felt it was too difficult tuning that task. So instead I regressed one single gaussian heatmap for all classes, so I can identify the location of each sea lion by detecting peaks without considering their classes, much alike the first step in some object detection methods, in which they propose a set of class agnostic bounding boxes before knowing their classes.\n<img src=\"http://i.imgur.com/7Lx31GB.jpg\" alt=\"heatmap\" title=\"\">\n<img src=\"http://i.imgur.com/p8VNf98.jpg\" alt=\"peak_detection\" title=\"\">\nTo classify the detected sea lions, an auxiliary semantic segmentation task is added to the same Unet and trained together that output a 6 class binary mask, the ground truth is generated by thresh-holding the gaussian heatmap. So after each sea lion was detected by peak detection, an argmax operation will be applied to that same location on the 6 class binary mask to decide which class it belongs to.\n<img src=\"http://i.imgur.com/0IJWu4B.png\" alt=\"6 class masks\" title=\"\">\nMy final submission was made by re-scaling all the test images to 0.6x original size, later I made some attempts try to generate more results with different re-scaling factors and ensemble them, I work with my team mate @hx364 on different scales, but it was still too slow.</p>\n\n<p>Meanwhile I made some experiments to see if I can estimate an optimal scaling factor for each testing image. The main idea is to find a proper scaling factor with template matching: I first hand picked about 400 template patches from training images of sea lions of class 2 and 3 that I thought has a \"good\" scale, then I rotate them to all 360 degrees and add to the template collection to cover different rotations. A VGG was used to extract features for all the templates, and I averaged all the features and divided by their covariances to make an \"exemplar sea lion feature\".\n<img src=\"http://i.imgur.com/1uKN4Mu.png\" alt=\"templates\" title=\"\">\nSince I've already detected some sea lions with the 0.6x method, I center croped all of the detections that belong to class 2,3 within in the same image, and re-scale each cropped patch with a series of scaling factors (1.0,  0.9, 0.8, 0.7, 0.6, 0.5, 0.4 )to make an image pyramid, and compute their mean cosine distance to the \"exemplar sea lion\" using VGG features, pick the lowest one and its corresponding scaling factor as the final results. \n<img src=\"http://i.imgur.com/JZ20JLd.png\" alt=\"estimated scales\" title=\"\">\nWith that approach I experimentally estimated the scaling factor for some of the test images, but I wasn't able to finish it before the deadline, so I am not sure if it really works but I am still very interested in exploring that method further.</p>",
  "messages": [
    {
      "id": "197177",
      "postDate": "06/28/2017 23:40:54",
      "content": "<p>Congratulations to the winners! \nI really enjoy the time I spend on  that competition, though for a while I almost <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/33788#190898\">gave up finishing it</a> .\nMy approach is kind of similar to some already released approaches, I really learned a lot by reading those posts, so I decide to share my approach here. In general it is a combination of gaussian heatmap regression and semantic segmentation. At the beginning I tried to regress 5 heatmaps for each individual class with Unet, later I felt it was too difficult tuning that task. So instead I regressed one single gaussian heatmap for all classes, so I can identify the location of each sea lion by detecting peaks without considering their classes, much alike the first step in some object detection methods, in which they propose a set of class agnostic bounding boxes before knowing their classes.\n<img src=\"http://i.imgur.com/7Lx31GB.jpg\" alt=\"heatmap\" title=\"\">\n<img src=\"http://i.imgur.com/p8VNf98.jpg\" alt=\"peak_detection\" title=\"\">\nTo classify the detected sea lions, an auxiliary semantic segmentation task is added to the same Unet and trained together that output a 6 class binary mask, the ground truth is generated by thresh-holding the gaussian heatmap. So after each sea lion was detected by peak detection, an argmax operation will be applied to that same location on the 6 class binary mask to decide which class it belongs to.\n<img src=\"http://i.imgur.com/0IJWu4B.png\" alt=\"6 class masks\" title=\"\">\nMy final submission was made by re-scaling all the test images to 0.6x original size, later I made some attempts try to generate more results with different re-scaling factors and ensemble them, I work with my team mate @hx364 on different scales, but it was still too slow.</p>\n\n<p>Meanwhile I made some experiments to see if I can estimate an optimal scaling factor for each testing image. The main idea is to find a proper scaling factor with template matching: I first hand picked about 400 template patches from training images of sea lions of class 2 and 3 that I thought has a \"good\" scale, then I rotate them to all 360 degrees and add to the template collection to cover different rotations. A VGG was used to extract features for all the templates, and I averaged all the features and divided by their covariances to make an \"exemplar sea lion feature\".\n<img src=\"http://i.imgur.com/1uKN4Mu.png\" alt=\"templates\" title=\"\">\nSince I've already detected some sea lions with the 0.6x method, I center croped all of the detections that belong to class 2,3 within in the same image, and re-scale each cropped patch with a series of scaling factors (1.0,  0.9, 0.8, 0.7, 0.6, 0.5, 0.4 )to make an image pyramid, and compute their mean cosine distance to the \"exemplar sea lion\" using VGG features, pick the lowest one and its corresponding scaling factor as the final results. \n<img src=\"http://i.imgur.com/JZ20JLd.png\" alt=\"estimated scales\" title=\"\">\nWith that approach I experimentally estimated the scaling factor for some of the test images, but I wasn't able to finish it before the deadline, so I am not sure if it really works but I am still very interested in exploring that method further.</p>",
      "rawMarkdown": "Congratulations to the winners! \nI really enjoy the time I spend on  that competition, though for a while I almost [gave up finishing it][1] .\nMy approach is kind of similar to some already released approaches, I really learned a lot by reading those posts, so I decide to share my approach here. In general it is a combination of gaussian heatmap regression and semantic segmentation. At the beginning I tried to regress 5 heatmaps for each individual class with Unet, later I felt it was too difficult tuning that task. So instead I regressed one single gaussian heatmap for all classes, so I can identify the location of each sea lion by detecting peaks without considering their classes, much alike the first step in some object detection methods, in which they propose a set of class agnostic bounding boxes before knowing their classes.\n![heatmap][2]\n![peak_detection][3]\nTo classify the detected sea lions, an auxiliary semantic segmentation task is added to the same Unet and trained together that output a 6 class binary mask, the ground truth is generated by thresh-holding the gaussian heatmap. So after each sea lion was detected by peak detection, an argmax operation will be applied to that same location on the 6 class binary mask to decide which class it belongs to.\n![6 class masks][4]\nMy final submission was made by re-scaling all the test images to 0.6x original size, later I made some attempts try to generate more results with different re-scaling factors and ensemble them, I work with my team mate @hx364 on different scales, but it was still too slow.\n\nMeanwhile I made some experiments to see if I can estimate an optimal scaling factor for each testing image. The main idea is to find a proper scaling factor with template matching: I first hand picked about 400 template patches from training images of sea lions of class 2 and 3 that I thought has a \"good\" scale, then I rotate them to all 360 degrees and add to the template collection to cover different rotations. A VGG was used to extract features for all the templates, and I averaged all the features and divided by their covariances to make an \"exemplar sea lion feature\".\n![templates][5]\nSince I've already detected some sea lions with the 0.6x method, I center croped all of the detections that belong to class 2,3 within in the same image, and re-scale each cropped patch with a series of scaling factors (1.0,  0.9, 0.8, 0.7, 0.6, 0.5, 0.4 )to make an image pyramid, and compute their mean cosine distance to the \"exemplar sea lion\" using VGG features, pick the lowest one and its corresponding scaling factor as the final results. \n![estimated scales][6]\nWith that approach I experimentally estimated the scaling factor for some of the test images, but I wasn't able to finish it before the deadline, so I am not sure if it really works but I am still very interested in exploring that method further.\n\n\n  [1]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/33788#190898\n  [2]: http://i.imgur.com/7Lx31GB.jpg\n  [3]: http://i.imgur.com/p8VNf98.jpg\n  [4]: http://i.imgur.com/0IJWu4B.png\n  [5]: http://i.imgur.com/1uKN4Mu.png\n  [6]: http://i.imgur.com/JZ20JLd.png",
      "votes": null
    },
    {
      "id": "197184",
      "postDate": "06/29/2017 00:51:14",
      "content": "<p>Nice work! Would you mind to sure your U-Net code? </p>",
      "rawMarkdown": "Nice work! Would you mind to sure your U-Net code?",
      "votes": null
    },
    {
      "id": "197187",
      "postDate": "06/29/2017 00:54:09",
      "content": "<p>Thank you, I will push my project to github this weekend</p>",
      "rawMarkdown": "Thank you, I will push my project to github this weekend",
      "votes": null
    },
    {
      "id": "197191",
      "postDate": "06/29/2017 01:13:23",
      "content": "<p>So basically you had 2 losses: L2 loss for density map regression and softmax loss on 6 classes for segmentation? <br>\nDid you weight these losses or simply summed them up?</p>",
      "rawMarkdown": "So basically you had 2 losses: L2 loss for density map regression and softmax loss on 6 classes for segmentation?   \nDid you weight these losses or simply summed them up?",
      "votes": null
    },
    {
      "id": "197195",
      "postDate": "06/29/2017 01:23:42",
      "content": "<p>Yes, I had 2 losses, the first one is L2 loss, the other is dice coefficient loss, I up weighted the dice coeff loss by a factor of 50, since I chose very large peak value in heatmap (500), so the L2 loss is relatively very large, while the dice coeff loss is always less than 1.0 . I doubt if  the overall results will be better I trained them separately, but when train them together, the heatmap do have a better looking than training it individually, maybe the multi-task learning  did help the Unet to converge towards better direction.</p>",
      "rawMarkdown": "Yes, I had 2 losses, the first one is L2 loss, the other is dice coefficient loss, I up weighted the dice coeff loss by a factor of 50, since I chose very large peak value in heatmap (500), so the L2 loss is relatively very large, while the dice coeff loss is always less than 1.0 . I doubt if  the overall results will be better I trained them separately, but when train them together, the heatmap do have a better looking than training it individually, maybe the multi-task learning  did help the Unet to converge towards better direction.",
      "votes": null
    },
    {
      "id": "209697",
      "postDate": "08/03/2017 04:27:07",
      "content": "<p>Do you happen to have a link to the github?</p>",
      "rawMarkdown": "Do you happen to have a link to the github?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 197184,
      "author_name": "",
      "author_url": "",
      "post_date": "06/29/2017 00:51:14",
      "content": "<p>Nice work! Would you mind to sure your U-Net code? </p>",
      "votes": null,
      "replies": [
        {
          "id": 197187,
          "author_name": "lzhang57",
          "author_url": "",
          "post_date": "06/29/2017 00:54:09",
          "content": "<p>Thank you, I will push my project to github this weekend</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209697,
          "author_name": "kblansit",
          "author_url": "",
          "post_date": "08/03/2017 04:27:07",
          "content": "<p>Do you happen to have a link to the github?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 197191,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "06/29/2017 01:13:23",
      "content": "<p>So basically you had 2 losses: L2 loss for density map regression and softmax loss on 6 classes for segmentation? <br>\nDid you weight these losses or simply summed them up?</p>",
      "votes": null,
      "replies": [
        {
          "id": 197195,
          "author_name": "lzhang57",
          "author_url": "",
          "post_date": "06/29/2017 01:23:42",
          "content": "<p>Yes, I had 2 losses, the first one is L2 loss, the other is dice coefficient loss, I up weighted the dice coeff loss by a factor of 50, since I chose very large peak value in heatmap (500), so the L2 loss is relatively very large, while the dice coeff loss is always less than 1.0 . I doubt if  the overall results will be better I trained them separately, but when train them together, the heatmap do have a better looking than training it individually, maybe the multi-task learning  did help the Unet to converge towards better direction.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "197177": "Congratulations to the winners! \nI really enjoy the time I spend on  that competition, though for a while I almost [gave up finishing it][1] .\nMy approach is kind of similar to some already released approaches, I really learned a lot by reading those posts, so I decide to share my approach here. In general it is a combination of gaussian heatmap regression and semantic segmentation. At the beginning I tried to regress 5 heatmaps for each individual class with Unet, later I felt it was too difficult tuning that task. So instead I regressed one single gaussian heatmap for all classes, so I can identify the location of each sea lion by detecting peaks without considering their classes, much alike the first step in some object detection methods, in which they propose a set of class agnostic bounding boxes before knowing their classes.\n![heatmap][2]\n![peak_detection][3]\nTo classify the detected sea lions, an auxiliary semantic segmentation task is added to the same Unet and trained together that output a 6 class binary mask, the ground truth is generated by thresh-holding the gaussian heatmap. So after each sea lion was detected by peak detection, an argmax operation will be applied to that same location on the 6 class binary mask to decide which class it belongs to.\n![6 class masks][4]\nMy final submission was made by re-scaling all the test images to 0.6x original size, later I made some attempts try to generate more results with different re-scaling factors and ensemble them, I work with my team mate @hx364 on different scales, but it was still too slow.\n\nMeanwhile I made some experiments to see if I can estimate an optimal scaling factor for each testing image. The main idea is to find a proper scaling factor with template matching: I first hand picked about 400 template patches from training images of sea lions of class 2 and 3 that I thought has a \"good\" scale, then I rotate them to all 360 degrees and add to the template collection to cover different rotations. A VGG was used to extract features for all the templates, and I averaged all the features and divided by their covariances to make an \"exemplar sea lion feature\".\n![templates][5]\nSince I've already detected some sea lions with the 0.6x method, I center croped all of the detections that belong to class 2,3 within in the same image, and re-scale each cropped patch with a series of scaling factors (1.0,  0.9, 0.8, 0.7, 0.6, 0.5, 0.4 )to make an image pyramid, and compute their mean cosine distance to the \"exemplar sea lion\" using VGG features, pick the lowest one and its corresponding scaling factor as the final results. \n![estimated scales][6]\nWith that approach I experimentally estimated the scaling factor for some of the test images, but I wasn't able to finish it before the deadline, so I am not sure if it really works but I am still very interested in exploring that method further.\n\n\n  [1]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/33788#190898\n  [2]: http://i.imgur.com/7Lx31GB.jpg\n  [3]: http://i.imgur.com/p8VNf98.jpg\n  [4]: http://i.imgur.com/0IJWu4B.png\n  [5]: http://i.imgur.com/1uKN4Mu.png\n  [6]: http://i.imgur.com/JZ20JLd.png",
    "197184": "Nice work! Would you mind to sure your U-Net code?",
    "197187": "Thank you, I will push my project to github this weekend",
    "197191": "So basically you had 2 losses: L2 loss for density map regression and softmax loss on 6 classes for segmentation?   \nDid you weight these losses or simply summed them up?",
    "197195": "Yes, I had 2 losses, the first one is L2 loss, the other is dice coefficient loss, I up weighted the dice coeff loss by a factor of 50, since I chose very large peak value in heatmap (500), so the L2 loss is relatively very large, while the dice coeff loss is always less than 1.0 . I doubt if  the overall results will be better I trained them separately, but when train them together, the heatmap do have a better looking than training it individually, maybe the multi-task learning  did help the Unet to converge towards better direction.",
    "209697": "Do you happen to have a link to the github?"
  },
  "source": "meta"
}