{
  "id": 35462,
  "title": "AdaBoost, SSD: object detection for lions counting",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/writeups/harshml-adaboost-ssd-object-detection-for-lions-co",
  "author_name": "",
  "post_date": "2017-06-29T06:55:29.437Z",
  "votes": 23,
  "comment_count": 14,
  "views": 1,
  "content": "<p>We made a bet with my friend, how far we can go in this challenge (and two month of life have blown away). The problem looks like a type of crowd counting task. Recent solutions prepare density maps for dotted crowd annotations, train network to predict them, and then sum the density in each pixel to obtain the final number. However, I've decided to use object detection approaches, and my friend started with segmentation...</p>\n\n<hr>\n\n<p>My path:</p>\n\n<p>1) Classic ml - AdaBoost</p>\n\n<p>2) SSD on full images (since it is faster than faster rcnn)</p>\n\n<p>3) SSD on tiles</p>\n\n<p>For object detection boxes is needed, so decided to use squares around the dots coordinates (size was selected manually, OpenCV <code>inRange</code> did the work with dots coordinates extraction).</p>\n\n<p>As for now DL is a mainstream with state of art results, but around 3 years ago, <em>the best off-the-shelf classifier</em> <strong>AdaBoost</strong> crushes the leaderboards, so cannot skip to evaluate it. One of the fastest and successful AdaBoost pipelines is boosted decision forest with aggregated channel features (thx to Piotr Dollar for this milestone in object detection). Thus trained 3 models: adult, subadult males and pups. This gives ~20 on public LB <em>(it took around of an hour to evaluate one model on all test set, on CPU)</em>. The recall is almost 100%, however, due to lions rotations, models learned to find circle-shaped objects, so there were a lot of false positives on stones.</p>\n\n<p>DL's time. Single Shot Multibox Detector was selected for the speed reasons and good accuracy. Scaled the images to fit the GPU memory (~1200x800) and run with VGG backbone out of the box. DL features give ~18 on LB, mostly due to better (than mean values) work on females and juveniles, testing time is also ~1 hour, but on GPU. Surprisingly, AdaBoost with ACF did well on adult, subadult males, the difference was ~0.1-0.3 compared to SSD. The issues:</p>\n\n<ul>\n<li><p>pups detected badly</p></li>\n<li><p>can't distinguish between females and juveniles</p></li>\n</ul>\n\n<p>What tried to beat this (without any success):</p>\n\n<ul>\n<li><p>Better features - construct a hyper feature from conv3_3 &amp; conv4_3 &amp; conv5_3 to integrate the details from shallow layer and context from deeper.</p></li>\n<li><p>Learn the image scale - average pool of feature maps + fc128 + fc16 and concatenate these 16 learned channels with feature maps for final prediction. So, I also did nothing with the scale problem.</p></li>\n<li><p>Tried multitask loss - added regression output to SSD classification and localization. Regression produced more or less reasonable values for adult, subadult males, but for classes with huge deviation in number of lions per image obtained <code>nan</code>.  Looks like need to spend more time here =), and use small tiles instead of full images.</p></li>\n</ul>\n\n<p>So, gave up with sophisticated methods, tile the train data. Used huge tiles 1000x1000 for speed reasons and to avoid border effects. It helped to detect pups better, gives ~17 on LB (during training and testing images were resized to 3000 by the shortest side, then tiles extracted, did flip augmentation), runs ~8 hours on GPU:\n<img src=\"http://i67.tinypic.com/juhicw.png\" alt=\"Picture 1\" title=\"\">\n<img src=\"http://i63.tinypic.com/nv76ns.png\" alt=\"Picture 2\" title=\"\">\nThe rest 1.5 point is stacking models (spent ~100 submissions while doing this).</p>\n\n<p>Thanks for everyone who did involve in this competition. It was fun!</p>",
  "messages": [
    {
      "id": "197138",
      "postDate": "06/28/2017 22:04:04",
      "content": "<p>We made a bet with my friend, how far we can go in this challenge (and two month of life have blown away). The problem looks like a type of crowd counting task. Recent solutions prepare density maps for dotted crowd annotations, train network to predict them, and then sum the density in each pixel to obtain the final number. However, I've decided to use object detection approaches, and my friend started with segmentation...</p>\n\n<hr>\n\n<p>My path:</p>\n\n<p>1) Classic ml - AdaBoost</p>\n\n<p>2) SSD on full images (since it is faster than faster rcnn)</p>\n\n<p>3) SSD on tiles</p>\n\n<p>For object detection boxes is needed, so decided to use squares around the dots coordinates (size was selected manually, OpenCV <code>inRange</code> did the work with dots coordinates extraction).</p>\n\n<p>As for now DL is a mainstream with state of art results, but around 3 years ago, <em>the best off-the-shelf classifier</em> <strong>AdaBoost</strong> crushes the leaderboards, so cannot skip to evaluate it. One of the fastest and successful AdaBoost pipelines is boosted decision forest with aggregated channel features (thx to Piotr Dollar for this milestone in object detection). Thus trained 3 models: adult, subadult males and pups. This gives ~20 on public LB <em>(it took around of an hour to evaluate one model on all test set, on CPU)</em>. The recall is almost 100%, however, due to lions rotations, models learned to find circle-shaped objects, so there were a lot of false positives on stones.</p>\n\n<p>DL's time. Single Shot Multibox Detector was selected for the speed reasons and good accuracy. Scaled the images to fit the GPU memory (~1200x800) and run with VGG backbone out of the box. DL features give ~18 on LB, mostly due to better (than mean values) work on females and juveniles, testing time is also ~1 hour, but on GPU. Surprisingly, AdaBoost with ACF did well on adult, subadult males, the difference was ~0.1-0.3 compared to SSD. The issues:</p>\n\n<ul>\n<li><p>pups detected badly</p></li>\n<li><p>can't distinguish between females and juveniles</p></li>\n</ul>\n\n<p>What tried to beat this (without any success):</p>\n\n<ul>\n<li><p>Better features - construct a hyper feature from conv3_3 &amp; conv4_3 &amp; conv5_3 to integrate the details from shallow layer and context from deeper.</p></li>\n<li><p>Learn the image scale - average pool of feature maps + fc128 + fc16 and concatenate these 16 learned channels with feature maps for final prediction. So, I also did nothing with the scale problem.</p></li>\n<li><p>Tried multitask loss - added regression output to SSD classification and localization. Regression produced more or less reasonable values for adult, subadult males, but for classes with huge deviation in number of lions per image obtained <code>nan</code>.  Looks like need to spend more time here =), and use small tiles instead of full images.</p></li>\n</ul>\n\n<p>So, gave up with sophisticated methods, tile the train data. Used huge tiles 1000x1000 for speed reasons and to avoid border effects. It helped to detect pups better, gives ~17 on LB (during training and testing images were resized to 3000 by the shortest side, then tiles extracted, did flip augmentation), runs ~8 hours on GPU:\n<img src=\"http://i67.tinypic.com/juhicw.png\" alt=\"Picture 1\" title=\"\">\n<img src=\"http://i63.tinypic.com/nv76ns.png\" alt=\"Picture 2\" title=\"\">\nThe rest 1.5 point is stacking models (spent ~100 submissions while doing this).</p>\n\n<p>Thanks for everyone who did involve in this competition. It was fun!</p>",
      "rawMarkdown": "We made a bet with my friend, how far we can go in this challenge (and two month of life have blown away). The problem looks like a type of crowd counting task. Recent solutions prepare density maps for dotted crowd annotations, train network to predict them, and then sum the density in each pixel to obtain the final number. However, I've decided to use object detection approaches, and my friend started with segmentation...\n___\nMy path:\n\n1) Classic ml - AdaBoost\n\n2) SSD on full images (since it is faster than faster rcnn)\n\n3) SSD on tiles\n\nFor object detection boxes is needed, so decided to use squares around the dots coordinates (size was selected manually, OpenCV `inRange` did the work with dots coordinates extraction).\n\nAs for now DL is a mainstream with state of art results, but around 3 years ago, *the best off-the-shelf classifier* **AdaBoost** crushes the leaderboards, so cannot skip to evaluate it. One of the fastest and successful AdaBoost pipelines is boosted decision forest with aggregated channel features (thx to Piotr Dollar for this milestone in object detection). Thus trained 3 models: adult, subadult males and pups. This gives ~20 on public LB *(it took around of an hour to evaluate one model on all test set, on CPU)*. The recall is almost 100%, however, due to lions rotations, models learned to find circle-shaped objects, so there were a lot of false positives on stones.\n\nDL's time. Single Shot Multibox Detector was selected for the speed reasons and good accuracy. Scaled the images to fit the GPU memory (~1200x800) and run with VGG backbone out of the box. DL features give ~18 on LB, mostly due to better (than mean values) work on females and juveniles, testing time is also ~1 hour, but on GPU. Surprisingly, AdaBoost with ACF did well on adult, subadult males, the difference was ~0.1-0.3 compared to SSD. The issues:\n\n* pups detected badly\n\n* can't distinguish between females and juveniles\n\nWhat tried to beat this (without any success):\n\n* Better features - construct a hyper feature from conv3_3 &amp; conv4_3 &amp; conv5_3 to integrate the details from shallow layer and context from deeper.\n\n* Learn the image scale - average pool of feature maps + fc128 + fc16 and concatenate these 16 learned channels with feature maps for final prediction. So, I also did nothing with the scale problem.\n\n* Tried multitask loss - added regression output to SSD classification and localization. Regression produced more or less reasonable values for adult, subadult males, but for classes with huge deviation in number of lions per image obtained `nan`.  Looks like need to spend more time here =), and use small tiles instead of full images.\n\nSo, gave up with sophisticated methods, tile the train data. Used huge tiles 1000x1000 for speed reasons and to avoid border effects. It helped to detect pups better, gives ~17 on LB (during training and testing images were resized to 3000 by the shortest side, then tiles extracted, did flip augmentation), runs ~8 hours on GPU:\n![Picture 1][1]\n![Picture 2][2]\nThe rest 1.5 point is stacking models (spent ~100 submissions while doing this).\n\nThanks for everyone who did involve in this competition. It was fun!\n\n  [1]: http://i67.tinypic.com/juhicw.png\n  [2]: http://i63.tinypic.com/nv76ns.png",
      "votes": null
    },
    {
      "id": "197152",
      "postDate": "06/28/2017 22:35:00",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "197160",
      "postDate": "06/28/2017 22:51:20",
      "content": "<ol>\n<li>Did you make any tricks to distinguish between so similar classes? i.e. something like 'finegrained classification'?</li>\n<li>How SSD performs with small objects and highly overlapped objects? did you do some tuning on this?</li>\n</ol>",
      "rawMarkdown": "1. Did you make any tricks to distinguish between so similar classes? i.e. something like 'finegrained classification'?\n2. How SSD performs with small objects and highly overlapped objects? did you do some tuning on this?",
      "votes": null
    },
    {
      "id": "197194",
      "postDate": "06/29/2017 01:20:50",
      "content": "<p>I was wondering how good could object detection do. Your result is beyond my imagination. Thanks for sharing, good job.</p>",
      "rawMarkdown": "I was wondering how good could object detection do. Your result is beyond my imagination. Thanks for sharing, good job.",
      "votes": null
    },
    {
      "id": "197263",
      "postDate": "06/29/2017 06:51:47",
      "content": "<p>I can't distinguish with my eyes between those similar classes, so switched to DL for better features, but found there similar problems. So I haven't tried one-vs-all or other tricks. My guess that cross-entropy loss doesn't good for this case, since it forces the rest classes to have near 0 score, and network cannot learn well if classes have similar features. I've tried hinge loss to simplify the training, but with no success =), probably due to wrong integration in SSD pipeline.</p>\n\n<p>Small objects is a pain for all OD. Usually hyper feature should help, but it didn't, or more sophisticated classifier (didn't help), also tried to learn predictors on different layers - conv3 - conv5, but results were similar. So looks like it's not a feature representation problem (it's strong enough), check the pups:\n<img src=\"http://i65.tinypic.com/2646had.jpg\" alt=\"pups\" title=\"\">\nThus SSD here performs as everything else (faster rcnn), but faster.</p>\n\n<p>From my experience highly overlapped objects here is one of the main issues with OD approach, they merges during nms and creates troubles during matching. Tried to add 6'th class: female + near pup, but without success. I played a lot with boxes sizes and nms thresholds, last submissions have only 2 box sizes during training: huge for males, smaller for the rest 3 classes to mach appropriately. During inference nms first 4 classes at first, then scale down boxes for females and juveniles, nms them with pups.</p>",
      "rawMarkdown": "I can't distinguish with my eyes between those similar classes, so switched to DL for better features, but found there similar problems. So I haven't tried one-vs-all or other tricks. My guess that cross-entropy loss doesn't good for this case, since it forces the rest classes to have near 0 score, and network cannot learn well if classes have similar features. I've tried hinge loss to simplify the training, but with no success =), probably due to wrong integration in SSD pipeline.\n\nSmall objects is a pain for all OD. Usually hyper feature should help, but it didn't, or more sophisticated classifier (didn't help), also tried to learn predictors on different layers - conv3 - conv5, but results were similar. So looks like it's not a feature representation problem (it's strong enough), check the pups:\n![pups][1]\nThus SSD here performs as everything else (faster rcnn), but faster.\n\nFrom my experience highly overlapped objects here is one of the main issues with OD approach, they merges during nms and creates troubles during matching. Tried to add 6'th class: female + near pup, but without success. I played a lot with boxes sizes and nms thresholds, last submissions have only 2 box sizes during training: huge for males, smaller for the rest 3 classes to mach appropriately. During inference nms first 4 classes at first, then scale down boxes for females and juveniles, nms them with pups.\n  [1]: http://i65.tinypic.com/2646had.jpg",
      "votes": null
    },
    {
      "id": "197276",
      "postDate": "06/29/2017 07:19:51",
      "content": "<p>Thanks for sharing! Question what do you mean by spending 100 submissions for stacking models? </p>",
      "rawMarkdown": "Thanks for sharing! Question what do you mean by spending 100 submissions for stacking models?",
      "votes": null
    },
    {
      "id": "197286",
      "postDate": "06/29/2017 07:49:33",
      "content": "<ol>\n<li>What is 'hyper feature' ?</li>\n<li>What NMS algorithm are you using something like groupRectangles or groupRectangles_meanshift from opencv or something more sophisticated?</li>\n</ol>",
      "rawMarkdown": "1. What is 'hyper feature' ?\n2. What NMS algorithm are you using something like groupRectangles or groupRectangles_meanshift from opencv or something more sophisticated?",
      "votes": null
    },
    {
      "id": "197302",
      "postDate": "06/29/2017 08:35:05",
      "content": "<p>Thank you, that's interesting. Can you elaborate a bit on AdaBoost pipeline you used (details/links)?</p>",
      "rawMarkdown": "Thank you, that's interesting. Can you elaborate a bit on AdaBoost pipeline you used (details/links)?",
      "votes": null
    },
    {
      "id": "197526",
      "postDate": "06/29/2017 18:50:19",
      "content": "<p>Hyper feature - it's a concatenated feature maps from different network layers: shallow for fine details, mid for class specific responses, deep for scene context, e.g. (conv3_3 &amp; maxpool) concatenated with conv4_3 concatenated with (conv5_3 &amp; deconvolution), after that 1x1 convolutions can be used to integrate the features and reduce the width (also features from each layer need to have similar scale before concatenation, so BN or scale/normalization layers are added). Nice example can be found the in paper <a href=\"https://arxiv.org/pdf/1604.00600.pdf\">\"HyperNet: Towards Accurate Region Proposal Generation and Joint Object Detection\"</a> by Intel guys.</p>\n\n<p>Since we have confidence for each rectangle, they can be suppressed using this cue: rectangle b, with confidence smaller than rectangle a, is suppressed if intersection_area(a, b) / min(a.area(), b.area()) &gt; threshold (0.3 is used), taken from: <a href=\"https://github.com/pdollar/toolbox/blob/master/detector/bbNms.m\">Piotr Dollar's toolbox</a>.</p>",
      "rawMarkdown": "Hyper feature - it's a concatenated feature maps from different network layers: shallow for fine details, mid for class specific responses, deep for scene context, e.g. (conv3_3 &amp; maxpool) concatenated with conv4_3 concatenated with (conv5_3 &amp; deconvolution), after that 1x1 convolutions can be used to integrate the features and reduce the width (also features from each layer need to have similar scale before concatenation, so BN or scale/normalization layers are added). Nice example can be found the in paper [\"HyperNet: Towards Accurate Region Proposal Generation and Joint Object Detection\"][1] by Intel guys.\n\nSince we have confidence for each rectangle, they can be suppressed using this cue: rectangle b, with confidence smaller than rectangle a, is suppressed if intersection_area(a, b) / min(a.area(), b.area()) &gt; threshold (0.3 is used), taken from: [Piotr Dollar's toolbox][2].\n\n\n  [1]: https://arxiv.org/pdf/1604.00600.pdf\n  [2]: https://github.com/pdollar/toolbox/blob/master/detector/bbNms.m",
      "votes": null
    },
    {
      "id": "197529",
      "postDate": "06/29/2017 18:54:26",
      "content": "<p>There is famous <a href=\"https://github.com/pdollar/toolbox\">Piotr Dollar's toolbox</a> for object detection approach: sliding window for localization + AdaBoost for classification. I've used it as is, RealBoost with 1024-2048 decision trees.</p>",
      "rawMarkdown": "There is famous [Piotr Dollar's toolbox][1] for object detection approach: sliding window for localization + AdaBoost for classification. I've used it as is, RealBoost with 1024-2048 decision trees.\n\n\n  [1]: https://github.com/pdollar/toolbox",
      "votes": null
    },
    {
      "id": "197532",
      "postDate": "06/29/2017 18:58:05",
      "content": "<p>I mean that tried the same models (didn't train new), but tested on images, resized to different size, also added different extra amount to pups, juveniles number, thus evaluate no new ideas, just tried to fit the LB.</p>",
      "rawMarkdown": "I mean that tried the same models (didn't train new), but tested on images, resized to different size, also added different extra amount to pups, juveniles number, thus evaluate no new ideas, just tried to fit the LB.",
      "votes": null
    },
    {
      "id": "197819",
      "postDate": "06/30/2017 09:07:32",
      "content": "<p>I also use objection detection.but my model can not classify sealion very well, which almost predict all sealions as adult and female.how about you?</p>",
      "rawMarkdown": "I also use objection detection.but my model can not classify sealion very well, which almost predict all sealions as adult and female.how about you?",
      "votes": null
    },
    {
      "id": "198032",
      "postDate": "06/30/2017 19:48:37",
      "content": "<p>I don't know how well it works, but network stably predicts adult, subadult males and pups. For the rest two classes scale is a problem that was unrivaled for me, and yes females predicted more often than juveniles.</p>",
      "rawMarkdown": "I don't know how well it works, but network stably predicts adult, subadult males and pups. For the rest two classes scale is a problem that was unrivaled for me, and yes females predicted more often than juveniles.",
      "votes": null
    },
    {
      "id": "198312",
      "postDate": "07/01/2017 21:04:35",
      "content": "<p>Seems in semantic segmentation field it's called 'hyper column' <a href=\"https://arxiv.org/pdf/1411.5752.pdf\">https://arxiv.org/pdf/1411.5752.pdf</a></p>",
      "rawMarkdown": "Seems in semantic segmentation field it's called 'hyper column' https://arxiv.org/pdf/1411.5752.pdf",
      "votes": null
    },
    {
      "id": "198791",
      "postDate": "07/03/2017 18:17:22",
      "content": "<p>Yep.</p>",
      "rawMarkdown": "Yep.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 197152,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "06/28/2017 22:35:00",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 197160,
      "author_name": "mrgloom",
      "author_url": "",
      "post_date": "06/28/2017 22:51:20",
      "content": "<ol>\n<li>Did you make any tricks to distinguish between so similar classes? i.e. something like 'finegrained classification'?</li>\n<li>How SSD performs with small objects and highly overlapped objects? did you do some tuning on this?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 197263,
          "author_name": "harshml",
          "author_url": "",
          "post_date": "06/29/2017 06:51:47",
          "content": "<p>I can't distinguish with my eyes between those similar classes, so switched to DL for better features, but found there similar problems. So I haven't tried one-vs-all or other tricks. My guess that cross-entropy loss doesn't good for this case, since it forces the rest classes to have near 0 score, and network cannot learn well if classes have similar features. I've tried hinge loss to simplify the training, but with no success =), probably due to wrong integration in SSD pipeline.</p>\n\n<p>Small objects is a pain for all OD. Usually hyper feature should help, but it didn't, or more sophisticated classifier (didn't help), also tried to learn predictors on different layers - conv3 - conv5, but results were similar. So looks like it's not a feature representation problem (it's strong enough), check the pups:\n<img src=\"http://i65.tinypic.com/2646had.jpg\" alt=\"pups\" title=\"\">\nThus SSD here performs as everything else (faster rcnn), but faster.</p>\n\n<p>From my experience highly overlapped objects here is one of the main issues with OD approach, they merges during nms and creates troubles during matching. Tried to add 6'th class: female + near pup, but without success. I played a lot with boxes sizes and nms thresholds, last submissions have only 2 box sizes during training: huge for males, smaller for the rest 3 classes to mach appropriately. During inference nms first 4 classes at first, then scale down boxes for females and juveniles, nms them with pups.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 197286,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/29/2017 07:49:33",
          "content": "<ol>\n<li>What is 'hyper feature' ?</li>\n<li>What NMS algorithm are you using something like groupRectangles or groupRectangles_meanshift from opencv or something more sophisticated?</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 197526,
          "author_name": "harshml",
          "author_url": "",
          "post_date": "06/29/2017 18:50:19",
          "content": "<p>Hyper feature - it's a concatenated feature maps from different network layers: shallow for fine details, mid for class specific responses, deep for scene context, e.g. (conv3_3 &amp; maxpool) concatenated with conv4_3 concatenated with (conv5_3 &amp; deconvolution), after that 1x1 convolutions can be used to integrate the features and reduce the width (also features from each layer need to have similar scale before concatenation, so BN or scale/normalization layers are added). Nice example can be found the in paper <a href=\"https://arxiv.org/pdf/1604.00600.pdf\">\"HyperNet: Towards Accurate Region Proposal Generation and Joint Object Detection\"</a> by Intel guys.</p>\n\n<p>Since we have confidence for each rectangle, they can be suppressed using this cue: rectangle b, with confidence smaller than rectangle a, is suppressed if intersection_area(a, b) / min(a.area(), b.area()) &gt; threshold (0.3 is used), taken from: <a href=\"https://github.com/pdollar/toolbox/blob/master/detector/bbNms.m\">Piotr Dollar's toolbox</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 198312,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "07/01/2017 21:04:35",
          "content": "<p>Seems in semantic segmentation field it's called 'hyper column' <a href=\"https://arxiv.org/pdf/1411.5752.pdf\">https://arxiv.org/pdf/1411.5752.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 198791,
          "author_name": "harshml",
          "author_url": "",
          "post_date": "07/03/2017 18:17:22",
          "content": "<p>Yep.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 197194,
      "author_name": "outrunner",
      "author_url": "",
      "post_date": "06/29/2017 01:20:50",
      "content": "<p>I was wondering how good could object detection do. Your result is beyond my imagination. Thanks for sharing, good job.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 197276,
      "author_name": "bfortuner",
      "author_url": "",
      "post_date": "06/29/2017 07:19:51",
      "content": "<p>Thanks for sharing! Question what do you mean by spending 100 submissions for stacking models? </p>",
      "votes": null,
      "replies": [
        {
          "id": 197532,
          "author_name": "harshml",
          "author_url": "",
          "post_date": "06/29/2017 18:58:05",
          "content": "<p>I mean that tried the same models (didn't train new), but tested on images, resized to different size, also added different extra amount to pups, juveniles number, thus evaluate no new ideas, just tried to fit the LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 197302,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "06/29/2017 08:35:05",
      "content": "<p>Thank you, that's interesting. Can you elaborate a bit on AdaBoost pipeline you used (details/links)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 197529,
          "author_name": "harshml",
          "author_url": "",
          "post_date": "06/29/2017 18:54:26",
          "content": "<p>There is famous <a href=\"https://github.com/pdollar/toolbox\">Piotr Dollar's toolbox</a> for object detection approach: sliding window for localization + AdaBoost for classification. I've used it as is, RealBoost with 1024-2048 decision trees.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 197819,
      "author_name": "whatsname",
      "author_url": "",
      "post_date": "06/30/2017 09:07:32",
      "content": "<p>I also use objection detection.but my model can not classify sealion very well, which almost predict all sealions as adult and female.how about you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 198032,
          "author_name": "harshml",
          "author_url": "",
          "post_date": "06/30/2017 19:48:37",
          "content": "<p>I don't know how well it works, but network stably predicts adult, subadult males and pups. For the rest two classes scale is a problem that was unrivaled for me, and yes females predicted more often than juveniles.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "197138": "We made a bet with my friend, how far we can go in this challenge (and two month of life have blown away). The problem looks like a type of crowd counting task. Recent solutions prepare density maps for dotted crowd annotations, train network to predict them, and then sum the density in each pixel to obtain the final number. However, I've decided to use object detection approaches, and my friend started with segmentation...\n___\nMy path:\n\n1) Classic ml - AdaBoost\n\n2) SSD on full images (since it is faster than faster rcnn)\n\n3) SSD on tiles\n\nFor object detection boxes is needed, so decided to use squares around the dots coordinates (size was selected manually, OpenCV `inRange` did the work with dots coordinates extraction).\n\nAs for now DL is a mainstream with state of art results, but around 3 years ago, *the best off-the-shelf classifier* **AdaBoost** crushes the leaderboards, so cannot skip to evaluate it. One of the fastest and successful AdaBoost pipelines is boosted decision forest with aggregated channel features (thx to Piotr Dollar for this milestone in object detection). Thus trained 3 models: adult, subadult males and pups. This gives ~20 on public LB *(it took around of an hour to evaluate one model on all test set, on CPU)*. The recall is almost 100%, however, due to lions rotations, models learned to find circle-shaped objects, so there were a lot of false positives on stones.\n\nDL's time. Single Shot Multibox Detector was selected for the speed reasons and good accuracy. Scaled the images to fit the GPU memory (~1200x800) and run with VGG backbone out of the box. DL features give ~18 on LB, mostly due to better (than mean values) work on females and juveniles, testing time is also ~1 hour, but on GPU. Surprisingly, AdaBoost with ACF did well on adult, subadult males, the difference was ~0.1-0.3 compared to SSD. The issues:\n\n* pups detected badly\n\n* can't distinguish between females and juveniles\n\nWhat tried to beat this (without any success):\n\n* Better features - construct a hyper feature from conv3_3 &amp; conv4_3 &amp; conv5_3 to integrate the details from shallow layer and context from deeper.\n\n* Learn the image scale - average pool of feature maps + fc128 + fc16 and concatenate these 16 learned channels with feature maps for final prediction. So, I also did nothing with the scale problem.\n\n* Tried multitask loss - added regression output to SSD classification and localization. Regression produced more or less reasonable values for adult, subadult males, but for classes with huge deviation in number of lions per image obtained `nan`.  Looks like need to spend more time here =), and use small tiles instead of full images.\n\nSo, gave up with sophisticated methods, tile the train data. Used huge tiles 1000x1000 for speed reasons and to avoid border effects. It helped to detect pups better, gives ~17 on LB (during training and testing images were resized to 3000 by the shortest side, then tiles extracted, did flip augmentation), runs ~8 hours on GPU:\n![Picture 1][1]\n![Picture 2][2]\nThe rest 1.5 point is stacking models (spent ~100 submissions while doing this).\n\nThanks for everyone who did involve in this competition. It was fun!\n\n  [1]: http://i67.tinypic.com/juhicw.png\n  [2]: http://i63.tinypic.com/nv76ns.png",
    "197152": "Thanks for sharing!",
    "197160": "1. Did you make any tricks to distinguish between so similar classes? i.e. something like 'finegrained classification'?\n2. How SSD performs with small objects and highly overlapped objects? did you do some tuning on this?",
    "197194": "I was wondering how good could object detection do. Your result is beyond my imagination. Thanks for sharing, good job.",
    "197263": "I can't distinguish with my eyes between those similar classes, so switched to DL for better features, but found there similar problems. So I haven't tried one-vs-all or other tricks. My guess that cross-entropy loss doesn't good for this case, since it forces the rest classes to have near 0 score, and network cannot learn well if classes have similar features. I've tried hinge loss to simplify the training, but with no success =), probably due to wrong integration in SSD pipeline.\n\nSmall objects is a pain for all OD. Usually hyper feature should help, but it didn't, or more sophisticated classifier (didn't help), also tried to learn predictors on different layers - conv3 - conv5, but results were similar. So looks like it's not a feature representation problem (it's strong enough), check the pups:\n![pups][1]\nThus SSD here performs as everything else (faster rcnn), but faster.\n\nFrom my experience highly overlapped objects here is one of the main issues with OD approach, they merges during nms and creates troubles during matching. Tried to add 6'th class: female + near pup, but without success. I played a lot with boxes sizes and nms thresholds, last submissions have only 2 box sizes during training: huge for males, smaller for the rest 3 classes to mach appropriately. During inference nms first 4 classes at first, then scale down boxes for females and juveniles, nms them with pups.\n  [1]: http://i65.tinypic.com/2646had.jpg",
    "197276": "Thanks for sharing! Question what do you mean by spending 100 submissions for stacking models?",
    "197286": "1. What is 'hyper feature' ?\n2. What NMS algorithm are you using something like groupRectangles or groupRectangles_meanshift from opencv or something more sophisticated?",
    "197302": "Thank you, that's interesting. Can you elaborate a bit on AdaBoost pipeline you used (details/links)?",
    "197526": "Hyper feature - it's a concatenated feature maps from different network layers: shallow for fine details, mid for class specific responses, deep for scene context, e.g. (conv3_3 &amp; maxpool) concatenated with conv4_3 concatenated with (conv5_3 &amp; deconvolution), after that 1x1 convolutions can be used to integrate the features and reduce the width (also features from each layer need to have similar scale before concatenation, so BN or scale/normalization layers are added). Nice example can be found the in paper [\"HyperNet: Towards Accurate Region Proposal Generation and Joint Object Detection\"][1] by Intel guys.\n\nSince we have confidence for each rectangle, they can be suppressed using this cue: rectangle b, with confidence smaller than rectangle a, is suppressed if intersection_area(a, b) / min(a.area(), b.area()) &gt; threshold (0.3 is used), taken from: [Piotr Dollar's toolbox][2].\n\n\n  [1]: https://arxiv.org/pdf/1604.00600.pdf\n  [2]: https://github.com/pdollar/toolbox/blob/master/detector/bbNms.m",
    "197529": "There is famous [Piotr Dollar's toolbox][1] for object detection approach: sliding window for localization + AdaBoost for classification. I've used it as is, RealBoost with 1024-2048 decision trees.\n\n\n  [1]: https://github.com/pdollar/toolbox",
    "197532": "I mean that tried the same models (didn't train new), but tested on images, resized to different size, also added different extra amount to pups, juveniles number, thus evaluate no new ideas, just tried to fit the LB.",
    "197819": "I also use objection detection.but my model can not classify sealion very well, which almost predict all sealions as adult and female.how about you?",
    "198032": "I don't know how well it works, but network stably predicts adult, subadult males and pups. For the rest two classes scale is a problem that was unrivaled for me, and yes females predicted more often than juveniles.",
    "198312": "Seems in semantic segmentation field it's called 'hyper column' https://arxiv.org/pdf/1411.5752.pdf",
    "198791": "Yep."
  },
  "source": "meta"
}