{
  "id": 35545,
  "title": "6th solution, objection detection + cnn classification",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/writeups/leilei-weiwei-6th-solution-objection-detection-cnn",
  "author_name": "",
  "post_date": "2017-06-30T06:39:09.390Z",
  "votes": 10,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks to @threeplusone for your coordinates.</p>\n\n<p>My solution consists of two parts, firstly using faster rcnn to get candidate box, and then using cnns to classify.</p>\n\n<p>Two faster rcnns(AN and BN) are trained. AN is trained through multi – scale input ,zooming in each picture two to four times randomly. BN is trained through single – scale input,zooming in each picture twice. Because The scale of the same age group of sealions in different picture changes greatly, for example, in Test ,the adult’s size vary from ~100 pixels to ~300 piexls. When sealions are too big, my single-scale network does not completely cover the whole sealion (mainly adult male),but it has less false positives. Because in second stage I need to use box’s size, I train the muti-scale faster rcnn . It completely covers the sealions, but it has more false positives.</p>\n\n<p>And then I train a CNN to reduce false positives.</p>\n\n<p>Finally, I train about a dozen cnns to classify candidate boxes from AN. I choose best four them lastly and average their results.</p>\n\n<p>Four cnns are as follows,</p>\n\n<p>1)72x72 input, vgg-like cnn</p>\n\n<p>2)72x72 and 144x144 muti - input network</p>\n\n<p>3)72x72 and 72x72(108x108 patch scaled to 72x72) muti - input network</p>\n\n<p>4)same as net3, but output five classes(no background)</p>\n\n<p>In test stage, I will scale each picture at a certain scale(scale = target-scale/ max_box_length , max_box_length is the longest edge of all the candidates from BN in a picture, target-scale is about 78-88). The candidate boxes' coordinates are scaled at the same scale.</p>",
  "messages": [
    {
      "id": "197760",
      "postDate": "06/30/2017 06:12:54",
      "content": "<p>Thanks to @threeplusone for your coordinates.</p>\n\n<p>My solution consists of two parts, firstly using faster rcnn to get candidate box, and then using cnns to classify.</p>\n\n<p>Two faster rcnns(AN and BN) are trained. AN is trained through multi – scale input ,zooming in each picture two to four times randomly. BN is trained through single – scale input,zooming in each picture twice. Because The scale of the same age group of sealions in different picture changes greatly, for example, in Test ,the adult’s size vary from ~100 pixels to ~300 piexls. When sealions are too big, my single-scale network does not completely cover the whole sealion (mainly adult male),but it has less false positives. Because in second stage I need to use box’s size, I train the muti-scale faster rcnn . It completely covers the sealions, but it has more false positives.</p>\n\n<p>And then I train a CNN to reduce false positives.</p>\n\n<p>Finally, I train about a dozen cnns to classify candidate boxes from AN. I choose best four them lastly and average their results.</p>\n\n<p>Four cnns are as follows,</p>\n\n<p>1)72x72 input, vgg-like cnn</p>\n\n<p>2)72x72 and 144x144 muti - input network</p>\n\n<p>3)72x72 and 72x72(108x108 patch scaled to 72x72) muti - input network</p>\n\n<p>4)same as net3, but output five classes(no background)</p>\n\n<p>In test stage, I will scale each picture at a certain scale(scale = target-scale/ max_box_length , max_box_length is the longest edge of all the candidates from BN in a picture, target-scale is about 78-88). The candidate boxes' coordinates are scaled at the same scale.</p>",
      "rawMarkdown": "Thanks to @threeplusone for your coordinates.\n\nMy solution consists of two parts, firstly using faster rcnn to get candidate box, and then using cnns to classify.\n\nTwo faster rcnns(AN and BN) are trained. AN is trained through multi – scale input ,zooming in each picture two to four times randomly. BN is trained through single – scale input,zooming in each picture twice. Because The scale of the same age group of sealions in different picture changes greatly, for example, in Test ,the adult’s size vary from ~100 pixels to ~300 piexls. When sealions are too big, my single-scale network does not completely cover the whole sealion (mainly adult male),but it has less false positives. Because in second stage I need to use box’s size, I train the muti-scale faster rcnn . It completely covers the sealions, but it has more false positives.\n\nAnd then I train a CNN to reduce false positives.\n\nFinally, I train about a dozen cnns to classify candidate boxes from AN. I choose best four them lastly and average their results.\n\nFour cnns are as follows,\n\n1)72x72 input, vgg-like cnn\n\n2)72x72 and 144x144 muti - input network\n\n3)72x72 and 72x72(108x108 patch scaled to 72x72) muti - input network\n\n4)same as net3, but output five classes(no background)\n\nIn test stage, I will scale each picture at a certain scale(scale = target-scale/ max_box_length , max_box_length is the longest edge of all the candidates from BN in a picture, target-scale is about 78-88). The candidate boxes' coordinates are scaled at the same scale.",
      "votes": null
    },
    {
      "id": "197877",
      "postDate": "06/30/2017 12:03:32",
      "content": "<p>Thanks. That is, your solution does not use BN model for region proposals, only AN? And on the second stage you fed to CNNs not exact candidate boxes, but crops of predefined size (72, 108, 144) centered on region proposals?</p>",
      "rawMarkdown": "Thanks. That is, your solution does not use BN model for region proposals, only AN? And on the second stage you fed to CNNs not exact candidate boxes, but crops of predefined size (72, 108, 144) centered on region proposals?",
      "votes": null
    },
    {
      "id": "197883",
      "postDate": "06/30/2017 12:17:24",
      "content": "<p>1) your solution does not use BN model for region proposals, only AN?</p>\n\n<p>I use AN only to adjust picture size, then use cnn to classify proposals from BN.</p>\n\n<p>2)on the second stage you fed to CNNs not exact candidate boxes, but crops of predefined size (72, 108, 144) centered on region proposals?</p>\n\n<p>Yes,the crop is centered on region proposals, and extracted after scaling picture</p>",
      "rawMarkdown": "1) your solution does not use BN model for region proposals, only AN?\n\nI use AN only to adjust picture size, then use cnn to classify proposals from BN.\n\n2)on the second stage you fed to CNNs not exact candidate boxes, but crops of predefined size (72, 108, 144) centered on region proposals?\n\nYes,the crop is centered on region proposals, and extracted after scaling picture",
      "votes": null
    },
    {
      "id": "216618",
      "postDate": "08/27/2017 01:36:12",
      "content": "<p>Hi yingxiaowei,</p>\n\n<p>Would you let me know kind of ground truth did you provide to train the network? Is the ground truth present as bounding box coordinate using in Faster RCNN ( [class, x_left, x_right, y_left, y_right] )? If yes, what size of the GT boxes did you use for different class respectively?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi yingxiaowei,\n\nWould you let me know kind of ground truth did you provide to train the network? Is the ground truth present as bounding box coordinate using in Faster RCNN ( [class, x_left, x_right, y_left, y_right] )? If yes, what size of the GT boxes did you use for different class respectively?\n\nThanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 197877,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "06/30/2017 12:03:32",
      "content": "<p>Thanks. That is, your solution does not use BN model for region proposals, only AN? And on the second stage you fed to CNNs not exact candidate boxes, but crops of predefined size (72, 108, 144) centered on region proposals?</p>",
      "votes": null,
      "replies": [
        {
          "id": 197883,
          "author_name": "whatsname",
          "author_url": "",
          "post_date": "06/30/2017 12:17:24",
          "content": "<p>1) your solution does not use BN model for region proposals, only AN?</p>\n\n<p>I use AN only to adjust picture size, then use cnn to classify proposals from BN.</p>\n\n<p>2)on the second stage you fed to CNNs not exact candidate boxes, but crops of predefined size (72, 108, 144) centered on region proposals?</p>\n\n<p>Yes,the crop is centered on region proposals, and extracted after scaling picture</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 216618,
      "author_name": "robinrorobin",
      "author_url": "",
      "post_date": "08/27/2017 01:36:12",
      "content": "<p>Hi yingxiaowei,</p>\n\n<p>Would you let me know kind of ground truth did you provide to train the network? Is the ground truth present as bounding box coordinate using in Faster RCNN ( [class, x_left, x_right, y_left, y_right] )? If yes, what size of the GT boxes did you use for different class respectively?</p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "197760": "Thanks to @threeplusone for your coordinates.\n\nMy solution consists of two parts, firstly using faster rcnn to get candidate box, and then using cnns to classify.\n\nTwo faster rcnns(AN and BN) are trained. AN is trained through multi – scale input ,zooming in each picture two to four times randomly. BN is trained through single – scale input,zooming in each picture twice. Because The scale of the same age group of sealions in different picture changes greatly, for example, in Test ,the adult’s size vary from ~100 pixels to ~300 piexls. When sealions are too big, my single-scale network does not completely cover the whole sealion (mainly adult male),but it has less false positives. Because in second stage I need to use box’s size, I train the muti-scale faster rcnn . It completely covers the sealions, but it has more false positives.\n\nAnd then I train a CNN to reduce false positives.\n\nFinally, I train about a dozen cnns to classify candidate boxes from AN. I choose best four them lastly and average their results.\n\nFour cnns are as follows,\n\n1)72x72 input, vgg-like cnn\n\n2)72x72 and 144x144 muti - input network\n\n3)72x72 and 72x72(108x108 patch scaled to 72x72) muti - input network\n\n4)same as net3, but output five classes(no background)\n\nIn test stage, I will scale each picture at a certain scale(scale = target-scale/ max_box_length , max_box_length is the longest edge of all the candidates from BN in a picture, target-scale is about 78-88). The candidate boxes' coordinates are scaled at the same scale.",
    "197877": "Thanks. That is, your solution does not use BN model for region proposals, only AN? And on the second stage you fed to CNNs not exact candidate boxes, but crops of predefined size (72, 108, 144) centered on region proposals?",
    "197883": "1) your solution does not use BN model for region proposals, only AN?\n\nI use AN only to adjust picture size, then use cnn to classify proposals from BN.\n\n2)on the second stage you fed to CNNs not exact candidate boxes, but crops of predefined size (72, 108, 144) centered on region proposals?\n\nYes,the crop is centered on region proposals, and extracted after scaling picture",
    "216618": "Hi yingxiaowei,\n\nWould you let me know kind of ground truth did you provide to train the network? Is the ground truth present as bounding box coordinate using in Faster RCNN ( [class, x_left, x_right, y_left, y_right] )? If yes, what size of the GT boxes did you use for different class respectively?\n\nThanks!"
  },
  "source": "meta"
}