{
  "id": 34974,
  "title": "Object Detection vs Image Segmentation Approaches",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/34974",
  "author_name": "",
  "post_date": "2017-06-19T14:17:05.519236600Z",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I was curious if anyone has successfully implemented an object detection method such as Faster R-CNN, or if everyone is working with image segmentation methods such as U-nets or <a href=\"https://arxiv.org/abs/1608.06197\">the density estimator</a>. Perhaps both an object detection method could be used with a density estimator that could help count areas with high concentration of sea lions? Any ideas on how to get passed manually labeling bounding boxes for an object detection method?</p>",
  "messages": [
    {
      "id": "194160",
      "postDate": "06/19/2017 14:17:05",
      "content": "<p>I was curious if anyone has successfully implemented an object detection method such as Faster R-CNN, or if everyone is working with image segmentation methods such as U-nets or <a href=\"https://arxiv.org/abs/1608.06197\">the density estimator</a>. Perhaps both an object detection method could be used with a density estimator that could help count areas with high concentration of sea lions? Any ideas on how to get passed manually labeling bounding boxes for an object detection method?</p>",
      "rawMarkdown": "I was curious if anyone has successfully implemented an object detection method such as Faster R-CNN, or if everyone is working with image segmentation methods such as U-nets or [the density estimator][1]. Perhaps both an object detection method could be used with a density estimator that could help count areas with high concentration of sea lions? Any ideas on how to get passed manually labeling bounding boxes for an object detection method?\n\n\n  [1]: https://arxiv.org/abs/1608.06197",
      "votes": null
    },
    {
      "id": "194196",
      "postDate": "06/19/2017 16:16:36",
      "content": "<p>You can download this <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/31472\">file</a>, and then: </p>\n\n<pre><code>import pandas as pd\n\nsize = [30, 30, 25, 40, 12]\ndf1 = pd.read_csv('coords.csv', dtype={'image_id': object})\ndf2 = df1\n\n\nf = open('frcnn_in.txt', 'w')\nfor index, row in df1.iterrows():\n    x1 = row.x - size[row.klass]\n    x2 = row.x + size[row.klass]\n    y1 = row.y - size[row.klass]\n    y2 = row.y + size[row.klass]\n    line = '../kaggle_data/Train/' + str(row.image_id) + '.jpg,' + str(x1) +  ',' + str(y1) + ','  + str(x2) +  ',' + str(y2) + ',' + str(row.klass)  + '\\n'\n    f.write(line) \n\nf.close()  \n</code></pre>\n\n<p>This file can be used for <a href=\"https://github.com/yhenon/keras-frcnn\">Keras-FRCNN</a></p>",
      "rawMarkdown": "You can download this [file][1], and then: \n\n    import pandas as pd\n    \n    size = [30, 30, 25, 40, 12]\n    df1 = pd.read_csv('coords.csv', dtype={'image_id': object})\n    df2 = df1\n    \n    \n    f = open('frcnn_in.txt', 'w')\n    for index, row in df1.iterrows():\n        x1 = row.x - size[row.klass]\n        x2 = row.x + size[row.klass]\n        y1 = row.y - size[row.klass]\n        y2 = row.y + size[row.klass]\n        line = '../kaggle_data/Train/' + str(row.image_id) + '.jpg,' + str(x1) +  ',' + str(y1) + ','  + str(x2) +  ',' + str(y2) + ',' + str(row.klass)  + '\\n'\n        f.write(line) \n        \n    f.close()  \n\nThis file can be used for [Keras-FRCNN][2]\n\n\n  [1]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/31472\n  [2]: https://github.com/yhenon/keras-frcnn",
      "votes": null
    },
    {
      "id": "194249",
      "postDate": "06/19/2017 21:15:21",
      "content": "<p>@MechCoder did you try this approach? I am getting only Exceptions, if you faced the similar issue please share. </p>\n\n<p>Thanks in advance</p>\n\n<p><img src=\"https://user-images.githubusercontent.com/9487316/27304276-1f41abaa-555b-11e7-80cf-8ab256160780.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "MechCoder did you try this approach? I am getting only Exceptions, if you faced the similar issue please share. \n\nThanks in advance\n\n![enter image description here][1]\n\n\n  [1]: https://user-images.githubusercontent.com/9487316/27304276-1f41abaa-555b-11e7-80cf-8ab256160780.png",
      "votes": null
    },
    {
      "id": "194280",
      "postDate": "06/20/2017 00:06:52",
      "content": "<p>@mechcoder thanks for the help, </p>\n\n<p>for the box's sizes however, how did you determine the values of [30, 30, 25, 40, 12]? Also which class does each correspond to?</p>",
      "rawMarkdown": "mechcoder thanks for the help, \n\nfor the box's sizes however, how did you determine the values of [30, 30, 25, 40, 12]? Also which class does each correspond to?",
      "votes": null
    },
    {
      "id": "194308",
      "postDate": "06/20/2017 03:08:44",
      "content": "<p>@Daft, sorry if there is any bug the code. I tried this approach before and I cannot find the exact code. I didn't get any exception in training but all these small objects are ignored in testing.....</p>",
      "rawMarkdown": "Daft, sorry if there is any bug the code. I tried this approach before and I cannot find the exact code. I didn't get any exception in training but all these small objects are ignored in testing.....",
      "votes": null
    },
    {
      "id": "194309",
      "postDate": "06/20/2017 03:11:27",
      "content": "<p>@LivingProgram, check this: <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/33253#183882\">https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/33253#183882</a></p>",
      "rawMarkdown": "LivingProgram, check this: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/33253#183882",
      "votes": null
    },
    {
      "id": "194381",
      "postDate": "06/20/2017 09:50:56",
      "content": "<p>@mechcoder,\nI see so these are the delta values from the center: deltas = {'pups': {'x': -12, 'y': 12}, 'juveniles': {'x': -40, 'y': 40}, 'subadult_males': {'x': -30, 'y': 30}, 'adult_females': {'x': -25, 'y': 25}, 'adult_males': {'x': -30, 'y': 30}}. Thanks for the clarification.</p>",
      "rawMarkdown": "mechcoder,\nI see so these are the delta values from the center: deltas = {'pups': {'x': -12, 'y': 12}, 'juveniles': {'x': -40, 'y': 40}, 'subadult_males': {'x': -30, 'y': 30}, 'adult_females': {'x': -25, 'y': 25}, 'adult_males': {'x': -30, 'y': 30}}. Thanks for the clarification.",
      "votes": null
    },
    {
      "id": "194386",
      "postDate": "06/20/2017 10:10:06",
      "content": "<p>You'll need to threshold the negative values. \nSince the bounding boxes are exceeded the original image.</p>",
      "rawMarkdown": "You'll need to threshold the negative values. \nSince the bounding boxes are exceeded the original image.",
      "votes": null
    },
    {
      "id": "194429",
      "postDate": "06/20/2017 14:24:10",
      "content": "<p>@Eugene Liu, you made a great point. I guess adding background objects is necessary here.</p>",
      "rawMarkdown": "Eugene Liu, you made a great point. I guess adding background objects is necessary here.",
      "votes": null
    },
    {
      "id": "195176",
      "postDate": "06/22/2017 22:07:00",
      "content": "<p>because I believe your suggestion is a good approach to this problem, here are some of my 2 cents on this...</p>\n\n<p>In implementing Faster RCNN I found challenges to pre-processing the images, especially because the raw images are very large.  I am using a P100 with 16 GB ram and I still needed to chop-up images in order to train the CNN.  I began with the approach given by MechCoder, where bounding boxes are automatically generated for each dotted sealion.</p>\n\n<p>It's a bit of a mess because the blob detection algorithm used to find dotted sealions is not perfect.</p>\n\n<p>I am currently stuck with the issue that it takes about 5 seconds to perform an inference and record the results for each image in the Test set.  5 sec x 19,000 images = a very long time to generate a submission .csv.</p>\n\n<p>Also, there are a few tune-able parameters in faster RCNN.  Its my first time using this algorithm and I am still learning how to fine-tune the hyperparameters.</p>\n\n<p>To your point about performing the first pass with a \"density estimator\", I this this is an excellent idea.  Potentially something like SegNet or an FCN would do the trick.  However, I believe Faster RCNN already has a built-in step to detect object proposals before classifying them as foreground or background.  It would seem that this is similar to finding areas of high sealion density before performing object detection, just like you suggest.  Do you agree?</p>",
      "rawMarkdown": "because I believe your suggestion is a good approach to this problem, here are some of my 2 cents on this...\n\nIn implementing Faster RCNN I found challenges to pre-processing the images, especially because the raw images are very large.  I am using a P100 with 16 GB ram and I still needed to chop-up images in order to train the CNN.  I began with the approach given by MechCoder, where bounding boxes are automatically generated for each dotted sealion.\n\nIt's a bit of a mess because the blob detection algorithm used to find dotted sealions is not perfect.\n\nI am currently stuck with the issue that it takes about 5 seconds to perform an inference and record the results for each image in the Test set.  5 sec x 19,000 images = a very long time to generate a submission .csv.\n\nAlso, there are a few tune-able parameters in faster RCNN.  Its my first time using this algorithm and I am still learning how to fine-tune the hyperparameters.\n\nTo your point about performing the first pass with a \"density estimator\", I this this is an excellent idea.  Potentially something like SegNet or an FCN would do the trick.  However, I believe Faster RCNN already has a built-in step to detect object proposals before classifying them as foreground or background.  It would seem that this is similar to finding areas of high sealion density before performing object detection, just like you suggest.  Do you agree?",
      "votes": null
    },
    {
      "id": "195237",
      "postDate": "06/23/2017 02:28:11",
      "content": "<p>@jeffalltogether ,\nThanks for the very informative explanation of your current model, and I agree with a lot of your points.</p>\n\n<ul>\n<li>the blob detection algorithm is definetly not the best, but I manually labeled dots for some of the rather \"wrong images\" earlier in the competition (over <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/32857\">here</a>) so I don't think I need to worry about that</li>\n<li>I have no clue whether or not automatically generated bounding boxes are accurate (only way to tell is through testing both methods but time is constrained), but I feel accurate data is important so I have invested time to manually label the boxes of many of the sea lions</li>\n<li>I haven't run anything but it appears other are having trouble with <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35150\">long computation times</a></li>\n<li>It's also my first time and I found <a href=\"https://github.com/yhenon/keras-frcnn\">this keras implementation</a> of faster rcnn very helpful</li>\n<li>As for your last point, I agree with what you said about Faster RCNN having a built in step (the region proposal network) to propose object regions before actually classifying them. But I disagree that it is similar to finding areas of high sealion density before performing object detection. If I'm not mistaken I think the region proposal network (RPN) is a sliding window over the output of the feature extractor, and since it has to compute anchor boxes for every single window over the entire feature extractor, it has to compute a lot of stuff! So when you have such massive images, you have lots of windows, lots of anchors and the processing time is a lot longer. So that's why maybe something like the density estimator or <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/34671\">binary tile estimator</a> can skim the image quickly at first so that only a few of your tiles are actually fed to the feature extractor and eventually the RPN, thus cutting down on the computation time spent on the RPN. </li>\n</ul>\n\n<p>So far that's my theory, but I haven't run a model yet, perhaps if when you tested on an image you could record how much time the network spends on the feature extractor step, RPN step,  and the classification step to find the bottleneck? What do you think?</p>",
      "rawMarkdown": "jeffalltogether ,\nThanks for the very informative explanation of your current model, and I agree with a lot of your points.\n\n* the blob detection algorithm is definetly not the best, but I manually labeled dots for some of the rather \"wrong images\" earlier in the competition (over [here][1]) so I don't think I need to worry about that\n* I have no clue whether or not automatically generated bounding boxes are accurate (only way to tell is through testing both methods but time is constrained), but I feel accurate data is important so I have invested time to manually label the boxes of many of the sea lions\n* I haven't run anything but it appears other are having trouble with [long computation times][2]\n* It's also my first time and I found [this keras implementation][3] of faster rcnn very helpful\n* As for your last point, I agree with what you said about Faster RCNN having a built in step (the region proposal network) to propose object regions before actually classifying them. But I disagree that it is similar to finding areas of high sealion density before performing object detection. If I'm not mistaken I think the region proposal network (RPN) is a sliding window over the output of the feature extractor, and since it has to compute anchor boxes for every single window over the entire feature extractor, it has to compute a lot of stuff! So when you have such massive images, you have lots of windows, lots of anchors and the processing time is a lot longer. So that's why maybe something like the density estimator or [binary tile estimator][4] can skim the image quickly at first so that only a few of your tiles are actually fed to the feature extractor and eventually the RPN, thus cutting down on the computation time spent on the RPN. \n\nSo far that's my theory, but I haven't run a model yet, perhaps if when you tested on an image you could record how much time the network spends on the feature extractor step, RPN step,  and the classification step to find the bottleneck? What do you think?\n\n\n  [1]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/32857\n  [2]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35150\n  [3]: https://github.com/yhenon/keras-frcnn\n  [4]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/34671",
      "votes": null
    },
    {
      "id": "195440",
      "postDate": "06/23/2017 16:02:10",
      "content": "<p>Thank for the comments.  I understand what you mean about the RPN and its raster approach.\nI guess I still have the same question that you have.  Would adding a pre-processing step to eliminate areas of very low sea lion probability improve the accuracy of a target-detection classifier.  </p>\n\n<p>This idea is a lot like medical diagnostic testing where the first diagnostic test has a very high sensitivity at the expense of specificity, so that no disease + people are missed in the initial screening.  Patients with + screens then go on to other tests with high specificity. </p>\n\n<p>Based on this paradigm, the pre-processing density estimator could be set to have a high sensitivity, then the second object detection classifier set to high specificity.</p>",
      "rawMarkdown": "Thank for the comments.  I understand what you mean about the RPN and its raster approach.\nI guess I still have the same question that you have.  Would adding a pre-processing step to eliminate areas of very low sea lion probability improve the accuracy of a target-detection classifier.  \n\nThis idea is a lot like medical diagnostic testing where the first diagnostic test has a very high sensitivity at the expense of specificity, so that no disease + people are missed in the initial screening.  Patients with + screens then go on to other tests with high specificity. \n\nBased on this paradigm, the pre-processing density estimator could be set to have a high sensitivity, then the second object detection classifier set to high specificity.",
      "votes": null
    },
    {
      "id": "195503",
      "postDate": "06/23/2017 19:07:04",
      "content": "<p>@jeffalltogether, Exactly that would be great! </p>",
      "rawMarkdown": "jeffalltogether, Exactly that would be great!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 194196,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "06/19/2017 16:16:36",
      "content": "<p>You can download this <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/31472\">file</a>, and then: </p>\n\n<pre><code>import pandas as pd\n\nsize = [30, 30, 25, 40, 12]\ndf1 = pd.read_csv('coords.csv', dtype={'image_id': object})\ndf2 = df1\n\n\nf = open('frcnn_in.txt', 'w')\nfor index, row in df1.iterrows():\n    x1 = row.x - size[row.klass]\n    x2 = row.x + size[row.klass]\n    y1 = row.y - size[row.klass]\n    y2 = row.y + size[row.klass]\n    line = '../kaggle_data/Train/' + str(row.image_id) + '.jpg,' + str(x1) +  ',' + str(y1) + ','  + str(x2) +  ',' + str(y2) + ',' + str(row.klass)  + '\\n'\n    f.write(line) \n\nf.close()  \n</code></pre>\n\n<p>This file can be used for <a href=\"https://github.com/yhenon/keras-frcnn\">Keras-FRCNN</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 194249,
          "author_name": "syeddanish",
          "author_url": "",
          "post_date": "06/19/2017 21:15:21",
          "content": "<p>@MechCoder did you try this approach? I am getting only Exceptions, if you faced the similar issue please share. </p>\n\n<p>Thanks in advance</p>\n\n<p><img src=\"https://user-images.githubusercontent.com/9487316/27304276-1f41abaa-555b-11e7-80cf-8ab256160780.png\" alt=\"enter image description here\" title=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194280,
          "author_name": "livingprogram",
          "author_url": "",
          "post_date": "06/20/2017 00:06:52",
          "content": "<p>@mechcoder thanks for the help, </p>\n\n<p>for the box's sizes however, how did you determine the values of [30, 30, 25, 40, 12]? Also which class does each correspond to?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194308,
          "author_name": "shujian",
          "author_url": "",
          "post_date": "06/20/2017 03:08:44",
          "content": "<p>@Daft, sorry if there is any bug the code. I tried this approach before and I cannot find the exact code. I didn't get any exception in training but all these small objects are ignored in testing.....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194309,
          "author_name": "shujian",
          "author_url": "",
          "post_date": "06/20/2017 03:11:27",
          "content": "<p>@LivingProgram, check this: <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/33253#183882\">https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/33253#183882</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194381,
          "author_name": "livingprogram",
          "author_url": "",
          "post_date": "06/20/2017 09:50:56",
          "content": "<p>@mechcoder,\nI see so these are the delta values from the center: deltas = {'pups': {'x': -12, 'y': 12}, 'juveniles': {'x': -40, 'y': 40}, 'subadult_males': {'x': -30, 'y': 30}, 'adult_females': {'x': -25, 'y': 25}, 'adult_males': {'x': -30, 'y': 30}}. Thanks for the clarification.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194386,
          "author_name": "",
          "author_url": "",
          "post_date": "06/20/2017 10:10:06",
          "content": "<p>You'll need to threshold the negative values. \nSince the bounding boxes are exceeded the original image.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194429,
          "author_name": "shujian",
          "author_url": "",
          "post_date": "06/20/2017 14:24:10",
          "content": "<p>@Eugene Liu, you made a great point. I guess adding background objects is necessary here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 195176,
      "author_name": "jeffalltogether",
      "author_url": "",
      "post_date": "06/22/2017 22:07:00",
      "content": "<p>because I believe your suggestion is a good approach to this problem, here are some of my 2 cents on this...</p>\n\n<p>In implementing Faster RCNN I found challenges to pre-processing the images, especially because the raw images are very large.  I am using a P100 with 16 GB ram and I still needed to chop-up images in order to train the CNN.  I began with the approach given by MechCoder, where bounding boxes are automatically generated for each dotted sealion.</p>\n\n<p>It's a bit of a mess because the blob detection algorithm used to find dotted sealions is not perfect.</p>\n\n<p>I am currently stuck with the issue that it takes about 5 seconds to perform an inference and record the results for each image in the Test set.  5 sec x 19,000 images = a very long time to generate a submission .csv.</p>\n\n<p>Also, there are a few tune-able parameters in faster RCNN.  Its my first time using this algorithm and I am still learning how to fine-tune the hyperparameters.</p>\n\n<p>To your point about performing the first pass with a \"density estimator\", I this this is an excellent idea.  Potentially something like SegNet or an FCN would do the trick.  However, I believe Faster RCNN already has a built-in step to detect object proposals before classifying them as foreground or background.  It would seem that this is similar to finding areas of high sealion density before performing object detection, just like you suggest.  Do you agree?</p>",
      "votes": null,
      "replies": [
        {
          "id": 195237,
          "author_name": "livingprogram",
          "author_url": "",
          "post_date": "06/23/2017 02:28:11",
          "content": "<p>@jeffalltogether ,\nThanks for the very informative explanation of your current model, and I agree with a lot of your points.</p>\n\n<ul>\n<li>the blob detection algorithm is definetly not the best, but I manually labeled dots for some of the rather \"wrong images\" earlier in the competition (over <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/32857\">here</a>) so I don't think I need to worry about that</li>\n<li>I have no clue whether or not automatically generated bounding boxes are accurate (only way to tell is through testing both methods but time is constrained), but I feel accurate data is important so I have invested time to manually label the boxes of many of the sea lions</li>\n<li>I haven't run anything but it appears other are having trouble with <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35150\">long computation times</a></li>\n<li>It's also my first time and I found <a href=\"https://github.com/yhenon/keras-frcnn\">this keras implementation</a> of faster rcnn very helpful</li>\n<li>As for your last point, I agree with what you said about Faster RCNN having a built in step (the region proposal network) to propose object regions before actually classifying them. But I disagree that it is similar to finding areas of high sealion density before performing object detection. If I'm not mistaken I think the region proposal network (RPN) is a sliding window over the output of the feature extractor, and since it has to compute anchor boxes for every single window over the entire feature extractor, it has to compute a lot of stuff! So when you have such massive images, you have lots of windows, lots of anchors and the processing time is a lot longer. So that's why maybe something like the density estimator or <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/34671\">binary tile estimator</a> can skim the image quickly at first so that only a few of your tiles are actually fed to the feature extractor and eventually the RPN, thus cutting down on the computation time spent on the RPN. </li>\n</ul>\n\n<p>So far that's my theory, but I haven't run a model yet, perhaps if when you tested on an image you could record how much time the network spends on the feature extractor step, RPN step,  and the classification step to find the bottleneck? What do you think?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195440,
          "author_name": "jeffalltogether",
          "author_url": "",
          "post_date": "06/23/2017 16:02:10",
          "content": "<p>Thank for the comments.  I understand what you mean about the RPN and its raster approach.\nI guess I still have the same question that you have.  Would adding a pre-processing step to eliminate areas of very low sea lion probability improve the accuracy of a target-detection classifier.  </p>\n\n<p>This idea is a lot like medical diagnostic testing where the first diagnostic test has a very high sensitivity at the expense of specificity, so that no disease + people are missed in the initial screening.  Patients with + screens then go on to other tests with high specificity. </p>\n\n<p>Based on this paradigm, the pre-processing density estimator could be set to have a high sensitivity, then the second object detection classifier set to high specificity.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195503,
          "author_name": "livingprogram",
          "author_url": "",
          "post_date": "06/23/2017 19:07:04",
          "content": "<p>@jeffalltogether, Exactly that would be great! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "194160": "I was curious if anyone has successfully implemented an object detection method such as Faster R-CNN, or if everyone is working with image segmentation methods such as U-nets or [the density estimator][1]. Perhaps both an object detection method could be used with a density estimator that could help count areas with high concentration of sea lions? Any ideas on how to get passed manually labeling bounding boxes for an object detection method?\n\n\n  [1]: https://arxiv.org/abs/1608.06197",
    "194196": "You can download this [file][1], and then: \n\n    import pandas as pd\n    \n    size = [30, 30, 25, 40, 12]\n    df1 = pd.read_csv('coords.csv', dtype={'image_id': object})\n    df2 = df1\n    \n    \n    f = open('frcnn_in.txt', 'w')\n    for index, row in df1.iterrows():\n        x1 = row.x - size[row.klass]\n        x2 = row.x + size[row.klass]\n        y1 = row.y - size[row.klass]\n        y2 = row.y + size[row.klass]\n        line = '../kaggle_data/Train/' + str(row.image_id) + '.jpg,' + str(x1) +  ',' + str(y1) + ','  + str(x2) +  ',' + str(y2) + ',' + str(row.klass)  + '\\n'\n        f.write(line) \n        \n    f.close()  \n\nThis file can be used for [Keras-FRCNN][2]\n\n\n  [1]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/31472\n  [2]: https://github.com/yhenon/keras-frcnn",
    "194249": "MechCoder did you try this approach? I am getting only Exceptions, if you faced the similar issue please share. \n\nThanks in advance\n\n![enter image description here][1]\n\n\n  [1]: https://user-images.githubusercontent.com/9487316/27304276-1f41abaa-555b-11e7-80cf-8ab256160780.png",
    "194280": "mechcoder thanks for the help, \n\nfor the box's sizes however, how did you determine the values of [30, 30, 25, 40, 12]? Also which class does each correspond to?",
    "194308": "Daft, sorry if there is any bug the code. I tried this approach before and I cannot find the exact code. I didn't get any exception in training but all these small objects are ignored in testing.....",
    "194309": "LivingProgram, check this: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/33253#183882",
    "194381": "mechcoder,\nI see so these are the delta values from the center: deltas = {'pups': {'x': -12, 'y': 12}, 'juveniles': {'x': -40, 'y': 40}, 'subadult_males': {'x': -30, 'y': 30}, 'adult_females': {'x': -25, 'y': 25}, 'adult_males': {'x': -30, 'y': 30}}. Thanks for the clarification.",
    "194386": "You'll need to threshold the negative values. \nSince the bounding boxes are exceeded the original image.",
    "194429": "Eugene Liu, you made a great point. I guess adding background objects is necessary here.",
    "195176": "because I believe your suggestion is a good approach to this problem, here are some of my 2 cents on this...\n\nIn implementing Faster RCNN I found challenges to pre-processing the images, especially because the raw images are very large.  I am using a P100 with 16 GB ram and I still needed to chop-up images in order to train the CNN.  I began with the approach given by MechCoder, where bounding boxes are automatically generated for each dotted sealion.\n\nIt's a bit of a mess because the blob detection algorithm used to find dotted sealions is not perfect.\n\nI am currently stuck with the issue that it takes about 5 seconds to perform an inference and record the results for each image in the Test set.  5 sec x 19,000 images = a very long time to generate a submission .csv.\n\nAlso, there are a few tune-able parameters in faster RCNN.  Its my first time using this algorithm and I am still learning how to fine-tune the hyperparameters.\n\nTo your point about performing the first pass with a \"density estimator\", I this this is an excellent idea.  Potentially something like SegNet or an FCN would do the trick.  However, I believe Faster RCNN already has a built-in step to detect object proposals before classifying them as foreground or background.  It would seem that this is similar to finding areas of high sealion density before performing object detection, just like you suggest.  Do you agree?",
    "195237": "jeffalltogether ,\nThanks for the very informative explanation of your current model, and I agree with a lot of your points.\n\n* the blob detection algorithm is definetly not the best, but I manually labeled dots for some of the rather \"wrong images\" earlier in the competition (over [here][1]) so I don't think I need to worry about that\n* I have no clue whether or not automatically generated bounding boxes are accurate (only way to tell is through testing both methods but time is constrained), but I feel accurate data is important so I have invested time to manually label the boxes of many of the sea lions\n* I haven't run anything but it appears other are having trouble with [long computation times][2]\n* It's also my first time and I found [this keras implementation][3] of faster rcnn very helpful\n* As for your last point, I agree with what you said about Faster RCNN having a built in step (the region proposal network) to propose object regions before actually classifying them. But I disagree that it is similar to finding areas of high sealion density before performing object detection. If I'm not mistaken I think the region proposal network (RPN) is a sliding window over the output of the feature extractor, and since it has to compute anchor boxes for every single window over the entire feature extractor, it has to compute a lot of stuff! So when you have such massive images, you have lots of windows, lots of anchors and the processing time is a lot longer. So that's why maybe something like the density estimator or [binary tile estimator][4] can skim the image quickly at first so that only a few of your tiles are actually fed to the feature extractor and eventually the RPN, thus cutting down on the computation time spent on the RPN. \n\nSo far that's my theory, but I haven't run a model yet, perhaps if when you tested on an image you could record how much time the network spends on the feature extractor step, RPN step,  and the classification step to find the bottleneck? What do you think?\n\n\n  [1]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/32857\n  [2]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35150\n  [3]: https://github.com/yhenon/keras-frcnn\n  [4]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/34671",
    "195440": "Thank for the comments.  I understand what you mean about the RPN and its raster approach.\nI guess I still have the same question that you have.  Would adding a pre-processing step to eliminate areas of very low sea lion probability improve the accuracy of a target-detection classifier.  \n\nThis idea is a lot like medical diagnostic testing where the first diagnostic test has a very high sensitivity at the expense of specificity, so that no disease + people are missed in the initial screening.  Patients with + screens then go on to other tests with high specificity. \n\nBased on this paradigm, the pre-processing density estimator could be set to have a high sensitivity, then the second object detection classifier set to high specificity.",
    "195503": "jeffalltogether, Exactly that would be great!"
  },
  "source": "meta"
}