{
  "id": 16959,
  "title": "Success Of Whale Detection",
  "url": "/competitions/noaa-right-whale-recognition/discussion/16959",
  "author_name": "",
  "post_date": "2015-10-11T23:12:07.203Z",
  "votes": null,
  "comment_count": 2,
  "views": 1563,
  "content": "<p>I'm curious what other peoples success rate has been with whale detection/cropping. I've attempted to train a HAAR Cascade detector using OpenCV with 707 positive images and 2000+ negative images. However, still, the results are subpar, with 50% of the detections missing the whale.</p>\n\n<p>To be able to detect in a reasonable time and use OpenCV, I had to downsample to 256 pixel width and convert to grayscale. Any thoughts on whether this might be the source of error?</p>",
  "messages": [
    {
      "id": "95786",
      "postDate": "10/11/2015 23:12:07",
      "content": "<p>I'm curious what other peoples success rate has been with whale detection/cropping. I've attempted to train a HAAR Cascade detector using OpenCV with 707 positive images and 2000+ negative images. However, still, the results are subpar, with 50% of the detections missing the whale.</p>\n\n<p>To be able to detect in a reasonable time and use OpenCV, I had to downsample to 256 pixel width and convert to grayscale. Any thoughts on whether this might be the source of error?</p>",
      "rawMarkdown": "I'm curious what other peoples success rate has been with whale detection/cropping. I've attempted to train a HAAR Cascade detector using OpenCV with 707 positive images and 2000+ negative images. However, still, the results are subpar, with 50% of the detections missing the whale.\r\n\r\nTo be able to detect in a reasonable time and use OpenCV, I had to downsample to 256 pixel width and convert to grayscale. Any thoughts on whether this might be the source of error?",
      "votes": null
    },
    {
      "id": "95993",
      "postDate": "10/13/2015 14:59:35",
      "content": "<p>Okay, after further some debugging, I've been able to improve on the system. However, I'm experiencing some issues if anyone would like to chime in and discuss.</p>\n\n<p>I've found that a large source of my error was that at 256 pixel width, the cascade classifier was finding 8+ instances of whales PER image. Clearly, the majority of the 'found' whales were false positives, and was severely affecting the output. However, downspampling to a 128 pixel width provides MUCH more accurate classifications, with an average of 1 whale per image (correct), and around 75% of the photos landing on the whale. It also seems that the majority of false positives can be attributed to instances where multiple whales were found in an image.</p>\n\n<p>The issue now is that the haar cascade seems to be focusing on either the entire whale, or focused on water with the head near the edge of the frame (but still in view). This looks like it's proving to be a problem, because any attempts to beat 6.1 logloss (uniform probability) have failed with the training and test images extracted by my classifier. I believe the issue is that the photos are not focusing enough on the head and calosity patterns of the whales.</p>\n\n<p>Any intuitions on why the number of false positives significantly decreases when using the haar cascade classifier on a smaller/more downsampled image? Or how to combat a lack of focus on calosity patterns/the head of a whale during detection, rather than merely including it in a large region (which defeats the purpose of detection.) </p>",
      "rawMarkdown": "Okay, after further some debugging, I've been able to improve on the system. However, I'm experiencing some issues if anyone would like to chime in and discuss.\r\n\r\nI've found that a large source of my error was that at 256 pixel width, the cascade classifier was finding 8+ instances of whales PER image. Clearly, the majority of the 'found' whales were false positives, and was severely affecting the output. However, downspampling to a 128 pixel width provides MUCH more accurate classifications, with an average of 1 whale per image (correct), and around 75% of the photos landing on the whale. It also seems that the majority of false positives can be attributed to instances where multiple whales were found in an image.\r\n\r\nThe issue now is that the haar cascade seems to be focusing on either the entire whale, or focused on water with the head near the edge of the frame (but still in view). This looks like it's proving to be a problem, because any attempts to beat 6.1 logloss (uniform probability) have failed with the training and test images extracted by my classifier. I believe the issue is that the photos are not focusing enough on the head and calosity patterns of the whales.\r\n\r\nAny intuitions on why the number of false positives significantly decreases when using the haar cascade classifier on a smaller/more downsampled image? Or how to combat a lack of focus on calosity patterns/the head of a whale during detection, rather than merely including it in a large region (which defeats the purpose of detection.)",
      "votes": null
    },
    {
      "id": "96575",
      "postDate": "10/18/2015 16:52:21",
      "content": "<p>You can average the center of the instances found and then put a box around the center, this works for me with a simple edge detector. </p>",
      "rawMarkdown": "You can average the center of the instances found and then put a box around the center, this works for me with a simple edge detector.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 95993,
      "author_name": "mrtwiggy",
      "author_url": "",
      "post_date": "10/13/2015 14:59:35",
      "content": "<p>Okay, after further some debugging, I've been able to improve on the system. However, I'm experiencing some issues if anyone would like to chime in and discuss.</p>\n\n<p>I've found that a large source of my error was that at 256 pixel width, the cascade classifier was finding 8+ instances of whales PER image. Clearly, the majority of the 'found' whales were false positives, and was severely affecting the output. However, downspampling to a 128 pixel width provides MUCH more accurate classifications, with an average of 1 whale per image (correct), and around 75% of the photos landing on the whale. It also seems that the majority of false positives can be attributed to instances where multiple whales were found in an image.</p>\n\n<p>The issue now is that the haar cascade seems to be focusing on either the entire whale, or focused on water with the head near the edge of the frame (but still in view). This looks like it's proving to be a problem, because any attempts to beat 6.1 logloss (uniform probability) have failed with the training and test images extracted by my classifier. I believe the issue is that the photos are not focusing enough on the head and calosity patterns of the whales.</p>\n\n<p>Any intuitions on why the number of false positives significantly decreases when using the haar cascade classifier on a smaller/more downsampled image? Or how to combat a lack of focus on calosity patterns/the head of a whale during detection, rather than merely including it in a large region (which defeats the purpose of detection.) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96575,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "10/18/2015 16:52:21",
      "content": "<p>You can average the center of the instances found and then put a box around the center, this works for me with a simple edge detector. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "95786": "I'm curious what other peoples success rate has been with whale detection/cropping. I've attempted to train a HAAR Cascade detector using OpenCV with 707 positive images and 2000+ negative images. However, still, the results are subpar, with 50% of the detections missing the whale.\r\n\r\nTo be able to detect in a reasonable time and use OpenCV, I had to downsample to 256 pixel width and convert to grayscale. Any thoughts on whether this might be the source of error?",
    "95993": "Okay, after further some debugging, I've been able to improve on the system. However, I'm experiencing some issues if anyone would like to chime in and discuss.\r\n\r\nI've found that a large source of my error was that at 256 pixel width, the cascade classifier was finding 8+ instances of whales PER image. Clearly, the majority of the 'found' whales were false positives, and was severely affecting the output. However, downspampling to a 128 pixel width provides MUCH more accurate classifications, with an average of 1 whale per image (correct), and around 75% of the photos landing on the whale. It also seems that the majority of false positives can be attributed to instances where multiple whales were found in an image.\r\n\r\nThe issue now is that the haar cascade seems to be focusing on either the entire whale, or focused on water with the head near the edge of the frame (but still in view). This looks like it's proving to be a problem, because any attempts to beat 6.1 logloss (uniform probability) have failed with the training and test images extracted by my classifier. I believe the issue is that the photos are not focusing enough on the head and calosity patterns of the whales.\r\n\r\nAny intuitions on why the number of false positives significantly decreases when using the haar cascade classifier on a smaller/more downsampled image? Or how to combat a lack of focus on calosity patterns/the head of a whale during detection, rather than merely including it in a large region (which defeats the purpose of detection.)",
    "96575": "You can average the center of the instances found and then put a box around the center, this works for me with a simple edge detector."
  },
  "source": "meta"
}