{
  "id": 17473,
  "title": "Finding the Whale by Histogram Similarity",
  "url": "/competitions/noaa-right-whale-recognition/discussion/17473",
  "author_name": "",
  "post_date": "2015-11-18T20:25:27.850Z",
  "votes": 15,
  "comment_count": 14,
  "views": 2731,
  "content": "<p>Hi, I've written a script to extract the whale by differentiating it's histogram. Results are not perfect (foam and splashes messes with the identification) but I think it generally works. </p>\n\n<p>Hope this helps!</p>\n\n<p><a href=\"http://eduardofv.com/2015/11/18/detecting-whales-kaggle-right-whale-recognition-challenge/\">Blog Post</a></p>\n\n<p><a href=\"https://github.com/eduardofv/whale_detector\">Github</a></p>\n\n<p><img src=\"https://raw.githubusercontent.com/eduardofv/whale_detector/master/w_7489.jpg.areas.jpg\" alt=\"enter image description here\" title>\n<img src=\"https://raw.githubusercontent.com/eduardofv/whale_detector/master/w_7489.jpg.extract.jpg\" alt=\"enter image description here\" title=\"Extract\"></p>",
  "messages": [
    {
      "id": "99053",
      "postDate": "11/18/2015 20:25:27",
      "content": "<p>Hi, I've written a script to extract the whale by differentiating it's histogram. Results are not perfect (foam and splashes messes with the identification) but I think it generally works. </p>\n\n<p>Hope this helps!</p>\n\n<p><a href=\"http://eduardofv.com/2015/11/18/detecting-whales-kaggle-right-whale-recognition-challenge/\">Blog Post</a></p>\n\n<p><a href=\"https://github.com/eduardofv/whale_detector\">Github</a></p>\n\n<p><img src=\"https://raw.githubusercontent.com/eduardofv/whale_detector/master/w_7489.jpg.areas.jpg\" alt=\"enter image description here\" title>\n<img src=\"https://raw.githubusercontent.com/eduardofv/whale_detector/master/w_7489.jpg.extract.jpg\" alt=\"enter image description here\" title=\"Extract\"></p>",
      "rawMarkdown": "Hi, I've written a script to extract the whale by differentiating it's histogram. Results are not perfect (foam and splashes messes with the identification) but I think it generally works. \r\n\r\nHope this helps!\r\n\r\n[Blog Post][1]\r\n\r\n[Github][2]\r\n\r\n![enter image description here][3]\r\n![enter image description here][4]\r\n\r\n\r\n  [1]: http://eduardofv.com/2015/11/18/detecting-whales-kaggle-right-whale-recognition-challenge/\r\n  [2]: https://github.com/eduardofv/whale_detector\r\n  [3]: https://raw.githubusercontent.com/eduardofv/whale_detector/master/w_7489.jpg.areas.jpg\r\n  [4]: https://raw.githubusercontent.com/eduardofv/whale_detector/master/w_7489.jpg.extract.jpg \"Extract\"",
      "votes": null
    },
    {
      "id": "99068",
      "postDate": "11/19/2015 01:48:06",
      "content": "<p>Tnx for sharing. This is very interesting.</p>",
      "rawMarkdown": "Tnx for sharing. This is very interesting.",
      "votes": null
    },
    {
      "id": "99224",
      "postDate": "11/21/2015 12:42:35",
      "content": "<p>Hi @eduardofv!</p>\n\n<p>I run some evaluations on your method. It looks promising! Almost all images have selected pixels in the head region. And overall, your method is able to get rid of ~90% of the image. </p>\n\n<p>So it's very good as a preprocessing step, thinking that the actual whale head is always (~99%) selected.</p>",
      "rawMarkdown": "Hi @eduardofv!\r\n\r\n I run some evaluations on your method. It looks promising! Almost all images have selected pixels in the head region. And overall, your method is able to get rid of ~90% of the image. \r\n\r\nSo it's very good as a preprocessing step, thinking that the actual whale head is always (~99%) selected.",
      "votes": null
    },
    {
      "id": "99289",
      "postDate": "11/22/2015 18:32:03",
      "content": "<p>Great visoft! Thanks for running those tests!</p>",
      "rawMarkdown": "Great visoft! Thanks for running those tests!",
      "votes": null
    },
    {
      "id": "99329",
      "postDate": "11/23/2015 13:27:55",
      "content": "<p>Brilliant. Here are some statistics derived from my processing of the entire test set:   </p>\n\n<pre><code>Average original  image size 2348 x 3523   \nAverage extracted image size 1120 x 1418   \nAverage extracted area / original area 19%  \n</code></pre>\n\n<p>I don't mean to quibble, your work is extraordinary, but I find many of the animals' noses are right up against one edge and sometimes clipped a little bit which may reduce the amount of information available for identification. One thought I had was along these lines:   </p>\n\n<pre><code>by = by - 40 if by &gt; 40 else 0    \nbh = bh + 80 if bh &lt; img.shape[0] - by else img.shape[0] - by   \nbx = bx - 40 if bx &gt; 40 else 0   \nbw = bw + 80 if bw &lt; img.shape[1] - bx else img.shape[1] - bx   \n\nextract = org[ by:(by+bh), bx:(bx+bw) ]\n</code></pre>",
      "rawMarkdown": "Brilliant. Here are some statistics derived from my processing of the entire test set:   \r\n\r\n    Average original  image size 2348 x 3523   \r\n    Average extracted image size 1120 x 1418   \r\n    Average extracted area / original area 19%  \r\n\r\nI don't mean to quibble, your work is extraordinary, but I find many of the animals' noses are right up against one edge and sometimes clipped a little bit which may reduce the amount of information available for identification. One thought I had was along these lines:   \r\n\r\n    by = by - 40 if by > 40 else 0    \r\n    bh = bh + 80 if bh < img.shape[0] - by else img.shape[0] - by   \r\n    bx = bx - 40 if bx > 40 else 0   \r\n    bw = bw + 80 if bw < img.shape[1] - bx else img.shape[1] - bx   \r\n       \r\n    extract = org[ by:(by+bh), bx:(bx+bw) ]",
      "votes": null
    },
    {
      "id": "99333",
      "postDate": "11/23/2015 14:49:03",
      "content": "<p>@eduardofv - this is great. Thanks for sharing!</p>",
      "rawMarkdown": "eduardofv - this is great. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "99345",
      "postDate": "11/23/2015 16:26:19",
      "content": "<p>Hi George, thanks for sharing. You are right, as a matter of fact on my extracts I'm adding a 10px &quot;padding&quot; around the ROI to consider these cases. Also, I'm selecting a &quot;minimal area rectangle&quot; around the ROI and rotating it to make the final extract. It gives a smaller and tighter extract. It still doesn't give the orientation of the whale but helps a bit. I didn't posted it because makes a larger post and I'm not really sure how much it helps. But in short, yes, padding may give a better extract.</p>\n\n<p>Edit: Also, there are some parameters that may be tweaked, like the minimum similarity, the BW thresholding, or the amount of erode/dilate.</p>",
      "rawMarkdown": "Hi George, thanks for sharing. You are right, as a matter of fact on my extracts I'm adding a 10px \"padding\" around the ROI to consider these cases. Also, I'm selecting a \"minimal area rectangle\" around the ROI and rotating it to make the final extract. It gives a smaller and tighter extract. It still doesn't give the orientation of the whale but helps a bit. I didn't posted it because makes a larger post and I'm not really sure how much it helps. But in short, yes, padding may give a better extract.\r\n\r\nEdit: Also, there are some parameters that may be tweaked, like the minimum similarity, the BW thresholding, or the amount of erode/dilate.",
      "votes": null
    },
    {
      "id": "99384",
      "postDate": "11/24/2015 02:18:48",
      "content": "<p>Have you tried applying CLAHE before running your algo?  </p>\n\n<pre><code>def clahe(rgb):\n    claheizer = cv2.createCLAHE()\n    lab = cv2.cvtColor(rgb.astype('uint8'), cv2.COLOR_RGB2Lab)\n    lab[:,:,0] = claheizer.apply(lab[:,:,0])\n    rgb = cv2.cvtColor(lab.astype('uint8'), cv2.COLOR_Lab2RGB)\n    return rgb\n</code></pre>",
      "rawMarkdown": "Have you tried applying CLAHE before running your algo?  \r\n\r\n    def clahe(rgb):\r\n        claheizer = cv2.createCLAHE()\r\n        lab = cv2.cvtColor(rgb.astype('uint8'), cv2.COLOR_RGB2Lab)\r\n        lab[:,:,0] = claheizer.apply(lab[:,:,0])\r\n        rgb = cv2.cvtColor(lab.astype('uint8'), cv2.COLOR_Lab2RGB)\r\n        return rgb",
      "votes": null
    },
    {
      "id": "99385",
      "postDate": "11/24/2015 03:00:28",
      "content": "<p>Nope, I tried a simple histogram equalization and gave worse results. I&#180;ll try CLAHE as soon as I have a chance (I've never used it before).</p>",
      "rawMarkdown": "Nope, I tried a simple histogram equalization and gave worse results. I´ll try CLAHE as soon as I have a chance (I've never used it before).",
      "votes": null
    },
    {
      "id": "99386",
      "postDate": "11/24/2015 03:05:46",
      "content": "<p>It makes all the images look similar (example attached).  </p>\n\n<p>Look forward to checking out your code.</p>",
      "rawMarkdown": "It makes all the images look similar (example attached).  \r\n\r\nLook forward to checking out your code.",
      "votes": null
    },
    {
      "id": "99394",
      "postDate": "11/24/2015 07:20:02",
      "content": "<p>Thank you very much for your code.</p>\n\n<p>I have got a basic question : you present the processing on the image w_7489, but I can't find it in the data base imgs. Is it another image that you found elsewhere ?</p>\n\n<p>Delphine</p>",
      "rawMarkdown": "Thank you very much for your code.\r\n\r\nI have got a basic question : you present the processing on the image w_7489, but I can't find it in the data base imgs. Is it another image that you found elsewhere ?\r\n\r\nDelphine",
      "votes": null
    },
    {
      "id": "99407",
      "postDate": "11/24/2015 12:14:24",
      "content": "<p>@Delphine, </p>\n\n<p>Image 7489 was omitted from the original zip file and must be downloaded separately from <a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition/data\">here</a>.</p>",
      "rawMarkdown": "Delphine, \r\n\r\nImage 7489 was omitted from the original zip file and must be downloaded separately from [here](https://www.kaggle.com/c/noaa-right-whale-recognition/data).",
      "votes": null
    },
    {
      "id": "99411",
      "postDate": "11/24/2015 13:52:01",
      "content": "<p>@dietCoke I made a quick test with a few images with CLAHE and found that it detects more areas and generally the detection on the whale is less tight. Anyway, I've updated the code on Github with your code but commented the actual call.</p>\n\n<p><img src=\"https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_7489.png\" alt=\"enter image description here\" title>\n<img src=\"https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_1.png\" alt=\"enter image description here\" title>\n<img src=\"https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_2.png\" alt=\"enter image description here\" title>\n<img src=\"https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_4.png\" alt=\"enter image description here\" title></p>",
      "rawMarkdown": "dietCoke I made a quick test with a few images with CLAHE and found that it detects more areas and generally the detection on the whale is less tight. Anyway, I've updated the code on Github with your code but commented the actual call.\r\n\r\n![enter image description here][1]\r\n![enter image description here][2]\r\n![enter image description here][3]\r\n![enter image description here][4]\r\n\r\n\r\n  [1]: https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_7489.png\r\n  [2]: https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_1.png\r\n  [3]: https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_2.png\r\n  [4]: https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_4.png",
      "votes": null
    },
    {
      "id": "99412",
      "postDate": "11/24/2015 14:18:28",
      "content": "<p>@Kevin Burnham\nThanks, I hadn't seen it.</p>",
      "rawMarkdown": "Kevin Burnham\r\nThanks, I hadn't seen it.",
      "votes": null
    },
    {
      "id": "1103966",
      "postDate": "12/06/2020 13:19:57",
      "content": "<p>Interesting application</p>",
      "rawMarkdown": "Interesting application",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1103966,
      "author_name": "amritpal333",
      "author_url": "",
      "post_date": "12/06/2020 13:19:57",
      "content": "<p>Interesting application</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99068,
      "author_name": "vinhnguyen",
      "author_url": "",
      "post_date": "11/19/2015 01:48:06",
      "content": "<p>Tnx for sharing. This is very interesting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99224,
      "author_name": "visoft",
      "author_url": "",
      "post_date": "11/21/2015 12:42:35",
      "content": "<p>Hi @eduardofv!</p>\n\n<p>I run some evaluations on your method. It looks promising! Almost all images have selected pixels in the head region. And overall, your method is able to get rid of ~90% of the image. </p>\n\n<p>So it's very good as a preprocessing step, thinking that the actual whale head is always (~99%) selected.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99289,
      "author_name": "eduardofv",
      "author_url": "",
      "post_date": "11/22/2015 18:32:03",
      "content": "<p>Great visoft! Thanks for running those tests!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99329,
      "author_name": "grfiv4",
      "author_url": "",
      "post_date": "11/23/2015 13:27:55",
      "content": "<p>Brilliant. Here are some statistics derived from my processing of the entire test set:   </p>\n\n<pre><code>Average original  image size 2348 x 3523   \nAverage extracted image size 1120 x 1418   \nAverage extracted area / original area 19%  \n</code></pre>\n\n<p>I don't mean to quibble, your work is extraordinary, but I find many of the animals' noses are right up against one edge and sometimes clipped a little bit which may reduce the amount of information available for identification. One thought I had was along these lines:   </p>\n\n<pre><code>by = by - 40 if by &gt; 40 else 0    \nbh = bh + 80 if bh &lt; img.shape[0] - by else img.shape[0] - by   \nbx = bx - 40 if bx &gt; 40 else 0   \nbw = bw + 80 if bw &lt; img.shape[1] - bx else img.shape[1] - bx   \n\nextract = org[ by:(by+bh), bx:(bx+bw) ]\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99333,
      "author_name": "juliussimonelli",
      "author_url": "",
      "post_date": "11/23/2015 14:49:03",
      "content": "<p>@eduardofv - this is great. Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99345,
      "author_name": "eduardofv",
      "author_url": "",
      "post_date": "11/23/2015 16:26:19",
      "content": "<p>Hi George, thanks for sharing. You are right, as a matter of fact on my extracts I'm adding a 10px &quot;padding&quot; around the ROI to consider these cases. Also, I'm selecting a &quot;minimal area rectangle&quot; around the ROI and rotating it to make the final extract. It gives a smaller and tighter extract. It still doesn't give the orientation of the whale but helps a bit. I didn't posted it because makes a larger post and I'm not really sure how much it helps. But in short, yes, padding may give a better extract.</p>\n\n<p>Edit: Also, there are some parameters that may be tweaked, like the minimum similarity, the BW thresholding, or the amount of erode/dilate.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99384,
      "author_name": "dietcoke",
      "author_url": "",
      "post_date": "11/24/2015 02:18:48",
      "content": "<p>Have you tried applying CLAHE before running your algo?  </p>\n\n<pre><code>def clahe(rgb):\n    claheizer = cv2.createCLAHE()\n    lab = cv2.cvtColor(rgb.astype('uint8'), cv2.COLOR_RGB2Lab)\n    lab[:,:,0] = claheizer.apply(lab[:,:,0])\n    rgb = cv2.cvtColor(lab.astype('uint8'), cv2.COLOR_Lab2RGB)\n    return rgb\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99385,
      "author_name": "eduardofv",
      "author_url": "",
      "post_date": "11/24/2015 03:00:28",
      "content": "<p>Nope, I tried a simple histogram equalization and gave worse results. I&#180;ll try CLAHE as soon as I have a chance (I've never used it before).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99386,
      "author_name": "dietcoke",
      "author_url": "",
      "post_date": "11/24/2015 03:05:46",
      "content": "<p>It makes all the images look similar (example attached).  </p>\n\n<p>Look forward to checking out your code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99394,
      "author_name": "massenet",
      "author_url": "",
      "post_date": "11/24/2015 07:20:02",
      "content": "<p>Thank you very much for your code.</p>\n\n<p>I have got a basic question : you present the processing on the image w_7489, but I can't find it in the data base imgs. Is it another image that you found elsewhere ?</p>\n\n<p>Delphine</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99407,
      "author_name": "kburnham",
      "author_url": "",
      "post_date": "11/24/2015 12:14:24",
      "content": "<p>@Delphine, </p>\n\n<p>Image 7489 was omitted from the original zip file and must be downloaded separately from <a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition/data\">here</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99411,
      "author_name": "eduardofv",
      "author_url": "",
      "post_date": "11/24/2015 13:52:01",
      "content": "<p>@dietCoke I made a quick test with a few images with CLAHE and found that it detects more areas and generally the detection on the whale is less tight. Anyway, I've updated the code on Github with your code but commented the actual call.</p>\n\n<p><img src=\"https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_7489.png\" alt=\"enter image description here\" title>\n<img src=\"https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_1.png\" alt=\"enter image description here\" title>\n<img src=\"https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_2.png\" alt=\"enter image description here\" title>\n<img src=\"https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_4.png\" alt=\"enter image description here\" title></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99412,
      "author_name": "massenet",
      "author_url": "",
      "post_date": "11/24/2015 14:18:28",
      "content": "<p>@Kevin Burnham\nThanks, I hadn't seen it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "99053": "Hi, I've written a script to extract the whale by differentiating it's histogram. Results are not perfect (foam and splashes messes with the identification) but I think it generally works. \r\n\r\nHope this helps!\r\n\r\n[Blog Post][1]\r\n\r\n[Github][2]\r\n\r\n![enter image description here][3]\r\n![enter image description here][4]\r\n\r\n\r\n  [1]: http://eduardofv.com/2015/11/18/detecting-whales-kaggle-right-whale-recognition-challenge/\r\n  [2]: https://github.com/eduardofv/whale_detector\r\n  [3]: https://raw.githubusercontent.com/eduardofv/whale_detector/master/w_7489.jpg.areas.jpg\r\n  [4]: https://raw.githubusercontent.com/eduardofv/whale_detector/master/w_7489.jpg.extract.jpg \"Extract\"",
    "99068": "Tnx for sharing. This is very interesting.",
    "99224": "Hi @eduardofv!\r\n\r\n I run some evaluations on your method. It looks promising! Almost all images have selected pixels in the head region. And overall, your method is able to get rid of ~90% of the image. \r\n\r\nSo it's very good as a preprocessing step, thinking that the actual whale head is always (~99%) selected.",
    "99289": "Great visoft! Thanks for running those tests!",
    "99329": "Brilliant. Here are some statistics derived from my processing of the entire test set:   \r\n\r\n    Average original  image size 2348 x 3523   \r\n    Average extracted image size 1120 x 1418   \r\n    Average extracted area / original area 19%  \r\n\r\nI don't mean to quibble, your work is extraordinary, but I find many of the animals' noses are right up against one edge and sometimes clipped a little bit which may reduce the amount of information available for identification. One thought I had was along these lines:   \r\n\r\n    by = by - 40 if by > 40 else 0    \r\n    bh = bh + 80 if bh < img.shape[0] - by else img.shape[0] - by   \r\n    bx = bx - 40 if bx > 40 else 0   \r\n    bw = bw + 80 if bw < img.shape[1] - bx else img.shape[1] - bx   \r\n       \r\n    extract = org[ by:(by+bh), bx:(bx+bw) ]",
    "99333": "eduardofv - this is great. Thanks for sharing!",
    "99345": "Hi George, thanks for sharing. You are right, as a matter of fact on my extracts I'm adding a 10px \"padding\" around the ROI to consider these cases. Also, I'm selecting a \"minimal area rectangle\" around the ROI and rotating it to make the final extract. It gives a smaller and tighter extract. It still doesn't give the orientation of the whale but helps a bit. I didn't posted it because makes a larger post and I'm not really sure how much it helps. But in short, yes, padding may give a better extract.\r\n\r\nEdit: Also, there are some parameters that may be tweaked, like the minimum similarity, the BW thresholding, or the amount of erode/dilate.",
    "99384": "Have you tried applying CLAHE before running your algo?  \r\n\r\n    def clahe(rgb):\r\n        claheizer = cv2.createCLAHE()\r\n        lab = cv2.cvtColor(rgb.astype('uint8'), cv2.COLOR_RGB2Lab)\r\n        lab[:,:,0] = claheizer.apply(lab[:,:,0])\r\n        rgb = cv2.cvtColor(lab.astype('uint8'), cv2.COLOR_Lab2RGB)\r\n        return rgb",
    "99385": "Nope, I tried a simple histogram equalization and gave worse results. I´ll try CLAHE as soon as I have a chance (I've never used it before).",
    "99386": "It makes all the images look similar (example attached).  \r\n\r\nLook forward to checking out your code.",
    "99394": "Thank you very much for your code.\r\n\r\nI have got a basic question : you present the processing on the image w_7489, but I can't find it in the data base imgs. Is it another image that you found elsewhere ?\r\n\r\nDelphine",
    "99407": "Delphine, \r\n\r\nImage 7489 was omitted from the original zip file and must be downloaded separately from [here](https://www.kaggle.com/c/noaa-right-whale-recognition/data).",
    "99411": "dietCoke I made a quick test with a few images with CLAHE and found that it detects more areas and generally the detection on the whale is less tight. Anyway, I've updated the code on Github with your code but commented the actual call.\r\n\r\n![enter image description here][1]\r\n![enter image description here][2]\r\n![enter image description here][3]\r\n![enter image description here][4]\r\n\r\n\r\n  [1]: https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_7489.png\r\n  [2]: https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_1.png\r\n  [3]: https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_2.png\r\n  [4]: https://raw.githubusercontent.com/eduardofv/whale_detector/master/clahe.ss.w_4.png",
    "99412": "Kevin Burnham\r\nThanks, I hadn't seen it.",
    "1103966": "Interesting application"
  },
  "source": "meta"
}