{
  "id": 16187,
  "title": "Right whale identification game",
  "url": "/competitions/noaa-right-whale-recognition/discussion/16187",
  "author_name": "",
  "post_date": "2015-08-28T22:31:58.530Z",
  "votes": 1,
  "comment_count": 8,
  "views": 2081,
  "content": "<p><a href=\"http://www.neaq.org/education_and_activities/games_and_activities/online_games/right_whale_identification_games.php\">http://www.neaq.org/education_and_activities/games_and_activities/online_games/right_whale_identification_games.php</a></p>\n\n<p>Humans identify right whales by the white/grey callosities on their heads. I think deep learning can work for this but it looks non-trivial to crop out the relevant area of the photo.</p>",
  "messages": [
    {
      "id": "90737",
      "postDate": "08/28/2015 22:31:58",
      "content": "<p><a href=\"http://www.neaq.org/education_and_activities/games_and_activities/online_games/right_whale_identification_games.php\">http://www.neaq.org/education_and_activities/games_and_activities/online_games/right_whale_identification_games.php</a></p>\n\n<p>Humans identify right whales by the white/grey callosities on their heads. I think deep learning can work for this but it looks non-trivial to crop out the relevant area of the photo.</p>",
      "rawMarkdown": "http://www.neaq.org/education_and_activities/games_and_activities/online_games/right_whale_identification_games.php\r\n\r\nHumans identify right whales by the white/grey callosities on their heads. I think deep learning can work for this but it looks non-trivial to crop out the relevant area of the photo.",
      "votes": null
    },
    {
      "id": "90800",
      "postDate": "08/29/2015 02:28:05",
      "content": "<p>Here's some guidance on how you can automate the detection part:\n<a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales\">https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales</a></p>",
      "rawMarkdown": "Here's some guidance on how you can automate the detection part:\r\nhttps://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales",
      "votes": null
    },
    {
      "id": "90825",
      "postDate": "08/29/2015 07:05:18",
      "content": "<blockquote>\n  <p>Here's some guidance on how you can automate the detection part: <a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales\">https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales</a></p>\n</blockquote>\n\n<p>In the above guideline, it seems that external data (manual annotations) is needed to train a detector. Isn't it against the rules? If such annotations are essential, I think they should be provided as part of the dataset, otherwise some might have more or better annotations than others. </p>",
      "rawMarkdown": "> Here's some guidance on how you can automate the detection part: https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales\r\n\r\nIn the above guideline, it seems that external data (manual annotations) is needed to train a detector. Isn't it against the rules? If such annotations are essential, I think they should be provided as part of the dataset, otherwise some might have more or better annotations than others. \r\n\r\n  [1]: https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales",
      "votes": null
    },
    {
      "id": "90851",
      "postDate": "08/29/2015 14:22:40",
      "content": "<p>Hand labeling hundreds of images to build a face detector sounds extremely tedious. Can we hire Amazon Mechanical Turks to do it?</p>",
      "rawMarkdown": "Hand labeling hundreds of images to build a face detector sounds extremely tedious. Can we hire Amazon Mechanical Turks to do it?",
      "votes": null
    },
    {
      "id": "90906",
      "postDate": "08/30/2015 00:46:30",
      "content": "<p>Hi James, Plankton,\nI'll let Christin (from NOAA) comment on providing annotations or using external data for annotations.</p>\n\n<p>The example demonstrates subsampling the training set to manually annotate them as annotations are not provided.\nSince classifying whale from water is easier than recognizing the exact whale, you likely won't need to annotate the entire dataset, just a small subset.\nHow small a subset would depend on what clever image processing, feature extracting and machine learning techniques you use.\nThe example shows (1) annotating few images and train a classifier to 'detect' (2) use your annotator to annotate, crop and train a classifier to recognize.\nThis is one suggested approach, it is possible that there are alternative approaches that are much more efficient.</p>",
      "rawMarkdown": "Hi James, Plankton,\r\nI'll let Christin (from NOAA) comment on providing annotations or using external data for annotations.\r\n\r\nThe example demonstrates subsampling the training set to manually annotate them as annotations are not provided.\r\nSince classifying whale from water is easier than recognizing the exact whale, you likely won't need to annotate the entire dataset, just a small subset.\r\nHow small a subset would depend on what clever image processing, feature extracting and machine learning techniques you use.\r\nThe example shows (1) annotating few images and train a classifier to 'detect' (2) use your annotator to annotate, crop and train a classifier to recognize.\r\nThis is one suggested approach, it is possible that there are alternative approaches that are much more efficient.",
      "votes": null
    },
    {
      "id": "90918",
      "postDate": "08/30/2015 01:39:13",
      "content": "<p>[quote=Plankton;90825]\nIn the above guideline, it seems that external data (manual annotations) is needed to train a detector. Isn't it against the rules?\n[/quote]</p>\n\n<p>You can hand label, hand-crop, and hand-preprocesses the <em>training</em> images. </p>\n\n<p>You <strong>can't</strong> manually do anything to the <em>test</em> images.</p>\n\n<p>See the admin's response to this question:</p>\n\n<p><a href=\"https://www.kaggle.com/c/datasciencebowl/forums/t/12587/manual-vs-auto-feature-selection/64743\">https://www.kaggle.com/c/datasciencebowl/forums/t/12587/manual-vs-auto-feature-selection/64743</a></p>\n\n<p>It would have been easier to have the test and train images in two separate folders, so as not to risk manually adjusting a test images. </p>\n\n<p>Maybe someone can write a script . . .</p>",
      "rawMarkdown": "[quote=Plankton;90825]\r\nIn the above guideline, it seems that external data (manual annotations) is needed to train a detector. Isn't it against the rules?\r\n[/quote]\r\n\r\nYou can hand label, hand-crop, and hand-preprocesses the *training* images. \r\n\r\nYou **can't** manually do anything to the *test* images.\r\n\r\nSee the admin's response to this question:\r\n\r\nhttps://www.kaggle.com/c/datasciencebowl/forums/t/12587/manual-vs-auto-feature-selection/64743\r\n\r\nIt would have been easier to have the test and train images in two separate folders, so as not to risk manually adjusting a test images. \r\n\r\nMaybe someone can write a script . . .",
      "votes": null
    },
    {
      "id": "90919",
      "postDate": "08/30/2015 01:45:30",
      "content": "<p>In the Plankton competition, the <a href=\"https://www.kaggle.com/c/datasciencebowl/rules\">rules were explicit</a> that &quot;Semi-supervised learning is permitted.&quot;</p>\n\n<p>I'm expecting this to be an important clarification for this contest.</p>\n\n<p>Can we use the test images as part of the model building?</p>",
      "rawMarkdown": "In the Plankton competition, the [rules were explicit][1] that \"Semi-supervised learning is permitted.\"\r\n\r\nI'm expecting this to be an important clarification for this contest.\r\n\r\nCan we use the test images as part of the model building?\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/datasciencebowl/rules",
      "votes": null
    },
    {
      "id": "90921",
      "postDate": "08/30/2015 02:22:22",
      "content": "<p>I've manually cropped 1,000 images. It's difficult because some of the photos are bad, and many of the photos don't clearly show the part of the whale that we think will be most useful for identification (between the eyes and the end of the mouth). Dropping these 1,000 cropped images into a convnet produces nothing better than using a constant probability distribution for each siting. Some clever preprocessing will be needed to come up with a decent model.</p>\n\n<p>One thing I've noticed in looking at 1,000 whale photos is that some whales have highly distinctive landmarks; one frequently photographed whale has 4 circular callosities below its eyes, for example. If you see those you definitely know which whale it is. But those landmarks are not visible on every photo and it may be hard to train an algorithm to recognize them. I'll be watching the leaderboard closely to see if someone comes up with something good.</p>",
      "rawMarkdown": "I've manually cropped 1,000 images. It's difficult because some of the photos are bad, and many of the photos don't clearly show the part of the whale that we think will be most useful for identification (between the eyes and the end of the mouth). Dropping these 1,000 cropped images into a convnet produces nothing better than using a constant probability distribution for each siting. Some clever preprocessing will be needed to come up with a decent model.\r\n\r\nOne thing I've noticed in looking at 1,000 whale photos is that some whales have highly distinctive landmarks; one frequently photographed whale has 4 circular callosities below its eyes, for example. If you see those you definitely know which whale it is. But those landmarks are not visible on every photo and it may be hard to train an algorithm to recognize them. I'll be watching the leaderboard closely to see if someone comes up with something good.",
      "votes": null
    },
    {
      "id": "91070",
      "postDate": "08/31/2015 18:20:04",
      "content": "<p>Hi all,</p>\n\n<p>Yes, @inversion is correct. </p>\n\n<ul>\n<li>You can hand label, hand-crop, and hand-preprocesses the training images.</li>\n<li>You can't manually do anything to the test images.</li>\n<li>As a rule of thumb, if you are given a brand new test image, your algorithm should be able to make predictions without requiring any human annotations. </li>\n</ul>\n\n<p>EDIT: Additional rules added to permit semi-supervised learning (as stated above)</p>",
      "rawMarkdown": "Hi all,\r\n\r\nYes, @inversion is correct. \r\n\r\n - You can hand label, hand-crop, and hand-preprocesses the training images.\r\n - You can't manually do anything to the test images.\r\n - As a rule of thumb, if you are given a brand new test image, your algorithm should be able to make predictions without requiring any human annotations. \r\n\r\nEDIT: Additional rules added to permit semi-supervised learning (as stated above)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 90800,
      "author_name": "shashankprasanna",
      "author_url": "",
      "post_date": "08/29/2015 02:28:05",
      "content": "<p>Here's some guidance on how you can automate the detection part:\n<a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales\">https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90825,
      "author_name": "thuyenvn",
      "author_url": "",
      "post_date": "08/29/2015 07:05:18",
      "content": "<blockquote>\n  <p>Here's some guidance on how you can automate the detection part: <a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales\">https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales</a></p>\n</blockquote>\n\n<p>In the above guideline, it seems that external data (manual annotations) is needed to train a detector. Isn't it against the rules? If such annotations are essential, I think they should be provided as part of the dataset, otherwise some might have more or better annotations than others. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90851,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "08/29/2015 14:22:40",
      "content": "<p>Hand labeling hundreds of images to build a face detector sounds extremely tedious. Can we hire Amazon Mechanical Turks to do it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90906,
      "author_name": "shashankprasanna",
      "author_url": "",
      "post_date": "08/30/2015 00:46:30",
      "content": "<p>Hi James, Plankton,\nI'll let Christin (from NOAA) comment on providing annotations or using external data for annotations.</p>\n\n<p>The example demonstrates subsampling the training set to manually annotate them as annotations are not provided.\nSince classifying whale from water is easier than recognizing the exact whale, you likely won't need to annotate the entire dataset, just a small subset.\nHow small a subset would depend on what clever image processing, feature extracting and machine learning techniques you use.\nThe example shows (1) annotating few images and train a classifier to 'detect' (2) use your annotator to annotate, crop and train a classifier to recognize.\nThis is one suggested approach, it is possible that there are alternative approaches that are much more efficient.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90918,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "08/30/2015 01:39:13",
      "content": "<p>[quote=Plankton;90825]\nIn the above guideline, it seems that external data (manual annotations) is needed to train a detector. Isn't it against the rules?\n[/quote]</p>\n\n<p>You can hand label, hand-crop, and hand-preprocesses the <em>training</em> images. </p>\n\n<p>You <strong>can't</strong> manually do anything to the <em>test</em> images.</p>\n\n<p>See the admin's response to this question:</p>\n\n<p><a href=\"https://www.kaggle.com/c/datasciencebowl/forums/t/12587/manual-vs-auto-feature-selection/64743\">https://www.kaggle.com/c/datasciencebowl/forums/t/12587/manual-vs-auto-feature-selection/64743</a></p>\n\n<p>It would have been easier to have the test and train images in two separate folders, so as not to risk manually adjusting a test images. </p>\n\n<p>Maybe someone can write a script . . .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90919,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "08/30/2015 01:45:30",
      "content": "<p>In the Plankton competition, the <a href=\"https://www.kaggle.com/c/datasciencebowl/rules\">rules were explicit</a> that &quot;Semi-supervised learning is permitted.&quot;</p>\n\n<p>I'm expecting this to be an important clarification for this contest.</p>\n\n<p>Can we use the test images as part of the model building?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90921,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "08/30/2015 02:22:22",
      "content": "<p>I've manually cropped 1,000 images. It's difficult because some of the photos are bad, and many of the photos don't clearly show the part of the whale that we think will be most useful for identification (between the eyes and the end of the mouth). Dropping these 1,000 cropped images into a convnet produces nothing better than using a constant probability distribution for each siting. Some clever preprocessing will be needed to come up with a decent model.</p>\n\n<p>One thing I've noticed in looking at 1,000 whale photos is that some whales have highly distinctive landmarks; one frequently photographed whale has 4 circular callosities below its eyes, for example. If you see those you definitely know which whale it is. But those landmarks are not visible on every photo and it may be hard to train an algorithm to recognize them. I'll be watching the leaderboard closely to see if someone comes up with something good.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91070,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "08/31/2015 18:20:04",
      "content": "<p>Hi all,</p>\n\n<p>Yes, @inversion is correct. </p>\n\n<ul>\n<li>You can hand label, hand-crop, and hand-preprocesses the training images.</li>\n<li>You can't manually do anything to the test images.</li>\n<li>As a rule of thumb, if you are given a brand new test image, your algorithm should be able to make predictions without requiring any human annotations. </li>\n</ul>\n\n<p>EDIT: Additional rules added to permit semi-supervised learning (as stated above)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "90737": "http://www.neaq.org/education_and_activities/games_and_activities/online_games/right_whale_identification_games.php\r\n\r\nHumans identify right whales by the white/grey callosities on their heads. I think deep learning can work for this but it looks non-trivial to crop out the relevant area of the photo.",
    "90800": "Here's some guidance on how you can automate the detection part:\r\nhttps://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales",
    "90825": "> Here's some guidance on how you can automate the detection part: https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales\r\n\r\nIn the above guideline, it seems that external data (manual annotations) is needed to train a detector. Isn't it against the rules? If such annotations are essential, I think they should be provided as part of the dataset, otherwise some might have more or better annotations than others. \r\n\r\n  [1]: https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales",
    "90851": "Hand labeling hundreds of images to build a face detector sounds extremely tedious. Can we hire Amazon Mechanical Turks to do it?",
    "90906": "Hi James, Plankton,\r\nI'll let Christin (from NOAA) comment on providing annotations or using external data for annotations.\r\n\r\nThe example demonstrates subsampling the training set to manually annotate them as annotations are not provided.\r\nSince classifying whale from water is easier than recognizing the exact whale, you likely won't need to annotate the entire dataset, just a small subset.\r\nHow small a subset would depend on what clever image processing, feature extracting and machine learning techniques you use.\r\nThe example shows (1) annotating few images and train a classifier to 'detect' (2) use your annotator to annotate, crop and train a classifier to recognize.\r\nThis is one suggested approach, it is possible that there are alternative approaches that are much more efficient.",
    "90918": "[quote=Plankton;90825]\r\nIn the above guideline, it seems that external data (manual annotations) is needed to train a detector. Isn't it against the rules?\r\n[/quote]\r\n\r\nYou can hand label, hand-crop, and hand-preprocesses the *training* images. \r\n\r\nYou **can't** manually do anything to the *test* images.\r\n\r\nSee the admin's response to this question:\r\n\r\nhttps://www.kaggle.com/c/datasciencebowl/forums/t/12587/manual-vs-auto-feature-selection/64743\r\n\r\nIt would have been easier to have the test and train images in two separate folders, so as not to risk manually adjusting a test images. \r\n\r\nMaybe someone can write a script . . .",
    "90919": "In the Plankton competition, the [rules were explicit][1] that \"Semi-supervised learning is permitted.\"\r\n\r\nI'm expecting this to be an important clarification for this contest.\r\n\r\nCan we use the test images as part of the model building?\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/datasciencebowl/rules",
    "90921": "I've manually cropped 1,000 images. It's difficult because some of the photos are bad, and many of the photos don't clearly show the part of the whale that we think will be most useful for identification (between the eyes and the end of the mouth). Dropping these 1,000 cropped images into a convnet produces nothing better than using a constant probability distribution for each siting. Some clever preprocessing will be needed to come up with a decent model.\r\n\r\nOne thing I've noticed in looking at 1,000 whale photos is that some whales have highly distinctive landmarks; one frequently photographed whale has 4 circular callosities below its eyes, for example. If you see those you definitely know which whale it is. But those landmarks are not visible on every photo and it may be hard to train an algorithm to recognize them. I'll be watching the leaderboard closely to see if someone comes up with something good.",
    "91070": "Hi all,\r\n\r\nYes, @inversion is correct. \r\n\r\n - You can hand label, hand-crop, and hand-preprocesses the training images.\r\n - You can't manually do anything to the test images.\r\n - As a rule of thumb, if you are given a brand new test image, your algorithm should be able to make predictions without requiring any human annotations. \r\n\r\nEDIT: Additional rules added to permit semi-supervised learning (as stated above)"
  },
  "source": "meta"
}