{
  "id": 16954,
  "title": "A few questions",
  "url": "/competitions/noaa-right-whale-recognition/discussion/16954",
  "author_name": "",
  "post_date": "2015-10-11T14:40:21.527Z",
  "votes": null,
  "comment_count": 3,
  "views": 1128,
  "content": "<p>Hi everyone, </p>\n\n<p>I stumbled across this project a few days ago and I find it really fascinating. It's extra special for me because my wife and I always make sure to visit the New England Aquarium when we are in the Boston area. I've been a computer programmer for 18+ years working on mostly large web applications, but have found the concept of data science to be fascinating and something I want to explore further. </p>\n\n<p>This is my first image processing and machine learning project, so please bear with me. I have a few questions: </p>\n\n<p>For Step 1 in the &quot;Creating a &quot;Face Detector&quot; for Whales&quot; - for annotating the images, how big / small or broad / precise should the annotations be ? Should they include just the whale's head? Or the whole whale body? I annotated the 500 images from the subset, however, my detector isn't detecting any whales when I feed it an image from the regular set - or it detects three+ random boxes on the page. Did I use too many images when I built my classifier? I'll attach two examples of how I annotated images from the initial subset.  </p>\n\n<p>I'm a little concerned by this output when I run trainCascadeObjectDetector: </p>\n\n<p><em>Automatically setting ObjectTrainingSize to [ 32, 41 ]</em></p>\n\n<p>From the documentation:</p>\n\n<p><em>Before training, the function resizes the positive and negative samples to ObjectTrainingSize in pixels.</em></p>\n\n<p>This doesn't sound good at all. </p>\n\n<p><em>For optimal detection accuracy, specify an object training size close to the expected size of the object in the image.</em></p>\n\n<p>I'm confused about this as per my previous question, should we be limiting the size and what we annotate when we annotate a whale in the training set? For instance, just annotate a 100x100 square around the head of the whale? Maybe we should pass</p>\n\n<p><code>'ObjectTrainingSize','Auto'</code></p>\n\n<p>to trainCascadeObjectDetector ? </p>\n\n<p>Steps 1-4 seem pretty straight forward - thanks to Nina Chen, and then this is where I get a little confused. Now that we have a detector and have run it on the entire dataset. What does that return exactly? </p>\n\n<pre><code>bbox = step(WhaleDetectorMdl,img);\n</code></pre>\n\n<p>seems to be a 3x4 matrix. Probably because I have 3 annotations on this given image, so assuming a correct image would have only 1 annotation it would be a 1x4 matrix. Would this be (x1, y1, height, width)? </p>\n\n<p>Now what do we do? :) I assume we pick a machine learning algorithm, and feed it both the training.csv file and corresponding set of images referenced by tranining.csv to tell the algorithm &quot;Hey, For image w_9450.jpg, you can find whale_67614 at [1000, 500, 1000, 100, 120] - now learn!&quot;</p>\n\n<p>Should we start by looking at Classification algorithms? </p>\n\n<p>And finally - last question, once we've built this beautiful learned machine, which images do we run it on to generate a solution to submit for evaluation? The entire 11,468 images? From train.csv we already know the solution to 4,528 of the images, do we include those in our dataset? And if so, do we override what the machine learning algorithm has told us about the known image, and put 1 for the corresponding whale and 0 for all of the others to improve our evaluation? </p>",
  "messages": [
    {
      "id": "95757",
      "postDate": "10/11/2015 14:40:21",
      "content": "<p>Hi everyone, </p>\n\n<p>I stumbled across this project a few days ago and I find it really fascinating. It's extra special for me because my wife and I always make sure to visit the New England Aquarium when we are in the Boston area. I've been a computer programmer for 18+ years working on mostly large web applications, but have found the concept of data science to be fascinating and something I want to explore further. </p>\n\n<p>This is my first image processing and machine learning project, so please bear with me. I have a few questions: </p>\n\n<p>For Step 1 in the &quot;Creating a &quot;Face Detector&quot; for Whales&quot; - for annotating the images, how big / small or broad / precise should the annotations be ? Should they include just the whale's head? Or the whole whale body? I annotated the 500 images from the subset, however, my detector isn't detecting any whales when I feed it an image from the regular set - or it detects three+ random boxes on the page. Did I use too many images when I built my classifier? I'll attach two examples of how I annotated images from the initial subset.  </p>\n\n<p>I'm a little concerned by this output when I run trainCascadeObjectDetector: </p>\n\n<p><em>Automatically setting ObjectTrainingSize to [ 32, 41 ]</em></p>\n\n<p>From the documentation:</p>\n\n<p><em>Before training, the function resizes the positive and negative samples to ObjectTrainingSize in pixels.</em></p>\n\n<p>This doesn't sound good at all. </p>\n\n<p><em>For optimal detection accuracy, specify an object training size close to the expected size of the object in the image.</em></p>\n\n<p>I'm confused about this as per my previous question, should we be limiting the size and what we annotate when we annotate a whale in the training set? For instance, just annotate a 100x100 square around the head of the whale? Maybe we should pass</p>\n\n<p><code>'ObjectTrainingSize','Auto'</code></p>\n\n<p>to trainCascadeObjectDetector ? </p>\n\n<p>Steps 1-4 seem pretty straight forward - thanks to Nina Chen, and then this is where I get a little confused. Now that we have a detector and have run it on the entire dataset. What does that return exactly? </p>\n\n<pre><code>bbox = step(WhaleDetectorMdl,img);\n</code></pre>\n\n<p>seems to be a 3x4 matrix. Probably because I have 3 annotations on this given image, so assuming a correct image would have only 1 annotation it would be a 1x4 matrix. Would this be (x1, y1, height, width)? </p>\n\n<p>Now what do we do? :) I assume we pick a machine learning algorithm, and feed it both the training.csv file and corresponding set of images referenced by tranining.csv to tell the algorithm &quot;Hey, For image w_9450.jpg, you can find whale_67614 at [1000, 500, 1000, 100, 120] - now learn!&quot;</p>\n\n<p>Should we start by looking at Classification algorithms? </p>\n\n<p>And finally - last question, once we've built this beautiful learned machine, which images do we run it on to generate a solution to submit for evaluation? The entire 11,468 images? From train.csv we already know the solution to 4,528 of the images, do we include those in our dataset? And if so, do we override what the machine learning algorithm has told us about the known image, and put 1 for the corresponding whale and 0 for all of the others to improve our evaluation? </p>",
      "rawMarkdown": "Hi everyone, \r\n\r\nI stumbled across this project a few days ago and I find it really fascinating. It's extra special for me because my wife and I always make sure to visit the New England Aquarium when we are in the Boston area. I've been a computer programmer for 18+ years working on mostly large web applications, but have found the concept of data science to be fascinating and something I want to explore further. \r\n\r\nThis is my first image processing and machine learning project, so please bear with me. I have a few questions: \r\n\r\nFor Step 1 in the \"Creating a \"Face Detector\" for Whales\" - for annotating the images, how big / small or broad / precise should the annotations be ? Should they include just the whale's head? Or the whole whale body? I annotated the 500 images from the subset, however, my detector isn't detecting any whales when I feed it an image from the regular set - or it detects three+ random boxes on the page. Did I use too many images when I built my classifier? I'll attach two examples of how I annotated images from the initial subset.  \r\n\r\nI'm a little concerned by this output when I run trainCascadeObjectDetector: \r\n\r\n*Automatically setting ObjectTrainingSize to [ 32, 41 ]*\r\n\r\nFrom the documentation:\r\n\r\n*Before training, the function resizes the positive and negative samples to ObjectTrainingSize in pixels.*\r\n\r\nThis doesn't sound good at all. \r\n\r\n *For optimal detection accuracy, specify an object training size close to the expected size of the object in the image.*\r\n\r\nI'm confused about this as per my previous question, should we be limiting the size and what we annotate when we annotate a whale in the training set? For instance, just annotate a 100x100 square around the head of the whale? Maybe we should pass\r\n\r\n`'ObjectTrainingSize','Auto'`\r\n\r\nto trainCascadeObjectDetector ? \r\n\r\nSteps 1-4 seem pretty straight forward - thanks to Nina Chen, and then this is where I get a little confused. Now that we have a detector and have run it on the entire dataset. What does that return exactly? \r\n\r\n    bbox = step(WhaleDetectorMdl,img);\r\n\r\nseems to be a 3x4 matrix. Probably because I have 3 annotations on this given image, so assuming a correct image would have only 1 annotation it would be a 1x4 matrix. Would this be (x1, y1, height, width)? \r\n\r\nNow what do we do? :) I assume we pick a machine learning algorithm, and feed it both the training.csv file and corresponding set of images referenced by tranining.csv to tell the algorithm \"Hey, For image w_9450.jpg, you can find whale_67614 at [1000, 500, 1000, 100, 120] - now learn!\"\r\n\r\nShould we start by looking at Classification algorithms? \r\n\r\nAnd finally - last question, once we've built this beautiful learned machine, which images do we run it on to generate a solution to submit for evaluation? The entire 11,468 images? From train.csv we already know the solution to 4,528 of the images, do we include those in our dataset? And if so, do we override what the machine learning algorithm has told us about the known image, and put 1 for the corresponding whale and 0 for all of the others to improve our evaluation?",
      "votes": null
    },
    {
      "id": "95761",
      "postDate": "10/11/2015 15:33:54",
      "content": "<p>For the size of the annotations, it depends! Likely, you will be cropping out the 'whale detected' part of the image. This essentially 'removes' water pixels/not useful information. So the only information your classifier/recognizer will need are the important parts that distinguish whales from each other. There is quite a bit of useful information on this in the sticky threads, but it seems that the heads of whales are the most important part as they have identifying marks (calosities?) on their heads. But you might be able to find useful information with the whole whale body.</p>\n\n<p>I've been unable to get MatPlot yet, so my help is going to be limited. However, when you say you fed in 500 annotated images, does that mean you also got negatives? I believe you will need to feed in 'negative' annotated images where you are cropping out a piece of background that doesn't have a whale in it. </p>\n\n<p>As for which images to run your final classifier/recognizer on and submit for. If you download the sample_submission, you can iterate through the 'Image' column and use each of those images to run through your recognizer and put in the probabilities for submission.</p>",
      "rawMarkdown": "For the size of the annotations, it depends! Likely, you will be cropping out the 'whale detected' part of the image. This essentially 'removes' water pixels/not useful information. So the only information your classifier/recognizer will need are the important parts that distinguish whales from each other. There is quite a bit of useful information on this in the sticky threads, but it seems that the heads of whales are the most important part as they have identifying marks (calosities?) on their heads. But you might be able to find useful information with the whole whale body.\r\n\r\nI've been unable to get MatPlot yet, so my help is going to be limited. However, when you say you fed in 500 annotated images, does that mean you also got negatives? I believe you will need to feed in 'negative' annotated images where you are cropping out a piece of background that doesn't have a whale in it. \r\n\r\nAs for which images to run your final classifier/recognizer on and submit for. If you download the sample_submission, you can iterate through the 'Image' column and use each of those images to run through your recognizer and put in the probabilities for submission.",
      "votes": null
    },
    {
      "id": "97977",
      "postDate": "11/02/2015 09:06:47",
      "content": "<p>Hi!</p>\n\n<p>I have a lot of interest on learning in this field. I've been reading the questions of Travis and it will be great if somebody can give some answers.</p>\n\n<p>I think that we need to give a solution to the 11,468 images.</p>",
      "rawMarkdown": "Hi!\r\n\r\nI have a lot of interest on learning in this field. I've been reading the questions of Travis and it will be great if somebody can give some answers.\r\n\r\nI think that we need to give a solution to the 11,468 images.",
      "votes": null
    },
    {
      "id": "98848",
      "postDate": "11/15/2015 02:20:39",
      "content": "<p>hi,\nmy question is how we add the machine learning algorithm to these codes?</p>\n\n<p>and i think Travis's guess about bbox is right</p>",
      "rawMarkdown": "hi,\r\nmy question is how we add the machine learning algorithm to these codes?\r\n\r\nand i think Travis's guess about bbox is right",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 95761,
      "author_name": "mrtwiggy",
      "author_url": "",
      "post_date": "10/11/2015 15:33:54",
      "content": "<p>For the size of the annotations, it depends! Likely, you will be cropping out the 'whale detected' part of the image. This essentially 'removes' water pixels/not useful information. So the only information your classifier/recognizer will need are the important parts that distinguish whales from each other. There is quite a bit of useful information on this in the sticky threads, but it seems that the heads of whales are the most important part as they have identifying marks (calosities?) on their heads. But you might be able to find useful information with the whole whale body.</p>\n\n<p>I've been unable to get MatPlot yet, so my help is going to be limited. However, when you say you fed in 500 annotated images, does that mean you also got negatives? I believe you will need to feed in 'negative' annotated images where you are cropping out a piece of background that doesn't have a whale in it. </p>\n\n<p>As for which images to run your final classifier/recognizer on and submit for. If you download the sample_submission, you can iterate through the 'Image' column and use each of those images to run through your recognizer and put in the probabilities for submission.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 97977,
      "author_name": "pabloalicante",
      "author_url": "",
      "post_date": "11/02/2015 09:06:47",
      "content": "<p>Hi!</p>\n\n<p>I have a lot of interest on learning in this field. I've been reading the questions of Travis and it will be great if somebody can give some answers.</p>\n\n<p>I think that we need to give a solution to the 11,468 images.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 98848,
      "author_name": "poizona",
      "author_url": "",
      "post_date": "11/15/2015 02:20:39",
      "content": "<p>hi,\nmy question is how we add the machine learning algorithm to these codes?</p>\n\n<p>and i think Travis's guess about bbox is right</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "95757": "Hi everyone, \r\n\r\nI stumbled across this project a few days ago and I find it really fascinating. It's extra special for me because my wife and I always make sure to visit the New England Aquarium when we are in the Boston area. I've been a computer programmer for 18+ years working on mostly large web applications, but have found the concept of data science to be fascinating and something I want to explore further. \r\n\r\nThis is my first image processing and machine learning project, so please bear with me. I have a few questions: \r\n\r\nFor Step 1 in the \"Creating a \"Face Detector\" for Whales\" - for annotating the images, how big / small or broad / precise should the annotations be ? Should they include just the whale's head? Or the whole whale body? I annotated the 500 images from the subset, however, my detector isn't detecting any whales when I feed it an image from the regular set - or it detects three+ random boxes on the page. Did I use too many images when I built my classifier? I'll attach two examples of how I annotated images from the initial subset.  \r\n\r\nI'm a little concerned by this output when I run trainCascadeObjectDetector: \r\n\r\n*Automatically setting ObjectTrainingSize to [ 32, 41 ]*\r\n\r\nFrom the documentation:\r\n\r\n*Before training, the function resizes the positive and negative samples to ObjectTrainingSize in pixels.*\r\n\r\nThis doesn't sound good at all. \r\n\r\n *For optimal detection accuracy, specify an object training size close to the expected size of the object in the image.*\r\n\r\nI'm confused about this as per my previous question, should we be limiting the size and what we annotate when we annotate a whale in the training set? For instance, just annotate a 100x100 square around the head of the whale? Maybe we should pass\r\n\r\n`'ObjectTrainingSize','Auto'`\r\n\r\nto trainCascadeObjectDetector ? \r\n\r\nSteps 1-4 seem pretty straight forward - thanks to Nina Chen, and then this is where I get a little confused. Now that we have a detector and have run it on the entire dataset. What does that return exactly? \r\n\r\n    bbox = step(WhaleDetectorMdl,img);\r\n\r\nseems to be a 3x4 matrix. Probably because I have 3 annotations on this given image, so assuming a correct image would have only 1 annotation it would be a 1x4 matrix. Would this be (x1, y1, height, width)? \r\n\r\nNow what do we do? :) I assume we pick a machine learning algorithm, and feed it both the training.csv file and corresponding set of images referenced by tranining.csv to tell the algorithm \"Hey, For image w_9450.jpg, you can find whale_67614 at [1000, 500, 1000, 100, 120] - now learn!\"\r\n\r\nShould we start by looking at Classification algorithms? \r\n\r\nAnd finally - last question, once we've built this beautiful learned machine, which images do we run it on to generate a solution to submit for evaluation? The entire 11,468 images? From train.csv we already know the solution to 4,528 of the images, do we include those in our dataset? And if so, do we override what the machine learning algorithm has told us about the known image, and put 1 for the corresponding whale and 0 for all of the others to improve our evaluation?",
    "95761": "For the size of the annotations, it depends! Likely, you will be cropping out the 'whale detected' part of the image. This essentially 'removes' water pixels/not useful information. So the only information your classifier/recognizer will need are the important parts that distinguish whales from each other. There is quite a bit of useful information on this in the sticky threads, but it seems that the heads of whales are the most important part as they have identifying marks (calosities?) on their heads. But you might be able to find useful information with the whole whale body.\r\n\r\nI've been unable to get MatPlot yet, so my help is going to be limited. However, when you say you fed in 500 annotated images, does that mean you also got negatives? I believe you will need to feed in 'negative' annotated images where you are cropping out a piece of background that doesn't have a whale in it. \r\n\r\nAs for which images to run your final classifier/recognizer on and submit for. If you download the sample_submission, you can iterate through the 'Image' column and use each of those images to run through your recognizer and put in the probabilities for submission.",
    "97977": "Hi!\r\n\r\nI have a lot of interest on learning in this field. I've been reading the questions of Travis and it will be great if somebody can give some answers.\r\n\r\nI think that we need to give a solution to the 11,468 images.",
    "98848": "hi,\r\nmy question is how we add the machine learning algorithm to these codes?\r\n\r\nand i think Travis's guess about bbox is right"
  },
  "source": "meta"
}