{
  "id": 30900,
  "title": "Approach to the problem",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/30900",
  "author_name": "",
  "post_date": "2017-03-30T21:30:03.388381500Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Clearly this is a unique competition, which is great, but also means I'm quite lost on how to start...is this a supervised learning problem? I believe it is, but I'm struggling to see how to approach it.</p>\n\n<p>Lets see, first, we are given several pictures with 5 types of predictors...we know how many of each predictor is in each training picture, but we don't know which is which (well, we do know looking at the dotted images, but I don't think the model should be trained looking at those images because the testing images wont have the colored dots) - right?</p>\n\n<p>Second, given, for example, image 1 - we have 2 adult_males, 20 subadult_males and 12 juveniles - thinking on how to translate this data into a \"data frame\", should we split the image in 3 and fit 3 models that predict each one of the classes?</p>\n\n<p>Or maybe, are we supposed to somehow, go trough the dotted images, find the colored dots, get the coordinates around that dot, cut the same image on the train folder using those coordinates, associate the correct class to each cut image, put all images on a data frame with its class and then fit a model on it?</p>\n\n<p>then go trough the same cutting process on the test set set and predict each one of the cut images...</p>\n\n<p>Im not sure if that makes any sense, what do you guys think?</p>",
  "messages": [
    {
      "id": "171643",
      "postDate": "03/30/2017 21:30:03",
      "content": "<p>Clearly this is a unique competition, which is great, but also means I'm quite lost on how to start...is this a supervised learning problem? I believe it is, but I'm struggling to see how to approach it.</p>\n\n<p>Lets see, first, we are given several pictures with 5 types of predictors...we know how many of each predictor is in each training picture, but we don't know which is which (well, we do know looking at the dotted images, but I don't think the model should be trained looking at those images because the testing images wont have the colored dots) - right?</p>\n\n<p>Second, given, for example, image 1 - we have 2 adult_males, 20 subadult_males and 12 juveniles - thinking on how to translate this data into a \"data frame\", should we split the image in 3 and fit 3 models that predict each one of the classes?</p>\n\n<p>Or maybe, are we supposed to somehow, go trough the dotted images, find the colored dots, get the coordinates around that dot, cut the same image on the train folder using those coordinates, associate the correct class to each cut image, put all images on a data frame with its class and then fit a model on it?</p>\n\n<p>then go trough the same cutting process on the test set set and predict each one of the cut images...</p>\n\n<p>Im not sure if that makes any sense, what do you guys think?</p>",
      "rawMarkdown": "Clearly this is a unique competition, which is great, but also means I'm quite lost on how to start...is this a supervised learning problem? I believe it is, but I'm struggling to see how to approach it.\n\nLets see, first, we are given several pictures with 5 types of predictors...we know how many of each predictor is in each training picture, but we don't know which is which (well, we do know looking at the dotted images, but I don't think the model should be trained looking at those images because the testing images wont have the colored dots) - right?\n\nSecond, given, for example, image 1 - we have 2 adult_males, 20 subadult_males and 12 juveniles - thinking on how to translate this data into a \"data frame\", should we split the image in 3 and fit 3 models that predict each one of the classes?\n\nOr maybe, are we supposed to somehow, go trough the dotted images, find the colored dots, get the coordinates around that dot, cut the same image on the train folder using those coordinates, associate the correct class to each cut image, put all images on a data frame with its class and then fit a model on it?\n\nthen go trough the same cutting process on the test set set and predict each one of the cut images...\n\nIm not sure if that makes any sense, what do you guys think?",
      "votes": null
    },
    {
      "id": "171803",
      "postDate": "03/31/2017 14:37:54",
      "content": "<p>I think you're on the right track. Finding the dots and cutting out the images is the approach I've started with. It seems to me that they labeled these images with a simple paint program, and then resaved them as jpegs (definitely not ideal). So the first hurdle is being able to actually obtain the labeled data.</p>\n\n<p>This makes training further difficult since (unless our dot abstraction is perfect) we're starting out with slightly noisy data. Either way, I think extracting the dots can be done. I'm getting pretty good results on most images, but I still have a handful that are just downright failing.</p>",
      "rawMarkdown": "I think you're on the right track. Finding the dots and cutting out the images is the approach I've started with. It seems to me that they labeled these images with a simple paint program, and then resaved them as jpegs (definitely not ideal). So the first hurdle is being able to actually obtain the labeled data.\n\nThis makes training further difficult since (unless our dot abstraction is perfect) we're starting out with slightly noisy data. Either way, I think extracting the dots can be done. I'm getting pretty good results on most images, but I still have a handful that are just downright failing.",
      "votes": null
    },
    {
      "id": "171808",
      "postDate": "03/31/2017 14:40:52",
      "content": "<p>thanks for your reply...after I wrote that, I realized that we'll need to cut the testing images as well to isolate the lions and apply the model...but we, clearly, have no dots on the testing data. how should we approach them? </p>",
      "rawMarkdown": "thanks for your reply...after I wrote that, I realized that we'll need to cut the testing images as well to isolate the lions and apply the model...but we, clearly, have no dots on the testing data. how should we approach them?",
      "votes": null
    },
    {
      "id": "171933",
      "postDate": "04/01/2017 01:05:56",
      "content": "<p>The Training set is broken into two sets.  There is a dotted set showing the location of each lion, and then the exact same photo without the dots.</p>\n\n<p>My suggestion would be to use the dotted set to grab the coordinates, then use the coordinates to cut a small photo of a sea lion out of the non-dotted picture.  These photos then can be rotated, reflected, and transposed to create a fairly large training set of photos.</p>\n\n<p>My expectation is that the easiest approach would then be to use a CNN to pass a window over the photos and try to predict if you see each type of sea lion in the window.</p>",
      "rawMarkdown": "The Training set is broken into two sets.  There is a dotted set showing the location of each lion, and then the exact same photo without the dots.\n\nMy suggestion would be to use the dotted set to grab the coordinates, then use the coordinates to cut a small photo of a sea lion out of the non-dotted picture.  These photos then can be rotated, reflected, and transposed to create a fairly large training set of photos.\n\nMy expectation is that the easiest approach would then be to use a CNN to pass a window over the photos and try to predict if you see each type of sea lion in the window.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 171803,
      "author_name": "grimmace",
      "author_url": "",
      "post_date": "03/31/2017 14:37:54",
      "content": "<p>I think you're on the right track. Finding the dots and cutting out the images is the approach I've started with. It seems to me that they labeled these images with a simple paint program, and then resaved them as jpegs (definitely not ideal). So the first hurdle is being able to actually obtain the labeled data.</p>\n\n<p>This makes training further difficult since (unless our dot abstraction is perfect) we're starting out with slightly noisy data. Either way, I think extracting the dots can be done. I'm getting pretty good results on most images, but I still have a handful that are just downright failing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 171808,
          "author_name": "dmenin",
          "author_url": "",
          "post_date": "03/31/2017 14:40:52",
          "content": "<p>thanks for your reply...after I wrote that, I realized that we'll need to cut the testing images as well to isolate the lions and apply the model...but we, clearly, have no dots on the testing data. how should we approach them? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 171933,
          "author_name": "londondan",
          "author_url": "",
          "post_date": "04/01/2017 01:05:56",
          "content": "<p>The Training set is broken into two sets.  There is a dotted set showing the location of each lion, and then the exact same photo without the dots.</p>\n\n<p>My suggestion would be to use the dotted set to grab the coordinates, then use the coordinates to cut a small photo of a sea lion out of the non-dotted picture.  These photos then can be rotated, reflected, and transposed to create a fairly large training set of photos.</p>\n\n<p>My expectation is that the easiest approach would then be to use a CNN to pass a window over the photos and try to predict if you see each type of sea lion in the window.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "171643": "Clearly this is a unique competition, which is great, but also means I'm quite lost on how to start...is this a supervised learning problem? I believe it is, but I'm struggling to see how to approach it.\n\nLets see, first, we are given several pictures with 5 types of predictors...we know how many of each predictor is in each training picture, but we don't know which is which (well, we do know looking at the dotted images, but I don't think the model should be trained looking at those images because the testing images wont have the colored dots) - right?\n\nSecond, given, for example, image 1 - we have 2 adult_males, 20 subadult_males and 12 juveniles - thinking on how to translate this data into a \"data frame\", should we split the image in 3 and fit 3 models that predict each one of the classes?\n\nOr maybe, are we supposed to somehow, go trough the dotted images, find the colored dots, get the coordinates around that dot, cut the same image on the train folder using those coordinates, associate the correct class to each cut image, put all images on a data frame with its class and then fit a model on it?\n\nthen go trough the same cutting process on the test set set and predict each one of the cut images...\n\nIm not sure if that makes any sense, what do you guys think?",
    "171803": "I think you're on the right track. Finding the dots and cutting out the images is the approach I've started with. It seems to me that they labeled these images with a simple paint program, and then resaved them as jpegs (definitely not ideal). So the first hurdle is being able to actually obtain the labeled data.\n\nThis makes training further difficult since (unless our dot abstraction is perfect) we're starting out with slightly noisy data. Either way, I think extracting the dots can be done. I'm getting pretty good results on most images, but I still have a handful that are just downright failing.",
    "171808": "thanks for your reply...after I wrote that, I realized that we'll need to cut the testing images as well to isolate the lions and apply the model...but we, clearly, have no dots on the testing data. how should we approach them?",
    "171933": "The Training set is broken into two sets.  There is a dotted set showing the location of each lion, and then the exact same photo without the dots.\n\nMy suggestion would be to use the dotted set to grab the coordinates, then use the coordinates to cut a small photo of a sea lion out of the non-dotted picture.  These photos then can be rotated, reflected, and transposed to create a fairly large training set of photos.\n\nMy expectation is that the easiest approach would then be to use a CNN to pass a window over the photos and try to predict if you see each type of sea lion in the window."
  },
  "source": "meta"
}