{
  "id": 31556,
  "title": "GIGO: a good pipeline starts with good data",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/31556",
  "author_name": "",
  "post_date": "2017-04-13T07:33:32.357279200Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I've just learned how to use deep networks and TF, and this is my first attempt at an image classification. As a newbie I cannot wonder at the poor quality of the data. Even a human classifier would have problems with such data, and I am certain that the accuracy on the labeled data is pretty low, way lower that it could and it should be achieved.</p>\n\n<p>Of course as a data science competition, the quality of the data is just part of the challenge. I am talking here from the perspective of the scientist that collects and uses the results from the data. After all, the winning solution will be data limited.</p>\n\n<p>Just a few suggestions:</p>\n\n<ul>\n<li>Fix the overexposure problem\n<img src=\"http://i.imgur.com/7Hf5lKP.png\" alt=\"enter image description here\" title=\"\">\nDo you see the big male in the bottom left corner? Me neither.</li>\n<li>Consistent and apropiate altitude\n<img src=\"http://i.imgur.com/s9hMXpv.png\" alt=\"enter image description here\" title=\"\">\nBoth these images are 100% crop. I am not sure how to decide between classes when the size plays a big role in classification. Sure a scaled down car still looks like a car, but also a scale model car looks like a car. If you have to decide between scale models cars and real cars you will have a bad time.\n<img src=\"http://i.imgur.com/HR0qPrc.jpg\" alt=\"enter image description here\" title=\"\"></li>\n<li>Use infrared cameras\n<img src=\"http://i.imgur.com/lyQxZQx.jpg\" alt=\"enter image description here\" title=\"\">\nNeed I say more?</li>\n</ul>",
  "messages": [
    {
      "id": "174920",
      "postDate": "04/13/2017 07:33:32",
      "content": "<p>I've just learned how to use deep networks and TF, and this is my first attempt at an image classification. As a newbie I cannot wonder at the poor quality of the data. Even a human classifier would have problems with such data, and I am certain that the accuracy on the labeled data is pretty low, way lower that it could and it should be achieved.</p>\n\n<p>Of course as a data science competition, the quality of the data is just part of the challenge. I am talking here from the perspective of the scientist that collects and uses the results from the data. After all, the winning solution will be data limited.</p>\n\n<p>Just a few suggestions:</p>\n\n<ul>\n<li>Fix the overexposure problem\n<img src=\"http://i.imgur.com/7Hf5lKP.png\" alt=\"enter image description here\" title=\"\">\nDo you see the big male in the bottom left corner? Me neither.</li>\n<li>Consistent and apropiate altitude\n<img src=\"http://i.imgur.com/s9hMXpv.png\" alt=\"enter image description here\" title=\"\">\nBoth these images are 100% crop. I am not sure how to decide between classes when the size plays a big role in classification. Sure a scaled down car still looks like a car, but also a scale model car looks like a car. If you have to decide between scale models cars and real cars you will have a bad time.\n<img src=\"http://i.imgur.com/HR0qPrc.jpg\" alt=\"enter image description here\" title=\"\"></li>\n<li>Use infrared cameras\n<img src=\"http://i.imgur.com/lyQxZQx.jpg\" alt=\"enter image description here\" title=\"\">\nNeed I say more?</li>\n</ul>",
      "rawMarkdown": "I've just learned how to use deep networks and TF, and this is my first attempt at an image classification. As a newbie I cannot wonder at the poor quality of the data. Even a human classifier would have problems with such data, and I am certain that the accuracy on the labeled data is pretty low, way lower that it could and it should be achieved.\n\nOf course as a data science competition, the quality of the data is just part of the challenge. I am talking here from the perspective of the scientist that collects and uses the results from the data. After all, the winning solution will be data limited.\n\nJust a few suggestions:\n\n - Fix the overexposure problem\n![enter image description here][1]\nDo you see the big male in the bottom left corner? Me neither.\n - Consistent and apropiate altitude\n![enter image description here][2]\nBoth these images are 100% crop. I am not sure how to decide between classes when the size plays a big role in classification. Sure a scaled down car still looks like a car, but also a scale model car looks like a car. If you have to decide between scale models cars and real cars you will have a bad time.\n![enter image description here][3]\n - Use infrared cameras\n![enter image description here][4]\nNeed I say more?\n\n\n  [1]: http://i.imgur.com/7Hf5lKP.png\n  [2]: http://i.imgur.com/s9hMXpv.png\n  [3]: http://i.imgur.com/HR0qPrc.jpg\n  [4]: http://i.imgur.com/lyQxZQx.jpg",
      "votes": null
    },
    {
      "id": "175033",
      "postDate": "04/13/2017 14:54:59",
      "content": "<p>Agree, I don't understand why they didn't use infrared cameras, it's obvious...If we don't succeed to analyse our images, we change the hardware/image acquisition. In our case this usually solves the problem...</p>",
      "rawMarkdown": "Agree, I don't understand why they didn't use infrared cameras, it's obvious...If we don't succeed to analyse our images, we change the hardware/image acquisition. In our case this usually solves the problem...",
      "votes": null
    },
    {
      "id": "175051",
      "postDate": "04/13/2017 16:02:37",
      "content": "<p>After the DSTL challenge, the idea of multi-band imaging really is useful =). However, keep in mind this is the NOAA we're talking about here--not exactly the most well funded government institution. We live in an era where a significant portion of the population believes climate change is a hoax invented by the Chinese, and that the NOAA has cooked their numbers. In essence, it's not like they're a for-profit tech company like Google or Microsoft. So we just gotta make due with what we have =). An important part of the job of a data scientist is data preparation.</p>\n\n<blockquote>\n  <p>Both these images are 100% crop. I am not sure how to decide between classes when the size plays a big role in classification. </p>\n</blockquote>\n\n<p>The same way you did it with the image of the man in the car above -- relative sizes =). One method: your CNN will have to have a receptive field large enough to be able to discern between the relative sizes of your sea lions. Another method: Post processing. If your cnn can produce pixel estimates of the volume or length of your sea lions, you can compare their sizes and use that in your false positive reduction pipeline to determine which ones are adult males vs subadult males, since those are the categories that are closes together. There are <a href=\"https://github.com/felixlaumon/deform-conv\">other tricks out there</a> as well.</p>\n\n<p>That said, I agree there are a bunch of issues with the data :P Based off the interaction with the contest organizers thus far, I'm confident they will continue to address at least a few of them.</p>",
      "rawMarkdown": "After the DSTL challenge, the idea of multi-band imaging really is useful =). However, keep in mind this is the NOAA we're talking about here--not exactly the most well funded government institution. We live in an era where a significant portion of the population believes climate change is a hoax invented by the Chinese, and that the NOAA has cooked their numbers. In essence, it's not like they're a for-profit tech company like Google or Microsoft. So we just gotta make due with what we have =). An important part of the job of a data scientist is data preparation.\n\n> Both these images are 100% crop. I am not sure how to decide between classes when the size plays a big role in classification. \n\nThe same way you did it with the image of the man in the car above -- relative sizes =). One method: your CNN will have to have a receptive field large enough to be able to discern between the relative sizes of your sea lions. Another method: Post processing. If your cnn can produce pixel estimates of the volume or length of your sea lions, you can compare their sizes and use that in your false positive reduction pipeline to determine which ones are adult males vs subadult males, since those are the categories that are closes together. There are [other tricks out there][1] as well.\n\nThat said, I agree there are a bunch of issues with the data :P Based off the interaction with the contest organizers thus far, I'm confident they will continue to address at least a few of them.\n\n\n  [1]: https://github.com/felixlaumon/deform-conv",
      "votes": null
    },
    {
      "id": "175245",
      "postDate": "04/14/2017 15:04:53",
      "content": "<p>I am working at integrating the relative size into the model. I am looking at this right now: <a href=\"http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.655.7926&amp;rep=rep1&amp;type=pdf\">http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.655.7926&amp;rep=rep1&amp;type=pdf</a></p>\n\n<p>I would love to hear different approaches.</p>",
      "rawMarkdown": "I am working at integrating the relative size into the model. I am looking at this right now: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.655.7926&rep=rep1&type=pdf\n\nI would love to hear different approaches.",
      "votes": null
    },
    {
      "id": "186368",
      "postDate": "05/27/2017 20:13:20",
      "content": "<p>This link is broken - did you get anything interesting out of this paper?</p>",
      "rawMarkdown": "This link is broken - did you get anything interesting out of this paper?",
      "votes": null
    },
    {
      "id": "186387",
      "postDate": "05/27/2017 22:05:23",
      "content": "<p>This was supposed to be my first image based competition. After many hours of work I officially exited the competition yesterday without submitting anything. I couldn't get anything out of the full images.</p>",
      "rawMarkdown": "This was supposed to be my first image based competition. After many hours of work I officially exited the competition yesterday without submitting anything. I couldn't get anything out of the full images.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 175033,
      "author_name": "snowformatics",
      "author_url": "",
      "post_date": "04/13/2017 14:54:59",
      "content": "<p>Agree, I don't understand why they didn't use infrared cameras, it's obvious...If we don't succeed to analyse our images, we change the hardware/image acquisition. In our case this usually solves the problem...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 175051,
      "author_name": "authman",
      "author_url": "",
      "post_date": "04/13/2017 16:02:37",
      "content": "<p>After the DSTL challenge, the idea of multi-band imaging really is useful =). However, keep in mind this is the NOAA we're talking about here--not exactly the most well funded government institution. We live in an era where a significant portion of the population believes climate change is a hoax invented by the Chinese, and that the NOAA has cooked their numbers. In essence, it's not like they're a for-profit tech company like Google or Microsoft. So we just gotta make due with what we have =). An important part of the job of a data scientist is data preparation.</p>\n\n<blockquote>\n  <p>Both these images are 100% crop. I am not sure how to decide between classes when the size plays a big role in classification. </p>\n</blockquote>\n\n<p>The same way you did it with the image of the man in the car above -- relative sizes =). One method: your CNN will have to have a receptive field large enough to be able to discern between the relative sizes of your sea lions. Another method: Post processing. If your cnn can produce pixel estimates of the volume or length of your sea lions, you can compare their sizes and use that in your false positive reduction pipeline to determine which ones are adult males vs subadult males, since those are the categories that are closes together. There are <a href=\"https://github.com/felixlaumon/deform-conv\">other tricks out there</a> as well.</p>\n\n<p>That said, I agree there are a bunch of issues with the data :P Based off the interaction with the contest organizers thus far, I'm confident they will continue to address at least a few of them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 175245,
          "author_name": "radustoicescu",
          "author_url": "",
          "post_date": "04/14/2017 15:04:53",
          "content": "<p>I am working at integrating the relative size into the model. I am looking at this right now: <a href=\"http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.655.7926&amp;rep=rep1&amp;type=pdf\">http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.655.7926&amp;rep=rep1&amp;type=pdf</a></p>\n\n<p>I would love to hear different approaches.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 186368,
          "author_name": "jinkos",
          "author_url": "",
          "post_date": "05/27/2017 20:13:20",
          "content": "<p>This link is broken - did you get anything interesting out of this paper?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 186387,
          "author_name": "radustoicescu",
          "author_url": "",
          "post_date": "05/27/2017 22:05:23",
          "content": "<p>This was supposed to be my first image based competition. After many hours of work I officially exited the competition yesterday without submitting anything. I couldn't get anything out of the full images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "174920": "I've just learned how to use deep networks and TF, and this is my first attempt at an image classification. As a newbie I cannot wonder at the poor quality of the data. Even a human classifier would have problems with such data, and I am certain that the accuracy on the labeled data is pretty low, way lower that it could and it should be achieved.\n\nOf course as a data science competition, the quality of the data is just part of the challenge. I am talking here from the perspective of the scientist that collects and uses the results from the data. After all, the winning solution will be data limited.\n\nJust a few suggestions:\n\n - Fix the overexposure problem\n![enter image description here][1]\nDo you see the big male in the bottom left corner? Me neither.\n - Consistent and apropiate altitude\n![enter image description here][2]\nBoth these images are 100% crop. I am not sure how to decide between classes when the size plays a big role in classification. Sure a scaled down car still looks like a car, but also a scale model car looks like a car. If you have to decide between scale models cars and real cars you will have a bad time.\n![enter image description here][3]\n - Use infrared cameras\n![enter image description here][4]\nNeed I say more?\n\n\n  [1]: http://i.imgur.com/7Hf5lKP.png\n  [2]: http://i.imgur.com/s9hMXpv.png\n  [3]: http://i.imgur.com/HR0qPrc.jpg\n  [4]: http://i.imgur.com/lyQxZQx.jpg",
    "175033": "Agree, I don't understand why they didn't use infrared cameras, it's obvious...If we don't succeed to analyse our images, we change the hardware/image acquisition. In our case this usually solves the problem...",
    "175051": "After the DSTL challenge, the idea of multi-band imaging really is useful =). However, keep in mind this is the NOAA we're talking about here--not exactly the most well funded government institution. We live in an era where a significant portion of the population believes climate change is a hoax invented by the Chinese, and that the NOAA has cooked their numbers. In essence, it's not like they're a for-profit tech company like Google or Microsoft. So we just gotta make due with what we have =). An important part of the job of a data scientist is data preparation.\n\n> Both these images are 100% crop. I am not sure how to decide between classes when the size plays a big role in classification. \n\nThe same way you did it with the image of the man in the car above -- relative sizes =). One method: your CNN will have to have a receptive field large enough to be able to discern between the relative sizes of your sea lions. Another method: Post processing. If your cnn can produce pixel estimates of the volume or length of your sea lions, you can compare their sizes and use that in your false positive reduction pipeline to determine which ones are adult males vs subadult males, since those are the categories that are closes together. There are [other tricks out there][1] as well.\n\nThat said, I agree there are a bunch of issues with the data :P Based off the interaction with the contest organizers thus far, I'm confident they will continue to address at least a few of them.\n\n\n  [1]: https://github.com/felixlaumon/deform-conv",
    "175245": "I am working at integrating the relative size into the model. I am looking at this right now: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.655.7926&rep=rep1&type=pdf\n\nI would love to hear different approaches.",
    "186368": "This link is broken - did you get anything interesting out of this paper?",
    "186387": "This was supposed to be my first image based competition. After many hours of work I officially exited the competition yesterday without submitting anything. I couldn't get anything out of the full images."
  },
  "source": "meta"
}