{
  "id": 16359,
  "title": "What stops people from detecting whales by hand?",
  "url": "/competitions/noaa-right-whale-recognition/discussion/16359",
  "author_name": "",
  "post_date": "2015-09-07T05:42:33.767Z",
  "votes": null,
  "comment_count": 3,
  "views": 1713,
  "content": "<p>As far as I understood, a huge part of the competition is to build a system which will automatically detect the face of the whale on the picture. So what stops people from detecting the faces by hand (takes like 30 seconds per picture) and then feeding these nicely selected faces to an algorithm? It will not take a lot of time to do this, by outsourcing the work.</p>",
  "messages": [
    {
      "id": "91752",
      "postDate": "09/07/2015 05:42:33",
      "content": "<p>As far as I understood, a huge part of the competition is to build a system which will automatically detect the face of the whale on the picture. So what stops people from detecting the faces by hand (takes like 30 seconds per picture) and then feeding these nicely selected faces to an algorithm? It will not take a lot of time to do this, by outsourcing the work.</p>",
      "rawMarkdown": "As far as I understood, a huge part of the competition is to build a system which will automatically detect the face of the whale on the picture. So what stops people from detecting the faces by hand (takes like 30 seconds per picture) and then feeding these nicely selected faces to an algorithm? It will not take a lot of time to do this, by outsourcing the work.",
      "votes": null
    },
    {
      "id": "91754",
      "postDate": "09/07/2015 06:59:00",
      "content": "<p>[quote=SalvadorDali;91752]</p>\n\n<p>As far as I understood, a huge part of the competition is to build a system which will automatically detect the face of the whale on the picture. So what stops people from detecting the faces by hand (takes like 30 seconds per picture) and then feeding these nicely selected faces to an algorithm? It will not take a lot of time to do this, by outsourcing the work.</p>\n\n<p>[/quote]</p>\n\n<p>It would not be possible to publish that work and win the prize, since it would be clear that part of the pipeline was missing and the result could not be replicated without it.</p>\n\n<p>So the only purpose would be to get a higher leaderboard position.</p>\n\n<p>In addition to costing money, posts on crowd-sourcing sites are usually public, so there would be evidence of cheating which if discovered will get even the few kaggle rank/tier improvements possible with this exploit rendered null and void. Result: lost cash for the outsourcing, wasted time, nothing useful about ML learned by competing.</p>\n\n<p>You could I suppose outsource privately. In which case yes you might get away with it. It still does seem a lot of effort to go to for no financial reward.</p>",
      "rawMarkdown": "[quote=SalvadorDali;91752]\r\n\r\nAs far as I understood, a huge part of the competition is to build a system which will automatically detect the face of the whale on the picture. So what stops people from detecting the faces by hand (takes like 30 seconds per picture) and then feeding these nicely selected faces to an algorithm? It will not take a lot of time to do this, by outsourcing the work.\r\n\r\n[/quote]\r\n\r\nIt would not be possible to publish that work and win the prize, since it would be clear that part of the pipeline was missing and the result could not be replicated without it.\r\n\r\nSo the only purpose would be to get a higher leaderboard position.\r\n\r\nIn addition to costing money, posts on crowd-sourcing sites are usually public, so there would be evidence of cheating which if discovered will get even the few kaggle rank/tier improvements possible with this exploit rendered null and void. Result: lost cash for the outsourcing, wasted time, nothing useful about ML learned by competing.\r\n\r\nYou could I suppose outsource privately. In which case yes you might get away with it. It still does seem a lot of effort to go to for no financial reward.",
      "votes": null
    },
    {
      "id": "99678",
      "postDate": "11/28/2015 04:38:39",
      "content": "<p>That is a valid approach but seems there are hard limits on performance as well as high cost.</p>",
      "rawMarkdown": "That is a valid approach but seems there are hard limits on performance as well as high cost.",
      "votes": null
    },
    {
      "id": "100453",
      "postDate": "12/07/2015 17:10:11",
      "content": "<p>Salvador</p>\n\n<p>I think the answer is cost to scale.</p>\n\n<p>The task at hand is to take an image and match it against a reference database of say 5k known whales. So for each image an expert would need to spend 2 weeks to identify a particular whale. Expert would not be someone working on lower wages but some scientist who gets paid $1k per week to do accurately. Hence an accurate identification using humans would be prohibitively expensive.</p>\n\n<p>With an algorithm however you could do classify 1k target images in a matter of seconds for $1 variable cost. </p>\n\n<p>Let me provide an example. A relative of mine automated the fingerprint identification process for FBI using Machine Learning. Prior to that system FBI used to use a manual review for matching a target fingerprint against a reference data base of 2 million fingerprints. So the fingerprint analyst used to take an image and sift through tens of thousands of fingerprints to compare against. If you are the analyst you would not be able to do a good job. If you see fingerprints all day you can go pretty crazy. In any case it much quicker for you to spend a month doing nothing but reading magazines and sip coffee and come back and report a 'no match' without actually doing any efforts. </p>\n\n<p>Instead they a used Fourier transforms based fingerprint matching algorithm to automate this. Now the algo can generate a short list of handful of prints - ten or so for manual review. This is done in a matter of seconds. What a human could not do accurately in a month a computer can do in seconds. </p>\n\n<p>Again this is at a very small variable cost. The main cost is the fixed cost of man, data, and machine. They are trying to do a similar thing here for identifying whales instead of human fingerprints. </p>\n\n<p>Hope that helps.</p>",
      "rawMarkdown": "Salvador\r\n\r\nI think the answer is cost to scale.\r\n\r\nThe task at hand is to take an image and match it against a reference database of say 5k known whales. So for each image an expert would need to spend 2 weeks to identify a particular whale. Expert would not be someone working on lower wages but some scientist who gets paid $1k per week to do accurately. Hence an accurate identification using humans would be prohibitively expensive.\r\n\r\nWith an algorithm however you could do classify 1k target images in a matter of seconds for $1 variable cost. \r\n\r\nLet me provide an example. A relative of mine automated the fingerprint identification process for FBI using Machine Learning. Prior to that system FBI used to use a manual review for matching a target fingerprint against a reference data base of 2 million fingerprints. So the fingerprint analyst used to take an image and sift through tens of thousands of fingerprints to compare against. If you are the analyst you would not be able to do a good job. If you see fingerprints all day you can go pretty crazy. In any case it much quicker for you to spend a month doing nothing but reading magazines and sip coffee and come back and report a 'no match' without actually doing any efforts. \r\n\r\nInstead they a used Fourier transforms based fingerprint matching algorithm to automate this. Now the algo can generate a short list of handful of prints - ten or so for manual review. This is done in a matter of seconds. What a human could not do accurately in a month a computer can do in seconds. \r\n\r\nAgain this is at a very small variable cost. The main cost is the fixed cost of man, data, and machine. They are trying to do a similar thing here for identifying whales instead of human fingerprints. \r\n\r\nHope that helps.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 91754,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "09/07/2015 06:59:00",
      "content": "<p>[quote=SalvadorDali;91752]</p>\n\n<p>As far as I understood, a huge part of the competition is to build a system which will automatically detect the face of the whale on the picture. So what stops people from detecting the faces by hand (takes like 30 seconds per picture) and then feeding these nicely selected faces to an algorithm? It will not take a lot of time to do this, by outsourcing the work.</p>\n\n<p>[/quote]</p>\n\n<p>It would not be possible to publish that work and win the prize, since it would be clear that part of the pipeline was missing and the result could not be replicated without it.</p>\n\n<p>So the only purpose would be to get a higher leaderboard position.</p>\n\n<p>In addition to costing money, posts on crowd-sourcing sites are usually public, so there would be evidence of cheating which if discovered will get even the few kaggle rank/tier improvements possible with this exploit rendered null and void. Result: lost cash for the outsourcing, wasted time, nothing useful about ML learned by competing.</p>\n\n<p>You could I suppose outsource privately. In which case yes you might get away with it. It still does seem a lot of effort to go to for no financial reward.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 99678,
      "author_name": "michaeltconway",
      "author_url": "",
      "post_date": "11/28/2015 04:38:39",
      "content": "<p>That is a valid approach but seems there are hard limits on performance as well as high cost.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 100453,
      "author_name": "veerkar",
      "author_url": "",
      "post_date": "12/07/2015 17:10:11",
      "content": "<p>Salvador</p>\n\n<p>I think the answer is cost to scale.</p>\n\n<p>The task at hand is to take an image and match it against a reference database of say 5k known whales. So for each image an expert would need to spend 2 weeks to identify a particular whale. Expert would not be someone working on lower wages but some scientist who gets paid $1k per week to do accurately. Hence an accurate identification using humans would be prohibitively expensive.</p>\n\n<p>With an algorithm however you could do classify 1k target images in a matter of seconds for $1 variable cost. </p>\n\n<p>Let me provide an example. A relative of mine automated the fingerprint identification process for FBI using Machine Learning. Prior to that system FBI used to use a manual review for matching a target fingerprint against a reference data base of 2 million fingerprints. So the fingerprint analyst used to take an image and sift through tens of thousands of fingerprints to compare against. If you are the analyst you would not be able to do a good job. If you see fingerprints all day you can go pretty crazy. In any case it much quicker for you to spend a month doing nothing but reading magazines and sip coffee and come back and report a 'no match' without actually doing any efforts. </p>\n\n<p>Instead they a used Fourier transforms based fingerprint matching algorithm to automate this. Now the algo can generate a short list of handful of prints - ten or so for manual review. This is done in a matter of seconds. What a human could not do accurately in a month a computer can do in seconds. </p>\n\n<p>Again this is at a very small variable cost. The main cost is the fixed cost of man, data, and machine. They are trying to do a similar thing here for identifying whales instead of human fingerprints. </p>\n\n<p>Hope that helps.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "91752": "As far as I understood, a huge part of the competition is to build a system which will automatically detect the face of the whale on the picture. So what stops people from detecting the faces by hand (takes like 30 seconds per picture) and then feeding these nicely selected faces to an algorithm? It will not take a lot of time to do this, by outsourcing the work.",
    "91754": "[quote=SalvadorDali;91752]\r\n\r\nAs far as I understood, a huge part of the competition is to build a system which will automatically detect the face of the whale on the picture. So what stops people from detecting the faces by hand (takes like 30 seconds per picture) and then feeding these nicely selected faces to an algorithm? It will not take a lot of time to do this, by outsourcing the work.\r\n\r\n[/quote]\r\n\r\nIt would not be possible to publish that work and win the prize, since it would be clear that part of the pipeline was missing and the result could not be replicated without it.\r\n\r\nSo the only purpose would be to get a higher leaderboard position.\r\n\r\nIn addition to costing money, posts on crowd-sourcing sites are usually public, so there would be evidence of cheating which if discovered will get even the few kaggle rank/tier improvements possible with this exploit rendered null and void. Result: lost cash for the outsourcing, wasted time, nothing useful about ML learned by competing.\r\n\r\nYou could I suppose outsource privately. In which case yes you might get away with it. It still does seem a lot of effort to go to for no financial reward.",
    "99678": "That is a valid approach but seems there are hard limits on performance as well as high cost.",
    "100453": "Salvador\r\n\r\nI think the answer is cost to scale.\r\n\r\nThe task at hand is to take an image and match it against a reference database of say 5k known whales. So for each image an expert would need to spend 2 weeks to identify a particular whale. Expert would not be someone working on lower wages but some scientist who gets paid $1k per week to do accurately. Hence an accurate identification using humans would be prohibitively expensive.\r\n\r\nWith an algorithm however you could do classify 1k target images in a matter of seconds for $1 variable cost. \r\n\r\nLet me provide an example. A relative of mine automated the fingerprint identification process for FBI using Machine Learning. Prior to that system FBI used to use a manual review for matching a target fingerprint against a reference data base of 2 million fingerprints. So the fingerprint analyst used to take an image and sift through tens of thousands of fingerprints to compare against. If you are the analyst you would not be able to do a good job. If you see fingerprints all day you can go pretty crazy. In any case it much quicker for you to spend a month doing nothing but reading magazines and sip coffee and come back and report a 'no match' without actually doing any efforts. \r\n\r\nInstead they a used Fourier transforms based fingerprint matching algorithm to automate this. Now the algo can generate a short list of handful of prints - ten or so for manual review. This is done in a matter of seconds. What a human could not do accurately in a month a computer can do in seconds. \r\n\r\nAgain this is at a very small variable cost. The main cost is the fixed cost of man, data, and machine. They are trying to do a similar thing here for identifying whales instead of human fingerprints. \r\n\r\nHope that helps."
  },
  "source": "meta"
}