{
  "id": 31424,
  "title": "Incorrect labels & numbers",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/31424",
  "author_name": "",
  "post_date": "2017-04-10T20:52:49.494500800Z",
  "votes": 35,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Dots are incorrect in some cases: 259 - red dot in the sea, 200 - a lot of pups marked with red dots, should be green, 941 -  a lot of sea lions wasn't marked.</p>\n\n<p>It looks like dots was counted manually and filled in train.csv In some cases it leads to dramatic errors, e.g. according to train.csv there are 5 juveniles on 593, in fact 58!!! I kindly ask organizers to check numbers of sea lions on private LB data as it could leads to inadequate competition results.</p>",
  "messages": [
    {
      "id": "174265",
      "postDate": "04/10/2017 20:52:49",
      "content": "<p>Dots are incorrect in some cases: 259 - red dot in the sea, 200 - a lot of pups marked with red dots, should be green, 941 -  a lot of sea lions wasn't marked.</p>\n\n<p>It looks like dots was counted manually and filled in train.csv In some cases it leads to dramatic errors, e.g. according to train.csv there are 5 juveniles on 593, in fact 58!!! I kindly ask organizers to check numbers of sea lions on private LB data as it could leads to inadequate competition results.</p>",
      "rawMarkdown": "Dots are incorrect in some cases: 259 - red dot in the sea, 200 - a lot of pups marked with red dots, should be green, 941 -  a lot of sea lions wasn't marked.\n\nIt looks like dots was counted manually and filled in train.csv In some cases it leads to dramatic errors, e.g. according to train.csv there are 5 juveniles on 593, in fact 58!!! I kindly ask organizers to check numbers of sea lions on private LB data as it could leads to inadequate competition results.",
      "votes": null
    },
    {
      "id": "174276",
      "postDate": "04/10/2017 21:37:59",
      "content": "<p>This is such a fun competition that I concur. It'd be unfortunate for someone to get 5th place, then a couple weeks later we find out they're actually 1st place, lol.</p>\n\n<p>Some recommendations:</p>\n\n<ol>\n<li>There are redacted portions in the traindotted images, as well as in the train images. Contest officials should announce publicly that sea lions located in redacted regions, irrespective if found in train <em>or</em> traindotted, should not count toward totals (e.g. not just redactions in traindotted).</li>\n<li>Since we are provided with ground truth train.csv files, a couple of people should post their actual computer-vision dot counts. Let's crowd source compare what we <em>believe</em> should be the correct answers for the training set so that we're all working towards an accurate target. This should be OK + sanctioned by current contest rules, since they already provide us with a train.csv file...</li>\n<li>After cross validating our counts (#2 above) against each other, whatever the determining kernel is should be open sourced + shared with the contest officials so that they can ensure accurate, automated counts for the private leader board as well.</li>\n</ol>\n\n<p>The above will help all participants + organizers tremendously ^_^;</p>\n\n<p>While it might not fix drawn issues, such as:</p>\n\n<blockquote>\n  <p>259.jpg - red dot in the sea\n  200.jpg - a lot of pups marked with red dots, should be green\n  941.jpg - a lot of sea lions wasn't marked</p>\n</blockquote>\n\n<p>It <em>will</em> at least correct current our train.csv and private leaderboard, so stuff like: </p>\n\n<blockquote>\n  <p>593.jpg - there are 5 juveniles on 593, in fact 58</p>\n</blockquote>\n\n<p>BTW, 259, 200, and 941 are Not currently in the list of 57 bad train ids, so the findings by @nekrasov are indeed novel.</p>",
      "rawMarkdown": "This is such a fun competition that I concur. It'd be unfortunate for someone to get 5th place, then a couple weeks later we find out they're actually 1st place, lol.\n\nSome recommendations:\n\n 1. There are redacted portions in the traindotted images, as well as in the train images. Contest officials should announce publicly that sea lions located in redacted regions, irrespective if found in train _or_ traindotted, should not count toward totals (e.g. not just redactions in traindotted).\n 2. Since we are provided with ground truth train.csv files, a couple of people should post their actual computer-vision dot counts. Let's crowd source compare what we _believe_ should be the correct answers for the training set so that we're all working towards an accurate target. This should be OK + sanctioned by current contest rules, since they already provide us with a train.csv file...\n 3. After cross validating our counts (#2 above) against each other, whatever the determining kernel is should be open sourced + shared with the contest officials so that they can ensure accurate, automated counts for the private leader board as well.\n\nThe above will help all participants + organizers tremendously ^_^;\n\nWhile it might not fix drawn issues, such as:\n\n> 259.jpg - red dot in the sea\n> 200.jpg - a lot of pups marked with red dots, should be green\n> 941.jpg - a lot of sea lions wasn't marked\n\nIt _will_ at least correct current our train.csv and private leaderboard, so stuff like: \n\n> 593.jpg - there are 5 juveniles on 593, in fact 58\n\nBTW, 259, 200, and 941 are Not currently in the list of 57 bad train ids, so the findings by @nekrasov are indeed novel.",
      "votes": null
    },
    {
      "id": "174401",
      "postDate": "04/11/2017 10:19:12",
      "content": "<p>I share the same opinion as @authman.</p>\n\n<p>From a machine learning point of view, actual images with colored point annotations are worthless. Not only this doubles the memory requirements of the training set (because for every image there is a copy with annotations), people need to spend time parsing the coordinates of these annotations instead of working on the real problem of detecting sea lions. The vast majority of current kernels and posts deal with the problem of recovering the true coordinates. In addition, noise introduced by an incorrect detection of the annotations will affect the performance of further learning algorithms.</p>\n\n<p>In my mind, the hosts should either release the true coordinates of the annotations in form of a text document, or publish a kernel to parse these coordinates from annotated images. For the latter, all competitors would work on the base of the same ground truth.</p>",
      "rawMarkdown": "I share the same opinion as @authman.\n\nFrom a machine learning point of view, actual images with colored point annotations are worthless. Not only this doubles the memory requirements of the training set (because for every image there is a copy with annotations), people need to spend time parsing the coordinates of these annotations instead of working on the real problem of detecting sea lions. The vast majority of current kernels and posts deal with the problem of recovering the true coordinates. In addition, noise introduced by an incorrect detection of the annotations will affect the performance of further learning algorithms.\n\nIn my mind, the hosts should either release the true coordinates of the annotations in form of a text document, or publish a kernel to parse these coordinates from annotated images. For the latter, all competitors would work on the base of the same ground truth.",
      "votes": null
    },
    {
      "id": "175142",
      "postDate": "04/14/2017 00:58:43",
      "content": "<p>I also agree with this. Any updates?</p>",
      "rawMarkdown": "I also agree with this. Any updates?",
      "votes": null
    },
    {
      "id": "175429",
      "postDate": "04/15/2017 10:42:10",
      "content": "<p>I found another incorrect. I can see many pups in <code>695.jpg</code>.<br>\nHowever, it is written that there are only 2 pups by <code>train.csv</code>.</p>",
      "rawMarkdown": "I found another incorrect. I can see many pups in `695.jpg`.<br>\nHowever, it is written that there are only 2 pups by `train.csv`.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 174276,
      "author_name": "authman",
      "author_url": "",
      "post_date": "04/10/2017 21:37:59",
      "content": "<p>This is such a fun competition that I concur. It'd be unfortunate for someone to get 5th place, then a couple weeks later we find out they're actually 1st place, lol.</p>\n\n<p>Some recommendations:</p>\n\n<ol>\n<li>There are redacted portions in the traindotted images, as well as in the train images. Contest officials should announce publicly that sea lions located in redacted regions, irrespective if found in train <em>or</em> traindotted, should not count toward totals (e.g. not just redactions in traindotted).</li>\n<li>Since we are provided with ground truth train.csv files, a couple of people should post their actual computer-vision dot counts. Let's crowd source compare what we <em>believe</em> should be the correct answers for the training set so that we're all working towards an accurate target. This should be OK + sanctioned by current contest rules, since they already provide us with a train.csv file...</li>\n<li>After cross validating our counts (#2 above) against each other, whatever the determining kernel is should be open sourced + shared with the contest officials so that they can ensure accurate, automated counts for the private leader board as well.</li>\n</ol>\n\n<p>The above will help all participants + organizers tremendously ^_^;</p>\n\n<p>While it might not fix drawn issues, such as:</p>\n\n<blockquote>\n  <p>259.jpg - red dot in the sea\n  200.jpg - a lot of pups marked with red dots, should be green\n  941.jpg - a lot of sea lions wasn't marked</p>\n</blockquote>\n\n<p>It <em>will</em> at least correct current our train.csv and private leaderboard, so stuff like: </p>\n\n<blockquote>\n  <p>593.jpg - there are 5 juveniles on 593, in fact 58</p>\n</blockquote>\n\n<p>BTW, 259, 200, and 941 are Not currently in the list of 57 bad train ids, so the findings by @nekrasov are indeed novel.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 174401,
      "author_name": "depthfirstsearch",
      "author_url": "",
      "post_date": "04/11/2017 10:19:12",
      "content": "<p>I share the same opinion as @authman.</p>\n\n<p>From a machine learning point of view, actual images with colored point annotations are worthless. Not only this doubles the memory requirements of the training set (because for every image there is a copy with annotations), people need to spend time parsing the coordinates of these annotations instead of working on the real problem of detecting sea lions. The vast majority of current kernels and posts deal with the problem of recovering the true coordinates. In addition, noise introduced by an incorrect detection of the annotations will affect the performance of further learning algorithms.</p>\n\n<p>In my mind, the hosts should either release the true coordinates of the annotations in form of a text document, or publish a kernel to parse these coordinates from annotated images. For the latter, all competitors would work on the base of the same ground truth.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 175142,
      "author_name": "subbytech",
      "author_url": "",
      "post_date": "04/14/2017 00:58:43",
      "content": "<p>I also agree with this. Any updates?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 175429,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "04/15/2017 10:42:10",
      "content": "<p>I found another incorrect. I can see many pups in <code>695.jpg</code>.<br>\nHowever, it is written that there are only 2 pups by <code>train.csv</code>.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "174265": "Dots are incorrect in some cases: 259 - red dot in the sea, 200 - a lot of pups marked with red dots, should be green, 941 -  a lot of sea lions wasn't marked.\n\nIt looks like dots was counted manually and filled in train.csv In some cases it leads to dramatic errors, e.g. according to train.csv there are 5 juveniles on 593, in fact 58!!! I kindly ask organizers to check numbers of sea lions on private LB data as it could leads to inadequate competition results.",
    "174276": "This is such a fun competition that I concur. It'd be unfortunate for someone to get 5th place, then a couple weeks later we find out they're actually 1st place, lol.\n\nSome recommendations:\n\n 1. There are redacted portions in the traindotted images, as well as in the train images. Contest officials should announce publicly that sea lions located in redacted regions, irrespective if found in train _or_ traindotted, should not count toward totals (e.g. not just redactions in traindotted).\n 2. Since we are provided with ground truth train.csv files, a couple of people should post their actual computer-vision dot counts. Let's crowd source compare what we _believe_ should be the correct answers for the training set so that we're all working towards an accurate target. This should be OK + sanctioned by current contest rules, since they already provide us with a train.csv file...\n 3. After cross validating our counts (#2 above) against each other, whatever the determining kernel is should be open sourced + shared with the contest officials so that they can ensure accurate, automated counts for the private leader board as well.\n\nThe above will help all participants + organizers tremendously ^_^;\n\nWhile it might not fix drawn issues, such as:\n\n> 259.jpg - red dot in the sea\n> 200.jpg - a lot of pups marked with red dots, should be green\n> 941.jpg - a lot of sea lions wasn't marked\n\nIt _will_ at least correct current our train.csv and private leaderboard, so stuff like: \n\n> 593.jpg - there are 5 juveniles on 593, in fact 58\n\nBTW, 259, 200, and 941 are Not currently in the list of 57 bad train ids, so the findings by @nekrasov are indeed novel.",
    "174401": "I share the same opinion as @authman.\n\nFrom a machine learning point of view, actual images with colored point annotations are worthless. Not only this doubles the memory requirements of the training set (because for every image there is a copy with annotations), people need to spend time parsing the coordinates of these annotations instead of working on the real problem of detecting sea lions. The vast majority of current kernels and posts deal with the problem of recovering the true coordinates. In addition, noise introduced by an incorrect detection of the annotations will affect the performance of further learning algorithms.\n\nIn my mind, the hosts should either release the true coordinates of the annotations in form of a text document, or publish a kernel to parse these coordinates from annotated images. For the latter, all competitors would work on the base of the same ground truth.",
    "175142": "I also agree with this. Any updates?",
    "175429": "I found another incorrect. I can see many pups in `695.jpg`.<br>\nHowever, it is written that there are only 2 pups by `train.csv`."
  },
  "source": "meta"
}