{
  "id": 31456,
  "title": "leader board mining",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/31456",
  "author_name": "",
  "post_date": "2017-04-11T13:25:58.463684200Z",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Will the organizers please clarify their policy on leader board mining?  Is anything going to be done to prevent leader board mining, or is this becoming a standard part of the game?  Please don't pretend that the problem doesn't exist.</p>",
  "messages": [
    {
      "id": "174426",
      "postDate": "04/11/2017 13:25:58",
      "content": "<p>Will the organizers please clarify their policy on leader board mining?  Is anything going to be done to prevent leader board mining, or is this becoming a standard part of the game?  Please don't pretend that the problem doesn't exist.</p>",
      "rawMarkdown": "Will the organizers please clarify their policy on leader board mining?  Is anything going to be done to prevent leader board mining, or is this becoming a standard part of the game?  Please don't pretend that the problem doesn't exist.",
      "votes": null
    },
    {
      "id": "174427",
      "postDate": "04/11/2017 13:29:48",
      "content": "<p>I would like a public position about the issue as well. Do I need to post a 0 LB entry to bring this up to light? I just need to upload it :-)</p>",
      "rawMarkdown": "I would like a public position about the issue as well. Do I need to post a 0 LB entry to bring this up to light? I just need to upload it :-)",
      "votes": null
    },
    {
      "id": "174432",
      "postDate": "04/11/2017 13:48:45",
      "content": "<p>@r4m0n You have my full support to share the LB ground truth.  It will save many man-hours of uncreative efforts which are clearly underway.</p>",
      "rawMarkdown": "r4m0n You have my full support to share the LB ground truth.  It will save many man-hours of uncreative efforts which are clearly underway.",
      "votes": null
    },
    {
      "id": "174539",
      "postDate": "04/11/2017 22:37:05",
      "content": "<p>Can you describe how that works? I can see one changing one prediction at a time, submitting 5 times a day will eventually reveal all the answers. But that will take forever to figure all answers. Is there is a smarter way?</p>",
      "rawMarkdown": "Can you describe how that works? I can see one changing one prediction at a time, submitting 5 times a day will eventually reveal all the answers. But that will take forever to figure all answers. Is there is a smarter way?",
      "votes": null
    },
    {
      "id": "174540",
      "postDate": "04/11/2017 22:40:00",
      "content": "<p>Check <a href=\"https://www.kaggle.com/olegtrott/data-science-bowl-2017/the-perfect-score-script\">this kernel by Oleg</a> for the general idea. I'll say the adaptation to multiclass logloss is reasonably trivial, and that I extracted the full truth with 86 submissions.</p>",
      "rawMarkdown": "Check [this kernel by Oleg][1] for the general idea. I'll say the adaptation to multiclass logloss is reasonably trivial, and that I extracted the full truth with 86 submissions.\n\n\n  [1]: https://www.kaggle.com/olegtrott/data-science-bowl-2017/the-perfect-score-script",
      "votes": null
    },
    {
      "id": "174575",
      "postDate": "04/12/2017 05:49:58",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "175534",
      "postDate": "04/16/2017 01:33:21",
      "content": "<p>Isn't the purpose of the second stage to prevent leader board mining?</p>",
      "rawMarkdown": "Isn't the purpose of the second stage to prevent leader board mining?",
      "votes": null
    },
    {
      "id": "175535",
      "postDate": "04/16/2017 01:35:45",
      "content": "<p>Mining doesn't affect the stage 2 and final LB,  but mining the stage 1 gives you 512 extra samples to validate with.</p>",
      "rawMarkdown": "Mining doesn't affect the stage 2 and final LB,  but mining the stage 1 gives you 512 extra samples to validate with.",
      "votes": null
    },
    {
      "id": "180651",
      "postDate": "05/06/2017 11:37:20",
      "content": "<p>here is my suggestion to solve the problem:</p>\n\n<ul>\n<li><p>the test dataset is divided into public leader board set and private leader board set.</p></li>\n<li><p>if you hack the public leader board set , you are going to get extra labels. Hence I think the organizer can disclosed the true labels  of this \" public leader board set \" in \"a some days before deadline\". If everyone can know the true labels, there is no advantage of hacking the system. So there are going to be three dataset:</p>\n\n<ul><li>Train set: for your own training and validation</li>\n<li>Public test set : for people to compare results  (and for some people who want to hacks) and test their algorithm. The ground truth of this dataset will be released in \"a some days before deadline\".</li>\n<li>Private test set: the real blackbox test set that is \"unknown\" to anyone</li></ul></li>\n</ul>",
      "rawMarkdown": "here is my suggestion to solve the problem:\n\n - the test dataset is divided into public leader board set and private leader board set.\n\n - if you hack the public leader board set , you are going to get extra labels. Hence I think the organizer can disclosed the true labels  of this \" public leader board set \" in \"a some days before deadline\". If everyone can know the true labels, there is no advantage of hacking the system. So there are going to be three dataset:\n\n     - Train set: for your own training and validation\n     - Public test set : for people to compare results  (and for some people who want to hacks) and test their algorithm. The ground truth of this dataset will be released in \"a some days before deadline\".\n     - Private test set: the real blackbox test set that is \"unknown\" to anyone",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 174427,
      "author_name": "mumech",
      "author_url": "",
      "post_date": "04/11/2017 13:29:48",
      "content": "<p>I would like a public position about the issue as well. Do I need to post a 0 LB entry to bring this up to light? I just need to upload it :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 174432,
      "author_name": "aaalgo",
      "author_url": "",
      "post_date": "04/11/2017 13:48:45",
      "content": "<p>@r4m0n You have my full support to share the LB ground truth.  It will save many man-hours of uncreative efforts which are clearly underway.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 174539,
      "author_name": "bastiaanbergman",
      "author_url": "",
      "post_date": "04/11/2017 22:37:05",
      "content": "<p>Can you describe how that works? I can see one changing one prediction at a time, submitting 5 times a day will eventually reveal all the answers. But that will take forever to figure all answers. Is there is a smarter way?</p>",
      "votes": null,
      "replies": [
        {
          "id": 174540,
          "author_name": "mumech",
          "author_url": "",
          "post_date": "04/11/2017 22:40:00",
          "content": "<p>Check <a href=\"https://www.kaggle.com/olegtrott/data-science-bowl-2017/the-perfect-score-script\">this kernel by Oleg</a> for the general idea. I'll say the adaptation to multiclass logloss is reasonably trivial, and that I extracted the full truth with 86 submissions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 174575,
      "author_name": "outlace",
      "author_url": "",
      "post_date": "04/12/2017 05:49:58",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 175534,
      "author_name": "matthewmasters",
      "author_url": "",
      "post_date": "04/16/2017 01:33:21",
      "content": "<p>Isn't the purpose of the second stage to prevent leader board mining?</p>",
      "votes": null,
      "replies": [
        {
          "id": 175535,
          "author_name": "mumech",
          "author_url": "",
          "post_date": "04/16/2017 01:35:45",
          "content": "<p>Mining doesn't affect the stage 2 and final LB,  but mining the stage 1 gives you 512 extra samples to validate with.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 180651,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/06/2017 11:37:20",
      "content": "<p>here is my suggestion to solve the problem:</p>\n\n<ul>\n<li><p>the test dataset is divided into public leader board set and private leader board set.</p></li>\n<li><p>if you hack the public leader board set , you are going to get extra labels. Hence I think the organizer can disclosed the true labels  of this \" public leader board set \" in \"a some days before deadline\". If everyone can know the true labels, there is no advantage of hacking the system. So there are going to be three dataset:</p>\n\n<ul><li>Train set: for your own training and validation</li>\n<li>Public test set : for people to compare results  (and for some people who want to hacks) and test their algorithm. The ground truth of this dataset will be released in \"a some days before deadline\".</li>\n<li>Private test set: the real blackbox test set that is \"unknown\" to anyone</li></ul></li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "174426": "Will the organizers please clarify their policy on leader board mining?  Is anything going to be done to prevent leader board mining, or is this becoming a standard part of the game?  Please don't pretend that the problem doesn't exist.",
    "174427": "I would like a public position about the issue as well. Do I need to post a 0 LB entry to bring this up to light? I just need to upload it :-)",
    "174432": "r4m0n You have my full support to share the LB ground truth.  It will save many man-hours of uncreative efforts which are clearly underway.",
    "174539": "Can you describe how that works? I can see one changing one prediction at a time, submitting 5 times a day will eventually reveal all the answers. But that will take forever to figure all answers. Is there is a smarter way?",
    "174540": "Check [this kernel by Oleg][1] for the general idea. I'll say the adaptation to multiclass logloss is reasonably trivial, and that I extracted the full truth with 86 submissions.\n\n\n  [1]: https://www.kaggle.com/olegtrott/data-science-bowl-2017/the-perfect-score-script",
    "174575": "",
    "175534": "Isn't the purpose of the second stage to prevent leader board mining?",
    "175535": "Mining doesn't affect the stage 2 and final LB,  but mining the stage 1 gives you 512 extra samples to validate with.",
    "180651": "here is my suggestion to solve the problem:\n\n - the test dataset is divided into public leader board set and private leader board set.\n\n - if you hack the public leader board set , you are going to get extra labels. Hence I think the organizer can disclosed the true labels  of this \" public leader board set \" in \"a some days before deadline\". If everyone can know the true labels, there is no advantage of hacking the system. So there are going to be three dataset:\n\n     - Train set: for your own training and validation\n     - Public test set : for people to compare results  (and for some people who want to hacks) and test their algorithm. The ground truth of this dataset will be released in \"a some days before deadline\".\n     - Private test set: the real blackbox test set that is \"unknown\" to anyone"
  },
  "source": "meta"
}