{
  "id": 16485,
  "title": "Important: Data Leakage, and Reset of Competition",
  "url": "/competitions/dato-native/discussion/16485",
  "author_name": "",
  "post_date": "2015-09-14T17:54:07.750Z",
  "votes": 11,
  "comment_count": 12,
  "views": 2913,
  "content": "<p>Most of you are aware of this already: the data provided in this dataset had a hidden &quot;feature&quot; which is in date stamp of the files (a.k.a. <a href=\"https://www.kaggle.com/wiki/Leakage\">leakage</a>). Fortunately, we do have more new data to reset this competition. We will swap the test data with a new version, and invalidate all previous submissions on the leaderboard. We will make an announcement both on the forum and in email, for you to re-download the new dataset. </p>\n\n<p>We know that some of you will be frustrated with this information and apologize for the inconvenience. By resetting the competition, we hope to keep the competition fair and challenging. Thank you for your understanding and good work so far!</p>",
  "messages": [
    {
      "id": "92528",
      "postDate": "09/14/2015 17:54:07",
      "content": "<p>Most of you are aware of this already: the data provided in this dataset had a hidden &quot;feature&quot; which is in date stamp of the files (a.k.a. <a href=\"https://www.kaggle.com/wiki/Leakage\">leakage</a>). Fortunately, we do have more new data to reset this competition. We will swap the test data with a new version, and invalidate all previous submissions on the leaderboard. We will make an announcement both on the forum and in email, for you to re-download the new dataset. </p>\n\n<p>We know that some of you will be frustrated with this information and apologize for the inconvenience. By resetting the competition, we hope to keep the competition fair and challenging. Thank you for your understanding and good work so far!</p>",
      "rawMarkdown": "Most of you are aware of this already: the data provided in this dataset had a hidden \"feature\" which is in date stamp of the files (a.k.a. [leakage][1]). Fortunately, we do have more new data to reset this competition. We will swap the test data with a new version, and invalidate all previous submissions on the leaderboard. We will make an announcement both on the forum and in email, for you to re-download the new dataset. \r\n\r\nWe know that some of you will be frustrated with this information and apologize for the inconvenience. By resetting the competition, we hope to keep the competition fair and challenging. Thank you for your understanding and good work so far!\r\n\r\n  [1]: https://www.kaggle.com/wiki/Leakage",
      "votes": null
    },
    {
      "id": "92529",
      "postDate": "09/14/2015 17:56:12",
      "content": "<p>Great to hear! </p>\n\n<p>Will there be any extension to the deadlines?</p>",
      "rawMarkdown": "Great to hear! \r\n\r\nWill there be any extension to the deadlines?",
      "votes": null
    },
    {
      "id": "92530",
      "postDate": "09/14/2015 17:58:23",
      "content": "<p>Great, no need to fight against 0.9999. </p>",
      "rawMarkdown": "Great, no need to fight against 0.9999.",
      "votes": null
    },
    {
      "id": "92532",
      "postDate": "09/14/2015 18:01:53",
      "content": "<p>Will you be adding the existing test set to the training set? If you don't then some will have an unfair advantage given that they essentially have a training set that is more than twice the size (as others have pointed out to me that many people would cheat if given the chance).</p>",
      "rawMarkdown": "Will you be adding the existing test set to the training set? If you don't then some will have an unfair advantage given that they essentially have a training set that is more than twice the size (as others have pointed out to me that many people would cheat if given the chance).",
      "votes": null
    },
    {
      "id": "92537",
      "postDate": "09/14/2015 18:14:39",
      "content": "<p>I second the request for an extension of the competition deadline.</p>\n\n<p>Edit: and the addition of the leaky test data to the training set.</p>",
      "rawMarkdown": "I second the request for an extension of the competition deadline.\r\n\r\nEdit: and the addition of the leaky test data to the training set.",
      "votes": null
    },
    {
      "id": "92547",
      "postDate": "09/14/2015 18:38:19",
      "content": "<p>Edit - already asked above by David.</p>",
      "rawMarkdown": "Edit - already asked above by David.",
      "votes": null
    },
    {
      "id": "92553",
      "postDate": "09/14/2015 18:53:10",
      "content": "<p>I smell so much (potential) cheating already, and the competition has not even restarted. @David point is just too strong, and not the only way to cheat.</p>",
      "rawMarkdown": "I smell so much (potential) cheating already, and the competition has not even restarted. @David point is just too strong, and not the only way to cheat.",
      "votes": null
    },
    {
      "id": "92561",
      "postDate": "09/14/2015 19:09:48",
      "content": "<p>[quote=Gabriel Barello;92537]</p>\n\n<p>I second the request for an extension of the competition deadline.</p>\n\n<p>Edit: and the addition of the leaky test data to the training set.</p>\n\n<p>[/quote]\nI wasn't asking for an extension - just if there would be one.</p>\n\n<p>If they are only swapping out new test data, then there likely isn't a need for an extension (although it is also a crippled contest). But if they swap out train and test, then I actually would hope for more time to allow scripts to re-run on the new data.\n(lol, one of my scripts never finished on the old data - clearly great code on my part)</p>",
      "rawMarkdown": "[quote=Gabriel Barello;92537]\r\n\r\nI second the request for an extension of the competition deadline.\r\n\r\nEdit: and the addition of the leaky test data to the training set.\r\n\r\n[/quote]\r\nI wasn't asking for an extension - just if there would be one.\r\n\r\nIf they are only swapping out new test data, then there likely isn't a need for an extension (although it is also a crippled contest). But if they swap out train and test, then I actually would hope for more time to allow scripts to re-run on the new data.\r\n(lol, one of my scripts never finished on the old data - clearly great code on my part)",
      "votes": null
    },
    {
      "id": "92563",
      "postDate": "09/14/2015 19:12:43",
      "content": "<p>[quote=NxGTR;92553]</p>\n\n<p>I smell so much (potential) cheating already, and the competition has not even restarted. @David point is just too strong, and not the only way to cheat.</p>\n\n<p>[/quote]\nHmm. If kaggle combines the current train and test as the new train, and release the new test data. I don't see how you can cheat anymore. But we can apply the old codes to the new train directly. So I think it is a solution. Or is it? :-)</p>",
      "rawMarkdown": "[quote=NxGTR;92553]\r\n\r\nI smell so much (potential) cheating already, and the competition has not even restarted. @David point is just too strong, and not the only way to cheat.\r\n\r\n[/quote]\r\nHmm. If kaggle combines the current train and test as the new train, and release the new test data. I don't see how you can cheat anymore. But we can apply the old codes to the new train directly. So I think it is a solution. Or is it? :-)",
      "votes": null
    },
    {
      "id": "92564",
      "postDate": "09/14/2015 19:15:13",
      "content": "<p>It is probably for the best to combine the htmls in train.csv and sampleSubmission.csv the <strong>new</strong> train data so there will be no chance of cheating. Or there is other ways to cheat as well ? <br>\nEdit: while I am editing my post, @LittleBoat already made the same suggestion 1 mins ago :) </p>",
      "rawMarkdown": "It is probably for the best to combine the htmls in train.csv and sampleSubmission.csv the **new** train data so there will be no chance of cheating. Or there is other ways to cheat as well ?  \r\nEdit: while I am editing my post, @LittleBoat already made the same suggestion 1 mins ago :)",
      "votes": null
    },
    {
      "id": "92637",
      "postDate": "09/15/2015 01:32:37",
      "content": "<p>David has a valid point ! Dataset is not kosher, not yet !</p>",
      "rawMarkdown": "David has a valid point ! Dataset is not kosher, not yet !",
      "votes": null
    },
    {
      "id": "92638",
      "postDate": "09/15/2015 01:50:26",
      "content": "<p>[quote=David McGarry;92532]</p>\n\n<p>Will you be adding the existing test set to the training set? If you don't then some will have an unfair advantage given that they essentially have a training set that is more than twice the size (as others have pointed out to me that many people would cheat if given the chance).</p>\n\n<p>[/quote]</p>\n\n<p>If I read correctly the data page description, all 0 to 4 directories are now in train and a new test has been released. So it seems to me Kaggle did the right thing.</p>",
      "rawMarkdown": "[quote=David McGarry;92532]\r\n\r\nWill you be adding the existing test set to the training set? If you don't then some will have an unfair advantage given that they essentially have a training set that is more than twice the size (as others have pointed out to me that many people would cheat if given the chance).\r\n\r\n[/quote]\r\n\r\nIf I read correctly the data page description, all 0 to 4 directories are now in train and a new test has been released. So it seems to me Kaggle did the right thing.",
      "votes": null
    },
    {
      "id": "92658",
      "postDate": "09/15/2015 07:27:40",
      "content": "<p>To all of those who would like some better competition preparation, you are welcome to participate to this topic, I listed already a few things I observed before in competition but I have not the same Kaggle encyclopedic knowledge as some of you;): <br>\n<a href=\"https://www.kaggle.com/forums/f/15/kaggle-forum/t/16282/a-checklist-for-competition-preparation-dataset-approach\">https://www.kaggle.com/forums/f/15/kaggle-forum/t/16282/a-checklist-for-competition-preparation-dataset-approach</a></p>\n\n<p>(please note it is note a way of trolling Kaggle or to say that it is bad and so on. I just want to be constructive here, as I learn a lot from Kaggle and all of you, it is my way to repay the favor)</p>",
      "rawMarkdown": "To all of those who would like some better competition preparation, you are welcome to participate to this topic, I listed already a few things I observed before in competition but I have not the same Kaggle encyclopedic knowledge as some of you;):  \r\nhttps://www.kaggle.com/forums/f/15/kaggle-forum/t/16282/a-checklist-for-competition-preparation-dataset-approach\r\n\r\n(please note it is note a way of trolling Kaggle or to say that it is bad and so on. I just want to be constructive here, as I learn a lot from Kaggle and all of you, it is my way to repay the favor)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 92529,
      "author_name": "omgponies",
      "author_url": "",
      "post_date": "09/14/2015 17:56:12",
      "content": "<p>Great to hear! </p>\n\n<p>Will there be any extension to the deadlines?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92530,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "09/14/2015 17:58:23",
      "content": "<p>Great, no need to fight against 0.9999. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92532,
      "author_name": "dmcgarry",
      "author_url": "",
      "post_date": "09/14/2015 18:01:53",
      "content": "<p>Will you be adding the existing test set to the training set? If you don't then some will have an unfair advantage given that they essentially have a training set that is more than twice the size (as others have pointed out to me that many people would cheat if given the chance).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92537,
      "author_name": "gbrello",
      "author_url": "",
      "post_date": "09/14/2015 18:14:39",
      "content": "<p>I second the request for an extension of the competition deadline.</p>\n\n<p>Edit: and the addition of the leaky test data to the training set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92547,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "09/14/2015 18:38:19",
      "content": "<p>Edit - already asked above by David.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92553,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "09/14/2015 18:53:10",
      "content": "<p>I smell so much (potential) cheating already, and the competition has not even restarted. @David point is just too strong, and not the only way to cheat.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92561,
      "author_name": "omgponies",
      "author_url": "",
      "post_date": "09/14/2015 19:09:48",
      "content": "<p>[quote=Gabriel Barello;92537]</p>\n\n<p>I second the request for an extension of the competition deadline.</p>\n\n<p>Edit: and the addition of the leaky test data to the training set.</p>\n\n<p>[/quote]\nI wasn't asking for an extension - just if there would be one.</p>\n\n<p>If they are only swapping out new test data, then there likely isn't a need for an extension (although it is also a crippled contest). But if they swap out train and test, then I actually would hope for more time to allow scripts to re-run on the new data.\n(lol, one of my scripts never finished on the old data - clearly great code on my part)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92563,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "09/14/2015 19:12:43",
      "content": "<p>[quote=NxGTR;92553]</p>\n\n<p>I smell so much (potential) cheating already, and the competition has not even restarted. @David point is just too strong, and not the only way to cheat.</p>\n\n<p>[/quote]\nHmm. If kaggle combines the current train and test as the new train, and release the new test data. I don't see how you can cheat anymore. But we can apply the old codes to the new train directly. So I think it is a solution. Or is it? :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92564,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "09/14/2015 19:15:13",
      "content": "<p>It is probably for the best to combine the htmls in train.csv and sampleSubmission.csv the <strong>new</strong> train data so there will be no chance of cheating. Or there is other ways to cheat as well ? <br>\nEdit: while I am editing my post, @LittleBoat already made the same suggestion 1 mins ago :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92637,
      "author_name": "zsaurk",
      "author_url": "",
      "post_date": "09/15/2015 01:32:37",
      "content": "<p>David has a valid point ! Dataset is not kosher, not yet !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92638,
      "author_name": "pierregutierrez",
      "author_url": "",
      "post_date": "09/15/2015 01:50:26",
      "content": "<p>[quote=David McGarry;92532]</p>\n\n<p>Will you be adding the existing test set to the training set? If you don't then some will have an unfair advantage given that they essentially have a training set that is more than twice the size (as others have pointed out to me that many people would cheat if given the chance).</p>\n\n<p>[/quote]</p>\n\n<p>If I read correctly the data page description, all 0 to 4 directories are now in train and a new test has been released. So it seems to me Kaggle did the right thing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92658,
      "author_name": "wildwizard",
      "author_url": "",
      "post_date": "09/15/2015 07:27:40",
      "content": "<p>To all of those who would like some better competition preparation, you are welcome to participate to this topic, I listed already a few things I observed before in competition but I have not the same Kaggle encyclopedic knowledge as some of you;): <br>\n<a href=\"https://www.kaggle.com/forums/f/15/kaggle-forum/t/16282/a-checklist-for-competition-preparation-dataset-approach\">https://www.kaggle.com/forums/f/15/kaggle-forum/t/16282/a-checklist-for-competition-preparation-dataset-approach</a></p>\n\n<p>(please note it is note a way of trolling Kaggle or to say that it is bad and so on. I just want to be constructive here, as I learn a lot from Kaggle and all of you, it is my way to repay the favor)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "92528": "Most of you are aware of this already: the data provided in this dataset had a hidden \"feature\" which is in date stamp of the files (a.k.a. [leakage][1]). Fortunately, we do have more new data to reset this competition. We will swap the test data with a new version, and invalidate all previous submissions on the leaderboard. We will make an announcement both on the forum and in email, for you to re-download the new dataset. \r\n\r\nWe know that some of you will be frustrated with this information and apologize for the inconvenience. By resetting the competition, we hope to keep the competition fair and challenging. Thank you for your understanding and good work so far!\r\n\r\n  [1]: https://www.kaggle.com/wiki/Leakage",
    "92529": "Great to hear! \r\n\r\nWill there be any extension to the deadlines?",
    "92530": "Great, no need to fight against 0.9999.",
    "92532": "Will you be adding the existing test set to the training set? If you don't then some will have an unfair advantage given that they essentially have a training set that is more than twice the size (as others have pointed out to me that many people would cheat if given the chance).",
    "92537": "I second the request for an extension of the competition deadline.\r\n\r\nEdit: and the addition of the leaky test data to the training set.",
    "92547": "Edit - already asked above by David.",
    "92553": "I smell so much (potential) cheating already, and the competition has not even restarted. @David point is just too strong, and not the only way to cheat.",
    "92561": "[quote=Gabriel Barello;92537]\r\n\r\nI second the request for an extension of the competition deadline.\r\n\r\nEdit: and the addition of the leaky test data to the training set.\r\n\r\n[/quote]\r\nI wasn't asking for an extension - just if there would be one.\r\n\r\nIf they are only swapping out new test data, then there likely isn't a need for an extension (although it is also a crippled contest). But if they swap out train and test, then I actually would hope for more time to allow scripts to re-run on the new data.\r\n(lol, one of my scripts never finished on the old data - clearly great code on my part)",
    "92563": "[quote=NxGTR;92553]\r\n\r\nI smell so much (potential) cheating already, and the competition has not even restarted. @David point is just too strong, and not the only way to cheat.\r\n\r\n[/quote]\r\nHmm. If kaggle combines the current train and test as the new train, and release the new test data. I don't see how you can cheat anymore. But we can apply the old codes to the new train directly. So I think it is a solution. Or is it? :-)",
    "92564": "It is probably for the best to combine the htmls in train.csv and sampleSubmission.csv the **new** train data so there will be no chance of cheating. Or there is other ways to cheat as well ?  \r\nEdit: while I am editing my post, @LittleBoat already made the same suggestion 1 mins ago :)",
    "92637": "David has a valid point ! Dataset is not kosher, not yet !",
    "92638": "[quote=David McGarry;92532]\r\n\r\nWill you be adding the existing test set to the training set? If you don't then some will have an unfair advantage given that they essentially have a training set that is more than twice the size (as others have pointed out to me that many people would cheat if given the chance).\r\n\r\n[/quote]\r\n\r\nIf I read correctly the data page description, all 0 to 4 directories are now in train and a new test has been released. So it seems to me Kaggle did the right thing.",
    "92658": "To all of those who would like some better competition preparation, you are welcome to participate to this topic, I listed already a few things I observed before in competition but I have not the same Kaggle encyclopedic knowledge as some of you;):  \r\nhttps://www.kaggle.com/forums/f/15/kaggle-forum/t/16282/a-checklist-for-competition-preparation-dataset-approach\r\n\r\n(please note it is note a way of trolling Kaggle or to say that it is bad and so on. I just want to be constructive here, as I learn a lot from Kaggle and all of you, it is my way to repay the favor)"
  },
  "source": "meta"
}