{
  "id": 16482,
  "title": "What if competition restarts?",
  "url": "/competitions/dato-native/discussion/16482",
  "author_name": "",
  "post_date": "2015-09-14T16:40:29.840Z",
  "votes": 4,
  "comment_count": 15,
  "views": 1715,
  "content": "<p>I want to ask in advance, Kaggle admins, what will happen with the pre-leak people scores if the competition restarts? I mean, there were (including me) people who worked a full month for it, its not like a 1-day leak or something....</p>",
  "messages": [
    {
      "id": "92511",
      "postDate": "09/14/2015 16:40:29",
      "content": "<p>I want to ask in advance, Kaggle admins, what will happen with the pre-leak people scores if the competition restarts? I mean, there were (including me) people who worked a full month for it, its not like a 1-day leak or something....</p>",
      "rawMarkdown": "I want to ask in advance, Kaggle admins, what will happen with the pre-leak people scores if the competition restarts? I mean, there were (including me) people who worked a full month for it, its not like a 1-day leak or something....",
      "votes": null
    },
    {
      "id": "92524",
      "postDate": "09/14/2015 17:35:40",
      "content": "<p>Totally agreed, some of us have spent a lot of time on this contest, only to see that a simple hack non-ML does much better. And we can not guarantee fairness with a randomly-reshuffled data either. I vote for a premature end to this contest (with an end date of 10th of Sept --- first instance of leak), and then start a brand new contest with new data. I have raised the same point in <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16449/can-lb-top-3-confirm-it-is-leakage-or-start-of-art-feature-engineering\">this thread</a></p>\n\n<p>Now that correct test labels are (almost) out, @NxGTR can you verify that your model also scores ~0.97 on entire test dataset?</p>",
      "rawMarkdown": "Totally agreed, some of us have spent a lot of time on this contest, only to see that a simple hack non-ML does much better. And we can not guarantee fairness with a randomly-reshuffled data either. I vote for a premature end to this contest (with an end date of 10th of Sept --- first instance of leak), and then start a brand new contest with new data. I have raised the same point in [this thread][1]\r\n\r\nNow that correct test labels are (almost) out, @NxGTR can you verify that your model also scores ~0.97 on entire test dataset?\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16449/can-lb-top-3-confirm-it-is-leakage-or-start-of-art-feature-engineering",
      "votes": null
    },
    {
      "id": "92533",
      "postDate": "09/14/2015 18:04:41",
      "content": "<p>Same here,  I devoted most of my time here in the last few weeks. I think the first one who found this leakage is worthy of reward, but I don't want this leakage to destroy all the players' hard work.</p>",
      "rawMarkdown": "Same here,  I devoted most of my time here in the last few weeks. I think the first one who found this leakage is worthy of reward, but I don't want this leakage to destroy all the players' hard work.",
      "votes": null
    },
    {
      "id": "92541",
      "postDate": "09/14/2015 18:23:54",
      "content": "<p>[quote=Sudeep Juvekar;92524]\nNow that correct test labels are (almost) out, @NxGTR can you verify that your model also scores ~0.97 on entire test dataset?\n[/quote]</p>\n\n<p>Oh, very interesting question, I haven't thought that, I can estimate the overall performance with the full dataset, but cant find out the private score. I will take a look.</p>",
      "rawMarkdown": "[quote=Sudeep Juvekar;92524]\r\nNow that correct test labels are (almost) out, @NxGTR can you verify that your model also scores ~0.97 on entire test dataset?\r\n[/quote]\r\n\r\nOh, very interesting question, I haven't thought that, I can estimate the overall performance with the full dataset, but cant find out the private score. I will take a look.",
      "votes": null
    },
    {
      "id": "92542",
      "postDate": "09/14/2015 18:26:57",
      "content": "<p>You can find Private LB score (or something close to it) as well, right? Suppose you create a 1/0 csv solution using leakage, called &quot;leakage.csv&quot;. Following simple code calculates auc_score on complete test dataset.</p>\n\n<pre><code>import pandas\nfrom sklearn import metrics\n\nleaked = pandas.read_csv(&quot;leakage.csv&quot;)\nmy_best = pandas.read_csv(&quot;my_best.csv&quot;)\nprint metrics.roc_auc_score(leaked.sponsored, my_best.sponsored)\n</code></pre>",
      "rawMarkdown": "You can find Private LB score (or something close to it) as well, right? Suppose you create a 1/0 csv solution using leakage, called \"leakage.csv\". Following simple code calculates auc_score on complete test dataset.\r\n\r\n    import pandas\r\n    from sklearn import metrics\r\n\r\n    leaked = pandas.read_csv(\"leakage.csv\")\r\n    my_best = pandas.read_csv(\"my_best.csv\")\r\n    print metrics.roc_auc_score(leaked.sponsored, my_best.sponsored)",
      "votes": null
    },
    {
      "id": "92546",
      "postDate": "09/14/2015 18:33:11",
      "content": "<p>nxgtr@feikyLinux:/stuff/Kaggle/Dato/CompareFull$ python comp.py </p>\n\n<p>0.974095222735</p>",
      "rawMarkdown": "nxgtr@feikyLinux:/stuff/Kaggle/Dato/CompareFull$ python comp.py \r\n\r\n0.974095222735",
      "votes": null
    },
    {
      "id": "92550",
      "postDate": "09/14/2015 18:46:02",
      "content": "<p>[quote=NxGTR;92546]</p>\n\n<p>nxgtr@feikyLinux:/stuff/Kaggle/Dato/CompareFull$ python comp.py </p>\n\n<p>0.974095222735</p>\n\n<p>[/quote]</p>\n\n<p>What a pity, they have decided to reset this contest.</p>",
      "rawMarkdown": "[quote=NxGTR;92546]\r\n\r\nnxgtr@feikyLinux:/stuff/Kaggle/Dato/CompareFull$ python comp.py \r\n\r\n0.974095222735\r\n\r\n[/quote]\r\n\r\nWhat a pity, they have decided to reset this contest.",
      "votes": null
    },
    {
      "id": "92551",
      "postDate": "09/14/2015 18:49:23",
      "content": "<p>[quote=Eric;92550]\nWhat a pity, they have decided to reset this contest.\n[/quote]</p>\n\n<p>Well, I think we all who worked 1 month in this will get a &quot;thank you for participate&quot; :D</p>",
      "rawMarkdown": "[quote=Eric;92550]\r\nWhat a pity, they have decided to reset this contest.\r\n[/quote]\r\n\r\nWell, I think we all who worked 1 month in this will get a \"thank you for participate\" :D",
      "votes": null
    },
    {
      "id": "92554",
      "postDate": "09/14/2015 18:53:50",
      "content": "<p>[quote=NxGTR;92551]</p>\n\n<p>[quote=Eric;92550]\nWhat a pity, they have decided to reset this contest.\n[/quote]</p>\n\n<p>Well, I think we all who worked 1 month in this will get a &quot;thank you for participate&quot; :D</p>\n\n<p>[/quote]</p>\n\n<p>Nah, I bet you will still get 0.97 for the new data set and resume the first place.</p>\n\n<p>unless, you haven't deleted the whole dato folder containing the source code, right?  </p>",
      "rawMarkdown": "[quote=NxGTR;92551]\r\n\r\n[quote=Eric;92550]\r\nWhat a pity, they have decided to reset this contest.\r\n[/quote]\r\n\r\nWell, I think we all who worked 1 month in this will get a \"thank you for participate\" :D\r\n\r\n[/quote]\r\n\r\nNah, I bet you will still get 0.97 for the new data set and resume the first place.\r\n\r\nunless, you haven't deleted the whole dato folder containing the source code, right?",
      "votes": null
    },
    {
      "id": "92555",
      "postDate": "09/14/2015 18:57:16",
      "content": "<p>[quote=rcarson;92554]\nunless, you haven't deleted the whole dato folder containing the source code, right? <br>\n[/quote]</p>\n\n<p>Or maybe I just hired some people to do manual tagging, damn, this will be more expensive than planned :/</p>",
      "rawMarkdown": "[quote=rcarson;92554]\r\nunless, you haven't deleted the whole dato folder containing the source code, right?  \r\n[/quote]\r\n\r\nOr maybe I just hired some people to do manual tagging, damn, this will be more expensive than planned :/",
      "votes": null
    },
    {
      "id": "92560",
      "postDate": "09/14/2015 19:08:03",
      "content": "<p>I wish they would just give us new test and train - that way your scripts should work the same, and everything is good.</p>\n\n<p>If they are only going to do the pseudo-reset, they should give NxGTR both prizes (assuming the GTC code scores the same), and the reset contest is just for points.</p>",
      "rawMarkdown": "I wish they would just give us new test and train - that way your scripts should work the same, and everything is good.\r\n\r\nIf they are only going to do the pseudo-reset, they should give NxGTR both prizes (assuming the GTC code scores the same), and the reset contest is just for points.",
      "votes": null
    },
    {
      "id": "92615",
      "postDate": "09/14/2015 22:23:26",
      "content": "<p>Looks like old test and train have been merged to create new training file.</p>",
      "rawMarkdown": "Looks like old test and train have been merged to create new training file.",
      "votes": null
    },
    {
      "id": "92630",
      "postDate": "09/15/2015 00:01:00",
      "content": "<p>Now it doesn't fit into memory! :)</p>",
      "rawMarkdown": "Now it doesn't fit into memory! :)",
      "votes": null
    },
    {
      "id": "92641",
      "postDate": "09/15/2015 02:24:52",
      "content": "<p>[quote=NxGTR;92555]</p>\n\n<p>[quote=rcarson;92554]\nunless, you haven't deleted the whole dato folder containing the source code, right? <br>\n[/quote]</p>\n\n<p>Or maybe I just hired some people to do manual tagging, damn, this will be more expensive than planned :/</p>\n\n<p>[/quote]</p>\n\n<p>our 0.94 model scores 0.96 now. I bet your workers will get 0.99 this time :P</p>",
      "rawMarkdown": "[quote=NxGTR;92555]\r\n\r\n[quote=rcarson;92554]\r\nunless, you haven't deleted the whole dato folder containing the source code, right?  \r\n[/quote]\r\n\r\nOr maybe I just hired some people to do manual tagging, damn, this will be more expensive than planned :/\r\n\r\n[/quote]\r\n\r\nour 0.94 model scores 0.96 now. I bet your workers will get 0.99 this time :P",
      "votes": null
    },
    {
      "id": "92643",
      "postDate": "09/15/2015 02:29:13",
      "content": "<p>[quote=rcarson;92641]\nour 0.94 model scores 0.96 now. I bet your workers will get 0.99 this time :P\n[/quote]</p>\n\n<p>I made a special <a href=\"https://www.kaggle.com/forums/f/15/kaggle-forum/t/16496/that-moment-when\">post</a>, just for you :D</p>",
      "rawMarkdown": "[quote=rcarson;92641]\r\nour 0.94 model scores 0.96 now. I bet your workers will get 0.99 this time :P\r\n[/quote]\r\n\r\nI made a special [post][1], just for you :D\r\n\r\n\r\n  [1]: https://www.kaggle.com/forums/f/15/kaggle-forum/t/16496/that-moment-when",
      "votes": null
    },
    {
      "id": "92672",
      "postDate": "09/15/2015 10:43:54",
      "content": "<p>The training set becomes so big</p>",
      "rawMarkdown": "The training set becomes so big",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 92524,
      "author_name": "sjuvekar",
      "author_url": "",
      "post_date": "09/14/2015 17:35:40",
      "content": "<p>Totally agreed, some of us have spent a lot of time on this contest, only to see that a simple hack non-ML does much better. And we can not guarantee fairness with a randomly-reshuffled data either. I vote for a premature end to this contest (with an end date of 10th of Sept --- first instance of leak), and then start a brand new contest with new data. I have raised the same point in <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16449/can-lb-top-3-confirm-it-is-leakage-or-start-of-art-feature-engineering\">this thread</a></p>\n\n<p>Now that correct test labels are (almost) out, @NxGTR can you verify that your model also scores ~0.97 on entire test dataset?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92533,
      "author_name": "chonglinsun",
      "author_url": "",
      "post_date": "09/14/2015 18:04:41",
      "content": "<p>Same here,  I devoted most of my time here in the last few weeks. I think the first one who found this leakage is worthy of reward, but I don't want this leakage to destroy all the players' hard work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92541,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "09/14/2015 18:23:54",
      "content": "<p>[quote=Sudeep Juvekar;92524]\nNow that correct test labels are (almost) out, @NxGTR can you verify that your model also scores ~0.97 on entire test dataset?\n[/quote]</p>\n\n<p>Oh, very interesting question, I haven't thought that, I can estimate the overall performance with the full dataset, but cant find out the private score. I will take a look.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92542,
      "author_name": "sjuvekar",
      "author_url": "",
      "post_date": "09/14/2015 18:26:57",
      "content": "<p>You can find Private LB score (or something close to it) as well, right? Suppose you create a 1/0 csv solution using leakage, called &quot;leakage.csv&quot;. Following simple code calculates auc_score on complete test dataset.</p>\n\n<pre><code>import pandas\nfrom sklearn import metrics\n\nleaked = pandas.read_csv(&quot;leakage.csv&quot;)\nmy_best = pandas.read_csv(&quot;my_best.csv&quot;)\nprint metrics.roc_auc_score(leaked.sponsored, my_best.sponsored)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92546,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "09/14/2015 18:33:11",
      "content": "<p>nxgtr@feikyLinux:/stuff/Kaggle/Dato/CompareFull$ python comp.py </p>\n\n<p>0.974095222735</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92550,
      "author_name": "chonglinsun",
      "author_url": "",
      "post_date": "09/14/2015 18:46:02",
      "content": "<p>[quote=NxGTR;92546]</p>\n\n<p>nxgtr@feikyLinux:/stuff/Kaggle/Dato/CompareFull$ python comp.py </p>\n\n<p>0.974095222735</p>\n\n<p>[/quote]</p>\n\n<p>What a pity, they have decided to reset this contest.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92551,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "09/14/2015 18:49:23",
      "content": "<p>[quote=Eric;92550]\nWhat a pity, they have decided to reset this contest.\n[/quote]</p>\n\n<p>Well, I think we all who worked 1 month in this will get a &quot;thank you for participate&quot; :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92554,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "09/14/2015 18:53:50",
      "content": "<p>[quote=NxGTR;92551]</p>\n\n<p>[quote=Eric;92550]\nWhat a pity, they have decided to reset this contest.\n[/quote]</p>\n\n<p>Well, I think we all who worked 1 month in this will get a &quot;thank you for participate&quot; :D</p>\n\n<p>[/quote]</p>\n\n<p>Nah, I bet you will still get 0.97 for the new data set and resume the first place.</p>\n\n<p>unless, you haven't deleted the whole dato folder containing the source code, right?  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92555,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "09/14/2015 18:57:16",
      "content": "<p>[quote=rcarson;92554]\nunless, you haven't deleted the whole dato folder containing the source code, right? <br>\n[/quote]</p>\n\n<p>Or maybe I just hired some people to do manual tagging, damn, this will be more expensive than planned :/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92560,
      "author_name": "omgponies",
      "author_url": "",
      "post_date": "09/14/2015 19:08:03",
      "content": "<p>I wish they would just give us new test and train - that way your scripts should work the same, and everything is good.</p>\n\n<p>If they are only going to do the pseudo-reset, they should give NxGTR both prizes (assuming the GTC code scores the same), and the reset contest is just for points.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92615,
      "author_name": "nigelcarpenter",
      "author_url": "",
      "post_date": "09/14/2015 22:23:26",
      "content": "<p>Looks like old test and train have been merged to create new training file.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92630,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "09/15/2015 00:01:00",
      "content": "<p>Now it doesn't fit into memory! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92641,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "09/15/2015 02:24:52",
      "content": "<p>[quote=NxGTR;92555]</p>\n\n<p>[quote=rcarson;92554]\nunless, you haven't deleted the whole dato folder containing the source code, right? <br>\n[/quote]</p>\n\n<p>Or maybe I just hired some people to do manual tagging, damn, this will be more expensive than planned :/</p>\n\n<p>[/quote]</p>\n\n<p>our 0.94 model scores 0.96 now. I bet your workers will get 0.99 this time :P</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92643,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "09/15/2015 02:29:13",
      "content": "<p>[quote=rcarson;92641]\nour 0.94 model scores 0.96 now. I bet your workers will get 0.99 this time :P\n[/quote]</p>\n\n<p>I made a special <a href=\"https://www.kaggle.com/forums/f/15/kaggle-forum/t/16496/that-moment-when\">post</a>, just for you :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92672,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "09/15/2015 10:43:54",
      "content": "<p>The training set becomes so big</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "92511": "I want to ask in advance, Kaggle admins, what will happen with the pre-leak people scores if the competition restarts? I mean, there were (including me) people who worked a full month for it, its not like a 1-day leak or something....",
    "92524": "Totally agreed, some of us have spent a lot of time on this contest, only to see that a simple hack non-ML does much better. And we can not guarantee fairness with a randomly-reshuffled data either. I vote for a premature end to this contest (with an end date of 10th of Sept --- first instance of leak), and then start a brand new contest with new data. I have raised the same point in [this thread][1]\r\n\r\nNow that correct test labels are (almost) out, @NxGTR can you verify that your model also scores ~0.97 on entire test dataset?\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16449/can-lb-top-3-confirm-it-is-leakage-or-start-of-art-feature-engineering",
    "92533": "Same here,  I devoted most of my time here in the last few weeks. I think the first one who found this leakage is worthy of reward, but I don't want this leakage to destroy all the players' hard work.",
    "92541": "[quote=Sudeep Juvekar;92524]\r\nNow that correct test labels are (almost) out, @NxGTR can you verify that your model also scores ~0.97 on entire test dataset?\r\n[/quote]\r\n\r\nOh, very interesting question, I haven't thought that, I can estimate the overall performance with the full dataset, but cant find out the private score. I will take a look.",
    "92542": "You can find Private LB score (or something close to it) as well, right? Suppose you create a 1/0 csv solution using leakage, called \"leakage.csv\". Following simple code calculates auc_score on complete test dataset.\r\n\r\n    import pandas\r\n    from sklearn import metrics\r\n\r\n    leaked = pandas.read_csv(\"leakage.csv\")\r\n    my_best = pandas.read_csv(\"my_best.csv\")\r\n    print metrics.roc_auc_score(leaked.sponsored, my_best.sponsored)",
    "92546": "nxgtr@feikyLinux:/stuff/Kaggle/Dato/CompareFull$ python comp.py \r\n\r\n0.974095222735",
    "92550": "[quote=NxGTR;92546]\r\n\r\nnxgtr@feikyLinux:/stuff/Kaggle/Dato/CompareFull$ python comp.py \r\n\r\n0.974095222735\r\n\r\n[/quote]\r\n\r\nWhat a pity, they have decided to reset this contest.",
    "92551": "[quote=Eric;92550]\r\nWhat a pity, they have decided to reset this contest.\r\n[/quote]\r\n\r\nWell, I think we all who worked 1 month in this will get a \"thank you for participate\" :D",
    "92554": "[quote=NxGTR;92551]\r\n\r\n[quote=Eric;92550]\r\nWhat a pity, they have decided to reset this contest.\r\n[/quote]\r\n\r\nWell, I think we all who worked 1 month in this will get a \"thank you for participate\" :D\r\n\r\n[/quote]\r\n\r\nNah, I bet you will still get 0.97 for the new data set and resume the first place.\r\n\r\nunless, you haven't deleted the whole dato folder containing the source code, right?",
    "92555": "[quote=rcarson;92554]\r\nunless, you haven't deleted the whole dato folder containing the source code, right?  \r\n[/quote]\r\n\r\nOr maybe I just hired some people to do manual tagging, damn, this will be more expensive than planned :/",
    "92560": "I wish they would just give us new test and train - that way your scripts should work the same, and everything is good.\r\n\r\nIf they are only going to do the pseudo-reset, they should give NxGTR both prizes (assuming the GTC code scores the same), and the reset contest is just for points.",
    "92615": "Looks like old test and train have been merged to create new training file.",
    "92630": "Now it doesn't fit into memory! :)",
    "92641": "[quote=NxGTR;92555]\r\n\r\n[quote=rcarson;92554]\r\nunless, you haven't deleted the whole dato folder containing the source code, right?  \r\n[/quote]\r\n\r\nOr maybe I just hired some people to do manual tagging, damn, this will be more expensive than planned :/\r\n\r\n[/quote]\r\n\r\nour 0.94 model scores 0.96 now. I bet your workers will get 0.99 this time :P",
    "92643": "[quote=rcarson;92641]\r\nour 0.94 model scores 0.96 now. I bet your workers will get 0.99 this time :P\r\n[/quote]\r\n\r\nI made a special [post][1], just for you :D\r\n\r\n\r\n  [1]: https://www.kaggle.com/forums/f/15/kaggle-forum/t/16496/that-moment-when",
    "92672": "The training set becomes so big"
  },
  "source": "meta"
}