{
  "id": 16449,
  "title": "Can LB top 3 confirm it is leakage or start of art feature engineering",
  "url": "/competitions/dato-native/discussion/16449",
  "author_name": "",
  "post_date": "2015-09-12T14:22:22.240Z",
  "votes": 5,
  "comment_count": 73,
  "views": 7254,
  "content": "<p>The near perfect result is pretty surprising to me, I  am just wondering if top three players are willing to tell it is label/tag leakage or not ?  I think there is no harm confirming this, after all, leakage is nothing related to violating competition rules.</p>\n\n<p>The dataset is quite interesting, but I don't want to spend time on a  competition with leakage again, the west nile virus one is a  nightmare for me, and the tube price too. </p>\n\n<p>Edit: sorry, just find a typo in the title, I guess there is no way to fix that lol</p>\n\n<p>Update: I didn't expect people to get so angry (a few at me, I noticed I got some downvote) and disappointed about the issue. I have mixed feelings for this. But I don't remember I got so upset by not knowing the leakage of tube price ID after someone point it out after that competition was <strong>over</strong>, that one has more than 1000 players, yet only a few complaints, so I guess fewer upset player. So now I understand why people don't want to come out and confirm anything, because they want people to be happy. Sorry if I made people disappointed or upset by posting this topic.</p>",
  "messages": [
    {
      "id": "92249",
      "postDate": "09/12/2015 14:22:22",
      "content": "<p>The near perfect result is pretty surprising to me, I  am just wondering if top three players are willing to tell it is label/tag leakage or not ?  I think there is no harm confirming this, after all, leakage is nothing related to violating competition rules.</p>\n\n<p>The dataset is quite interesting, but I don't want to spend time on a  competition with leakage again, the west nile virus one is a  nightmare for me, and the tube price too. </p>\n\n<p>Edit: sorry, just find a typo in the title, I guess there is no way to fix that lol</p>\n\n<p>Update: I didn't expect people to get so angry (a few at me, I noticed I got some downvote) and disappointed about the issue. I have mixed feelings for this. But I don't remember I got so upset by not knowing the leakage of tube price ID after someone point it out after that competition was <strong>over</strong>, that one has more than 1000 players, yet only a few complaints, so I guess fewer upset player. So now I understand why people don't want to come out and confirm anything, because they want people to be happy. Sorry if I made people disappointed or upset by posting this topic.</p>",
      "rawMarkdown": "The near perfect result is pretty surprising to me, I  am just wondering if top three players are willing to tell it is label/tag leakage or not ?  I think there is no harm confirming this, after all, leakage is nothing related to violating competition rules.\r\n\r\nThe dataset is quite interesting, but I don't want to spend time on a  competition with leakage again, the west nile virus one is a  nightmare for me, and the tube price too. \r\n\r\nEdit: sorry, just find a typo in the title, I guess there is no way to fix that lol\r\n\r\nUpdate: I didn't expect people to get so angry (a few at me, I noticed I got some downvote) and disappointed about the issue. I have mixed feelings for this. But I don't remember I got so upset by not knowing the leakage of tube price ID after someone point it out after that competition was **over**, that one has more than 1000 players, yet only a few complaints, so I guess fewer upset player. So now I understand why people don't want to come out and confirm anything, because they want people to be happy. Sorry if I made people disappointed or upset by posting this topic.",
      "votes": null
    },
    {
      "id": "92252",
      "postDate": "09/12/2015 15:00:08",
      "content": "<p>[quote=SkyLibrary;92249]</p>\n\n<p>The near perfect result is pretty surprising to me, I  am just wondering if top three players are willing to tell it is label/tag leakage or not ?  I think there is no harm confirming this, after all, leakage is nothing related to violating competition rules.</p>\n\n<p>The dataset is quite interesting, but I don't want to spend time on a  competition with leakage again, the west nile virus one is a  nightmare for me, and the tube price too. </p>\n\n<p>Edit: sorry, just find typo in title, guess this is no way to fix that lol</p>\n\n<p>[/quote]</p>\n\n<p>It would be unfair with the 2 above me if I tell you anything. Last time I posted high model in Otto competition and everyone was just behind me. Someone from top 10 even criticized me by messaging me personally that because of me they are losing rank on Leader board.  </p>\n\n<p>I hope you can understand that I don't want to be involve in controversy every time :)</p>",
      "rawMarkdown": "[quote=SkyLibrary;92249]\r\n\r\nThe near perfect result is pretty surprising to me, I  am just wondering if top three players are willing to tell it is label/tag leakage or not ?  I think there is no harm confirming this, after all, leakage is nothing related to violating competition rules.\r\n\r\nThe dataset is quite interesting, but I don't want to spend time on a  competition with leakage again, the west nile virus one is a  nightmare for me, and the tube price too. \r\n\r\nEdit: sorry, just find typo in title, guess this is no way to fix that lol\r\n\r\n[/quote]\r\n\r\nIt would be unfair with the 2 above me if I tell you anything. Last time I posted high model in Otto competition and everyone was just behind me. Someone from top 10 even criticized me by messaging me personally that because of me they are losing rank on Leader board.  \r\n\r\nI hope you can understand that I don't want to be involve in controversy every time :)",
      "votes": null
    },
    {
      "id": "92254",
      "postDate": "09/12/2015 15:30:35",
      "content": "<p>Just wait few weeks :)</p>",
      "rawMarkdown": "Just wait few weeks :)",
      "votes": null
    },
    {
      "id": "92257",
      "postDate": "09/12/2015 16:21:59",
      "content": "<p>It is pretty much a yes or no question, if you guys don't want to answer it, then I will assume the answer is yes.</p>",
      "rawMarkdown": "It is pretty much a yes or no question, if you guys don't want to answer it, then I will assume the answer is yes.",
      "votes": null
    },
    {
      "id": "92276",
      "postDate": "09/12/2015 19:50:26",
      "content": "<p>Why would they have to say anything? It is already a HUGE advantage for everybody to know that it is possible to achieve that score. I dropped the competition because I believe it is highly likely that this is an exercise in finding a deterministic rule in the dataset, and that treasure hunt is not worth my time. I could be wrong though. </p>",
      "rawMarkdown": "Why would they have to say anything? It is already a HUGE advantage for everybody to know that it is possible to achieve that score. I dropped the competition because I believe it is highly likely that this is an exercise in finding a deterministic rule in the dataset, and that treasure hunt is not worth my time. I could be wrong though.",
      "votes": null
    },
    {
      "id": "92280",
      "postDate": "09/12/2015 20:26:16",
      "content": "<p>I didn't say they should answer my question, but I have the right to post the question no matter they reply or not. If they do, it will confirm your 'highly likely', if not, then it is still highly likely :)</p>\n\n<p>[quote=Giulio;92276]</p>\n\n<p>Why would they have to say anything? It is already a HUGE advantage for everybody to know that it is possible to achieve that score. I dropped the competition because I believe it is highly likely that this is an exercise in finding a deterministic rule in the dataset, and that treasure hunt is not worth my time. I could be wrong though. </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I didn't say they should answer my question, but I have the right to post the question no matter they reply or not. If they do, it will confirm your 'highly likely', if not, then it is still highly likely :)\r\n\r\n[quote=Giulio;92276]\r\n\r\nWhy would they have to say anything? It is already a HUGE advantage for everybody to know that it is possible to achieve that score. I dropped the competition because I believe it is highly likely that this is an exercise in finding a deterministic rule in the dataset, and that treasure hunt is not worth my time. I could be wrong though. \r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "92285",
      "postDate": "09/12/2015 22:48:30",
      "content": "<p>My guess is:</p>\n\n<p>They are running an unsupervised algorithm over both training and test that generates relationships between elements in both datasets. Let's call this a vocabulary. This thus creates highly abstract relationships between all documents at word level, but the difference now is that this is across both train and test.</p>\n\n<p>Then for training, create (meaningful) document vectors from this vocabulary and start training to indicate sponsored / not sponsored. The difference is that you're not training on occurrences of individual features of a document (which can be meaningful or not), but you're training against the known vector space of all documents (with 1/3 of the data). </p>",
      "rawMarkdown": "My guess is:\r\n\r\nThey are running an unsupervised algorithm over both training and test that generates relationships between elements in both datasets. Let's call this a vocabulary. This thus creates highly abstract relationships between all documents at word level, but the difference now is that this is across both train and test.\r\n\r\nThen for training, create (meaningful) document vectors from this vocabulary and start training to indicate sponsored / not sponsored. The difference is that you're not training on occurrences of individual features of a document (which can be meaningful or not), but you're training against the known vector space of all documents (with 1/3 of the data).",
      "votes": null
    },
    {
      "id": "92294",
      "postDate": "09/13/2015 01:05:02",
      "content": "<p>[quote=DataGeek;92252]</p>\n\n<p>[quote=SkyLibrary;92249]</p>\n\n<p>The near perfect result is pretty surprising to me, I  am just wondering if top three players are willing to tell it is label/tag leakage or not ?  I think there is no harm confirming this, after all, leakage is nothing related to violating competition rules.</p>\n\n<p>The dataset is quite interesting, but I don't want to spend time on a  competition with leakage again, the west nile virus one is a  nightmare for me, and the tube price too. </p>\n\n<p>Edit: sorry, just find typo in title, guess this is no way to fix that lol</p>\n\n<p>[/quote]</p>\n\n<p>It would be unfair with the 2 above me if I tell you anything. Last time I posted high model in Otto competition and everyone was just behind me. Someone from top 10 even criticized me by messaging me personally that because of me they are losing rank on Leader board.  </p>\n\n<p>I hope you can understand that I don't want to be involve in controversy every time :)</p>\n\n<p>[/quote]</p>\n\n<p>I think more and more people are getting offended by this, and joining the 0.99 club. :)</p>",
      "rawMarkdown": "[quote=DataGeek;92252]\r\n\r\n[quote=SkyLibrary;92249]\r\n\r\nThe near perfect result is pretty surprising to me, I  am just wondering if top three players are willing to tell it is label/tag leakage or not ?  I think there is no harm confirming this, after all, leakage is nothing related to violating competition rules.\r\n\r\nThe dataset is quite interesting, but I don't want to spend time on a  competition with leakage again, the west nile virus one is a  nightmare for me, and the tube price too. \r\n\r\nEdit: sorry, just find typo in title, guess this is no way to fix that lol\r\n\r\n[/quote]\r\n\r\nIt would be unfair with the 2 above me if I tell you anything. Last time I posted high model in Otto competition and everyone was just behind me. Someone from top 10 even criticized me by messaging me personally that because of me they are losing rank on Leader board.  \r\n\r\nI hope you can understand that I don't want to be involve in controversy every time :)\r\n\r\n[/quote]\r\n\r\nI think more and more people are getting offended by this, and joining the 0.99 club. :)",
      "votes": null
    },
    {
      "id": "92299",
      "postDate": "09/13/2015 03:07:39",
      "content": "<p>After download the data and play with it for 5 hours, I am here to confirm the answer is yes.</p>",
      "rawMarkdown": "After download the data and play with it for 5 hours, I am here to confirm the answer is yes.",
      "votes": null
    },
    {
      "id": "92301",
      "postDate": "09/13/2015 04:32:35",
      "content": "<p>Well this is fun, and here I thought this would be the one competition that is actually like the old days since the data is too large for scripts. Silly me.</p>",
      "rawMarkdown": "Well this is fun, and here I thought this would be the one competition that is actually like the old days since the data is too large for scripts. Silly me.",
      "votes": null
    },
    {
      "id": "92303",
      "postDate": "09/13/2015 05:37:23",
      "content": "<p>Interesting. I wonder if Kaggle will put a hold on this or just let it be it.</p>",
      "rawMarkdown": "Interesting. I wonder if Kaggle will put a hold on this or just let it be it.",
      "votes": null
    },
    {
      "id": "92304",
      "postDate": "09/13/2015 05:55:33",
      "content": "<p>I am a noob in machine learning. I played with natural language processing algorithms. I achieved .76 score then I found that beside language processing, we can have better result if we used other features in the HTML files such as: javascript, meta tag attribute... </p>",
      "rawMarkdown": "I am a noob in machine learning. I played with natural language processing algorithms. I achieved .76 score then I found that beside language processing, we can have better result if we used other features in the HTML files such as: javascript, meta tag attribute...",
      "votes": null
    },
    {
      "id": "92305",
      "postDate": "09/13/2015 06:04:10",
      "content": "<p>So there is little space for the improvement after you applying the leakage information into your model.</p>",
      "rawMarkdown": "So there is little space for the improvement after you applying the leakage information into your model.",
      "votes": null
    },
    {
      "id": "92308",
      "postDate": "09/13/2015 06:46:58",
      "content": "<p>For those of you who think this competition doesn't make sense now -- it is really a whole different world in tuning your model, when you have one almost perfect predictor. Just try to go beyond the level this predictor gives you, it is challenging. </p>",
      "rawMarkdown": "For those of you who think this competition doesn't make sense now -- it is really a whole different world in tuning your model, when you have one almost perfect predictor. Just try to go beyond the level this predictor gives you, it is challenging.",
      "votes": null
    },
    {
      "id": "92313",
      "postDate": "09/13/2015 09:28:41",
      "content": "<p>[quote=Dmitry Ulyanov;92308]</p>\n\n<p>For those of you who think this competition doesn't make sense now -- it is really a whole different world in tuning your model, when you have one almost perfect predictor. Just try to go beyond the level this predictor gives you, it is challenging. </p>\n\n<p>[/quote]</p>\n\n<p>But now you must know this &quot;perfect predictor&quot;, which it's a leakage and not real data, in order to be successful in this competition.</p>\n\n<p>So if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.</p>",
      "rawMarkdown": "[quote=Dmitry Ulyanov;92308]\r\n\r\nFor those of you who think this competition doesn't make sense now -- it is really a whole different world in tuning your model, when you have one almost perfect predictor. Just try to go beyond the level this predictor gives you, it is challenging. \r\n\r\n[/quote]\r\n\r\nBut now you must know this \"perfect predictor\", which it's a leakage and not real data, in order to be successful in this competition.\r\n\r\nSo if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.",
      "votes": null
    },
    {
      "id": "92314",
      "postDate": "09/13/2015 09:29:01",
      "content": "<p>This is unfair to WNV or Cat competitors to say that leakage exploitation was the only determinant in victory. <br>\nMany did exploit it (even myself in WNV, not knowing it was, while in Cat I discarded it believing I would make a methodology mistake), and we did not all achieve top scores.</p>\n\n<p>Leakage exploitation still asks for talent and skills.\nAlso, nothing tells that what you call leakage, as I said before, is just a very strong association rule in the dataset and was present from the start.</p>",
      "rawMarkdown": "This is unfair to WNV or Cat competitors to say that leakage exploitation was the only determinant in victory.  \r\nMany did exploit it (even myself in WNV, not knowing it was, while in Cat I discarded it believing I would make a methodology mistake), and we did not all achieve top scores.\r\n\r\nLeakage exploitation still asks for talent and skills.\r\nAlso, nothing tells that what you call leakage, as I said before, is just a very strong association rule in the dataset and was present from the start.",
      "votes": null
    },
    {
      "id": "92321",
      "postDate": "09/13/2015 13:18:19",
      "content": "<p>My code that is looking at what I think is the leak is apparently very slow, lol.</p>",
      "rawMarkdown": "My code that is looking at what I think is the leak is apparently very slow, lol.",
      "votes": null
    },
    {
      "id": "92322",
      "postDate": "09/13/2015 13:27:26",
      "content": "<p>I thought that, since an ad is &quot;sponsored&quot;, you'd find some kind of link or script in the files across all the sponsored ones (tracking links), but it isn't as easy as that. Just filtering on _gac for example gave me 0.546 and non-sponsored ones contain the same thing.  But I reckon it must be somewhere along those lines...</p>",
      "rawMarkdown": "I thought that, since an ad is \"sponsored\", you'd find some kind of link or script in the files across all the sponsored ones (tracking links), but it isn't as easy as that. Just filtering on _gac for example gave me 0.546 and non-sponsored ones contain the same thing.  But I reckon it must be somewhere along those lines...",
      "votes": null
    },
    {
      "id": "92324",
      "postDate": "09/13/2015 13:52:50",
      "content": "<p>[quote=SkyLibrary;92299]</p>\n\n<p>After download the data and play with it for 5 hours, I am here to confirm the answer is yes.</p>\n\n<p>[/quote]\nSo this is a leakage, and you have found it?</p>",
      "rawMarkdown": "[quote=SkyLibrary;92299]\r\n\r\nAfter download the data and play with it for 5 hours, I am here to confirm the answer is yes.\r\n\r\n[/quote]\r\nSo this is a leakage, and you have found it?",
      "votes": null
    },
    {
      "id": "92326",
      "postDate": "09/13/2015 14:33:21",
      "content": "<p>[quote=clustifier;92313]</p>\n\n<p>So if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.</p>\n\n<p>[/quote]</p>\n\n<p>Even with this leak it is still fair for players.</p>\n\n<p>The real bad thing is that it is pointless now for anyone trying to win the Dato Creators Award with this leakage existing.</p>",
      "rawMarkdown": "[quote=clustifier;92313]\r\n\r\nSo if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.\r\n\r\n[/quote]\r\n\r\nEven with this leak it is still fair for players.\r\n\r\nThe real bad thing is that it is pointless now for anyone trying to win the Dato Creators Award with this leakage existing.",
      "votes": null
    },
    {
      "id": "92327",
      "postDate": "09/13/2015 14:38:01",
      "content": "<p>I am sure that there is a leakage that can be used as a &quot;golden feature&quot;, but only using this can not reach 1.0 on the public leader board (0.99805). Or, as pointed out on the other thread, there are some errors in this dataset so the &quot;golden feature&quot; is perfect but labels are wrong.</p>\n\n<p>I want to know what the admin is thinking about this since further explorer might be in vain.</p>",
      "rawMarkdown": "I am sure that there is a leakage that can be used as a \"golden feature\", but only using this can not reach 1.0 on the public leader board (0.99805). Or, as pointed out on the other thread, there are some errors in this dataset so the \"golden feature\" is perfect but labels are wrong.\r\n\r\nI want to know what the admin is thinking about this since further explorer might be in vain.",
      "votes": null
    },
    {
      "id": "92331",
      "postDate": "09/13/2015 14:59:08",
      "content": "<p>[quote=rcarson;92326]</p>\n\n<p>[quote=clustifier;92313]</p>\n\n<p>So if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.</p>\n\n<p>[/quote]</p>\n\n<p>Even with this leak it is still fair for players.</p>\n\n<p>The real bad thing is that it is pointless now for anyone trying to win the Dato Creators Award with this leakage existing.</p>\n\n<p>[/quote]</p>\n\n<p>Define &quot;fair&quot;. Sure, all players have the same opportunity to find and exploit leakage but this is supposed to be a machine learning competition. I would have never signed up if it was a race to find some mistake in the data. I'm now stuck in a demlina where I have to decide if I want to, (a) continue using machine learning to tackle a seemingly interesting problem, (b) give in and find the leakage or (c) say screw it, delete the large data files from my system and feel even worse about the post-scripts world of kaggle.</p>",
      "rawMarkdown": "[quote=rcarson;92326]\r\n\r\n[quote=clustifier;92313]\r\n\r\nSo if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.\r\n\r\n[/quote]\r\n\r\nEven with this leak it is still fair for players.\r\n\r\nThe real bad thing is that it is pointless now for anyone trying to win the Dato Creators Award with this leakage existing.\r\n\r\n\r\n[/quote]\r\n\r\nDefine \"fair\". Sure, all players have the same opportunity to find and exploit leakage but this is supposed to be a machine learning competition. I would have never signed up if it was a race to find some mistake in the data. I'm now stuck in a demlina where I have to decide if I want to, (a) continue using machine learning to tackle a seemingly interesting problem, (b) give in and find the leakage or (c) say screw it, delete the large data files from my system and feel even worse about the post-scripts world of kaggle.",
      "votes": null
    },
    {
      "id": "92332",
      "postDate": "09/13/2015 15:11:02",
      "content": "<p>I guess this has something to do with scripts inside html files.</p>\n\n<p>For example, the function of ads' charging. But I'm not familiar with javascript unfortunately, lol.</p>",
      "rawMarkdown": "I guess this has something to do with scripts inside html files.\r\n\r\nFor example, the function of ads' charging. But I'm not familiar with javascript unfortunately, lol.",
      "votes": null
    },
    {
      "id": "92333",
      "postDate": "09/13/2015 15:17:04",
      "content": "<p>Stumbleupon doesn't require any additions to html page.</p>",
      "rawMarkdown": "Stumbleupon doesn't require any additions to html page.",
      "votes": null
    },
    {
      "id": "92334",
      "postDate": "09/13/2015 15:23:01",
      "content": "<p>[quote=David McGarry;92331]</p>\n\n<p>[quote=rcarson;92326]</p>\n\n<p>[quote=clustifier;92313]</p>\n\n<p>So if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.</p>\n\n<p>[/quote]</p>\n\n<p>Even with this leak it is still fair for players.</p>\n\n<p>The real bad thing is that it is pointless now for anyone trying to win the Dato Creators Award with this leakage existing.</p>\n\n<p>[/quote]</p>\n\n<p>Define &quot;fair&quot;. Sure, all players have the same opportunity to find and exploit leakage but this is supposed to be a machine learning competition. I would have never signed up if it was a race to find some mistake in the data. I'm not stuck in a demlina where I have to decide if I want to, (a) continue using machine learning to tackle a seemingly interesting problem, (b) give in and find the leakage or (c) say screw it, delete the large data files from my system and feel even worse about the post-scripts world of kaggle.</p>\n\n<p>[/quote]</p>\n\n<p>On the contrary, I found it fun to look for the leak ( I haven't found it yet). </p>\n\n<p>A little bit off the topic, I think a great part of kaggle competitions is not about machine learning. It is more about data mining. Finding soft or strong leak (good or golden features) is usually more dominant, given there are so many good out-of-box tools like xgboost for the learning part.</p>\n\n<p>In my opinion, finding a leak is a great skill even in real world case. But I agree in this case, it is a mess that the leak is too strong..</p>",
      "rawMarkdown": "[quote=David McGarry;92331]\r\n\r\n[quote=rcarson;92326]\r\n\r\n[quote=clustifier;92313]\r\n\r\nSo if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.\r\n\r\n[/quote]\r\n\r\nEven with this leak it is still fair for players.\r\n\r\nThe real bad thing is that it is pointless now for anyone trying to win the Dato Creators Award with this leakage existing.\r\n\r\n\r\n[/quote]\r\n\r\nDefine \"fair\". Sure, all players have the same opportunity to find and exploit leakage but this is supposed to be a machine learning competition. I would have never signed up if it was a race to find some mistake in the data. I'm not stuck in a demlina where I have to decide if I want to, (a) continue using machine learning to tackle a seemingly interesting problem, (b) give in and find the leakage or (c) say screw it, delete the large data files from my system and feel even worse about the post-scripts world of kaggle.\r\n\r\n[/quote]\r\n\r\nOn the contrary, I found it fun to look for the leak ( I haven't found it yet). \r\n\r\nA little bit off the topic, I think a great part of kaggle competitions is not about machine learning. It is more about data mining. Finding soft or strong leak (good or golden features) is usually more dominant, given there are so many good out-of-box tools like xgboost for the learning part.\r\n\r\nIn my opinion, finding a leak is a great skill even in real world case. But I agree in this case, it is a mess that the leak is too strong..",
      "votes": null
    },
    {
      "id": "92336",
      "postDate": "09/13/2015 15:26:43",
      "content": "<p>I think i found a  leakage, but it doesn't improve my score. LOL.</p>",
      "rawMarkdown": "I think i found a  leakage, but it doesn't improve my score. LOL.",
      "votes": null
    },
    {
      "id": "92340",
      "postDate": "09/13/2015 16:35:16",
      "content": "<p>Ok, let's play this out as if it was a real project.</p>\n\n<p>Business guy: we think we could really benefit from being able to use ML to classify pages with native add content.</p>\n\n<p>Data Scientist: give me a week to look at the data and I will get back to you with a first evaluation of what I think is possible. I'll work with engineering to get some data.</p>\n\n<p>...a week later...</p>\n\n<p>(very excited, yet skeptical) Data Scientist: I think I found something in the data that allows a perfect prediction. This is XYZ feature.</p>\n\n<p>Business Guy: oh, well, that makes sense. Seems kind of obvious. Wonder what &quot;name of a peer/manager/director who is taking the heat&quot; thought when he came up with the idea of using ML for this... You know what, I don't think we need ML anymore.</p>\n\n<p>(disappointed) Data Scientist: well, I guess I'll go back to building that boring dashboard then. Too bad, I was so excited of finally using some of those ML skills I had learnt on Kaggle...</p>\n\n<p>My point, when a perfect predictor exists, there is no reason to use ML because it is not an ML problem. This should have never made it to be an ML project/competition. Now, will the finding of XYZ feature be of value to Dato? Who knows? I would be surprised if it does. Will it still be interesting to try to maximize performance on misslabeld observations? I mean, why not, it's just a problem like any other. But, wow, not the problem I thought I was signing up for.\nGlad I haven't made a submission yet...</p>",
      "rawMarkdown": "Ok, let's play this out as if it was a real project.\r\n\r\nBusiness guy: we think we could really benefit from being able to use ML to classify pages with native add content.\r\n\r\nData Scientist: give me a week to look at the data and I will get back to you with a first evaluation of what I think is possible. I'll work with engineering to get some data.\r\n\r\n...a week later...\r\n\r\n(very excited, yet skeptical) Data Scientist: I think I found something in the data that allows a perfect prediction. This is XYZ feature.\r\n\r\nBusiness Guy: oh, well, that makes sense. Seems kind of obvious. Wonder what \"name of a peer/manager/director who is taking the heat\" thought when he came up with the idea of using ML for this... You know what, I don't think we need ML anymore.\r\n\r\n(disappointed) Data Scientist: well, I guess I'll go back to building that boring dashboard then. Too bad, I was so excited of finally using some of those ML skills I had learnt on Kaggle...\r\n\r\n\r\nMy point, when a perfect predictor exists, there is no reason to use ML because it is not an ML problem. This should have never made it to be an ML project/competition. Now, will the finding of XYZ feature be of value to Dato? Who knows? I would be surprised if it does. Will it still be interesting to try to maximize performance on misslabeld observations? I mean, why not, it's just a problem like any other. But, wow, not the problem I thought I was signing up for.\r\nGlad I haven't made a submission yet...",
      "votes": null
    },
    {
      "id": "92341",
      "postDate": "09/13/2015 16:39:06",
      "content": "<p>[quote=David McGarry;92331]\n I'm now stuck in a demlina where I have to decide if I want to, (a) continue using machine learning to tackle a seemingly interesting problem, (b) give in and find the leakage or (c) say screw it, delete the large data files from my system and feel even worse about the post-scripts world of kaggle.</p>\n\n<p>[/quote]</p>\n\n<p>+1 for (c)</p>",
      "rawMarkdown": "[quote=David McGarry;92331]\r\n I'm now stuck in a demlina where I have to decide if I want to, (a) continue using machine learning to tackle a seemingly interesting problem, (b) give in and find the leakage or (c) say screw it, delete the large data files from my system and feel even worse about the post-scripts world of kaggle.\r\n\r\n[/quote]\r\n\r\n+1 for (c)",
      "votes": null
    },
    {
      "id": "92342",
      "postDate": "09/13/2015 16:41:24",
      "content": "<p>Giulio, it's much worse than that... much worse :)  one of those things: &quot;soon as you see it...&quot; </p>\n\n<p>What I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. I had been trying to improve from 0.951 before this without success. That by itself is already a very reasonable score I think. </p>",
      "rawMarkdown": "Giulio, it's much worse than that... much worse :)  one of those things: \"soon as you see it...\" \r\n\r\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. I had been trying to improve from 0.951 before this without success. That by itself is already a very reasonable score I think.",
      "votes": null
    },
    {
      "id": "92343",
      "postDate": "09/13/2015 16:44:13",
      "content": "<p>In the end, don't tell me it is a filename leak... ==</p>",
      "rawMarkdown": "In the end, don't tell me it is a filename leak... ==",
      "votes": null
    },
    {
      "id": "92344",
      "postDate": "09/13/2015 16:51:18",
      "content": "<p>[quote=Gerard Toonstra;92342]\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. [/quote]</p>\n\n<p>I am very sad for this competition... I can confirm 0.973x is achievable without losing my own standards.</p>",
      "rawMarkdown": "[quote=Gerard Toonstra;92342]\r\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. [/quote]\r\n\r\nI am very sad for this competition... I can confirm 0.973x is achievable without losing my own standards.",
      "votes": null
    },
    {
      "id": "92345",
      "postDate": "09/13/2015 16:56:13",
      "content": "<p>[quote=NxGTR;92344]</p>\n\n<p>[quote=Gerard Toonstra;92342]\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. [/quote]</p>\n\n<p>I am very sad for this competition... I can confirm 0.973x is achievable without losing my own standards.</p>\n\n<p>[/quote]</p>\n\n<p>I'm sad too,  especially because of time I spent on it in the last few weeks.</p>",
      "rawMarkdown": "[quote=NxGTR;92344]\r\n\r\n[quote=Gerard Toonstra;92342]\r\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. [/quote]\r\n\r\nI am very sad for this competition... I can confirm 0.973x is achievable without losing my own standards.\r\n\r\n[/quote]\r\n\r\nI'm sad too,  especially because of time I spent on it in the last few weeks.",
      "votes": null
    },
    {
      "id": "92348",
      "postDate": "09/13/2015 17:19:20",
      "content": "<p>[quote=NxGTR;92344]</p>\n\n<p>[quote=Gerard Toonstra;92342]\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. [/quote]</p>\n\n<p>I am very sad for this competition... I can confirm 0.973x is achievable without losing my own standards.</p>\n\n<p>[/quote]</p>\n\n<p>I'm looking forward to an explanation of your approach, parametrizations and feature choices, if you decide to write it up.</p>",
      "rawMarkdown": "[quote=NxGTR;92344]\r\n\r\n[quote=Gerard Toonstra;92342]\r\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. [/quote]\r\n\r\nI am very sad for this competition... I can confirm 0.973x is achievable without losing my own standards.\r\n\r\n[/quote]\r\n\r\nI'm looking forward to an explanation of your approach, parametrizations and feature choices, if you decide to write it up.",
      "votes": null
    },
    {
      "id": "92349",
      "postDate": "09/13/2015 19:07:04",
      "content": "<p>The date in the zip files looks suspect.</p>",
      "rawMarkdown": "The date in the zip files looks suspect.",
      "votes": null
    },
    {
      "id": "92350",
      "postDate": "09/13/2015 19:25:14",
      "content": "<p>yes, it looks like the 1's files were written only on June 24 or July 27th.</p>",
      "rawMarkdown": "yes, it looks like the 1's files were written only on June 24 or July 27th.",
      "votes": null
    },
    {
      "id": "92353",
      "postDate": "09/13/2015 19:55:58",
      "content": "<p>[quote=Lawrence Chernin;92350]</p>\n\n<p>yes, it looks like the 1's files were written only on June 24 or July 27th.</p>\n\n<p>[/quote]\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  </p>",
      "rawMarkdown": "[quote=Lawrence Chernin;92350]\r\n\r\nyes, it looks like the 1's files were written only on June 24 or July 27th.\r\n\r\n[/quote]\r\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.",
      "votes": null
    },
    {
      "id": "92355",
      "postDate": "09/13/2015 20:02:11",
      "content": "<p>[quote=Little Boat;92353]</p>\n\n<p>[quote=Lawrence Chernin;92350]</p>\n\n<p>yes, it looks like the 1's files were written only on June 24 or July 27th.</p>\n\n<p>[/quote]\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  </p>\n\n<p>[/quote]\nThere is nothing to remove :)</p>\n\n<blockquote>\n  <p>StumbleUpon Ads is a native advertising platform that does not use\n  traditional ad units such as banners or text links. This means that on\n  StumbleUpon your ad is, quite simply, your URL. Since your web page\n  itself is the ad unit, no additional creative assets or copy are\n  needed to drive traffic. We drive directly to your page when a user\n  clicks the Stumble button.</p>\n  \n  <p><a href=\"http://ads.stumbleupon.com/wp-content/uploads/2014/09/CampaignCreationGuide.pdf\">http://ads.stumbleupon.com/wp-content/uploads/2014/09/CampaignCreationGuide.pdf</a></p>\n</blockquote>",
      "rawMarkdown": "[quote=Little Boat;92353]\r\n\r\n[quote=Lawrence Chernin;92350]\r\n\r\nyes, it looks like the 1's files were written only on June 24 or July 27th.\r\n\r\n[/quote]\r\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  \r\n\r\n[/quote]\r\nThere is nothing to remove :)\r\n\r\n> StumbleUpon Ads is a native advertising platform that does not use\r\n> traditional ad units such as banners or text links. This means that on\r\n> StumbleUpon your ad is, quite simply, your URL. Since your web page\r\n> itself is the ad unit, no additional creative assets or copy are\r\n> needed to drive traffic. We drive directly to your page when a user\r\n> clicks the Stumble button.\r\n\r\n\r\n> http://ads.stumbleupon.com/wp-content/uploads/2014/09/CampaignCreationGuide.pdf",
      "votes": null
    },
    {
      "id": "92356",
      "postDate": "09/13/2015 20:05:35",
      "content": "<p>Hmm. Then I don't see how that would be the leakage. Anyway, if I am bored enough to find out if it is, I will post it here. But I am hoping someone else can do it for us. :-)</p>\n\n<p>[quote=Artem;92355]</p>\n\n<p>[quote=Little Boat;92353]</p>\n\n<p>[quote=Lawrence Chernin;92350]</p>\n\n<p>yes, it looks like the 1's files were written only on June 24 or July 27th.</p>\n\n<p>[/quote]\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  </p>\n\n<p>[/quote]\nThere is nothing to remove :)</p>\n\n<blockquote>\n  <p>StumbleUpon Ads is a native advertising platform that does not use\n  traditional ad units such as banners or text links. This means that on\n  StumbleUpon your ad is, quite simply, your URL. Since your web page\n  itself is the ad unit, no additional creative assets or copy are\n  needed to drive traffic. We drive directly to your page when a user\n  clicks the Stumble button.</p>\n</blockquote>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Hmm. Then I don't see how that would be the leakage. Anyway, if I am bored enough to find out if it is, I will post it here. But I am hoping someone else can do it for us. :-)\r\n\r\n[quote=Artem;92355]\r\n\r\n[quote=Little Boat;92353]\r\n\r\n[quote=Lawrence Chernin;92350]\r\n\r\nyes, it looks like the 1's files were written only on June 24 or July 27th.\r\n\r\n[/quote]\r\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  \r\n\r\n[/quote]\r\nThere is nothing to remove :)\r\n\r\n> StumbleUpon Ads is a native advertising platform that does not use\r\n> traditional ad units such as banners or text links. This means that on\r\n> StumbleUpon your ad is, quite simply, your URL. Since your web page\r\n> itself is the ad unit, no additional creative assets or copy are\r\n> needed to drive traffic. We drive directly to your page when a user\r\n> clicks the Stumble button.\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "92357",
      "postDate": "09/13/2015 20:11:46",
      "content": "<p>[quote=Little Boat;92356]</p>\n\n<p>Hmm. Then I don't see how that would be the leakage. Anyway, if I am bored enough to find out if it is, I will post it here. But I am hoping someone else can do it for us. :-)</p>\n\n<p>[quote=Artem;92355]</p>\n\n<p>[quote=Little Boat;92353]</p>\n\n<p>[quote=Lawrence Chernin;92350]</p>\n\n<p>yes, it looks like the 1's files were written only on June 24 or July 27th.</p>\n\n<p>[/quote]\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  </p>\n\n<p>[/quote]\nThere is nothing to remove :)</p>\n\n<blockquote>\n  <p>StumbleUpon Ads is a native advertising platform that does not use\n  traditional ad units such as banners or text links. This means that on\n  StumbleUpon your ad is, quite simply, your URL. Since your web page\n  itself is the ad unit, no additional creative assets or copy are\n  needed to drive traffic. We drive directly to your page when a user\n  clicks the Stumble button.</p>\n</blockquote>\n\n<p>[/quote]\n[/quote]\nI checked it locally and it looks like true.</p>\n\n<p>Also i found another &quot;leakage&quot; that can give 0.86+ without ml.</p>",
      "rawMarkdown": "[quote=Little Boat;92356]\r\n\r\nHmm. Then I don't see how that would be the leakage. Anyway, if I am bored enough to find out if it is, I will post it here. But I am hoping someone else can do it for us. :-)\r\n\r\n[quote=Artem;92355]\r\n\r\n[quote=Little Boat;92353]\r\n\r\n[quote=Lawrence Chernin;92350]\r\n\r\nyes, it looks like the 1's files were written only on June 24 or July 27th.\r\n\r\n[/quote]\r\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  \r\n\r\n[/quote]\r\nThere is nothing to remove :)\r\n\r\n> StumbleUpon Ads is a native advertising platform that does not use\r\n> traditional ad units such as banners or text links. This means that on\r\n> StumbleUpon your ad is, quite simply, your URL. Since your web page\r\n> itself is the ad unit, no additional creative assets or copy are\r\n> needed to drive traffic. We drive directly to your page when a user\r\n> clicks the Stumble button.\r\n\r\n[/quote]\r\n[/quote]\r\nI checked it locally and it looks like true.\r\n\r\n\r\nAlso i found another \"leakage\" that can give 0.86+ without ml.",
      "votes": null
    },
    {
      "id": "92358",
      "postDate": "09/13/2015 20:13:48",
      "content": "<p>There is definitely leakage and I don't see how this could be solved/repaired, short of giving us a completely new data set...</p>\n\n<p>What the final few % is I have no idea. Could be mislabelling. Could be LB noise. Could be the tiniest of edges we are now competing for.</p>",
      "rawMarkdown": "There is definitely leakage and I don't see how this could be solved/repaired, short of giving us a completely new data set...\r\n\r\nWhat the final few % is I have no idea. Could be mislabelling. Could be LB noise. Could be the tiniest of edges we are now competing for.",
      "votes": null
    },
    {
      "id": "92359",
      "postDate": "09/13/2015 20:14:50",
      "content": "<p>[quote=Lawrence Chernin;92350]</p>\n\n<p>yes, it looks like the 1's files were written only on June 24 or July 27th.</p>\n\n<p>[/quote]</p>\n\n<p>Well that's fun. Hopefully they clear the LB and release a new dataset (using new HTML files would be ideal, but at least renaming all of the existing HTML files would discourage all but the least honest to ignore the known values in the test set).</p>",
      "rawMarkdown": "[quote=Lawrence Chernin;92350]\r\n\r\nyes, it looks like the 1's files were written only on June 24 or July 27th.\r\n\r\n[/quote]\r\n\r\nWell that's fun. Hopefully they clear the LB and release a new dataset (using new HTML files would be ideal, but at least renaming all of the existing HTML files would discourage all but the least honest to ignore the known values in the test set).",
      "votes": null
    },
    {
      "id": "92362",
      "postDate": "09/13/2015 20:45:35",
      "content": "<p>Now I can confirm that the date is the leakage.</p>",
      "rawMarkdown": "Now I can confirm that the date is the leakage.",
      "votes": null
    },
    {
      "id": "92363",
      "postDate": "09/13/2015 20:46:47",
      "content": "<p>[quote=Triskelion;92358]</p>\n\n<p>What the final few % is I have no idea. Could be mislabelling. Could be LB noise. Could be the tiniest of edges we are now competing for.</p>\n\n<p>[/quote]</p>\n\n<p>How about the tiniest of of edges that is NOT worth competing for.</p>",
      "rawMarkdown": "[quote=Triskelion;92358]\r\n\r\nWhat the final few % is I have no idea. Could be mislabelling. Could be LB noise. Could be the tiniest of edges we are now competing for.\r\n\r\n[/quote]\r\n\r\nHow about the tiniest of of edges that is NOT worth competing for.",
      "votes": null
    },
    {
      "id": "92366",
      "postDate": "09/13/2015 21:25:58",
      "content": "<p>The one I was looking at was an unusual amount of the file name numbers showing up in other files, but hadn't looked hard at it yet - the code is slow.\nMaybe just genuinely random, but seems odd if you look at it.</p>",
      "rawMarkdown": "The one I was looking at was an unusual amount of the file name numbers showing up in other files, but hadn't looked hard at it yet - the code is slow.\r\nMaybe just genuinely random, but seems odd if you look at it.",
      "votes": null
    },
    {
      "id": "92373",
      "postDate": "09/13/2015 22:42:13",
      "content": "<p>A small suggestion to Kaggle is that, may be they can give some incentives to the player who first reports this type of leak (like 10% badge in the profile?) to detect this type of problems earlier.</p>",
      "rawMarkdown": "A small suggestion to Kaggle is that, may be they can give some incentives to the player who first reports this type of leak (like 10% badge in the profile?) to detect this type of problems earlier.",
      "votes": null
    },
    {
      "id": "92376",
      "postDate": "09/13/2015 23:00:46",
      "content": "<p>[quote=senbong;92373]</p>\n\n<p>A small suggestion to Kaggle is that, may be they can give some incentives to the player who first reports this type of leak (like 10% badge in the profile?) to detect this type of problems earlier.</p>\n\n<p>[/quote]</p>\n\n<p>Mmm... what about incentives to players whose workflow is not:</p>\n\n<p>1) Lets find some leaks</p>\n\n<p>2) Do actual ML</p>\n\n<p>:D</p>",
      "rawMarkdown": "[quote=senbong;92373]\r\n\r\nA small suggestion to Kaggle is that, may be they can give some incentives to the player who first reports this type of leak (like 10% badge in the profile?) to detect this type of problems earlier.\r\n\r\n[/quote]\r\n\r\nMmm... what about incentives to players whose workflow is not:\r\n\r\n1) Lets find some leaks\r\n\r\n2) Do actual ML\r\n\r\n:D",
      "votes": null
    },
    {
      "id": "92377",
      "postDate": "09/13/2015 23:11:06",
      "content": "<p>[quote=NxGTR;92376]</p>\n\n<p>[quote=senbong;92373]</p>\n\n<p>A small suggestion to Kaggle is that, may be they can give some incentives to the player who first reports this type of leak (like 10% badge in the profile?) to detect this type of problems earlier.</p>\n\n<p>[/quote]</p>\n\n<p>Mmm... what about incentives to players whose workflow is not:</p>\n\n<p>1) Lets find some leaks</p>\n\n<p>2) Do actual ML</p>\n\n<p>:D</p>\n\n<p>[/quote]</p>\n\n<p>In my opinion, detecting this type of problem earlier is very important. The worst case is that someone keeps this leak to the later stage and uses it to win a competition.</p>",
      "rawMarkdown": "[quote=NxGTR;92376]\r\n\r\n[quote=senbong;92373]\r\n\r\nA small suggestion to Kaggle is that, may be they can give some incentives to the player who first reports this type of leak (like 10% badge in the profile?) to detect this type of problems earlier.\r\n\r\n[/quote]\r\n\r\nMmm... what about incentives to players whose workflow is not:\r\n\r\n1) Lets find some leaks\r\n\r\n2) Do actual ML\r\n\r\n:D\r\n\r\n\r\n[/quote]\r\n\r\n\r\nIn my opinion, detecting this type of problem earlier is very important. The worst case is that someone keeps this leak to the later stage and uses it to win a competition.",
      "votes": null
    },
    {
      "id": "92379",
      "postDate": "09/13/2015 23:43:34",
      "content": "<p>If the competition is stopped and isn't rerun then what might be nice is the offer of some sort of small reward (swag perhaps?) for those who attained a good score without the use of the &quot;file_modified&quot; </p>\n\n<p>I only really entered this competition to try large scale analysis for the first time and learn from other people. It would be a shame for other people's work to go to waste simply due to a massive data leak.</p>",
      "rawMarkdown": "If the competition is stopped and isn't rerun then what might be nice is the offer of some sort of small reward (swag perhaps?) for those who attained a good score without the use of the \"file_modified\" \r\n\r\nI only really entered this competition to try large scale analysis for the first time and learn from other people. It would be a shame for other people's work to go to waste simply due to a massive data leak.",
      "votes": null
    },
    {
      "id": "92380",
      "postDate": "09/13/2015 23:45:10",
      "content": "<p>[quote=Giulio;92363]</p>\n\n<p>[quote=Triskelion;92358]</p>\n\n<p>What the final few % is I have no idea. Could be mislabelling. Could be LB noise. Could be the tiniest of edges we are now competing for.</p>\n\n<p>[/quote]</p>\n\n<p>How about the tiniest of of edges that is NOT worth competing for.</p>\n\n<p>[/quote]</p>\n\n<p>Looks to me the 10.000$ is still up for the taking. Kinda agree with your sentiment though.</p>",
      "rawMarkdown": "[quote=Giulio;92363]\r\n\r\n[quote=Triskelion;92358]\r\n\r\nWhat the final few % is I have no idea. Could be mislabelling. Could be LB noise. Could be the tiniest of edges we are now competing for.\r\n\r\n[/quote]\r\n\r\nHow about the tiniest of of edges that is NOT worth competing for.\r\n\r\n[/quote]\r\n\r\nLooks to me the 10.000$ is still up for the taking. Kinda agree with your sentiment though.",
      "votes": null
    },
    {
      "id": "92389",
      "postDate": "09/14/2015 03:55:03",
      "content": "<p>If the date is the leakage, how about the Kaggle admin add an additional restriction just for this contest that using this leakage in the model is invalid?  </p>\n\n<p>Actually, I don't think winning with this leakage has any value for the original purpose.  It's not even like a data mining technology to discover the leakage.</p>\n\n<p>And we should also pay respect to the participants who has already payed much of time to get a high ranking without using leakage information.</p>\n\n<p>Another good solution is to replace the original data set with a normal one.</p>\n\n<p>The worst solution is that you know it's harmful to the contest, and just let it go. </p>",
      "rawMarkdown": "If the date is the leakage, how about the Kaggle admin add an additional restriction just for this contest that using this leakage in the model is invalid?  \r\n\r\nActually, I don't think winning with this leakage has any value for the original purpose.  It's not even like a data mining technology to discover the leakage.\r\n\r\nAnd we should also pay respect to the participants who has already payed much of time to get a high ranking without using leakage information.\r\n\r\nAnother good solution is to replace the original data set with a normal one.\r\n\r\nThe worst solution is that you know it's harmful to the contest, and just let it go.",
      "votes": null
    },
    {
      "id": "92392",
      "postDate": "09/14/2015 04:38:22",
      "content": "<p>[quote=YcdoiT;92389]</p>\n\n<p>If the date is the leakage, how about the Kaggle admin add an additional restriction just for this contest that using this leakage in the model is invalid?  </p>\n\n<p>Actually, I don't think winning with this leakage has any value for the original purpose.  It's not even like a data mining technology to discover the leakage.</p>\n\n<p>And we should also pay respect to the participants who has already payed much of time to get a high ranking without using leakage information.</p>\n\n<p>Another good solution is to replace the original data set with a normal one.</p>\n\n<p>The worst solution is that you know it's harmful to the contest, and just let it go. </p>\n\n<p>[/quote]</p>\n\n<p>Actually it will be quite difficult to go through everyone's codes to see if they used date variable or not. Alternatively, they can make two leaderboards, one just before leakage and one after and at least award the ranks/points, whichever is maximum of the two LB's.</p>",
      "rawMarkdown": "[quote=YcdoiT;92389]\r\n\r\nIf the date is the leakage, how about the Kaggle admin add an additional restriction just for this contest that using this leakage in the model is invalid?  \r\n\r\nActually, I don't think winning with this leakage has any value for the original purpose.  It's not even like a data mining technology to discover the leakage.\r\n\r\nAnd we should also pay respect to the participants who has already payed much of time to get a high ranking without using leakage information.\r\n\r\nAnother good solution is to replace the original data set with a normal one.\r\n\r\nThe worst solution is that you know it's harmful to the contest, and just let it go. \r\n\r\n[/quote]\r\n\r\nActually it will be quite difficult to go through everyone's codes to see if they used date variable or not. Alternatively, they can make two leaderboards, one just before leakage and one after and at least award the ranks/points, whichever is maximum of the two LB's.",
      "votes": null
    },
    {
      "id": "92400",
      "postDate": "09/14/2015 05:23:38",
      "content": "<p>@Subhajit Mandal, yeah, it's hard to tell if a person using this leakage unless he/she want to get the money(Top 1).\nSo, maybe replacing total data set is a good idea.</p>",
      "rawMarkdown": "Subhajit Mandal, yeah, it's hard to tell if a person using this leakage unless he/she want to get the money(Top 1).\r\nSo, maybe replacing total data set is a good idea.",
      "votes": null
    },
    {
      "id": "92402",
      "postDate": "09/14/2015 05:28:11",
      "content": "<p>Yeah theoretically, that's the best idea. But downloading the whole data, and extracting them, and running all the models which take a large amount of time and resources would be very difficult.</p>",
      "rawMarkdown": "Yeah theoretically, that's the best idea. But downloading the whole data, and extracting them, and running all the models which take a large amount of time and resources would be very difficult.",
      "votes": null
    },
    {
      "id": "92405",
      "postDate": "09/14/2015 05:33:26",
      "content": "<p>[quote=Subhajit Mandal;92402]</p>\n\n<p>Yeah theoretically, that's the best idea. But downloading the whole data, and extracting them, and running all the models which take a large amount of time and resources would be very difficult.</p>\n\n<p>[/quote]</p>\n\n<p>I think a new test dataset should be enough. We can still use the downloaded data as the training data.</p>",
      "rawMarkdown": "[quote=Subhajit Mandal;92402]\r\n\r\nYeah theoretically, that's the best idea. But downloading the whole data, and extracting them, and running all the models which take a large amount of time and resources would be very difficult.\r\n\r\n[/quote]\r\n\r\nI think a new test dataset should be enough. We can still use the downloaded data as the training data.",
      "votes": null
    },
    {
      "id": "92408",
      "postDate": "09/14/2015 05:39:15",
      "content": "<p>It looked like this competition was organized very well: lots of data left for testing - fairer scoring.</p>\n\n<p>It is a unfortunate that labels were leaked and now it will have to be cancelled.</p>\n\n<p>Maybe organizers could save the competition, by coming up with more unseen test data and converting the data given so far to a &quot;train&quot; set?</p>",
      "rawMarkdown": "It looked like this competition was organized very well: lots of data left for testing - fairer scoring.\r\n\r\nIt is a unfortunate that labels were leaked and now it will have to be cancelled.\r\n\r\nMaybe organizers could save the competition, by coming up with more unseen test data and converting the data given so far to a \"train\" set?",
      "votes": null
    },
    {
      "id": "92424",
      "postDate": "09/14/2015 06:14:10",
      "content": "<p>[quote=YcdoiT;92400]</p>\n\n<p>@Subhajit Mandal, yeah, it's hard to tell if a person using this leakage unless he/she want to get the money(Top 1).\nSo, maybe replacing total data set is a good idea.</p>\n\n<p>Yeah, whatever, at least 0.99805 of the test set was leaked.\n[/quote]</p>",
      "rawMarkdown": "[quote=YcdoiT;92400]\r\n\r\n@Subhajit Mandal, yeah, it's hard to tell if a person using this leakage unless he/she want to get the money(Top 1).\r\nSo, maybe replacing total data set is a good idea.\r\n\r\nYeah, whatever, at least 0.99805 of the test set was leaked.\r\n[/quote]",
      "votes": null
    },
    {
      "id": "92425",
      "postDate": "09/14/2015 06:14:41",
      "content": "<p>[quote=YcdoiT;92400]</p>\n\n<p>@Subhajit Mandal, yeah, it's hard to tell if a person using this leakage unless he/she want to get the money(Top 1).\nSo, maybe replacing total data set is a good idea.</p>\n\n<p>[/quote]</p>\n\n<p>Yeah, whatever, at least 0.99805 of the test set was leaked.</p>",
      "rawMarkdown": "[quote=YcdoiT;92400]\r\n\r\n@Subhajit Mandal, yeah, it's hard to tell if a person using this leakage unless he/she want to get the money(Top 1).\r\nSo, maybe replacing total data set is a good idea.\r\n\r\n[/quote]\r\n\r\nYeah, whatever, at least 0.99805 of the test set was leaked.",
      "votes": null
    },
    {
      "id": "92476",
      "postDate": "09/14/2015 13:55:35",
      "content": "<p>I'll vote for a restart with a new test data set.</p>",
      "rawMarkdown": "I'll vote for a restart with a new test data set.",
      "votes": null
    },
    {
      "id": "92478",
      "postDate": "09/14/2015 14:15:46",
      "content": "<p>Long silence from Dato/Kaggle suggests that they don't have additional data. In any case, I would expect creating datasets of these sizes (especially from 3rd party websites like StumbleUpon) should take a long time.</p>\n\n<p>I would vote for a 2-stage approach like Netflix, GE (and many other competitions on Kaggle). Keep the end-date of stage 1 on Sept. 10th --- that was the day the leak was discovered. Start stage-2 with much bigger (and cleaner) datasets when they are ready.</p>",
      "rawMarkdown": "Long silence from Dato/Kaggle suggests that they don't have additional data. In any case, I would expect creating datasets of these sizes (especially from 3rd party websites like StumbleUpon) should take a long time.\r\n\r\nI would vote for a 2-stage approach like Netflix, GE (and many other competitions on Kaggle). Keep the end-date of stage 1 on Sept. 10th --- that was the day the leak was discovered. Start stage-2 with much bigger (and cleaner) datasets when they are ready.",
      "votes": null
    },
    {
      "id": "92486",
      "postDate": "09/14/2015 14:32:23",
      "content": "<p>I think that renaming all of the files and even jumbling the composition of the training/test sets would be sufficient. It's true that someone could still keep both sets of data and do an expensive file similarity to find the &quot;ground truth&quot; and cheat -- but anyone who spend the time to do that would be pathetic. Plus all prize eligible solutions need to be revealed so those couldn't exactly cheat. </p>",
      "rawMarkdown": "I think that renaming all of the files and even jumbling the composition of the training/test sets would be sufficient. It's true that someone could still keep both sets of data and do an expensive file similarity to find the \"ground truth\" and cheat -- but anyone who spend the time to do that would be pathetic. Plus all prize eligible solutions need to be revealed so those couldn't exactly cheat.",
      "votes": null
    },
    {
      "id": "92488",
      "postDate": "09/14/2015 14:37:50",
      "content": "<p>[quote=David McGarry;92486]</p>\n\n<p>I think that renaming all of the files and even jumbling the composition of the training/test sets would be sufficient. It's true that someone could still keep both sets of data and do an expensive file similarity to find the &quot;ground truth&quot; and cheat -- but anyone who spend the time to do that would be pathetic. Plus all prize eligible solutions need to be revealed so those couldn't exactly cheat. </p>\n\n<p>[/quote]</p>\n\n<p>File similarity is one thing, but that's not the only path to cheating. People could treat combined train+test dataset as a giant training dataset, overfit the model to it and achieve high score on LB. </p>",
      "rawMarkdown": "[quote=David McGarry;92486]\r\n\r\nI think that renaming all of the files and even jumbling the composition of the training/test sets would be sufficient. It's true that someone could still keep both sets of data and do an expensive file similarity to find the \"ground truth\" and cheat -- but anyone who spend the time to do that would be pathetic. Plus all prize eligible solutions need to be revealed so those couldn't exactly cheat. \r\n\r\n[/quote]\r\n\r\nFile similarity is one thing, but that's not the only path to cheating. People could treat combined train+test dataset as a giant training dataset, overfit the model to it and achieve high score on LB.",
      "votes": null
    },
    {
      "id": "92489",
      "postDate": "09/14/2015 14:38:55",
      "content": "<p>Hashing 300k files and then matching by hash doesn't seem expensive to me. Actually, they can even be beautifulsoupped in a couple of hours on a single 8-core machine.  At the same time, feature selection done on ground truth might be quite effective and can't be proven to be a cheat.</p>",
      "rawMarkdown": "Hashing 300k files and then matching by hash doesn't seem expensive to me. Actually, they can even be beautifulsoupped in a couple of hours on a single 8-core machine.  At the same time, feature selection done on ground truth might be quite effective and can't be proven to be a cheat.",
      "votes": null
    },
    {
      "id": "92492",
      "postDate": "09/14/2015 15:12:59",
      "content": "<p>Guys, people work for Kaggle also take weekends off. So I believe if they are gonna do something, they will start doing it today. So we should just be patient.</p>",
      "rawMarkdown": "Guys, people work for Kaggle also take weekends off. So I believe if they are gonna do something, they will start doing it today. So we should just be patient.",
      "votes": null
    },
    {
      "id": "92497",
      "postDate": "09/14/2015 15:45:36",
      "content": "<p>The question is whether there is any advantage for Kaggle and Dato to reset the competition. Dato hasn't put a lot of money on the the plate, which is indicative that this wasn't a must solve problem for them. Maybe more of a relatively cheap way to get some PR for their tool, but here I'm just speculating.\nAnyways, here are the few things that stood out to me so far, probably the only things I will remember about yet another flawed competition:</p>\n\n<p>-NxGTR's .974 performance without using the leak (great job!)</p>\n\n<p>-Dimitry's being the first finding the leak. I think, without him finding the problem, this might have gone unnoticed till the end. Great eye and I appreciate the honesty to put this out immediately instead of using it as a last day submission.</p>\n\n<p>-The sadness of seeing the LB full of new entrants as soon as the leak was made public. Like sharks attracted by blood in the water. Come on guys, there is more to life than a few Kaggle points!</p>\n\n<p>-Yet another flawed competition. I mean, I know that life sometimes throws curve balls, but it is interesting to observe how even companies that do DS &amp; ML for a living, can overlook huge flaws like the one we're now dealing with.</p>\n\n<p>Anyhow, however this goes, I hope everybody gets to enjoy the remainder of the competition. Good luck to all!</p>",
      "rawMarkdown": "The question is whether there is any advantage for Kaggle and Dato to reset the competition. Dato hasn't put a lot of money on the the plate, which is indicative that this wasn't a must solve problem for them. Maybe more of a relatively cheap way to get some PR for their tool, but here I'm just speculating.\r\nAnyways, here are the few things that stood out to me so far, probably the only things I will remember about yet another flawed competition:\r\n\r\n-NxGTR's .974 performance without using the leak (great job!)\r\n\r\n-Dimitry's being the first finding the leak. I think, without him finding the problem, this might have gone unnoticed till the end. Great eye and I appreciate the honesty to put this out immediately instead of using it as a last day submission.\r\n\r\n-The sadness of seeing the LB full of new entrants as soon as the leak was made public. Like sharks attracted by blood in the water. Come on guys, there is more to life than a few Kaggle points!\r\n\r\n-Yet another flawed competition. I mean, I know that life sometimes throws curve balls, but it is interesting to observe how even companies that do DS & ML for a living, can overlook huge flaws like the one we're now dealing with.\r\n\r\nAnyhow, however this goes, I hope everybody gets to enjoy the remainder of the competition. Good luck to all!",
      "votes": null
    },
    {
      "id": "92500",
      "postDate": "09/14/2015 15:53:54",
      "content": "<p>[quote=Ilya Nekhay;92489]</p>\n\n<p>Hashing 300k files and then matching by hash doesn't seem expensive to me. Actually, they can even be beautifulsoupped in a couple of hours on a single 8-core machine.  At the same time, feature selection done on ground truth might be quite effective and can't be proven to be a cheat.</p>\n\n<p>[/quote]</p>\n\n<p>True, although &quot;a couple of hours&quot; to cheat sounds expensive to me, but perhaps I don't understand the motivations for some joining the competition. Plus introducing some innocuous white noise to files would make hashing more difficult, but again still completely solvable if someone really wanted to cheat. As for feature selection, given the size of the training set I'm not sure there is really that much to gain unless people are developing massively overfit features. I suppose some would, which is unfortunate.</p>\n\n<p>It would nice if the competition could be salvaged. I'm not sure how possible that is if more data isn't available given that my idea was more-or-less de-bunked.</p>",
      "rawMarkdown": "[quote=Ilya Nekhay;92489]\r\n\r\nHashing 300k files and then matching by hash doesn't seem expensive to me. Actually, they can even be beautifulsoupped in a couple of hours on a single 8-core machine.  At the same time, feature selection done on ground truth might be quite effective and can't be proven to be a cheat.\r\n\r\n[/quote]\r\n\r\nTrue, although \"a couple of hours\" to cheat sounds expensive to me, but perhaps I don't understand the motivations for some joining the competition. Plus introducing some innocuous white noise to files would make hashing more difficult, but again still completely solvable if someone really wanted to cheat. As for feature selection, given the size of the training set I'm not sure there is really that much to gain unless people are developing massively overfit features. I suppose some would, which is unfortunate.\r\n\r\nIt would nice if the competition could be salvaged. I'm not sure how possible that is if more data isn't available given that my idea was more-or-less de-bunked.",
      "votes": null
    },
    {
      "id": "92501",
      "postDate": "09/14/2015 15:54:56",
      "content": "<p>[quote=Giulio;92497]</p>\n\n<p>-The sadness of seeing the LB full of new entrants as soon as the leak was made public. Like sharks attracted by blood in the water. Come on guys, there is more to life than a few Kaggle points!</p>\n\n<p>[/quote]</p>\n\n<p>While I agree in principle, I for one have never so far found nor exploited a leak in a kaggle contest, and for me it will be a nice (albeit small) exercise to do so (if the data is still up until I get a chance to download it all).</p>",
      "rawMarkdown": "[quote=Giulio;92497]\r\n\r\n-The sadness of seeing the LB full of new entrants as soon as the leak was made public. Like sharks attracted by blood in the water. Come on guys, there is more to life than a few Kaggle points!\r\n\r\n[/quote]\r\n\r\nWhile I agree in principle, I for one have never so far found nor exploited a leak in a kaggle contest, and for me it will be a nice (albeit small) exercise to do so (if the data is still up until I get a chance to download it all).",
      "votes": null
    },
    {
      "id": "92503",
      "postDate": "09/14/2015 16:02:31",
      "content": "<p>[quote=Giulio;92497]</p>\n\n<p>The question is whether there is any advantage for Kaggle and Dato to reset the competition. Dato hasn't put a lot of money on the the plate, which is indicative that this wasn't a must solve problem for them. Maybe more of a relatively cheap way to get some PR for their tool, but here I'm just speculating.\nAnyways, here are the few things that stood out to me so far, probably the only things I will remember about yet another flawed competition:</p>\n\n<p>-NxGTR's .974 performance without using the leak (great job!)</p>\n\n<p>-Dimitry's being the first finding the leak. I think, without him finding the problem, this might have gone unnoticed till the end. Great eye and I appreciate the honesty to put this out immediately instead of using it as a last day submission.</p>\n\n<p>-The sadness of seeing the LB full of new entrants as soon as the leak was made public. Like sharks attracted by blood in the water. Come on guys, there is more to life than a few Kaggle points!</p>\n\n<p>-Yet another flawed competition. I mean, I know that life sometimes throws curve balls, but it is interesting to observe how even companies that do DS &amp; ML for a living, can overlook huge flaws like the one we're now dealing with.</p>\n\n<p>Anyhow, however this goes, I hope everybody gets to enjoy the remainder of the competition. Good luck to all!</p>\n\n<p>[/quote]</p>\n\n<p>I don't think that .974 is possible without using the leak - these webpages are well-disguised.\n@NxGTR - can you confirm that you didn't use the leak somehow/accidentally? I find it hard to believe, since these could be any webpages.</p>",
      "rawMarkdown": "[quote=Giulio;92497]\r\n\r\nThe question is whether there is any advantage for Kaggle and Dato to reset the competition. Dato hasn't put a lot of money on the the plate, which is indicative that this wasn't a must solve problem for them. Maybe more of a relatively cheap way to get some PR for their tool, but here I'm just speculating.\r\nAnyways, here are the few things that stood out to me so far, probably the only things I will remember about yet another flawed competition:\r\n\r\n-NxGTR's .974 performance without using the leak (great job!)\r\n\r\n-Dimitry's being the first finding the leak. I think, without him finding the problem, this might have gone unnoticed till the end. Great eye and I appreciate the honesty to put this out immediately instead of using it as a last day submission.\r\n\r\n-The sadness of seeing the LB full of new entrants as soon as the leak was made public. Like sharks attracted by blood in the water. Come on guys, there is more to life than a few Kaggle points!\r\n\r\n-Yet another flawed competition. I mean, I know that life sometimes throws curve balls, but it is interesting to observe how even companies that do DS & ML for a living, can overlook huge flaws like the one we're now dealing with.\r\n\r\nAnyhow, however this goes, I hope everybody gets to enjoy the remainder of the competition. Good luck to all!\r\n\r\n[/quote]\r\n\r\nI don't think that .974 is possible without using the leak - these webpages are well-disguised.\r\n@NxGTR - can you confirm that you didn't use the leak somehow/accidentally? I find it hard to believe, since these could be any webpages.",
      "votes": null
    },
    {
      "id": "92505",
      "postDate": "09/14/2015 16:08:50",
      "content": "<p>[quote=bekbolatov;92503]</p>\n\n<p>I don't think that .974 is possible without using the leak - these webpages are well-disguised.\n@NxGTR - can you confirm that you didn't use the leak somehow/accidentally? I find it hard to believe, since these could be any webpages.</p>\n\n<p>[/quote]</p>\n\n<p>Apparently NxGTR's score is legit. 1) my dumb model is 0.95 already without the leak. 2) The leak is 0.99805 immediately. It's clear that he didn't know it at all.</p>",
      "rawMarkdown": "[quote=bekbolatov;92503]\r\n\r\nI don't think that .974 is possible without using the leak - these webpages are well-disguised.\r\n@NxGTR - can you confirm that you didn't use the leak somehow/accidentally? I find it hard to believe, since these could be any webpages.\r\n\r\n[/quote]\r\n\r\nApparently NxGTR's score is legit. 1) my dumb model is 0.95 already without the leak. 2) The leak is 0.99805 immediately. It's clear that he didn't know it at all.",
      "votes": null
    },
    {
      "id": "92506",
      "postDate": "09/14/2015 16:09:55",
      "content": "<p>I got to 0.951 only using the content of the files, so I'm pretty sure it's possible. I made some decisions that worsened my model it seems, so I fell short after getting these results.</p>",
      "rawMarkdown": "I got to 0.951 only using the content of the files, so I'm pretty sure it's possible. I made some decisions that worsened my model it seems, so I fell short after getting these results.",
      "votes": null
    },
    {
      "id": "92507",
      "postDate": "09/14/2015 16:18:31",
      "content": "<p>Thanks guys. Totally above my expectations.</p>",
      "rawMarkdown": "Thanks guys. Totally above my expectations.",
      "votes": null
    },
    {
      "id": "92509",
      "postDate": "09/14/2015 16:30:22",
      "content": "<p>Anyone who had actually work in this one will know that my score its just good and legit, no leaks. My second submission got me almost 0.96 already, so 0.97x is not unbelievable.</p>",
      "rawMarkdown": "Anyone who had actually work in this one will know that my score its just good and legit, no leaks. My second submission got me almost 0.96 already, so 0.97x is not unbelievable.",
      "votes": null
    },
    {
      "id": "92512",
      "postDate": "09/14/2015 16:43:37",
      "content": "<p>It's all good, NxGTR. I haven't had time to try it out yet.</p>\n\n<p>I was just getting familiar with the data and initially it seemed pretty hard to have a good separation like that - turns out it isn't.</p>\n\n<p>I haven't decided whether I want to put an effort into this competition - given the easy separation.</p>\n\n<p>This means that practically all of the Native ads are detectable :)</p>",
      "rawMarkdown": "It's all good, NxGTR. I haven't had time to try it out yet.\r\n\r\nI was just getting familiar with the data and initially it seemed pretty hard to have a good separation like that - turns out it isn't.\r\n\r\nI haven't decided whether I want to put an effort into this competition - given the easy separation.\r\n\r\nThis means that practically all of the Native ads are detectable :)",
      "votes": null
    },
    {
      "id": "92519",
      "postDate": "09/14/2015 17:02:44",
      "content": "<p>I got a 0.97064 submission score with a model based on HTML alone.</p>\n\n<p>Mixing my content based model with the metadata model yields a slightly worse submission score (0.99826) than the metadata model alone (0.99846).</p>\n\n<p>The metadata based models worked by injecting tags like <code>&lt;file-meta-###/&gt;</code>, so that I wouldn't have to use a separate algorithm to mix the features.</p>",
      "rawMarkdown": "I got a 0.97064 submission score with a model based on HTML alone.\r\n\r\nMixing my content based model with the metadata model yields a slightly worse submission score (0.99826) than the metadata model alone (0.99846).\r\n\r\nThe metadata based models worked by injecting tags like `<file-meta-###/>`, so that I wouldn't have to use a separate algorithm to mix the features.",
      "votes": null
    },
    {
      "id": "305511",
      "postDate": "03/29/2018 01:14:05",
      "content": "<p>Hey bro</p>",
      "rawMarkdown": "Hey bro",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 92252,
      "author_name": "thakurrajanand",
      "author_url": "",
      "post_date": "09/12/2015 15:00:08",
      "content": "<p>[quote=SkyLibrary;92249]</p>\n\n<p>The near perfect result is pretty surprising to me, I  am just wondering if top three players are willing to tell it is label/tag leakage or not ?  I think there is no harm confirming this, after all, leakage is nothing related to violating competition rules.</p>\n\n<p>The dataset is quite interesting, but I don't want to spend time on a  competition with leakage again, the west nile virus one is a  nightmare for me, and the tube price too. </p>\n\n<p>Edit: sorry, just find typo in title, guess this is no way to fix that lol</p>\n\n<p>[/quote]</p>\n\n<p>It would be unfair with the 2 above me if I tell you anything. Last time I posted high model in Otto competition and everyone was just behind me. Someone from top 10 even criticized me by messaging me personally that because of me they are losing rank on Leader board.  </p>\n\n<p>I hope you can understand that I don't want to be involve in controversy every time :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92254,
      "author_name": "rushter",
      "author_url": "",
      "post_date": "09/12/2015 15:30:35",
      "content": "<p>Just wait few weeks :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92257,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "09/12/2015 16:21:59",
      "content": "<p>It is pretty much a yes or no question, if you guys don't want to answer it, then I will assume the answer is yes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92276,
      "author_name": "adjgiulio",
      "author_url": "",
      "post_date": "09/12/2015 19:50:26",
      "content": "<p>Why would they have to say anything? It is already a HUGE advantage for everybody to know that it is possible to achieve that score. I dropped the competition because I believe it is highly likely that this is an exercise in finding a deterministic rule in the dataset, and that treasure hunt is not worth my time. I could be wrong though. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92280,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "09/12/2015 20:26:16",
      "content": "<p>I didn't say they should answer my question, but I have the right to post the question no matter they reply or not. If they do, it will confirm your 'highly likely', if not, then it is still highly likely :)</p>\n\n<p>[quote=Giulio;92276]</p>\n\n<p>Why would they have to say anything? It is already a HUGE advantage for everybody to know that it is possible to achieve that score. I dropped the competition because I believe it is highly likely that this is an exercise in finding a deterministic rule in the dataset, and that treasure hunt is not worth my time. I could be wrong though. </p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92285,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "09/12/2015 22:48:30",
      "content": "<p>My guess is:</p>\n\n<p>They are running an unsupervised algorithm over both training and test that generates relationships between elements in both datasets. Let's call this a vocabulary. This thus creates highly abstract relationships between all documents at word level, but the difference now is that this is across both train and test.</p>\n\n<p>Then for training, create (meaningful) document vectors from this vocabulary and start training to indicate sponsored / not sponsored. The difference is that you're not training on occurrences of individual features of a document (which can be meaningful or not), but you're training against the known vector space of all documents (with 1/3 of the data). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92294,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "09/13/2015 01:05:02",
      "content": "<p>[quote=DataGeek;92252]</p>\n\n<p>[quote=SkyLibrary;92249]</p>\n\n<p>The near perfect result is pretty surprising to me, I  am just wondering if top three players are willing to tell it is label/tag leakage or not ?  I think there is no harm confirming this, after all, leakage is nothing related to violating competition rules.</p>\n\n<p>The dataset is quite interesting, but I don't want to spend time on a  competition with leakage again, the west nile virus one is a  nightmare for me, and the tube price too. </p>\n\n<p>Edit: sorry, just find typo in title, guess this is no way to fix that lol</p>\n\n<p>[/quote]</p>\n\n<p>It would be unfair with the 2 above me if I tell you anything. Last time I posted high model in Otto competition and everyone was just behind me. Someone from top 10 even criticized me by messaging me personally that because of me they are losing rank on Leader board.  </p>\n\n<p>I hope you can understand that I don't want to be involve in controversy every time :)</p>\n\n<p>[/quote]</p>\n\n<p>I think more and more people are getting offended by this, and joining the 0.99 club. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92299,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "09/13/2015 03:07:39",
      "content": "<p>After download the data and play with it for 5 hours, I am here to confirm the answer is yes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92301,
      "author_name": "dmcgarry",
      "author_url": "",
      "post_date": "09/13/2015 04:32:35",
      "content": "<p>Well this is fun, and here I thought this would be the one competition that is actually like the old days since the data is too large for scripts. Silly me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92303,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "09/13/2015 05:37:23",
      "content": "<p>Interesting. I wonder if Kaggle will put a hold on this or just let it be it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92304,
      "author_name": "",
      "author_url": "",
      "post_date": "09/13/2015 05:55:33",
      "content": "<p>I am a noob in machine learning. I played with natural language processing algorithms. I achieved .76 score then I found that beside language processing, we can have better result if we used other features in the HTML files such as: javascript, meta tag attribute... </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92305,
      "author_name": "lijunjie",
      "author_url": "",
      "post_date": "09/13/2015 06:04:10",
      "content": "<p>So there is little space for the improvement after you applying the leakage information into your model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92308,
      "author_name": "dulyanov",
      "author_url": "",
      "post_date": "09/13/2015 06:46:58",
      "content": "<p>For those of you who think this competition doesn't make sense now -- it is really a whole different world in tuning your model, when you have one almost perfect predictor. Just try to go beyond the level this predictor gives you, it is challenging. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92313,
      "author_name": "clustifier",
      "author_url": "",
      "post_date": "09/13/2015 09:28:41",
      "content": "<p>[quote=Dmitry Ulyanov;92308]</p>\n\n<p>For those of you who think this competition doesn't make sense now -- it is really a whole different world in tuning your model, when you have one almost perfect predictor. Just try to go beyond the level this predictor gives you, it is challenging. </p>\n\n<p>[/quote]</p>\n\n<p>But now you must know this &quot;perfect predictor&quot;, which it's a leakage and not real data, in order to be successful in this competition.</p>\n\n<p>So if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92314,
      "author_name": "wildwizard",
      "author_url": "",
      "post_date": "09/13/2015 09:29:01",
      "content": "<p>This is unfair to WNV or Cat competitors to say that leakage exploitation was the only determinant in victory. <br>\nMany did exploit it (even myself in WNV, not knowing it was, while in Cat I discarded it believing I would make a methodology mistake), and we did not all achieve top scores.</p>\n\n<p>Leakage exploitation still asks for talent and skills.\nAlso, nothing tells that what you call leakage, as I said before, is just a very strong association rule in the dataset and was present from the start.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92321,
      "author_name": "omgponies",
      "author_url": "",
      "post_date": "09/13/2015 13:18:19",
      "content": "<p>My code that is looking at what I think is the leak is apparently very slow, lol.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92322,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "09/13/2015 13:27:26",
      "content": "<p>I thought that, since an ad is &quot;sponsored&quot;, you'd find some kind of link or script in the files across all the sponsored ones (tracking links), but it isn't as easy as that. Just filtering on _gac for example gave me 0.546 and non-sponsored ones contain the same thing.  But I reckon it must be somewhere along those lines...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92324,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "09/13/2015 13:52:50",
      "content": "<p>[quote=SkyLibrary;92299]</p>\n\n<p>After download the data and play with it for 5 hours, I am here to confirm the answer is yes.</p>\n\n<p>[/quote]\nSo this is a leakage, and you have found it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92326,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "09/13/2015 14:33:21",
      "content": "<p>[quote=clustifier;92313]</p>\n\n<p>So if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.</p>\n\n<p>[/quote]</p>\n\n<p>Even with this leak it is still fair for players.</p>\n\n<p>The real bad thing is that it is pointless now for anyone trying to win the Dato Creators Award with this leakage existing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92327,
      "author_name": "khyh00",
      "author_url": "",
      "post_date": "09/13/2015 14:38:01",
      "content": "<p>I am sure that there is a leakage that can be used as a &quot;golden feature&quot;, but only using this can not reach 1.0 on the public leader board (0.99805). Or, as pointed out on the other thread, there are some errors in this dataset so the &quot;golden feature&quot; is perfect but labels are wrong.</p>\n\n<p>I want to know what the admin is thinking about this since further explorer might be in vain.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92331,
      "author_name": "dmcgarry",
      "author_url": "",
      "post_date": "09/13/2015 14:59:08",
      "content": "<p>[quote=rcarson;92326]</p>\n\n<p>[quote=clustifier;92313]</p>\n\n<p>So if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.</p>\n\n<p>[/quote]</p>\n\n<p>Even with this leak it is still fair for players.</p>\n\n<p>The real bad thing is that it is pointless now for anyone trying to win the Dato Creators Award with this leakage existing.</p>\n\n<p>[/quote]</p>\n\n<p>Define &quot;fair&quot;. Sure, all players have the same opportunity to find and exploit leakage but this is supposed to be a machine learning competition. I would have never signed up if it was a race to find some mistake in the data. I'm now stuck in a demlina where I have to decide if I want to, (a) continue using machine learning to tackle a seemingly interesting problem, (b) give in and find the leakage or (c) say screw it, delete the large data files from my system and feel even worse about the post-scripts world of kaggle.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92332,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "09/13/2015 15:11:02",
      "content": "<p>I guess this has something to do with scripts inside html files.</p>\n\n<p>For example, the function of ads' charging. But I'm not familiar with javascript unfortunately, lol.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92333,
      "author_name": "rushter",
      "author_url": "",
      "post_date": "09/13/2015 15:17:04",
      "content": "<p>Stumbleupon doesn't require any additions to html page.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92334,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "09/13/2015 15:23:01",
      "content": "<p>[quote=David McGarry;92331]</p>\n\n<p>[quote=rcarson;92326]</p>\n\n<p>[quote=clustifier;92313]</p>\n\n<p>So if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.</p>\n\n<p>[/quote]</p>\n\n<p>Even with this leak it is still fair for players.</p>\n\n<p>The real bad thing is that it is pointless now for anyone trying to win the Dato Creators Award with this leakage existing.</p>\n\n<p>[/quote]</p>\n\n<p>Define &quot;fair&quot;. Sure, all players have the same opportunity to find and exploit leakage but this is supposed to be a machine learning competition. I would have never signed up if it was a race to find some mistake in the data. I'm not stuck in a demlina where I have to decide if I want to, (a) continue using machine learning to tackle a seemingly interesting problem, (b) give in and find the leakage or (c) say screw it, delete the large data files from my system and feel even worse about the post-scripts world of kaggle.</p>\n\n<p>[/quote]</p>\n\n<p>On the contrary, I found it fun to look for the leak ( I haven't found it yet). </p>\n\n<p>A little bit off the topic, I think a great part of kaggle competitions is not about machine learning. It is more about data mining. Finding soft or strong leak (good or golden features) is usually more dominant, given there are so many good out-of-box tools like xgboost for the learning part.</p>\n\n<p>In my opinion, finding a leak is a great skill even in real world case. But I agree in this case, it is a mess that the leak is too strong..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92336,
      "author_name": "rushter",
      "author_url": "",
      "post_date": "09/13/2015 15:26:43",
      "content": "<p>I think i found a  leakage, but it doesn't improve my score. LOL.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92340,
      "author_name": "adjgiulio",
      "author_url": "",
      "post_date": "09/13/2015 16:35:16",
      "content": "<p>Ok, let's play this out as if it was a real project.</p>\n\n<p>Business guy: we think we could really benefit from being able to use ML to classify pages with native add content.</p>\n\n<p>Data Scientist: give me a week to look at the data and I will get back to you with a first evaluation of what I think is possible. I'll work with engineering to get some data.</p>\n\n<p>...a week later...</p>\n\n<p>(very excited, yet skeptical) Data Scientist: I think I found something in the data that allows a perfect prediction. This is XYZ feature.</p>\n\n<p>Business Guy: oh, well, that makes sense. Seems kind of obvious. Wonder what &quot;name of a peer/manager/director who is taking the heat&quot; thought when he came up with the idea of using ML for this... You know what, I don't think we need ML anymore.</p>\n\n<p>(disappointed) Data Scientist: well, I guess I'll go back to building that boring dashboard then. Too bad, I was so excited of finally using some of those ML skills I had learnt on Kaggle...</p>\n\n<p>My point, when a perfect predictor exists, there is no reason to use ML because it is not an ML problem. This should have never made it to be an ML project/competition. Now, will the finding of XYZ feature be of value to Dato? Who knows? I would be surprised if it does. Will it still be interesting to try to maximize performance on misslabeld observations? I mean, why not, it's just a problem like any other. But, wow, not the problem I thought I was signing up for.\nGlad I haven't made a submission yet...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92341,
      "author_name": "adjgiulio",
      "author_url": "",
      "post_date": "09/13/2015 16:39:06",
      "content": "<p>[quote=David McGarry;92331]\n I'm now stuck in a demlina where I have to decide if I want to, (a) continue using machine learning to tackle a seemingly interesting problem, (b) give in and find the leakage or (c) say screw it, delete the large data files from my system and feel even worse about the post-scripts world of kaggle.</p>\n\n<p>[/quote]</p>\n\n<p>+1 for (c)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92342,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "09/13/2015 16:41:24",
      "content": "<p>Giulio, it's much worse than that... much worse :)  one of those things: &quot;soon as you see it...&quot; </p>\n\n<p>What I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. I had been trying to improve from 0.951 before this without success. That by itself is already a very reasonable score I think. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92343,
      "author_name": "senbong",
      "author_url": "",
      "post_date": "09/13/2015 16:44:13",
      "content": "<p>In the end, don't tell me it is a filename leak... ==</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92344,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "09/13/2015 16:51:18",
      "content": "<p>[quote=Gerard Toonstra;92342]\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. [/quote]</p>\n\n<p>I am very sad for this competition... I can confirm 0.973x is achievable without losing my own standards.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92345,
      "author_name": "chonglinsun",
      "author_url": "",
      "post_date": "09/13/2015 16:56:13",
      "content": "<p>[quote=NxGTR;92344]</p>\n\n<p>[quote=Gerard Toonstra;92342]\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. [/quote]</p>\n\n<p>I am very sad for this competition... I can confirm 0.973x is achievable without losing my own standards.</p>\n\n<p>[/quote]</p>\n\n<p>I'm sad too,  especially because of time I spent on it in the last few weeks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92348,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "09/13/2015 17:19:20",
      "content": "<p>[quote=NxGTR;92344]</p>\n\n<p>[quote=Gerard Toonstra;92342]\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. [/quote]</p>\n\n<p>I am very sad for this competition... I can confirm 0.973x is achievable without losing my own standards.</p>\n\n<p>[/quote]</p>\n\n<p>I'm looking forward to an explanation of your approach, parametrizations and feature choices, if you decide to write it up.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92349,
      "author_name": "gkoundry",
      "author_url": "",
      "post_date": "09/13/2015 19:07:04",
      "content": "<p>The date in the zip files looks suspect.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92350,
      "author_name": "lawrencechernin",
      "author_url": "",
      "post_date": "09/13/2015 19:25:14",
      "content": "<p>yes, it looks like the 1's files were written only on June 24 or July 27th.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92353,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "09/13/2015 19:55:58",
      "content": "<p>[quote=Lawrence Chernin;92350]</p>\n\n<p>yes, it looks like the 1's files were written only on June 24 or July 27th.</p>\n\n<p>[/quote]\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92355,
      "author_name": "rushter",
      "author_url": "",
      "post_date": "09/13/2015 20:02:11",
      "content": "<p>[quote=Little Boat;92353]</p>\n\n<p>[quote=Lawrence Chernin;92350]</p>\n\n<p>yes, it looks like the 1's files were written only on June 24 or July 27th.</p>\n\n<p>[/quote]\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  </p>\n\n<p>[/quote]\nThere is nothing to remove :)</p>\n\n<blockquote>\n  <p>StumbleUpon Ads is a native advertising platform that does not use\n  traditional ad units such as banners or text links. This means that on\n  StumbleUpon your ad is, quite simply, your URL. Since your web page\n  itself is the ad unit, no additional creative assets or copy are\n  needed to drive traffic. We drive directly to your page when a user\n  clicks the Stumble button.</p>\n  \n  <p><a href=\"http://ads.stumbleupon.com/wp-content/uploads/2014/09/CampaignCreationGuide.pdf\">http://ads.stumbleupon.com/wp-content/uploads/2014/09/CampaignCreationGuide.pdf</a></p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92356,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "09/13/2015 20:05:35",
      "content": "<p>Hmm. Then I don't see how that would be the leakage. Anyway, if I am bored enough to find out if it is, I will post it here. But I am hoping someone else can do it for us. :-)</p>\n\n<p>[quote=Artem;92355]</p>\n\n<p>[quote=Little Boat;92353]</p>\n\n<p>[quote=Lawrence Chernin;92350]</p>\n\n<p>yes, it looks like the 1's files were written only on June 24 or July 27th.</p>\n\n<p>[/quote]\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  </p>\n\n<p>[/quote]\nThere is nothing to remove :)</p>\n\n<blockquote>\n  <p>StumbleUpon Ads is a native advertising platform that does not use\n  traditional ad units such as banners or text links. This means that on\n  StumbleUpon your ad is, quite simply, your URL. Since your web page\n  itself is the ad unit, no additional creative assets or copy are\n  needed to drive traffic. We drive directly to your page when a user\n  clicks the Stumble button.</p>\n</blockquote>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92357,
      "author_name": "rushter",
      "author_url": "",
      "post_date": "09/13/2015 20:11:46",
      "content": "<p>[quote=Little Boat;92356]</p>\n\n<p>Hmm. Then I don't see how that would be the leakage. Anyway, if I am bored enough to find out if it is, I will post it here. But I am hoping someone else can do it for us. :-)</p>\n\n<p>[quote=Artem;92355]</p>\n\n<p>[quote=Little Boat;92353]</p>\n\n<p>[quote=Lawrence Chernin;92350]</p>\n\n<p>yes, it looks like the 1's files were written only on June 24 or July 27th.</p>\n\n<p>[/quote]\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  </p>\n\n<p>[/quote]\nThere is nothing to remove :)</p>\n\n<blockquote>\n  <p>StumbleUpon Ads is a native advertising platform that does not use\n  traditional ad units such as banners or text links. This means that on\n  StumbleUpon your ad is, quite simply, your URL. Since your web page\n  itself is the ad unit, no additional creative assets or copy are\n  needed to drive traffic. We drive directly to your page when a user\n  clicks the Stumble button.</p>\n</blockquote>\n\n<p>[/quote]\n[/quote]\nI checked it locally and it looks like true.</p>\n\n<p>Also i found another &quot;leakage&quot; that can give 0.86+ without ml.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92358,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "09/13/2015 20:13:48",
      "content": "<p>There is definitely leakage and I don't see how this could be solved/repaired, short of giving us a completely new data set...</p>\n\n<p>What the final few % is I have no idea. Could be mislabelling. Could be LB noise. Could be the tiniest of edges we are now competing for.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92359,
      "author_name": "dmcgarry",
      "author_url": "",
      "post_date": "09/13/2015 20:14:50",
      "content": "<p>[quote=Lawrence Chernin;92350]</p>\n\n<p>yes, it looks like the 1's files were written only on June 24 or July 27th.</p>\n\n<p>[/quote]</p>\n\n<p>Well that's fun. Hopefully they clear the LB and release a new dataset (using new HTML files would be ideal, but at least renaming all of the existing HTML files would discourage all but the least honest to ignore the known values in the test set).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92362,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "09/13/2015 20:45:35",
      "content": "<p>Now I can confirm that the date is the leakage.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92363,
      "author_name": "adjgiulio",
      "author_url": "",
      "post_date": "09/13/2015 20:46:47",
      "content": "<p>[quote=Triskelion;92358]</p>\n\n<p>What the final few % is I have no idea. Could be mislabelling. Could be LB noise. Could be the tiniest of edges we are now competing for.</p>\n\n<p>[/quote]</p>\n\n<p>How about the tiniest of of edges that is NOT worth competing for.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92366,
      "author_name": "omgponies",
      "author_url": "",
      "post_date": "09/13/2015 21:25:58",
      "content": "<p>The one I was looking at was an unusual amount of the file name numbers showing up in other files, but hadn't looked hard at it yet - the code is slow.\nMaybe just genuinely random, but seems odd if you look at it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92373,
      "author_name": "senbong",
      "author_url": "",
      "post_date": "09/13/2015 22:42:13",
      "content": "<p>A small suggestion to Kaggle is that, may be they can give some incentives to the player who first reports this type of leak (like 10% badge in the profile?) to detect this type of problems earlier.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92376,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "09/13/2015 23:00:46",
      "content": "<p>[quote=senbong;92373]</p>\n\n<p>A small suggestion to Kaggle is that, may be they can give some incentives to the player who first reports this type of leak (like 10% badge in the profile?) to detect this type of problems earlier.</p>\n\n<p>[/quote]</p>\n\n<p>Mmm... what about incentives to players whose workflow is not:</p>\n\n<p>1) Lets find some leaks</p>\n\n<p>2) Do actual ML</p>\n\n<p>:D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92377,
      "author_name": "senbong",
      "author_url": "",
      "post_date": "09/13/2015 23:11:06",
      "content": "<p>[quote=NxGTR;92376]</p>\n\n<p>[quote=senbong;92373]</p>\n\n<p>A small suggestion to Kaggle is that, may be they can give some incentives to the player who first reports this type of leak (like 10% badge in the profile?) to detect this type of problems earlier.</p>\n\n<p>[/quote]</p>\n\n<p>Mmm... what about incentives to players whose workflow is not:</p>\n\n<p>1) Lets find some leaks</p>\n\n<p>2) Do actual ML</p>\n\n<p>:D</p>\n\n<p>[/quote]</p>\n\n<p>In my opinion, detecting this type of problem earlier is very important. The worst case is that someone keeps this leak to the later stage and uses it to win a competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 305511,
          "author_name": "work4diapers",
          "author_url": "",
          "post_date": "03/29/2018 01:14:05",
          "content": "<p>Hey bro</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 92379,
      "author_name": "hillbert",
      "author_url": "",
      "post_date": "09/13/2015 23:43:34",
      "content": "<p>If the competition is stopped and isn't rerun then what might be nice is the offer of some sort of small reward (swag perhaps?) for those who attained a good score without the use of the &quot;file_modified&quot; </p>\n\n<p>I only really entered this competition to try large scale analysis for the first time and learn from other people. It would be a shame for other people's work to go to waste simply due to a massive data leak.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92380,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "09/13/2015 23:45:10",
      "content": "<p>[quote=Giulio;92363]</p>\n\n<p>[quote=Triskelion;92358]</p>\n\n<p>What the final few % is I have no idea. Could be mislabelling. Could be LB noise. Could be the tiniest of edges we are now competing for.</p>\n\n<p>[/quote]</p>\n\n<p>How about the tiniest of of edges that is NOT worth competing for.</p>\n\n<p>[/quote]</p>\n\n<p>Looks to me the 10.000$ is still up for the taking. Kinda agree with your sentiment though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92389,
      "author_name": "lijunjie",
      "author_url": "",
      "post_date": "09/14/2015 03:55:03",
      "content": "<p>If the date is the leakage, how about the Kaggle admin add an additional restriction just for this contest that using this leakage in the model is invalid?  </p>\n\n<p>Actually, I don't think winning with this leakage has any value for the original purpose.  It's not even like a data mining technology to discover the leakage.</p>\n\n<p>And we should also pay respect to the participants who has already payed much of time to get a high ranking without using leakage information.</p>\n\n<p>Another good solution is to replace the original data set with a normal one.</p>\n\n<p>The worst solution is that you know it's harmful to the contest, and just let it go. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92392,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "09/14/2015 04:38:22",
      "content": "<p>[quote=YcdoiT;92389]</p>\n\n<p>If the date is the leakage, how about the Kaggle admin add an additional restriction just for this contest that using this leakage in the model is invalid?  </p>\n\n<p>Actually, I don't think winning with this leakage has any value for the original purpose.  It's not even like a data mining technology to discover the leakage.</p>\n\n<p>And we should also pay respect to the participants who has already payed much of time to get a high ranking without using leakage information.</p>\n\n<p>Another good solution is to replace the original data set with a normal one.</p>\n\n<p>The worst solution is that you know it's harmful to the contest, and just let it go. </p>\n\n<p>[/quote]</p>\n\n<p>Actually it will be quite difficult to go through everyone's codes to see if they used date variable or not. Alternatively, they can make two leaderboards, one just before leakage and one after and at least award the ranks/points, whichever is maximum of the two LB's.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92400,
      "author_name": "lijunjie",
      "author_url": "",
      "post_date": "09/14/2015 05:23:38",
      "content": "<p>@Subhajit Mandal, yeah, it's hard to tell if a person using this leakage unless he/she want to get the money(Top 1).\nSo, maybe replacing total data set is a good idea.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92402,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "09/14/2015 05:28:11",
      "content": "<p>Yeah theoretically, that's the best idea. But downloading the whole data, and extracting them, and running all the models which take a large amount of time and resources would be very difficult.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92405,
      "author_name": "senbong",
      "author_url": "",
      "post_date": "09/14/2015 05:33:26",
      "content": "<p>[quote=Subhajit Mandal;92402]</p>\n\n<p>Yeah theoretically, that's the best idea. But downloading the whole data, and extracting them, and running all the models which take a large amount of time and resources would be very difficult.</p>\n\n<p>[/quote]</p>\n\n<p>I think a new test dataset should be enough. We can still use the downloaded data as the training data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92408,
      "author_name": "bekbolatov",
      "author_url": "",
      "post_date": "09/14/2015 05:39:15",
      "content": "<p>It looked like this competition was organized very well: lots of data left for testing - fairer scoring.</p>\n\n<p>It is a unfortunate that labels were leaked and now it will have to be cancelled.</p>\n\n<p>Maybe organizers could save the competition, by coming up with more unseen test data and converting the data given so far to a &quot;train&quot; set?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92424,
      "author_name": "chonglinsun",
      "author_url": "",
      "post_date": "09/14/2015 06:14:10",
      "content": "<p>[quote=YcdoiT;92400]</p>\n\n<p>@Subhajit Mandal, yeah, it's hard to tell if a person using this leakage unless he/she want to get the money(Top 1).\nSo, maybe replacing total data set is a good idea.</p>\n\n<p>Yeah, whatever, at least 0.99805 of the test set was leaked.\n[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92425,
      "author_name": "chonglinsun",
      "author_url": "",
      "post_date": "09/14/2015 06:14:41",
      "content": "<p>[quote=YcdoiT;92400]</p>\n\n<p>@Subhajit Mandal, yeah, it's hard to tell if a person using this leakage unless he/she want to get the money(Top 1).\nSo, maybe replacing total data set is a good idea.</p>\n\n<p>[/quote]</p>\n\n<p>Yeah, whatever, at least 0.99805 of the test set was leaked.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92476,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "09/14/2015 13:55:35",
      "content": "<p>I'll vote for a restart with a new test data set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92478,
      "author_name": "sjuvekar",
      "author_url": "",
      "post_date": "09/14/2015 14:15:46",
      "content": "<p>Long silence from Dato/Kaggle suggests that they don't have additional data. In any case, I would expect creating datasets of these sizes (especially from 3rd party websites like StumbleUpon) should take a long time.</p>\n\n<p>I would vote for a 2-stage approach like Netflix, GE (and many other competitions on Kaggle). Keep the end-date of stage 1 on Sept. 10th --- that was the day the leak was discovered. Start stage-2 with much bigger (and cleaner) datasets when they are ready.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92486,
      "author_name": "dmcgarry",
      "author_url": "",
      "post_date": "09/14/2015 14:32:23",
      "content": "<p>I think that renaming all of the files and even jumbling the composition of the training/test sets would be sufficient. It's true that someone could still keep both sets of data and do an expensive file similarity to find the &quot;ground truth&quot; and cheat -- but anyone who spend the time to do that would be pathetic. Plus all prize eligible solutions need to be revealed so those couldn't exactly cheat. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92488,
      "author_name": "sjuvekar",
      "author_url": "",
      "post_date": "09/14/2015 14:37:50",
      "content": "<p>[quote=David McGarry;92486]</p>\n\n<p>I think that renaming all of the files and even jumbling the composition of the training/test sets would be sufficient. It's true that someone could still keep both sets of data and do an expensive file similarity to find the &quot;ground truth&quot; and cheat -- but anyone who spend the time to do that would be pathetic. Plus all prize eligible solutions need to be revealed so those couldn't exactly cheat. </p>\n\n<p>[/quote]</p>\n\n<p>File similarity is one thing, but that's not the only path to cheating. People could treat combined train+test dataset as a giant training dataset, overfit the model to it and achieve high score on LB. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92489,
      "author_name": "ilyanekhay",
      "author_url": "",
      "post_date": "09/14/2015 14:38:55",
      "content": "<p>Hashing 300k files and then matching by hash doesn't seem expensive to me. Actually, they can even be beautifulsoupped in a couple of hours on a single 8-core machine.  At the same time, feature selection done on ground truth might be quite effective and can't be proven to be a cheat.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92492,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "09/14/2015 15:12:59",
      "content": "<p>Guys, people work for Kaggle also take weekends off. So I believe if they are gonna do something, they will start doing it today. So we should just be patient.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92497,
      "author_name": "adjgiulio",
      "author_url": "",
      "post_date": "09/14/2015 15:45:36",
      "content": "<p>The question is whether there is any advantage for Kaggle and Dato to reset the competition. Dato hasn't put a lot of money on the the plate, which is indicative that this wasn't a must solve problem for them. Maybe more of a relatively cheap way to get some PR for their tool, but here I'm just speculating.\nAnyways, here are the few things that stood out to me so far, probably the only things I will remember about yet another flawed competition:</p>\n\n<p>-NxGTR's .974 performance without using the leak (great job!)</p>\n\n<p>-Dimitry's being the first finding the leak. I think, without him finding the problem, this might have gone unnoticed till the end. Great eye and I appreciate the honesty to put this out immediately instead of using it as a last day submission.</p>\n\n<p>-The sadness of seeing the LB full of new entrants as soon as the leak was made public. Like sharks attracted by blood in the water. Come on guys, there is more to life than a few Kaggle points!</p>\n\n<p>-Yet another flawed competition. I mean, I know that life sometimes throws curve balls, but it is interesting to observe how even companies that do DS &amp; ML for a living, can overlook huge flaws like the one we're now dealing with.</p>\n\n<p>Anyhow, however this goes, I hope everybody gets to enjoy the remainder of the competition. Good luck to all!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92500,
      "author_name": "dmcgarry",
      "author_url": "",
      "post_date": "09/14/2015 15:53:54",
      "content": "<p>[quote=Ilya Nekhay;92489]</p>\n\n<p>Hashing 300k files and then matching by hash doesn't seem expensive to me. Actually, they can even be beautifulsoupped in a couple of hours on a single 8-core machine.  At the same time, feature selection done on ground truth might be quite effective and can't be proven to be a cheat.</p>\n\n<p>[/quote]</p>\n\n<p>True, although &quot;a couple of hours&quot; to cheat sounds expensive to me, but perhaps I don't understand the motivations for some joining the competition. Plus introducing some innocuous white noise to files would make hashing more difficult, but again still completely solvable if someone really wanted to cheat. As for feature selection, given the size of the training set I'm not sure there is really that much to gain unless people are developing massively overfit features. I suppose some would, which is unfortunate.</p>\n\n<p>It would nice if the competition could be salvaged. I'm not sure how possible that is if more data isn't available given that my idea was more-or-less de-bunked.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92501,
      "author_name": "tobycheese",
      "author_url": "",
      "post_date": "09/14/2015 15:54:56",
      "content": "<p>[quote=Giulio;92497]</p>\n\n<p>-The sadness of seeing the LB full of new entrants as soon as the leak was made public. Like sharks attracted by blood in the water. Come on guys, there is more to life than a few Kaggle points!</p>\n\n<p>[/quote]</p>\n\n<p>While I agree in principle, I for one have never so far found nor exploited a leak in a kaggle contest, and for me it will be a nice (albeit small) exercise to do so (if the data is still up until I get a chance to download it all).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92503,
      "author_name": "bekbolatov",
      "author_url": "",
      "post_date": "09/14/2015 16:02:31",
      "content": "<p>[quote=Giulio;92497]</p>\n\n<p>The question is whether there is any advantage for Kaggle and Dato to reset the competition. Dato hasn't put a lot of money on the the plate, which is indicative that this wasn't a must solve problem for them. Maybe more of a relatively cheap way to get some PR for their tool, but here I'm just speculating.\nAnyways, here are the few things that stood out to me so far, probably the only things I will remember about yet another flawed competition:</p>\n\n<p>-NxGTR's .974 performance without using the leak (great job!)</p>\n\n<p>-Dimitry's being the first finding the leak. I think, without him finding the problem, this might have gone unnoticed till the end. Great eye and I appreciate the honesty to put this out immediately instead of using it as a last day submission.</p>\n\n<p>-The sadness of seeing the LB full of new entrants as soon as the leak was made public. Like sharks attracted by blood in the water. Come on guys, there is more to life than a few Kaggle points!</p>\n\n<p>-Yet another flawed competition. I mean, I know that life sometimes throws curve balls, but it is interesting to observe how even companies that do DS &amp; ML for a living, can overlook huge flaws like the one we're now dealing with.</p>\n\n<p>Anyhow, however this goes, I hope everybody gets to enjoy the remainder of the competition. Good luck to all!</p>\n\n<p>[/quote]</p>\n\n<p>I don't think that .974 is possible without using the leak - these webpages are well-disguised.\n@NxGTR - can you confirm that you didn't use the leak somehow/accidentally? I find it hard to believe, since these could be any webpages.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92505,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "09/14/2015 16:08:50",
      "content": "<p>[quote=bekbolatov;92503]</p>\n\n<p>I don't think that .974 is possible without using the leak - these webpages are well-disguised.\n@NxGTR - can you confirm that you didn't use the leak somehow/accidentally? I find it hard to believe, since these could be any webpages.</p>\n\n<p>[/quote]</p>\n\n<p>Apparently NxGTR's score is legit. 1) my dumb model is 0.95 already without the leak. 2) The leak is 0.99805 immediately. It's clear that he didn't know it at all.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92506,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "09/14/2015 16:09:55",
      "content": "<p>I got to 0.951 only using the content of the files, so I'm pretty sure it's possible. I made some decisions that worsened my model it seems, so I fell short after getting these results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92507,
      "author_name": "bekbolatov",
      "author_url": "",
      "post_date": "09/14/2015 16:18:31",
      "content": "<p>Thanks guys. Totally above my expectations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92509,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "09/14/2015 16:30:22",
      "content": "<p>Anyone who had actually work in this one will know that my score its just good and legit, no leaks. My second submission got me almost 0.96 already, so 0.97x is not unbelievable.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92512,
      "author_name": "bekbolatov",
      "author_url": "",
      "post_date": "09/14/2015 16:43:37",
      "content": "<p>It's all good, NxGTR. I haven't had time to try it out yet.</p>\n\n<p>I was just getting familiar with the data and initially it seemed pretty hard to have a good separation like that - turns out it isn't.</p>\n\n<p>I haven't decided whether I want to put an effort into this competition - given the easy separation.</p>\n\n<p>This means that practically all of the Native ads are detectable :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92519,
      "author_name": "mortehu",
      "author_url": "",
      "post_date": "09/14/2015 17:02:44",
      "content": "<p>I got a 0.97064 submission score with a model based on HTML alone.</p>\n\n<p>Mixing my content based model with the metadata model yields a slightly worse submission score (0.99826) than the metadata model alone (0.99846).</p>\n\n<p>The metadata based models worked by injecting tags like <code>&lt;file-meta-###/&gt;</code>, so that I wouldn't have to use a separate algorithm to mix the features.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "92249": "The near perfect result is pretty surprising to me, I  am just wondering if top three players are willing to tell it is label/tag leakage or not ?  I think there is no harm confirming this, after all, leakage is nothing related to violating competition rules.\r\n\r\nThe dataset is quite interesting, but I don't want to spend time on a  competition with leakage again, the west nile virus one is a  nightmare for me, and the tube price too. \r\n\r\nEdit: sorry, just find a typo in the title, I guess there is no way to fix that lol\r\n\r\nUpdate: I didn't expect people to get so angry (a few at me, I noticed I got some downvote) and disappointed about the issue. I have mixed feelings for this. But I don't remember I got so upset by not knowing the leakage of tube price ID after someone point it out after that competition was **over**, that one has more than 1000 players, yet only a few complaints, so I guess fewer upset player. So now I understand why people don't want to come out and confirm anything, because they want people to be happy. Sorry if I made people disappointed or upset by posting this topic.",
    "92252": "[quote=SkyLibrary;92249]\r\n\r\nThe near perfect result is pretty surprising to me, I  am just wondering if top three players are willing to tell it is label/tag leakage or not ?  I think there is no harm confirming this, after all, leakage is nothing related to violating competition rules.\r\n\r\nThe dataset is quite interesting, but I don't want to spend time on a  competition with leakage again, the west nile virus one is a  nightmare for me, and the tube price too. \r\n\r\nEdit: sorry, just find typo in title, guess this is no way to fix that lol\r\n\r\n[/quote]\r\n\r\nIt would be unfair with the 2 above me if I tell you anything. Last time I posted high model in Otto competition and everyone was just behind me. Someone from top 10 even criticized me by messaging me personally that because of me they are losing rank on Leader board.  \r\n\r\nI hope you can understand that I don't want to be involve in controversy every time :)",
    "92254": "Just wait few weeks :)",
    "92257": "It is pretty much a yes or no question, if you guys don't want to answer it, then I will assume the answer is yes.",
    "92276": "Why would they have to say anything? It is already a HUGE advantage for everybody to know that it is possible to achieve that score. I dropped the competition because I believe it is highly likely that this is an exercise in finding a deterministic rule in the dataset, and that treasure hunt is not worth my time. I could be wrong though.",
    "92280": "I didn't say they should answer my question, but I have the right to post the question no matter they reply or not. If they do, it will confirm your 'highly likely', if not, then it is still highly likely :)\r\n\r\n[quote=Giulio;92276]\r\n\r\nWhy would they have to say anything? It is already a HUGE advantage for everybody to know that it is possible to achieve that score. I dropped the competition because I believe it is highly likely that this is an exercise in finding a deterministic rule in the dataset, and that treasure hunt is not worth my time. I could be wrong though. \r\n\r\n[/quote]",
    "92285": "My guess is:\r\n\r\nThey are running an unsupervised algorithm over both training and test that generates relationships between elements in both datasets. Let's call this a vocabulary. This thus creates highly abstract relationships between all documents at word level, but the difference now is that this is across both train and test.\r\n\r\nThen for training, create (meaningful) document vectors from this vocabulary and start training to indicate sponsored / not sponsored. The difference is that you're not training on occurrences of individual features of a document (which can be meaningful or not), but you're training against the known vector space of all documents (with 1/3 of the data).",
    "92294": "[quote=DataGeek;92252]\r\n\r\n[quote=SkyLibrary;92249]\r\n\r\nThe near perfect result is pretty surprising to me, I  am just wondering if top three players are willing to tell it is label/tag leakage or not ?  I think there is no harm confirming this, after all, leakage is nothing related to violating competition rules.\r\n\r\nThe dataset is quite interesting, but I don't want to spend time on a  competition with leakage again, the west nile virus one is a  nightmare for me, and the tube price too. \r\n\r\nEdit: sorry, just find typo in title, guess this is no way to fix that lol\r\n\r\n[/quote]\r\n\r\nIt would be unfair with the 2 above me if I tell you anything. Last time I posted high model in Otto competition and everyone was just behind me. Someone from top 10 even criticized me by messaging me personally that because of me they are losing rank on Leader board.  \r\n\r\nI hope you can understand that I don't want to be involve in controversy every time :)\r\n\r\n[/quote]\r\n\r\nI think more and more people are getting offended by this, and joining the 0.99 club. :)",
    "92299": "After download the data and play with it for 5 hours, I am here to confirm the answer is yes.",
    "92301": "Well this is fun, and here I thought this would be the one competition that is actually like the old days since the data is too large for scripts. Silly me.",
    "92303": "Interesting. I wonder if Kaggle will put a hold on this or just let it be it.",
    "92304": "I am a noob in machine learning. I played with natural language processing algorithms. I achieved .76 score then I found that beside language processing, we can have better result if we used other features in the HTML files such as: javascript, meta tag attribute...",
    "92305": "So there is little space for the improvement after you applying the leakage information into your model.",
    "92308": "For those of you who think this competition doesn't make sense now -- it is really a whole different world in tuning your model, when you have one almost perfect predictor. Just try to go beyond the level this predictor gives you, it is challenging.",
    "92313": "[quote=Dmitry Ulyanov;92308]\r\n\r\nFor those of you who think this competition doesn't make sense now -- it is really a whole different world in tuning your model, when you have one almost perfect predictor. Just try to go beyond the level this predictor gives you, it is challenging. \r\n\r\n[/quote]\r\n\r\nBut now you must know this \"perfect predictor\", which it's a leakage and not real data, in order to be successful in this competition.\r\n\r\nSo if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.",
    "92314": "This is unfair to WNV or Cat competitors to say that leakage exploitation was the only determinant in victory.  \r\nMany did exploit it (even myself in WNV, not knowing it was, while in Cat I discarded it believing I would make a methodology mistake), and we did not all achieve top scores.\r\n\r\nLeakage exploitation still asks for talent and skills.\r\nAlso, nothing tells that what you call leakage, as I said before, is just a very strong association rule in the dataset and was present from the start.",
    "92321": "My code that is looking at what I think is the leak is apparently very slow, lol.",
    "92322": "I thought that, since an ad is \"sponsored\", you'd find some kind of link or script in the files across all the sponsored ones (tracking links), but it isn't as easy as that. Just filtering on _gac for example gave me 0.546 and non-sponsored ones contain the same thing.  But I reckon it must be somewhere along those lines...",
    "92324": "[quote=SkyLibrary;92299]\r\n\r\nAfter download the data and play with it for 5 hours, I am here to confirm the answer is yes.\r\n\r\n[/quote]\r\nSo this is a leakage, and you have found it?",
    "92326": "[quote=clustifier;92313]\r\n\r\nSo if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.\r\n\r\n[/quote]\r\n\r\nEven with this leak it is still fair for players.\r\n\r\nThe real bad thing is that it is pointless now for anyone trying to win the Dato Creators Award with this leakage existing.",
    "92327": "I am sure that there is a leakage that can be used as a \"golden feature\", but only using this can not reach 1.0 on the public leader board (0.99805). Or, as pointed out on the other thread, there are some errors in this dataset so the \"golden feature\" is perfect but labels are wrong.\r\n\r\nI want to know what the admin is thinking about this since further explorer might be in vain.",
    "92331": "[quote=rcarson;92326]\r\n\r\n[quote=clustifier;92313]\r\n\r\nSo if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.\r\n\r\n[/quote]\r\n\r\nEven with this leak it is still fair for players.\r\n\r\nThe real bad thing is that it is pointless now for anyone trying to win the Dato Creators Award with this leakage existing.\r\n\r\n\r\n[/quote]\r\n\r\nDefine \"fair\". Sure, all players have the same opportunity to find and exploit leakage but this is supposed to be a machine learning competition. I would have never signed up if it was a race to find some mistake in the data. I'm now stuck in a demlina where I have to decide if I want to, (a) continue using machine learning to tackle a seemingly interesting problem, (b) give in and find the leakage or (c) say screw it, delete the large data files from my system and feel even worse about the post-scripts world of kaggle.",
    "92332": "I guess this has something to do with scripts inside html files.\r\n\r\nFor example, the function of ads' charging. But I'm not familiar with javascript unfortunately, lol.",
    "92333": "Stumbleupon doesn't require any additions to html page.",
    "92334": "[quote=David McGarry;92331]\r\n\r\n[quote=rcarson;92326]\r\n\r\n[quote=clustifier;92313]\r\n\r\nSo if it's a leakage, and according to @SkyLibrary it is, the fair thing to do is to restart this competition without this leakage.\r\n\r\n[/quote]\r\n\r\nEven with this leak it is still fair for players.\r\n\r\nThe real bad thing is that it is pointless now for anyone trying to win the Dato Creators Award with this leakage existing.\r\n\r\n\r\n[/quote]\r\n\r\nDefine \"fair\". Sure, all players have the same opportunity to find and exploit leakage but this is supposed to be a machine learning competition. I would have never signed up if it was a race to find some mistake in the data. I'm not stuck in a demlina where I have to decide if I want to, (a) continue using machine learning to tackle a seemingly interesting problem, (b) give in and find the leakage or (c) say screw it, delete the large data files from my system and feel even worse about the post-scripts world of kaggle.\r\n\r\n[/quote]\r\n\r\nOn the contrary, I found it fun to look for the leak ( I haven't found it yet). \r\n\r\nA little bit off the topic, I think a great part of kaggle competitions is not about machine learning. It is more about data mining. Finding soft or strong leak (good or golden features) is usually more dominant, given there are so many good out-of-box tools like xgboost for the learning part.\r\n\r\nIn my opinion, finding a leak is a great skill even in real world case. But I agree in this case, it is a mess that the leak is too strong..",
    "92336": "I think i found a  leakage, but it doesn't improve my score. LOL.",
    "92340": "Ok, let's play this out as if it was a real project.\r\n\r\nBusiness guy: we think we could really benefit from being able to use ML to classify pages with native add content.\r\n\r\nData Scientist: give me a week to look at the data and I will get back to you with a first evaluation of what I think is possible. I'll work with engineering to get some data.\r\n\r\n...a week later...\r\n\r\n(very excited, yet skeptical) Data Scientist: I think I found something in the data that allows a perfect prediction. This is XYZ feature.\r\n\r\nBusiness Guy: oh, well, that makes sense. Seems kind of obvious. Wonder what \"name of a peer/manager/director who is taking the heat\" thought when he came up with the idea of using ML for this... You know what, I don't think we need ML anymore.\r\n\r\n(disappointed) Data Scientist: well, I guess I'll go back to building that boring dashboard then. Too bad, I was so excited of finally using some of those ML skills I had learnt on Kaggle...\r\n\r\n\r\nMy point, when a perfect predictor exists, there is no reason to use ML because it is not an ML problem. This should have never made it to be an ML project/competition. Now, will the finding of XYZ feature be of value to Dato? Who knows? I would be surprised if it does. Will it still be interesting to try to maximize performance on misslabeld observations? I mean, why not, it's just a problem like any other. But, wow, not the problem I thought I was signing up for.\r\nGlad I haven't made a submission yet...",
    "92341": "[quote=David McGarry;92331]\r\n I'm now stuck in a demlina where I have to decide if I want to, (a) continue using machine learning to tackle a seemingly interesting problem, (b) give in and find the leakage or (c) say screw it, delete the large data files from my system and feel even worse about the post-scripts world of kaggle.\r\n\r\n[/quote]\r\n\r\n+1 for (c)",
    "92342": "Giulio, it's much worse than that... much worse :)  one of those things: \"soon as you see it...\" \r\n\r\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. I had been trying to improve from 0.951 before this without success. That by itself is already a very reasonable score I think.",
    "92343": "In the end, don't tell me it is a filename leak... ==",
    "92344": "[quote=Gerard Toonstra;92342]\r\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. [/quote]\r\n\r\nI am very sad for this competition... I can confirm 0.973x is achievable without losing my own standards.",
    "92345": "[quote=NxGTR;92344]\r\n\r\n[quote=Gerard Toonstra;92342]\r\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. [/quote]\r\n\r\nI am very sad for this competition... I can confirm 0.973x is achievable without losing my own standards.\r\n\r\n[/quote]\r\n\r\nI'm sad too,  especially because of time I spent on it in the last few weeks.",
    "92348": "[quote=NxGTR;92344]\r\n\r\n[quote=Gerard Toonstra;92342]\r\nWhat I am interested in is what NXGTR was using for his model before this happened. The real submissions up to 0.970 or so. [/quote]\r\n\r\nI am very sad for this competition... I can confirm 0.973x is achievable without losing my own standards.\r\n\r\n[/quote]\r\n\r\nI'm looking forward to an explanation of your approach, parametrizations and feature choices, if you decide to write it up.",
    "92349": "The date in the zip files looks suspect.",
    "92350": "yes, it looks like the 1's files were written only on June 24 or July 27th.",
    "92353": "[quote=Lawrence Chernin;92350]\r\n\r\nyes, it looks like the 1's files were written only on June 24 or July 27th.\r\n\r\n[/quote]\r\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.",
    "92355": "[quote=Little Boat;92353]\r\n\r\n[quote=Lawrence Chernin;92350]\r\n\r\nyes, it looks like the 1's files were written only on June 24 or July 27th.\r\n\r\n[/quote]\r\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  \r\n\r\n[/quote]\r\nThere is nothing to remove :)\r\n\r\n> StumbleUpon Ads is a native advertising platform that does not use\r\n> traditional ad units such as banners or text links. This means that on\r\n> StumbleUpon your ad is, quite simply, your URL. Since your web page\r\n> itself is the ad unit, no additional creative assets or copy are\r\n> needed to drive traffic. We drive directly to your page when a user\r\n> clicks the Stumble button.\r\n\r\n\r\n> http://ads.stumbleupon.com/wp-content/uploads/2014/09/CampaignCreationGuide.pdf",
    "92356": "Hmm. Then I don't see how that would be the leakage. Anyway, if I am bored enough to find out if it is, I will post it here. But I am hoping someone else can do it for us. :-)\r\n\r\n[quote=Artem;92355]\r\n\r\n[quote=Little Boat;92353]\r\n\r\n[quote=Lawrence Chernin;92350]\r\n\r\nyes, it looks like the 1's files were written only on June 24 or July 27th.\r\n\r\n[/quote]\r\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  \r\n\r\n[/quote]\r\nThere is nothing to remove :)\r\n\r\n> StumbleUpon Ads is a native advertising platform that does not use\r\n> traditional ad units such as banners or text links. This means that on\r\n> StumbleUpon your ad is, quite simply, your URL. Since your web page\r\n> itself is the ad unit, no additional creative assets or copy are\r\n> needed to drive traffic. We drive directly to your page when a user\r\n> clicks the Stumble button.\r\n\r\n[/quote]",
    "92357": "[quote=Little Boat;92356]\r\n\r\nHmm. Then I don't see how that would be the leakage. Anyway, if I am bored enough to find out if it is, I will post it here. But I am hoping someone else can do it for us. :-)\r\n\r\n[quote=Artem;92355]\r\n\r\n[quote=Little Boat;92353]\r\n\r\n[quote=Lawrence Chernin;92350]\r\n\r\nyes, it looks like the 1's files were written only on June 24 or July 27th.\r\n\r\n[/quote]\r\nIf that is the leakage, I think it might be because dato removed the 'sponsored' label in original html files or did something similar at the same time.  \r\n\r\n[/quote]\r\nThere is nothing to remove :)\r\n\r\n> StumbleUpon Ads is a native advertising platform that does not use\r\n> traditional ad units such as banners or text links. This means that on\r\n> StumbleUpon your ad is, quite simply, your URL. Since your web page\r\n> itself is the ad unit, no additional creative assets or copy are\r\n> needed to drive traffic. We drive directly to your page when a user\r\n> clicks the Stumble button.\r\n\r\n[/quote]\r\n[/quote]\r\nI checked it locally and it looks like true.\r\n\r\n\r\nAlso i found another \"leakage\" that can give 0.86+ without ml.",
    "92358": "There is definitely leakage and I don't see how this could be solved/repaired, short of giving us a completely new data set...\r\n\r\nWhat the final few % is I have no idea. Could be mislabelling. Could be LB noise. Could be the tiniest of edges we are now competing for.",
    "92359": "[quote=Lawrence Chernin;92350]\r\n\r\nyes, it looks like the 1's files were written only on June 24 or July 27th.\r\n\r\n[/quote]\r\n\r\nWell that's fun. Hopefully they clear the LB and release a new dataset (using new HTML files would be ideal, but at least renaming all of the existing HTML files would discourage all but the least honest to ignore the known values in the test set).",
    "92362": "Now I can confirm that the date is the leakage.",
    "92363": "[quote=Triskelion;92358]\r\n\r\nWhat the final few % is I have no idea. Could be mislabelling. Could be LB noise. Could be the tiniest of edges we are now competing for.\r\n\r\n[/quote]\r\n\r\nHow about the tiniest of of edges that is NOT worth competing for.",
    "92366": "The one I was looking at was an unusual amount of the file name numbers showing up in other files, but hadn't looked hard at it yet - the code is slow.\r\nMaybe just genuinely random, but seems odd if you look at it.",
    "92373": "A small suggestion to Kaggle is that, may be they can give some incentives to the player who first reports this type of leak (like 10% badge in the profile?) to detect this type of problems earlier.",
    "92376": "[quote=senbong;92373]\r\n\r\nA small suggestion to Kaggle is that, may be they can give some incentives to the player who first reports this type of leak (like 10% badge in the profile?) to detect this type of problems earlier.\r\n\r\n[/quote]\r\n\r\nMmm... what about incentives to players whose workflow is not:\r\n\r\n1) Lets find some leaks\r\n\r\n2) Do actual ML\r\n\r\n:D",
    "92377": "[quote=NxGTR;92376]\r\n\r\n[quote=senbong;92373]\r\n\r\nA small suggestion to Kaggle is that, may be they can give some incentives to the player who first reports this type of leak (like 10% badge in the profile?) to detect this type of problems earlier.\r\n\r\n[/quote]\r\n\r\nMmm... what about incentives to players whose workflow is not:\r\n\r\n1) Lets find some leaks\r\n\r\n2) Do actual ML\r\n\r\n:D\r\n\r\n\r\n[/quote]\r\n\r\n\r\nIn my opinion, detecting this type of problem earlier is very important. The worst case is that someone keeps this leak to the later stage and uses it to win a competition.",
    "92379": "If the competition is stopped and isn't rerun then what might be nice is the offer of some sort of small reward (swag perhaps?) for those who attained a good score without the use of the \"file_modified\" \r\n\r\nI only really entered this competition to try large scale analysis for the first time and learn from other people. It would be a shame for other people's work to go to waste simply due to a massive data leak.",
    "92380": "[quote=Giulio;92363]\r\n\r\n[quote=Triskelion;92358]\r\n\r\nWhat the final few % is I have no idea. Could be mislabelling. Could be LB noise. Could be the tiniest of edges we are now competing for.\r\n\r\n[/quote]\r\n\r\nHow about the tiniest of of edges that is NOT worth competing for.\r\n\r\n[/quote]\r\n\r\nLooks to me the 10.000$ is still up for the taking. Kinda agree with your sentiment though.",
    "92389": "If the date is the leakage, how about the Kaggle admin add an additional restriction just for this contest that using this leakage in the model is invalid?  \r\n\r\nActually, I don't think winning with this leakage has any value for the original purpose.  It's not even like a data mining technology to discover the leakage.\r\n\r\nAnd we should also pay respect to the participants who has already payed much of time to get a high ranking without using leakage information.\r\n\r\nAnother good solution is to replace the original data set with a normal one.\r\n\r\nThe worst solution is that you know it's harmful to the contest, and just let it go.",
    "92392": "[quote=YcdoiT;92389]\r\n\r\nIf the date is the leakage, how about the Kaggle admin add an additional restriction just for this contest that using this leakage in the model is invalid?  \r\n\r\nActually, I don't think winning with this leakage has any value for the original purpose.  It's not even like a data mining technology to discover the leakage.\r\n\r\nAnd we should also pay respect to the participants who has already payed much of time to get a high ranking without using leakage information.\r\n\r\nAnother good solution is to replace the original data set with a normal one.\r\n\r\nThe worst solution is that you know it's harmful to the contest, and just let it go. \r\n\r\n[/quote]\r\n\r\nActually it will be quite difficult to go through everyone's codes to see if they used date variable or not. Alternatively, they can make two leaderboards, one just before leakage and one after and at least award the ranks/points, whichever is maximum of the two LB's.",
    "92400": "Subhajit Mandal, yeah, it's hard to tell if a person using this leakage unless he/she want to get the money(Top 1).\r\nSo, maybe replacing total data set is a good idea.",
    "92402": "Yeah theoretically, that's the best idea. But downloading the whole data, and extracting them, and running all the models which take a large amount of time and resources would be very difficult.",
    "92405": "[quote=Subhajit Mandal;92402]\r\n\r\nYeah theoretically, that's the best idea. But downloading the whole data, and extracting them, and running all the models which take a large amount of time and resources would be very difficult.\r\n\r\n[/quote]\r\n\r\nI think a new test dataset should be enough. We can still use the downloaded data as the training data.",
    "92408": "It looked like this competition was organized very well: lots of data left for testing - fairer scoring.\r\n\r\nIt is a unfortunate that labels were leaked and now it will have to be cancelled.\r\n\r\nMaybe organizers could save the competition, by coming up with more unseen test data and converting the data given so far to a \"train\" set?",
    "92424": "[quote=YcdoiT;92400]\r\n\r\n@Subhajit Mandal, yeah, it's hard to tell if a person using this leakage unless he/she want to get the money(Top 1).\r\nSo, maybe replacing total data set is a good idea.\r\n\r\nYeah, whatever, at least 0.99805 of the test set was leaked.\r\n[/quote]",
    "92425": "[quote=YcdoiT;92400]\r\n\r\n@Subhajit Mandal, yeah, it's hard to tell if a person using this leakage unless he/she want to get the money(Top 1).\r\nSo, maybe replacing total data set is a good idea.\r\n\r\n[/quote]\r\n\r\nYeah, whatever, at least 0.99805 of the test set was leaked.",
    "92476": "I'll vote for a restart with a new test data set.",
    "92478": "Long silence from Dato/Kaggle suggests that they don't have additional data. In any case, I would expect creating datasets of these sizes (especially from 3rd party websites like StumbleUpon) should take a long time.\r\n\r\nI would vote for a 2-stage approach like Netflix, GE (and many other competitions on Kaggle). Keep the end-date of stage 1 on Sept. 10th --- that was the day the leak was discovered. Start stage-2 with much bigger (and cleaner) datasets when they are ready.",
    "92486": "I think that renaming all of the files and even jumbling the composition of the training/test sets would be sufficient. It's true that someone could still keep both sets of data and do an expensive file similarity to find the \"ground truth\" and cheat -- but anyone who spend the time to do that would be pathetic. Plus all prize eligible solutions need to be revealed so those couldn't exactly cheat.",
    "92488": "[quote=David McGarry;92486]\r\n\r\nI think that renaming all of the files and even jumbling the composition of the training/test sets would be sufficient. It's true that someone could still keep both sets of data and do an expensive file similarity to find the \"ground truth\" and cheat -- but anyone who spend the time to do that would be pathetic. Plus all prize eligible solutions need to be revealed so those couldn't exactly cheat. \r\n\r\n[/quote]\r\n\r\nFile similarity is one thing, but that's not the only path to cheating. People could treat combined train+test dataset as a giant training dataset, overfit the model to it and achieve high score on LB.",
    "92489": "Hashing 300k files and then matching by hash doesn't seem expensive to me. Actually, they can even be beautifulsoupped in a couple of hours on a single 8-core machine.  At the same time, feature selection done on ground truth might be quite effective and can't be proven to be a cheat.",
    "92492": "Guys, people work for Kaggle also take weekends off. So I believe if they are gonna do something, they will start doing it today. So we should just be patient.",
    "92497": "The question is whether there is any advantage for Kaggle and Dato to reset the competition. Dato hasn't put a lot of money on the the plate, which is indicative that this wasn't a must solve problem for them. Maybe more of a relatively cheap way to get some PR for their tool, but here I'm just speculating.\r\nAnyways, here are the few things that stood out to me so far, probably the only things I will remember about yet another flawed competition:\r\n\r\n-NxGTR's .974 performance without using the leak (great job!)\r\n\r\n-Dimitry's being the first finding the leak. I think, without him finding the problem, this might have gone unnoticed till the end. Great eye and I appreciate the honesty to put this out immediately instead of using it as a last day submission.\r\n\r\n-The sadness of seeing the LB full of new entrants as soon as the leak was made public. Like sharks attracted by blood in the water. Come on guys, there is more to life than a few Kaggle points!\r\n\r\n-Yet another flawed competition. I mean, I know that life sometimes throws curve balls, but it is interesting to observe how even companies that do DS & ML for a living, can overlook huge flaws like the one we're now dealing with.\r\n\r\nAnyhow, however this goes, I hope everybody gets to enjoy the remainder of the competition. Good luck to all!",
    "92500": "[quote=Ilya Nekhay;92489]\r\n\r\nHashing 300k files and then matching by hash doesn't seem expensive to me. Actually, they can even be beautifulsoupped in a couple of hours on a single 8-core machine.  At the same time, feature selection done on ground truth might be quite effective and can't be proven to be a cheat.\r\n\r\n[/quote]\r\n\r\nTrue, although \"a couple of hours\" to cheat sounds expensive to me, but perhaps I don't understand the motivations for some joining the competition. Plus introducing some innocuous white noise to files would make hashing more difficult, but again still completely solvable if someone really wanted to cheat. As for feature selection, given the size of the training set I'm not sure there is really that much to gain unless people are developing massively overfit features. I suppose some would, which is unfortunate.\r\n\r\nIt would nice if the competition could be salvaged. I'm not sure how possible that is if more data isn't available given that my idea was more-or-less de-bunked.",
    "92501": "[quote=Giulio;92497]\r\n\r\n-The sadness of seeing the LB full of new entrants as soon as the leak was made public. Like sharks attracted by blood in the water. Come on guys, there is more to life than a few Kaggle points!\r\n\r\n[/quote]\r\n\r\nWhile I agree in principle, I for one have never so far found nor exploited a leak in a kaggle contest, and for me it will be a nice (albeit small) exercise to do so (if the data is still up until I get a chance to download it all).",
    "92503": "[quote=Giulio;92497]\r\n\r\nThe question is whether there is any advantage for Kaggle and Dato to reset the competition. Dato hasn't put a lot of money on the the plate, which is indicative that this wasn't a must solve problem for them. Maybe more of a relatively cheap way to get some PR for their tool, but here I'm just speculating.\r\nAnyways, here are the few things that stood out to me so far, probably the only things I will remember about yet another flawed competition:\r\n\r\n-NxGTR's .974 performance without using the leak (great job!)\r\n\r\n-Dimitry's being the first finding the leak. I think, without him finding the problem, this might have gone unnoticed till the end. Great eye and I appreciate the honesty to put this out immediately instead of using it as a last day submission.\r\n\r\n-The sadness of seeing the LB full of new entrants as soon as the leak was made public. Like sharks attracted by blood in the water. Come on guys, there is more to life than a few Kaggle points!\r\n\r\n-Yet another flawed competition. I mean, I know that life sometimes throws curve balls, but it is interesting to observe how even companies that do DS & ML for a living, can overlook huge flaws like the one we're now dealing with.\r\n\r\nAnyhow, however this goes, I hope everybody gets to enjoy the remainder of the competition. Good luck to all!\r\n\r\n[/quote]\r\n\r\nI don't think that .974 is possible without using the leak - these webpages are well-disguised.\r\n@NxGTR - can you confirm that you didn't use the leak somehow/accidentally? I find it hard to believe, since these could be any webpages.",
    "92505": "[quote=bekbolatov;92503]\r\n\r\nI don't think that .974 is possible without using the leak - these webpages are well-disguised.\r\n@NxGTR - can you confirm that you didn't use the leak somehow/accidentally? I find it hard to believe, since these could be any webpages.\r\n\r\n[/quote]\r\n\r\nApparently NxGTR's score is legit. 1) my dumb model is 0.95 already without the leak. 2) The leak is 0.99805 immediately. It's clear that he didn't know it at all.",
    "92506": "I got to 0.951 only using the content of the files, so I'm pretty sure it's possible. I made some decisions that worsened my model it seems, so I fell short after getting these results.",
    "92507": "Thanks guys. Totally above my expectations.",
    "92509": "Anyone who had actually work in this one will know that my score its just good and legit, no leaks. My second submission got me almost 0.96 already, so 0.97x is not unbelievable.",
    "92512": "It's all good, NxGTR. I haven't had time to try it out yet.\r\n\r\nI was just getting familiar with the data and initially it seemed pretty hard to have a good separation like that - turns out it isn't.\r\n\r\nI haven't decided whether I want to put an effort into this competition - given the easy separation.\r\n\r\nThis means that practically all of the Native ads are detectable :)",
    "92519": "I got a 0.97064 submission score with a model based on HTML alone.\r\n\r\nMixing my content based model with the metadata model yields a slightly worse submission score (0.99826) than the metadata model alone (0.99846).\r\n\r\nThe metadata based models worked by injecting tags like `<file-meta-###/>`, so that I wouldn't have to use a separate algorithm to mix the features.",
    "305511": "Hey bro"
  },
  "source": "meta"
}