{
  "id": 45927,
  "title": "Leakage, Leakage, Leakages",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/45927",
  "author_name": "",
  "post_date": "2017-12-18T03:23:55.035382100Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Congrats to all of us for the finish. I'd like to list out all the leakages we have faced in this competition. </p>\n\n<p>(1) Train, test are not randomly split. As you knew that, all scores are zeros. See my screenshot.</p>\n\n<p>(2) Expiration date in members data</p>\n\n<p>(3) Mean scale to 0.0357 based on probing the LB score.</p>\n\n<p>And more... Can you share your findings here?</p>",
  "messages": [
    {
      "id": "259271",
      "postDate": "12/18/2017 03:23:55",
      "content": "<p>Congrats to all of us for the finish. I'd like to list out all the leakages we have faced in this competition. </p>\n\n<p>(1) Train, test are not randomly split. As you knew that, all scores are zeros. See my screenshot.</p>\n\n<p>(2) Expiration date in members data</p>\n\n<p>(3) Mean scale to 0.0357 based on probing the LB score.</p>\n\n<p>And more... Can you share your findings here?</p>",
      "rawMarkdown": "Congrats to all of us for the finish. I'd like to list out all the leakages we have faced in this competition. \n\n(1) Train, test are not randomly split. As you knew that, all scores are zeros. See my screenshot.\n\n(2) Expiration date in members data\n\n(3) Mean scale to 0.0357 based on probing the LB score.\n\nAnd more... Can you share your findings here?",
      "votes": null
    },
    {
      "id": "259274",
      "postDate": "12/18/2017 03:39:29",
      "content": "<p>One addition to (2) Expiration date in KKBox recommendation challenge members file. </p>\n\n<p>(4) Based on <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/45921#259258\">https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/45921#259258</a> I think the difference between members_v3 and members file in KKBox recommendation challenge should be treated as a leakage. </p>",
      "rawMarkdown": "One addition to (2) Expiration date in KKBox recommendation challenge members file. \n\n(4) Based on https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/45921#259258 I think the difference between members_v3 and members file in KKBox recommendation challenge should be treated as a leakage.",
      "votes": null
    },
    {
      "id": "259980",
      "postDate": "12/19/2017 10:58:41",
      "content": "<p>I don't consider this leakage, but the training_v2.csv labels and the labels generated by the scala program only agree on 40% of the cases.  Since the submission msnos are labeled using the scala file, models trained on the train_v2.csv file would be at a severe disadvantage.</p>\n\n<p>The scala file was public so using it was fair game.  However, I assumed that the provided training labels would be correct.</p>",
      "rawMarkdown": "I don't consider this leakage, but the training_v2.csv labels and the labels generated by the scala program only agree on 40% of the cases.  Since the submission msnos are labeled using the scala file, models trained on the train_v2.csv file would be at a severe disadvantage.\n\nThe scala file was public so using it was fair game.  However, I assumed that the provided training labels would be correct.",
      "votes": null
    },
    {
      "id": "260166",
      "postDate": "12/19/2017 18:52:19",
      "content": "<p>Agreed on all counts.  All leakage, and all were influential in improving the LB score for my solution (except #2 obviously, unless you count the members file data from the recommendation challenge, which is #4 that Hang mentioned).  Real world accuracy  would be quite less.</p>\n\n<p>To be fair, some of these leaks are common in any prediction competition with temporal elements.  The only way to prevent #3 for example would be to limit submissions of teams to just a few to prevent LB probing.  With 5 submissions a day and potentially multiple team members, LB probing is a fact of life.</p>\n\n<p>Quick question, did anyone have any success with finding any leakage in msno?  I thought I found some during EDA and started to get CV improvement using it but it never amounted to any LB improvement.</p>",
      "rawMarkdown": "Agreed on all counts.  All leakage, and all were influential in improving the LB score for my solution (except #2 obviously, unless you count the members file data from the recommendation challenge, which is #4 that Hang mentioned).  Real world accuracy  would be quite less.\n\nTo be fair, some of these leaks are common in any prediction competition with temporal elements.  The only way to prevent #3 for example would be to limit submissions of teams to just a few to prevent LB probing.  With 5 submissions a day and potentially multiple team members, LB probing is a fact of life.\n\nQuick question, did anyone have any success with finding any leakage in msno?  I thought I found some during EDA and started to get CV improvement using it but it never amounted to any LB improvement.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 259274,
      "author_name": "soundwaveli00",
      "author_url": "",
      "post_date": "12/18/2017 03:39:29",
      "content": "<p>One addition to (2) Expiration date in KKBox recommendation challenge members file. </p>\n\n<p>(4) Based on <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/45921#259258\">https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/45921#259258</a> I think the difference between members_v3 and members file in KKBox recommendation challenge should be treated as a leakage. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 259980,
      "author_name": "",
      "author_url": "",
      "post_date": "12/19/2017 10:58:41",
      "content": "<p>I don't consider this leakage, but the training_v2.csv labels and the labels generated by the scala program only agree on 40% of the cases.  Since the submission msnos are labeled using the scala file, models trained on the train_v2.csv file would be at a severe disadvantage.</p>\n\n<p>The scala file was public so using it was fair game.  However, I assumed that the provided training labels would be correct.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 260166,
      "author_name": "bryangregory",
      "author_url": "",
      "post_date": "12/19/2017 18:52:19",
      "content": "<p>Agreed on all counts.  All leakage, and all were influential in improving the LB score for my solution (except #2 obviously, unless you count the members file data from the recommendation challenge, which is #4 that Hang mentioned).  Real world accuracy  would be quite less.</p>\n\n<p>To be fair, some of these leaks are common in any prediction competition with temporal elements.  The only way to prevent #3 for example would be to limit submissions of teams to just a few to prevent LB probing.  With 5 submissions a day and potentially multiple team members, LB probing is a fact of life.</p>\n\n<p>Quick question, did anyone have any success with finding any leakage in msno?  I thought I found some during EDA and started to get CV improvement using it but it never amounted to any LB improvement.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "259271": "Congrats to all of us for the finish. I'd like to list out all the leakages we have faced in this competition. \n\n(1) Train, test are not randomly split. As you knew that, all scores are zeros. See my screenshot.\n\n(2) Expiration date in members data\n\n(3) Mean scale to 0.0357 based on probing the LB score.\n\nAnd more... Can you share your findings here?",
    "259274": "One addition to (2) Expiration date in KKBox recommendation challenge members file. \n\n(4) Based on https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/45921#259258 I think the difference between members_v3 and members file in KKBox recommendation challenge should be treated as a leakage.",
    "259980": "I don't consider this leakage, but the training_v2.csv labels and the labels generated by the scala program only agree on 40% of the cases.  Since the submission msnos are labeled using the scala file, models trained on the train_v2.csv file would be at a severe disadvantage.\n\nThe scala file was public so using it was fair game.  However, I assumed that the provided training labels would be correct.",
    "260166": "Agreed on all counts.  All leakage, and all were influential in improving the LB score for my solution (except #2 obviously, unless you count the members file data from the recommendation challenge, which is #4 that Hang mentioned).  Real world accuracy  would be quite less.\n\nTo be fair, some of these leaks are common in any prediction competition with temporal elements.  The only way to prevent #3 for example would be to limit submissions of teams to just a few to prevent LB probing.  With 5 submissions a day and potentially multiple team members, LB probing is a fact of life.\n\nQuick question, did anyone have any success with finding any leakage in msno?  I thought I found some during EDA and started to get CV improvement using it but it never amounted to any LB improvement."
  },
  "source": "meta"
}