{
  "id": 41939,
  "title": "0.00000？leak?",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/41939",
  "author_name": "",
  "post_date": "2017-10-25T13:19:19.944030200Z",
  "votes": 8,
  "comment_count": 11,
  "views": 0,
  "content": "<p>How come at this time so many challengers have built perfect models(or discover the same msno from somewhere?) to this game? And almost at the same time? Am I missing anything? Because of another ongoing WSDM game?Two games use same msno? In such case there is no need to build models to predict.</p>",
  "messages": [
    {
      "id": "235454",
      "postDate": "10/25/2017 13:19:19",
      "content": "<p>How come at this time so many challengers have built perfect models(or discover the same msno from somewhere?) to this game? And almost at the same time? Am I missing anything? Because of another ongoing WSDM game?Two games use same msno? In such case there is no need to build models to predict.</p>",
      "rawMarkdown": "How come at this time so many challengers have built perfect models(or discover the same msno from somewhere?) to this game? And almost at the same time? Am I missing anything? Because of another ongoing WSDM game?Two games use same msno? In such case there is no need to build models to predict.",
      "votes": null
    },
    {
      "id": "235461",
      "postDate": "10/25/2017 13:55:05",
      "content": "<p>0:617459    is 1\n617459:       is 0   </p>",
      "rawMarkdown": "0:617459    is 1\n617459:       is 0",
      "votes": null
    },
    {
      "id": "235471",
      "postDate": "10/25/2017 14:23:16",
      "content": "<p>Organizers forgot to mix data sets. Check the initial train set and you will see that all churners are placed at the beginning. The same story with the test set. Unfortunately, the whole competition turned out to be a waste of time.</p>",
      "rawMarkdown": "Organizers forgot to mix data sets. Check the initial train set and you will see that all churners are placed at the beginning. The same story with the test set. Unfortunately, the whole competition turned out to be a waste of time.",
      "votes": null
    },
    {
      "id": "235499",
      "postDate": "10/25/2017 15:36:23",
      "content": "<p>What do you all think of the huge discrepancy between the train and test distributions? <br>\n6.3% of the train users churned -- 63471/992931 <br>\n63.5% of the test users churned -- 617460/970960    </p>\n\n<p>Edit: After making some strategic submissions, I am fairly certain that only the final 40% of the rows are used in calculating the LB.  </p>\n\n<p>This indicates that roughly 8.9% of the LB test users churned -- 34884/388384</p>",
      "rawMarkdown": "What do you all think of the huge discrepancy between the train and test distributions?  \n6.3% of the train users churned -- 63471/992931  \n63.5% of the test users churned -- 617460/970960    \n\nEdit: After making some strategic submissions, I am fairly certain that only the final 40% of the rows are used in calculating the LB.  \n\nThis indicates that roughly 8.9% of the LB test users churned -- 34884/388384",
      "votes": null
    },
    {
      "id": "235504",
      "postDate": "10/25/2017 15:44:11",
      "content": "<p>I'll just leave it here:</p>\n\n<pre><code>df_test = pd.read_csv('sample_submission_zero.csv')\ndf_test.loc[:617459, 'is_churn'] = 1\ndf_test.to_csv('res2.csv', index=False)\n</code></pre>",
      "rawMarkdown": "I'll just leave it here:\n\n    df_test = pd.read_csv('sample_submission_zero.csv')\n    df_test.loc[:617459, 'is_churn'] = 1\n    df_test.to_csv('res2.csv', index=False)",
      "votes": null
    },
    {
      "id": "235508",
      "postDate": "10/25/2017 16:02:46",
      "content": "<p>The leaderboard is calculated with last 40% of the test data.  We don't know the score of the rest part</p>",
      "rawMarkdown": "The leaderboard is calculated with last 40% of the test data.  We don't know the score of the rest part",
      "votes": null
    },
    {
      "id": "235535",
      "postDate": "10/25/2017 17:13:34",
      "content": "<p>Its unfortunate.  I only just started playing around with this data, but I did learn some things about predicting churn.  Having 400 million rows of customer behavior to build models and learn from is a huge win</p>",
      "rawMarkdown": "Its unfortunate.  I only just started playing around with this data, but I did learn some things about predicting churn.  Having 400 million rows of customer behavior to build models and learn from is a huge win",
      "votes": null
    },
    {
      "id": "235540",
      "postDate": "10/25/2017 17:24:11",
      "content": "<p>Hey everyone,</p>\n\n<p>Thanks for the vibrant discussion! We've been monitoring this for the past few days, are aware of it, and are looking into it further. We don't like it when this happens just as much as you do - it's no fun for anyone :).</p>\n\n<p>Stand by!</p>",
      "rawMarkdown": "Hey everyone,\n\nThanks for the vibrant discussion! We've been monitoring this for the past few days, are aware of it, and are looking into it further. We don't like it when this happens just as much as you do - it's no fun for anyone :).\n\nStand by!",
      "votes": null
    },
    {
      "id": "235712",
      "postDate": "10/25/2017 22:56:52",
      "content": "<p>But it still part of the same dataset. If it is because of not sufling than it will still be useless</p>",
      "rawMarkdown": "But it still part of the same dataset. If it is because of not sufling than it will still be useless",
      "votes": null
    },
    {
      "id": "235746",
      "postDate": "10/26/2017 01:03:34",
      "content": "<p>So one can use the LB score information and try submitting some times to figure out the turning point of is_churn. Will finally find out only the top 617460 are churners and reach 0.00000 score.</p>",
      "rawMarkdown": "So one can use the LB score information and try submitting some times to figure out the turning point of is_churn. Will finally find out only the top 617460 are churners and reach 0.00000 score.",
      "votes": null
    },
    {
      "id": "236020",
      "postDate": "10/26/2017 14:29:51",
      "content": "<p>really bad thing! I invested a lot of time in my lstm ... :-(</p>",
      "rawMarkdown": "really bad thing! I invested a lot of time in my lstm ... :-(",
      "votes": null
    },
    {
      "id": "236097",
      "postDate": "10/26/2017 17:28:08",
      "content": "<p>Seems totally doable to automatically: </p>\n\n<ul>\n<li><p>Randomize datasets and assign new IDs as a standard preprocessing step for non-forecasting competitions. For good measure: Calculate randomness/variance of train and test target data.</p></li>\n<li><p>Make sure file creation dates are reset to a single date.</p></li>\n<li><p>Remove train-test duplicates.</p></li>\n<li><p>Run a logistic regression/decision tree on every single feature and discard those with evaluations above a certain threshold.</p></li>\n</ul>\n\n<p>What am I missing here? It is not like the first time that train or test data order was indicative of the target.</p>\n\n<p>With lots of competitors, any remaining leakage will be found, sure. But this type of order-ID leakage seems so avoidable.</p>\n\n<p>My priors say this competition will be reset soon and resume with a new data set.</p>",
      "rawMarkdown": "Seems totally doable to automatically: \n\n- Randomize datasets and assign new IDs as a standard preprocessing step for non-forecasting competitions. For good measure: Calculate randomness/variance of train and test target data.\n\n- Make sure file creation dates are reset to a single date.\n\n- Remove train-test duplicates.\n\n- Run a logistic regression/decision tree on every single feature and discard those with evaluations above a certain threshold.\n\nWhat am I missing here? It is not like the first time that train or test data order was indicative of the target.\n\nWith lots of competitors, any remaining leakage will be found, sure. But this type of order-ID leakage seems so avoidable.\n\nMy priors say this competition will be reset soon and resume with a new data set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 235461,
      "author_name": "mseasd",
      "author_url": "",
      "post_date": "10/25/2017 13:55:05",
      "content": "<p>0:617459    is 1\n617459:       is 0   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 235471,
      "author_name": "andreyjanzenx",
      "author_url": "",
      "post_date": "10/25/2017 14:23:16",
      "content": "<p>Organizers forgot to mix data sets. Check the initial train set and you will see that all churners are placed at the beginning. The same story with the test set. Unfortunately, the whole competition turned out to be a waste of time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 235535,
          "author_name": "autokad",
          "author_url": "",
          "post_date": "10/25/2017 17:13:34",
          "content": "<p>Its unfortunate.  I only just started playing around with this data, but I did learn some things about predicting churn.  Having 400 million rows of customer behavior to build models and learn from is a huge win</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 235746,
          "author_name": "swimming7",
          "author_url": "",
          "post_date": "10/26/2017 01:03:34",
          "content": "<p>So one can use the LB score information and try submitting some times to figure out the turning point of is_churn. Will finally find out only the top 617460 are churners and reach 0.00000 score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 235499,
      "author_name": "npa02012",
      "author_url": "",
      "post_date": "10/25/2017 15:36:23",
      "content": "<p>What do you all think of the huge discrepancy between the train and test distributions? <br>\n6.3% of the train users churned -- 63471/992931 <br>\n63.5% of the test users churned -- 617460/970960    </p>\n\n<p>Edit: After making some strategic submissions, I am fairly certain that only the final 40% of the rows are used in calculating the LB.  </p>\n\n<p>This indicates that roughly 8.9% of the LB test users churned -- 34884/388384</p>",
      "votes": null,
      "replies": [
        {
          "id": 235508,
          "author_name": "shichence",
          "author_url": "",
          "post_date": "10/25/2017 16:02:46",
          "content": "<p>The leaderboard is calculated with last 40% of the test data.  We don't know the score of the rest part</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 235712,
          "author_name": "muhammadalfiansyah",
          "author_url": "",
          "post_date": "10/25/2017 22:56:52",
          "content": "<p>But it still part of the same dataset. If it is because of not sufling than it will still be useless</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 235504,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "10/25/2017 15:44:11",
      "content": "<p>I'll just leave it here:</p>\n\n<pre><code>df_test = pd.read_csv('sample_submission_zero.csv')\ndf_test.loc[:617459, 'is_churn'] = 1\ndf_test.to_csv('res2.csv', index=False)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 235540,
      "author_name": "addisonhoward",
      "author_url": "",
      "post_date": "10/25/2017 17:24:11",
      "content": "<p>Hey everyone,</p>\n\n<p>Thanks for the vibrant discussion! We've been monitoring this for the past few days, are aware of it, and are looking into it further. We don't like it when this happens just as much as you do - it's no fun for anyone :).</p>\n\n<p>Stand by!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 236020,
      "author_name": "patrickklein",
      "author_url": "",
      "post_date": "10/26/2017 14:29:51",
      "content": "<p>really bad thing! I invested a lot of time in my lstm ... :-(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 236097,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "10/26/2017 17:28:08",
      "content": "<p>Seems totally doable to automatically: </p>\n\n<ul>\n<li><p>Randomize datasets and assign new IDs as a standard preprocessing step for non-forecasting competitions. For good measure: Calculate randomness/variance of train and test target data.</p></li>\n<li><p>Make sure file creation dates are reset to a single date.</p></li>\n<li><p>Remove train-test duplicates.</p></li>\n<li><p>Run a logistic regression/decision tree on every single feature and discard those with evaluations above a certain threshold.</p></li>\n</ul>\n\n<p>What am I missing here? It is not like the first time that train or test data order was indicative of the target.</p>\n\n<p>With lots of competitors, any remaining leakage will be found, sure. But this type of order-ID leakage seems so avoidable.</p>\n\n<p>My priors say this competition will be reset soon and resume with a new data set.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "235454": "How come at this time so many challengers have built perfect models(or discover the same msno from somewhere?) to this game? And almost at the same time? Am I missing anything? Because of another ongoing WSDM game?Two games use same msno? In such case there is no need to build models to predict.",
    "235461": "0:617459    is 1\n617459:       is 0",
    "235471": "Organizers forgot to mix data sets. Check the initial train set and you will see that all churners are placed at the beginning. The same story with the test set. Unfortunately, the whole competition turned out to be a waste of time.",
    "235499": "What do you all think of the huge discrepancy between the train and test distributions?  \n6.3% of the train users churned -- 63471/992931  \n63.5% of the test users churned -- 617460/970960    \n\nEdit: After making some strategic submissions, I am fairly certain that only the final 40% of the rows are used in calculating the LB.  \n\nThis indicates that roughly 8.9% of the LB test users churned -- 34884/388384",
    "235504": "I'll just leave it here:\n\n    df_test = pd.read_csv('sample_submission_zero.csv')\n    df_test.loc[:617459, 'is_churn'] = 1\n    df_test.to_csv('res2.csv', index=False)",
    "235508": "The leaderboard is calculated with last 40% of the test data.  We don't know the score of the rest part",
    "235535": "Its unfortunate.  I only just started playing around with this data, but I did learn some things about predicting churn.  Having 400 million rows of customer behavior to build models and learn from is a huge win",
    "235540": "Hey everyone,\n\nThanks for the vibrant discussion! We've been monitoring this for the past few days, are aware of it, and are looking into it further. We don't like it when this happens just as much as you do - it's no fun for anyone :).\n\nStand by!",
    "235712": "But it still part of the same dataset. If it is because of not sufling than it will still be useless",
    "235746": "So one can use the LB score information and try submitting some times to figure out the turning point of is_churn. Will finally find out only the top 617460 are churners and reach 0.00000 score.",
    "236020": "really bad thing! I invested a lot of time in my lstm ... :-(",
    "236097": "Seems totally doable to automatically: \n\n- Randomize datasets and assign new IDs as a standard preprocessing step for non-forecasting competitions. For good measure: Calculate randomness/variance of train and test target data.\n\n- Make sure file creation dates are reset to a single date.\n\n- Remove train-test duplicates.\n\n- Run a logistic regression/decision tree on every single feature and discard those with evaluations above a certain threshold.\n\nWhat am I missing here? It is not like the first time that train or test data order was indicative of the target.\n\nWith lots of competitors, any remaining leakage will be found, sure. But this type of order-ID leakage seems so avoidable.\n\nMy priors say this competition will be reset soon and resume with a new data set."
  },
  "source": "meta"
}