{
  "id": 42087,
  "title": "Announcement: New test data will be released",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/42087",
  "author_name": "",
  "post_date": "2017-10-26T17:29:19.864254500Z",
  "votes": 8,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi all, </p>\n\n<p>We're very sorry for the leak - but we will try our best to fix it. We're working closely with KKBox to pull new test data for a different month (possibly April). We will release the new data as soon as we can, but please understand it takes time to gather, clean and validate new data. We do not have an estimated time at this point. </p>\n\n<p>In the mean time, please continue to use the train data to build model, and the test data can be used as your validation set. </p>\n\n<p>Kaggle admin</p>",
  "messages": [
    {
      "id": "236098",
      "postDate": "10/26/2017 17:29:19",
      "content": "<p>Hi all, </p>\n\n<p>We're very sorry for the leak - but we will try our best to fix it. We're working closely with KKBox to pull new test data for a different month (possibly April). We will release the new data as soon as we can, but please understand it takes time to gather, clean and validate new data. We do not have an estimated time at this point. </p>\n\n<p>In the mean time, please continue to use the train data to build model, and the test data can be used as your validation set. </p>\n\n<p>Kaggle admin</p>",
      "rawMarkdown": "Hi all, \n\nWe're very sorry for the leak - but we will try our best to fix it. We're working closely with KKBox to pull new test data for a different month (possibly April). We will release the new data as soon as we can, but please understand it takes time to gather, clean and validate new data. We do not have an estimated time at this point. \n\nIn the mean time, please continue to use the train data to build model, and the test data can be used as your validation set. \n\nKaggle admin",
      "votes": null
    },
    {
      "id": "236296",
      "postDate": "10/27/2017 02:34:07",
      "content": "<p>Hi Wendy, is it possible to release the labels of the test data set for our validation purpose. Since the zero score is only calculated over 40% of the data set and according to the current approach of getting 0 score, there is a huge discrepancy between the train and test distributions: 63471/992931 = 6.3% of the train users churned and 617460/970960 = 63.5% of the test users churned, it would be helpful for us to know the correct labels of the whole test set. Thanks.</p>",
      "rawMarkdown": "Hi Wendy, is it possible to release the labels of the test data set for our validation purpose. Since the zero score is only calculated over 40% of the data set and according to the current approach of getting 0 score, there is a huge discrepancy between the train and test distributions: 63471/992931 = 6.3% of the train users churned and 617460/970960 = 63.5% of the test users churned, it would be helpful for us to know the correct labels of the whole test set. Thanks.",
      "votes": null
    },
    {
      "id": "236762",
      "postDate": "10/28/2017 03:47:21",
      "content": "<p>It would be very helpful to my attempt to release an additional month of both transaction history and userLogs for the test users.  Predicting churn from two months out is a very different task in some ways than predicting a churn that is right around the corner.  Thanks for your efforts.</p>",
      "rawMarkdown": "It would be very helpful to my attempt to release an additional month of both transaction history and userLogs for the test users.  Predicting churn from two months out is a very different task in some ways than predicting a churn that is right around the corner.  Thanks for your efforts.",
      "votes": null
    },
    {
      "id": "236879",
      "postDate": "10/28/2017 14:49:58",
      "content": "<p>Yes, we are aware of the difficulties of predicting churn two months ahead of time. We will add more user listening log and transaction log entries to keep the task the same as we announced.</p>",
      "rawMarkdown": "Yes, we are aware of the difficulties of predicting churn two months ahead of time. We will add more user listening log and transaction log entries to keep the task the same as we announced.",
      "votes": null
    },
    {
      "id": "236926",
      "postDate": "10/28/2017 18:18:50",
      "content": "<p>I would like to recommend you to pay special attention to the 'expiration_date' column in 'members' table because, as it follows from one of your previous messages, \"it is possible to contain future information, but it may not give you much useful information\". However, it was the most important feature in every model I have tried. Moreover, some people had the substantial discrepancy between the local cv and public leaderboard scores, which might be caused by that feature. </p>",
      "rawMarkdown": "I would like to recommend you to pay special attention to the 'expiration_date' column in 'members' table because, as it follows from one of your previous messages, \"it is possible to contain future information, but it may not give you much useful information\". However, it was the most important feature in every model I have tried. Moreover, some people had the substantial discrepancy between the local cv and public leaderboard scores, which might be caused by that feature.",
      "votes": null
    },
    {
      "id": "238032",
      "postDate": "10/31/2017 14:38:54",
      "content": "<p>The top 40% of users in the test set are used for the LB score. That is, we cant say anything about the is_churn predictions for the bottom 60% rows ( 1 --&gt; 582460). <br>\nSo rows 582461 --&gt; 617460 are churned users. <br>\n617461 --&gt; 970960 are not churned. <br>\n(617460-582460)/(970960-582460) = 35000/388500 = 9% of users churned. </p>",
      "rawMarkdown": "The top 40% of users in the test set are used for the LB score. That is, we cant say anything about the is_churn predictions for the bottom 60% rows ( 1 --&gt; 582460).  \nSo rows 582461 --&gt; 617460 are churned users.  \n617461 --&gt; 970960 are not churned.  \n(617460-582460)/(970960-582460) = 35000/388500 = 9% of users churned.",
      "votes": null
    },
    {
      "id": "238276",
      "postDate": "11/01/2017 02:50:33",
      "content": "<p>Make sense Nicolas. We calculated the churn ratio from zero submission with score 3.11 and got similar result for churn rate: 9%.</p>",
      "rawMarkdown": "Make sense Nicolas. We calculated the churn ratio from zero submission with score 3.11 and got similar result for churn rate: 9%.",
      "votes": null
    },
    {
      "id": "239821",
      "postDate": "11/04/2017 17:25:45",
      "content": "<p>Are there any updates about the new test set?\nI would really appreciate if you can provide us with a time estimate.\nThanks a lot. </p>",
      "rawMarkdown": "Are there any updates about the new test set?\nI would really appreciate if you can provide us with a time estimate.\nThanks a lot.",
      "votes": null
    },
    {
      "id": "240224",
      "postDate": "11/06/2017 00:54:09",
      "content": "<p>Hey there,</p>\n\n<p>We're still working on collecting more data! Thanks for your patience! No time estimate yet, but soon!</p>",
      "rawMarkdown": "Hey there,\n\nWe're still working on collecting more data! Thanks for your patience! No time estimate yet, but soon!",
      "votes": null
    },
    {
      "id": "240834",
      "postDate": "11/07/2017 13:37:17",
      "content": "<p>Have you added old test set to the training data? It's clear that participants who have old data set have an advantage (more data) than new participants.</p>",
      "rawMarkdown": "Have you added old test set to the training data? It's clear that participants who have old data set have an advantage (more data) than new participants.",
      "votes": null
    },
    {
      "id": "240838",
      "postDate": "11/07/2017 13:47:55",
      "content": "<p>It seems like some new files are meant to completely replace the old files while others are supplementary (e.g. logs - the new file is significantly smaller than the old one). Could you be more clear in your data description/instructions?</p>",
      "rawMarkdown": "It seems like some new files are meant to completely replace the old files while others are supplementary (e.g. logs - the new file is significantly smaller than the old one). Could you be more clear in your data description/instructions?",
      "votes": null
    },
    {
      "id": "240923",
      "postDate": "11/07/2017 16:44:05",
      "content": "<p>Will there be a dealine update when the new data is released?</p>",
      "rawMarkdown": "Will there be a dealine update when the new data is released?",
      "votes": null
    },
    {
      "id": "241028",
      "postDate": "11/07/2017 21:33:26",
      "content": "<p>Bryan,</p>\n\n<p>Yup! We'll give a full update on deadlines, etc with the new release. Thanks for your patience!</p>",
      "rawMarkdown": "Bryan,\n\nYup! We'll give a full update on deadlines, etc with the new release. Thanks for your patience!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 236296,
      "author_name": "henryvu",
      "author_url": "",
      "post_date": "10/27/2017 02:34:07",
      "content": "<p>Hi Wendy, is it possible to release the labels of the test data set for our validation purpose. Since the zero score is only calculated over 40% of the data set and according to the current approach of getting 0 score, there is a huge discrepancy between the train and test distributions: 63471/992931 = 6.3% of the train users churned and 617460/970960 = 63.5% of the test users churned, it would be helpful for us to know the correct labels of the whole test set. Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 238032,
          "author_name": "npa02012",
          "author_url": "",
          "post_date": "10/31/2017 14:38:54",
          "content": "<p>The top 40% of users in the test set are used for the LB score. That is, we cant say anything about the is_churn predictions for the bottom 60% rows ( 1 --&gt; 582460). <br>\nSo rows 582461 --&gt; 617460 are churned users. <br>\n617461 --&gt; 970960 are not churned. <br>\n(617460-582460)/(970960-582460) = 35000/388500 = 9% of users churned. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 238276,
          "author_name": "henryvu",
          "author_url": "",
          "post_date": "11/01/2017 02:50:33",
          "content": "<p>Make sense Nicolas. We calculated the churn ratio from zero submission with score 3.11 and got similar result for churn rate: 9%.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 236762,
      "author_name": "stravinsky",
      "author_url": "",
      "post_date": "10/28/2017 03:47:21",
      "content": "<p>It would be very helpful to my attempt to release an additional month of both transaction history and userLogs for the test users.  Predicting churn from two months out is a very different task in some ways than predicting a churn that is right around the corner.  Thanks for your efforts.</p>",
      "votes": null,
      "replies": [
        {
          "id": 236879,
          "author_name": "ardenkkbox",
          "author_url": "",
          "post_date": "10/28/2017 14:49:58",
          "content": "<p>Yes, we are aware of the difficulties of predicting churn two months ahead of time. We will add more user listening log and transaction log entries to keep the task the same as we announced.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 236926,
          "author_name": "andreyjanzenx",
          "author_url": "",
          "post_date": "10/28/2017 18:18:50",
          "content": "<p>I would like to recommend you to pay special attention to the 'expiration_date' column in 'members' table because, as it follows from one of your previous messages, \"it is possible to contain future information, but it may not give you much useful information\". However, it was the most important feature in every model I have tried. Moreover, some people had the substantial discrepancy between the local cv and public leaderboard scores, which might be caused by that feature. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 239821,
      "author_name": "raniamibrahim",
      "author_url": "",
      "post_date": "11/04/2017 17:25:45",
      "content": "<p>Are there any updates about the new test set?\nI would really appreciate if you can provide us with a time estimate.\nThanks a lot. </p>",
      "votes": null,
      "replies": [
        {
          "id": 240224,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "11/06/2017 00:54:09",
          "content": "<p>Hey there,</p>\n\n<p>We're still working on collecting more data! Thanks for your patience! No time estimate yet, but soon!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 240834,
      "author_name": "",
      "author_url": "",
      "post_date": "11/07/2017 13:37:17",
      "content": "<p>Have you added old test set to the training data? It's clear that participants who have old data set have an advantage (more data) than new participants.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 240838,
      "author_name": "simonsti",
      "author_url": "",
      "post_date": "11/07/2017 13:47:55",
      "content": "<p>It seems like some new files are meant to completely replace the old files while others are supplementary (e.g. logs - the new file is significantly smaller than the old one). Could you be more clear in your data description/instructions?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 240923,
      "author_name": "puremath86",
      "author_url": "",
      "post_date": "11/07/2017 16:44:05",
      "content": "<p>Will there be a dealine update when the new data is released?</p>",
      "votes": null,
      "replies": [
        {
          "id": 241028,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "11/07/2017 21:33:26",
          "content": "<p>Bryan,</p>\n\n<p>Yup! We'll give a full update on deadlines, etc with the new release. Thanks for your patience!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "236098": "Hi all, \n\nWe're very sorry for the leak - but we will try our best to fix it. We're working closely with KKBox to pull new test data for a different month (possibly April). We will release the new data as soon as we can, but please understand it takes time to gather, clean and validate new data. We do not have an estimated time at this point. \n\nIn the mean time, please continue to use the train data to build model, and the test data can be used as your validation set. \n\nKaggle admin",
    "236296": "Hi Wendy, is it possible to release the labels of the test data set for our validation purpose. Since the zero score is only calculated over 40% of the data set and according to the current approach of getting 0 score, there is a huge discrepancy between the train and test distributions: 63471/992931 = 6.3% of the train users churned and 617460/970960 = 63.5% of the test users churned, it would be helpful for us to know the correct labels of the whole test set. Thanks.",
    "236762": "It would be very helpful to my attempt to release an additional month of both transaction history and userLogs for the test users.  Predicting churn from two months out is a very different task in some ways than predicting a churn that is right around the corner.  Thanks for your efforts.",
    "236879": "Yes, we are aware of the difficulties of predicting churn two months ahead of time. We will add more user listening log and transaction log entries to keep the task the same as we announced.",
    "236926": "I would like to recommend you to pay special attention to the 'expiration_date' column in 'members' table because, as it follows from one of your previous messages, \"it is possible to contain future information, but it may not give you much useful information\". However, it was the most important feature in every model I have tried. Moreover, some people had the substantial discrepancy between the local cv and public leaderboard scores, which might be caused by that feature.",
    "238032": "The top 40% of users in the test set are used for the LB score. That is, we cant say anything about the is_churn predictions for the bottom 60% rows ( 1 --&gt; 582460).  \nSo rows 582461 --&gt; 617460 are churned users.  \n617461 --&gt; 970960 are not churned.  \n(617460-582460)/(970960-582460) = 35000/388500 = 9% of users churned.",
    "238276": "Make sense Nicolas. We calculated the churn ratio from zero submission with score 3.11 and got similar result for churn rate: 9%.",
    "239821": "Are there any updates about the new test set?\nI would really appreciate if you can provide us with a time estimate.\nThanks a lot.",
    "240224": "Hey there,\n\nWe're still working on collecting more data! Thanks for your patience! No time estimate yet, but soon!",
    "240834": "Have you added old test set to the training data? It's clear that participants who have old data set have an advantage (more data) than new participants.",
    "240838": "It seems like some new files are meant to completely replace the old files while others are supplementary (e.g. logs - the new file is significantly smaller than the old one). Could you be more clear in your data description/instructions?",
    "240923": "Will there be a dealine update when the new data is released?",
    "241028": "Bryan,\n\nYup! We'll give a full update on deadlines, etc with the new release. Thanks for your patience!"
  },
  "source": "meta"
}