{
  "id": 39672,
  "title": "Welcome!",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/39672",
  "author_name": "",
  "post_date": "2017-09-18T21:40:37.887982200Z",
  "votes": 18,
  "comment_count": 25,
  "views": 0,
  "content": "<p>On behalf of KKBox, WSDM planning staff, and Kaggle, I'd like to welcome you to this research competition. This is a series of two competition, the second one will launch in the coming weeks. </p>\n\n<p>Please ask any questions in this thread and we will answer them here. Good luck!</p>",
  "messages": [
    {
      "id": "222453",
      "postDate": "09/18/2017 21:40:37",
      "content": "<p>On behalf of KKBox, WSDM planning staff, and Kaggle, I'd like to welcome you to this research competition. This is a series of two competition, the second one will launch in the coming weeks. </p>\n\n<p>Please ask any questions in this thread and we will answer them here. Good luck!</p>",
      "rawMarkdown": "On behalf of KKBox, WSDM planning staff, and Kaggle, I'd like to welcome you to this research competition. This is a series of two competition, the second one will launch in the coming weeks. \n\nPlease ask any questions in this thread and we will answer them here. Good luck!",
      "votes": null
    },
    {
      "id": "222454",
      "postDate": "09/18/2017 21:58:38",
      "content": "<p>Thanks Wendy, Interesting competition! <br>\nJust for curiosity... is the KKBox benchmark a real thing? it performs very strage.</p>",
      "rawMarkdown": "Thanks Wendy, Interesting competition!  \nJust for curiosity... is the KKBox benchmark a real thing? it performs very strage.",
      "votes": null
    },
    {
      "id": "222505",
      "postDate": "09/19/2017 03:47:07",
      "content": "<p>KKBOX benchmark is used for validating the competition workflow and scoring mechanism. It is generated without any feature engineering and model selection process. Intuitively, it should only outperforms simple rule based models, e.g. guessing all 0s, 1s or random. </p>",
      "rawMarkdown": "KKBOX benchmark is used for validating the competition workflow and scoring mechanism. It is generated without any feature engineering and model selection process. Intuitively, it should only outperforms simple rule based models, e.g. guessing all 0s, 1s or random.",
      "votes": null
    },
    {
      "id": "222613",
      "postDate": "09/19/2017 12:43:17",
      "content": "<p>Hello, \nIs there needed to participate to this one to participate to the second competition ?\nThanks.</p>",
      "rawMarkdown": "Hello, \nIs there needed to participate to this one to participate to the second competition ?\nThanks.",
      "votes": null
    },
    {
      "id": "222628",
      "postDate": "09/19/2017 13:48:42",
      "content": "<p>Hi Wendy, both competitions do not award standard ranking points and tiers?</p>",
      "rawMarkdown": "Hi Wendy, both competitions do not award standard ranking points and tiers?",
      "votes": null
    },
    {
      "id": "222653",
      "postDate": "09/19/2017 15:54:34",
      "content": "<p>Hi Skanderbeg,</p>\n\n<p>No - each competition stands alone, you do not have to participate in both.</p>",
      "rawMarkdown": "Hi Skanderbeg,\n\nNo - each competition stands alone, you do not have to participate in both.",
      "votes": null
    },
    {
      "id": "222674",
      "postDate": "09/19/2017 17:03:27",
      "content": "<p>That is correct. </p>",
      "rawMarkdown": "That is correct.",
      "votes": null
    },
    {
      "id": "223667",
      "postDate": "09/23/2017 01:16:19",
      "content": "<p>Hi Wendy,  where I can take part in another task of music recommendation in WSDM CUP 2018?</p>",
      "rawMarkdown": "Hi Wendy,  where I can take part in another task of music recommendation in WSDM CUP 2018?",
      "votes": null
    },
    {
      "id": "225000",
      "postDate": "09/28/2017 03:13:55",
      "content": "<p>Hi Wendy \nI have a question about user ur5l+RJ9n6L7h96PqgBIOfgFxSM95YzhdrA8xS1NvRQ=,\nin train.csv ,this user is churn. but i see the detail transactions ,there will be a discrepany.\ncould you explain this phenomenon,\nthanks!</p>",
      "rawMarkdown": "Hi Wendy \nI have a question about user ur5l+RJ9n6L7h96PqgBIOfgFxSM95YzhdrA8xS1NvRQ=,\nin train.csv ,this user is churn. but i see the detail transactions ,there will be a discrepany.\ncould you explain this phenomenon,\nthanks!",
      "votes": null
    },
    {
      "id": "225016",
      "postDate": "09/28/2017 03:52:27",
      "content": "<p>Letting people see the detailed transaction data for users in train.csv is intentional.  Recall our target of prediction: A churn user is defined as a user who didn't make a membership renewal 30 days after his/her membership expires in March. I am quite certain that you will not find any transaction data after 2017-02-28. If you want to build a model to make a prediction, you must use history transaction data. The train.csv is a sample data set that tells you whether a user whose membership expired in Feb. made a service renewal within 30 day after the expiration date. Since the data set is imbalanced, we encourage participants to explore the two-year worth of transaction history to discover patterns of churn users.</p>",
      "rawMarkdown": "Letting people see the detailed transaction data for users in train.csv is intentional.  Recall our target of prediction: A churn user is defined as a user who didn't make a membership renewal 30 days after his/her membership expires in March. I am quite certain that you will not find any transaction data after 2017-02-28. If you want to build a model to make a prediction, you must use history transaction data. The train.csv is a sample data set that tells you whether a user whose membership expired in Feb. made a service renewal within 30 day after the expiration date. Since the data set is imbalanced, we encourage participants to explore the two-year worth of transaction history to discover patterns of churn users.",
      "votes": null
    },
    {
      "id": "225051",
      "postDate": "09/28/2017 05:14:29",
      "content": "<p>oh got it ！thanks Arden</p>",
      "rawMarkdown": "oh got it ！thanks Arden",
      "votes": null
    },
    {
      "id": "225105",
      "postDate": "09/28/2017 09:05:48",
      "content": "<p>\"Winners will present their findings at the WSDM conference February 6-8, 2018 in Los Angeles, CA\"</p>\n\n<p>What is the package provided for the winners? flight tickets, accomodation...</p>",
      "rawMarkdown": "\"Winners will present their findings at the WSDM conference February 6-8, 2018 in Los Angeles, CA\"\n\nWhat is the package provided for the winners? flight tickets, accomodation...",
      "votes": null
    },
    {
      "id": "227183",
      "postDate": "10/03/2017 22:01:49",
      "content": "<p>Hi @Wendy Kan! Can you please address theese questions to the organizers?</p>\n\n<p>Some clarification regarding the travel grants:</p>\n\n<p>Are master dregree students elegible?\nIs it necessary to have an advisor joining the team in order to be elegible?\nTo be elegible as advisor, is it necessary to have an teacher-student relationship with all the other teamates?\nStudents that have achieved their titles in the second semester of 2017 are elegible?</p>\n\n<p>Thank you in advance!</p>",
      "rawMarkdown": "Hi @Wendy Kan! Can you please address theese questions to the organizers?\n\nSome clarification regarding the travel grants:\n\nAre master dregree students elegible?\nIs it necessary to have an advisor joining the team in order to be elegible?\nTo be elegible as advisor, is it necessary to have an teacher-student relationship with all the other teamates?\nStudents that have achieved their titles in the second semester of 2017 are elegible?\n\nThank you in advance!",
      "votes": null
    },
    {
      "id": "228023",
      "postDate": "10/05/2017 18:21:48",
      "content": "<p>Please disregard, new to kaggle, and I misinterpreted the data.  Thanks!!</p>\n\n<p>Is there any way we can have the userLog data for the final period?  If the userLog data ended after the final transaction date, it was not included in the given data.  What this means in practice is that in the days to months immediately preceding the prediction of churn (it's variable), critical information is just dropped in a basically arbitrary fashion.</p>\n\n<p>I'm new to kaggle, so I understand if this sort of thing is against the rules or whatever.  I am not looking for future data, just up-to-date data.  My kernel graphs should explain what I mean.  (I hope!)</p>\n\n<p>Thanks for putting on this competition!</p>",
      "rawMarkdown": "Please disregard, new to kaggle, and I misinterpreted the data.  Thanks!!\n\nIs there any way we can have the userLog data for the final period?  If the userLog data ended after the final transaction date, it was not included in the given data.  What this means in practice is that in the days to months immediately preceding the prediction of churn (it's variable), critical information is just dropped in a basically arbitrary fashion.\n\nI'm new to kaggle, so I understand if this sort of thing is against the rules or whatever.  I am not looking for future data, just up-to-date data.  My kernel graphs should explain what I mean.  (I hope!)\n\nThanks for putting on this competition!",
      "votes": null
    },
    {
      "id": "228170",
      "postDate": "10/06/2017 02:32:30",
      "content": "<p>Hi,  raddar, </p>\n\n<p>We provided prizes for winners and travel grants for top 4 student teams that are ranked among top 10.   Flight tickets and accommodations are not included. \nFor teams already are winners are excluded from travel grant receivers. </p>",
      "rawMarkdown": "Hi,  raddar, \n\nWe provided prizes for winners and travel grants for top 4 student teams that are ranked among top 10.   Flight tickets and accommodations are not included. \nFor teams already are winners are excluded from travel grant receivers.",
      "votes": null
    },
    {
      "id": "228224",
      "postDate": "10/06/2017 06:19:27",
      "content": "<p>Hi, Alosio, \nThanks for your questions. The answers are listed below - </p>\n\n<p>Q: Are master dregree students elegible?\nA: Mater degree students are eligible.\nQ: Is it necessary to have an advisor joining the team in order to be eligible?\nA: No, it is not necessary \nQ: To be eligible as an advisor, is it necessary to have a teacher-student relationship with all the other teammates? \nA: No, it is not necessary, too. '\nQ: Students that have achieved their titles in the second semester of 2017 are eligible?\nA: \"Student participants\" means that the participant has not graduated while the competitions began.  </p>",
      "rawMarkdown": "Hi, Alosio, \nThanks for your questions. The answers are listed below - \n\nQ: Are master dregree students elegible?\nA: Mater degree students are eligible.\nQ: Is it necessary to have an advisor joining the team in order to be eligible?\nA: No, it is not necessary \nQ: To be eligible as an advisor, is it necessary to have a teacher-student relationship with all the other teammates? \nA: No, it is not necessary, too. '\nQ: Students that have achieved their titles in the second semester of 2017 are eligible?\nA: \"Student participants\" means that the participant has not graduated while the competitions began.",
      "votes": null
    },
    {
      "id": "229428",
      "postDate": "10/09/2017 15:46:01",
      "content": "<p>Thanks a lot!!!</p>",
      "rawMarkdown": "Thanks a lot!!!",
      "votes": null
    },
    {
      "id": "232166",
      "postDate": "10/17/2017 01:33:27",
      "content": "<p>I feel like this would have been a hell of a lot easier if all the data files were merged and the churns were clearly defined (0/1 or true/false). Especially since I'm doing analysis on a potato.</p>",
      "rawMarkdown": "I feel like this would have been a hell of a lot easier if all the data files were merged and the churns were clearly defined (0/1 or true/false). Especially since I'm doing analysis on a potato.",
      "votes": null
    },
    {
      "id": "232646",
      "postDate": "10/18/2017 02:01:34",
      "content": "<p>There is a very big dataset included in this competition (~30GB uncompressed).</p>\n\n<p>Can I extract the required info from this dataset using distributed computing (outside kaggle), and add that useful info as an additional dataset using the 'Upload a Dataset' option? So that everyone can use that useful information, and competition remains fair to even those who don't have access to distributed computing systems.</p>\n\n<p>I asked this question to Kaggle support, and they replied back with this:</p>\n\n<blockquote>\n  <p>The KKBox churn competition doesn’t allow external data, which this may be considered for this competition. However if you'd like a second opinion, I recommend asking this on the forums and we'll let the WSDM/KKBox team respond!</p>\n</blockquote>\n\n<p>@Wendy Kan, can you please advise us on that?</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "There is a very big dataset included in this competition (~30GB uncompressed).\n\nCan I extract the required info from this dataset using distributed computing (outside kaggle), and add that useful info as an additional dataset using the 'Upload a Dataset' option? So that everyone can use that useful information, and competition remains fair to even those who don't have access to distributed computing systems.\n\nI asked this question to Kaggle support, and they replied back with this:\n\n&gt; The KKBox churn competition doesn’t allow external data, which this may be considered for this competition. However if you'd like a second opinion, I recommend asking this on the forums and we'll let the WSDM/KKBox team respond!\n\n@Wendy Kan, can you please advise us on that?\n\nThanks.",
      "votes": null
    },
    {
      "id": "233006",
      "postDate": "10/19/2017 01:31:34",
      "content": "<p>Hi, below is a comment I left for a request by @lamthuy to open the training set generation scripts. </p>\n\n<p>'After spending so much time processing the transaction data for the users in the train set, I found out the majority (~90%) of them will be 'out of scope' based on the definition given in the competition. I don't know what's really happening in the test set. Also someone was asking the Log-loss question and how come the loss is not log(2) when submitting all 0.5 as the prediction. Then another question is the 'expiration date' feature and if it's futuristic. In the XGboost model, you can see it's 'the' most important factor for the prediction. It certainly looks like futuristic and should be removed from the data if so. When you have a model with so much difference in the 'validation' and final 'test' result, it points to high probabilities of issues with the data itself, such as they are very different distribution. These questions point to some serious issues of this competition. I hope more people upvote to catch the organizers' attention. If they are serious about it and try to get some good results please do respond to these serious concerns.'</p>\n\n<p>Below are the examples I am talking about for the first comment regarding the majority of the training data is actually out-of-scope for training purpose. As you see, all of them renew their membership right in Feb. This should be 'out-of-scope' by definition. And ~90% of users in the training set do that. With all due respect, if you are serious about the competition, please do address these concerns. Thanks.</p>\n\n<pre><code>                                       msno transaction_date membership_expire_date\n</code></pre>\n\n<p>0  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2016-11-16             2016-12-15\n1  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2016-12-15             2017-01-15\n2  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2017-01-15             2017-02-15\n3  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2017-02-15             2017-03-15\n                                            msno transaction_date membership_expire_date\n18  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-09-30             2016-11-19\n19  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-10-31             2016-12-19\n20  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-11-30             2017-01-19\n21  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-12-31             2017-02-19\n22  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2017-01-31             2017-03-19\n                                            msno transaction_date membership_expire_date\n56  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2016-10-15             2016-11-15\n57  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2016-11-15             2016-12-15\n58  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2016-12-15             2017-01-15\n59  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2017-01-15             2017-02-15\n60  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2017-02-15             2017-03-15</p>",
      "rawMarkdown": "Hi, below is a comment I left for a request by @lamthuy to open the training set generation scripts. \n\n'After spending so much time processing the transaction data for the users in the train set, I found out the majority (~90%) of them will be 'out of scope' based on the definition given in the competition. I don't know what's really happening in the test set. Also someone was asking the Log-loss question and how come the loss is not log(2) when submitting all 0.5 as the prediction. Then another question is the 'expiration date' feature and if it's futuristic. In the XGboost model, you can see it's 'the' most important factor for the prediction. It certainly looks like futuristic and should be removed from the data if so. When you have a model with so much difference in the 'validation' and final 'test' result, it points to high probabilities of issues with the data itself, such as they are very different distribution. These questions point to some serious issues of this competition. I hope more people upvote to catch the organizers' attention. If they are serious about it and try to get some good results please do respond to these serious concerns.'\n\nBelow are the examples I am talking about for the first comment regarding the majority of the training data is actually out-of-scope for training purpose. As you see, all of them renew their membership right in Feb. This should be 'out-of-scope' by definition. And ~90% of users in the training set do that. With all due respect, if you are serious about the competition, please do address these concerns. Thanks.\n\n                                           msno transaction_date membership_expire_date\n0  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2016-11-16             2016-12-15\n1  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2016-12-15             2017-01-15\n2  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2017-01-15             2017-02-15\n3  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2017-02-15             2017-03-15\n                                            msno transaction_date membership_expire_date\n18  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-09-30             2016-11-19\n19  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-10-31             2016-12-19\n20  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-11-30             2017-01-19\n21  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-12-31             2017-02-19\n22  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2017-01-31             2017-03-19\n                                            msno transaction_date membership_expire_date\n56  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2016-10-15             2016-11-15\n57  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2016-11-15             2016-12-15\n58  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2016-12-15             2017-01-15\n59  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2017-01-15             2017-02-15\n60  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2017-02-15             2017-03-15",
      "votes": null
    },
    {
      "id": "234008",
      "postDate": "10/22/2017 02:48:38",
      "content": "<p>For those transactions entries whose 'msno' and 'transaction_date' are the same, how can I tell the time order of them? Or which one is the final plan of the user on that day?</p>\n\n<p>Eg.\nxm6fmAfgZx1OYUXaJuHOObD0H2EAtIktv9NYIVlaTf4=\nThe user with 'msno' above have 4 transactions on date 20150311, but the expiration date is respectively 20150403, 20150410, 20150417, 20150424</p>",
      "rawMarkdown": "For those transactions entries whose 'msno' and 'transaction_date' are the same, how can I tell the time order of them? Or which one is the final plan of the user on that day?\n\nEg.\nxm6fmAfgZx1OYUXaJuHOObD0H2EAtIktv9NYIVlaTf4=\nThe user with 'msno' above have 4 transactions on date 20150311, but the expiration date is respectively 20150403, 20150410, 20150417, 20150424",
      "votes": null
    },
    {
      "id": "234011",
      "postDate": "10/22/2017 02:52:09",
      "content": "<p>Why is not 'payment_plan_days' equals to 'membership_expire_date' minus 'transaction_date'? Can you please give a formal definition?</p>\n\n<p>Besides, the 'plan_list_price' and 'actual _amount_paid' is also confusing. What do they mean?</p>",
      "rawMarkdown": "Why is not 'payment_plan_days' equals to 'membership_expire_date' minus 'transaction_date'? Can you please give a formal definition?\n\nBesides, the 'plan_list_price' and 'actual _amount_paid' is also confusing. What do they mean?",
      "votes": null
    },
    {
      "id": "234030",
      "postDate": "10/22/2017 03:56:04",
      "content": "<p>The result is ln(2), which is correct. I think you have regarded 'log' as log10 rather than loge.</p>",
      "rawMarkdown": "The result is ln(2), which is correct. I think you have regarded 'log' as log10 rather than loge.",
      "votes": null
    },
    {
      "id": "236658",
      "postDate": "10/27/2017 19:52:50",
      "content": "<p>This is the original post from the Log-loss question:</p>\n\n<p>\"Is the evaluation metric is actually Log-Loss? One way to check this is to set blindly is_churn = 0.5 for all msno in sample_submission_zero.csv for data submission. According to Log-Loss function, this will result to log(2) = 0.6931472 regardless the actual binary target yi values, either 0 or 1.</p>\n\n<p>I tried to submit this sample_submission_zero.csv with is_churn value is set at 0.5 for all rows and expecting to get a score of 0.6931472. I got a score of 1.70064 instead.\"</p>\n\n<p>ln(2) is 0.6931472 not 1.70064. I did not try to submit all 0.5 myself. But I am not mixing up between ln(2) and log10(2), log10(2) is 0.3010.</p>",
      "rawMarkdown": "This is the original post from the Log-loss question:\n\n\"Is the evaluation metric is actually Log-Loss? One way to check this is to set blindly is_churn = 0.5 for all msno in sample_submission_zero.csv for data submission. According to Log-Loss function, this will result to log(2) = 0.6931472 regardless the actual binary target yi values, either 0 or 1.\n\nI tried to submit this sample_submission_zero.csv with is_churn value is set at 0.5 for all rows and expecting to get a score of 0.6931472. I got a score of 1.70064 instead.\"\n\nln(2) is 0.6931472 not 1.70064. I did not try to submit all 0.5 myself. But I am not mixing up between ln(2) and log10(2), log10(2) is 0.3010.",
      "votes": null
    },
    {
      "id": "245892",
      "postDate": "11/20/2017 02:10:05",
      "content": "<p>Hello, Wendy as competition data changed we are now predicting churn of the members that have their latest expiration date within April. But what I've noticed in sample submission v2 is that 4454 members have their expiration date greater than 2017-04-31. Can you clarify the circumstances that determine which members to predict and which not to. Thanks</p>",
      "rawMarkdown": "Hello, Wendy as competition data changed we are now predicting churn of the members that have their latest expiration date within April. But what I've noticed in sample submission v2 is that 4454 members have their expiration date greater than 2017-04-31. Can you clarify the circumstances that determine which members to predict and which not to. Thanks",
      "votes": null
    },
    {
      "id": "252448",
      "postDate": "12/03/2017 00:05:51",
      "content": "<p>Is the presentation at the WSDM conference mandatory to receive the prize?</p>",
      "rawMarkdown": "Is the presentation at the WSDM conference mandatory to receive the prize?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 222454,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "09/18/2017 21:58:38",
      "content": "<p>Thanks Wendy, Interesting competition! <br>\nJust for curiosity... is the KKBox benchmark a real thing? it performs very strage.</p>",
      "votes": null,
      "replies": [
        {
          "id": 222505,
          "author_name": "ardenkkbox",
          "author_url": "",
          "post_date": "09/19/2017 03:47:07",
          "content": "<p>KKBOX benchmark is used for validating the competition workflow and scoring mechanism. It is generated without any feature engineering and model selection process. Intuitively, it should only outperforms simple rule based models, e.g. guessing all 0s, 1s or random. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 222613,
      "author_name": "skanderbeg",
      "author_url": "",
      "post_date": "09/19/2017 12:43:17",
      "content": "<p>Hello, \nIs there needed to participate to this one to participate to the second competition ?\nThanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 222653,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "09/19/2017 15:54:34",
          "content": "<p>Hi Skanderbeg,</p>\n\n<p>No - each competition stands alone, you do not have to participate in both.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 222628,
      "author_name": "cczaixian",
      "author_url": "",
      "post_date": "09/19/2017 13:48:42",
      "content": "<p>Hi Wendy, both competitions do not award standard ranking points and tiers?</p>",
      "votes": null,
      "replies": [
        {
          "id": 222674,
          "author_name": "wendykan",
          "author_url": "",
          "post_date": "09/19/2017 17:03:27",
          "content": "<p>That is correct. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 223667,
      "author_name": "bigheartc",
      "author_url": "",
      "post_date": "09/23/2017 01:16:19",
      "content": "<p>Hi Wendy,  where I can take part in another task of music recommendation in WSDM CUP 2018?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 225000,
      "author_name": "",
      "author_url": "",
      "post_date": "09/28/2017 03:13:55",
      "content": "<p>Hi Wendy \nI have a question about user ur5l+RJ9n6L7h96PqgBIOfgFxSM95YzhdrA8xS1NvRQ=,\nin train.csv ,this user is churn. but i see the detail transactions ,there will be a discrepany.\ncould you explain this phenomenon,\nthanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 225016,
          "author_name": "ardenkkbox",
          "author_url": "",
          "post_date": "09/28/2017 03:52:27",
          "content": "<p>Letting people see the detailed transaction data for users in train.csv is intentional.  Recall our target of prediction: A churn user is defined as a user who didn't make a membership renewal 30 days after his/her membership expires in March. I am quite certain that you will not find any transaction data after 2017-02-28. If you want to build a model to make a prediction, you must use history transaction data. The train.csv is a sample data set that tells you whether a user whose membership expired in Feb. made a service renewal within 30 day after the expiration date. Since the data set is imbalanced, we encourage participants to explore the two-year worth of transaction history to discover patterns of churn users.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225051,
          "author_name": "",
          "author_url": "",
          "post_date": "09/28/2017 05:14:29",
          "content": "<p>oh got it ！thanks Arden</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225105,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "09/28/2017 09:05:48",
      "content": "<p>\"Winners will present their findings at the WSDM conference February 6-8, 2018 in Los Angeles, CA\"</p>\n\n<p>What is the package provided for the winners? flight tickets, accomodation...</p>",
      "votes": null,
      "replies": [
        {
          "id": 228170,
          "author_name": "yianchenkkbox",
          "author_url": "",
          "post_date": "10/06/2017 02:32:30",
          "content": "<p>Hi,  raddar, </p>\n\n<p>We provided prizes for winners and travel grants for top 4 student teams that are ranked among top 10.   Flight tickets and accommodations are not included. \nFor teams already are winners are excluded from travel grant receivers. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 227183,
      "author_name": "aloisiodn",
      "author_url": "",
      "post_date": "10/03/2017 22:01:49",
      "content": "<p>Hi @Wendy Kan! Can you please address theese questions to the organizers?</p>\n\n<p>Some clarification regarding the travel grants:</p>\n\n<p>Are master dregree students elegible?\nIs it necessary to have an advisor joining the team in order to be elegible?\nTo be elegible as advisor, is it necessary to have an teacher-student relationship with all the other teamates?\nStudents that have achieved their titles in the second semester of 2017 are elegible?</p>\n\n<p>Thank you in advance!</p>",
      "votes": null,
      "replies": [
        {
          "id": 228224,
          "author_name": "yianchenkkbox",
          "author_url": "",
          "post_date": "10/06/2017 06:19:27",
          "content": "<p>Hi, Alosio, \nThanks for your questions. The answers are listed below - </p>\n\n<p>Q: Are master dregree students elegible?\nA: Mater degree students are eligible.\nQ: Is it necessary to have an advisor joining the team in order to be eligible?\nA: No, it is not necessary \nQ: To be eligible as an advisor, is it necessary to have a teacher-student relationship with all the other teammates? \nA: No, it is not necessary, too. '\nQ: Students that have achieved their titles in the second semester of 2017 are eligible?\nA: \"Student participants\" means that the participant has not graduated while the competitions began.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 229428,
          "author_name": "aloisiodn",
          "author_url": "",
          "post_date": "10/09/2017 15:46:01",
          "content": "<p>Thanks a lot!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 228023,
      "author_name": "stravinsky",
      "author_url": "",
      "post_date": "10/05/2017 18:21:48",
      "content": "<p>Please disregard, new to kaggle, and I misinterpreted the data.  Thanks!!</p>\n\n<p>Is there any way we can have the userLog data for the final period?  If the userLog data ended after the final transaction date, it was not included in the given data.  What this means in practice is that in the days to months immediately preceding the prediction of churn (it's variable), critical information is just dropped in a basically arbitrary fashion.</p>\n\n<p>I'm new to kaggle, so I understand if this sort of thing is against the rules or whatever.  I am not looking for future data, just up-to-date data.  My kernel graphs should explain what I mean.  (I hope!)</p>\n\n<p>Thanks for putting on this competition!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 232166,
      "author_name": "kelvinwop",
      "author_url": "",
      "post_date": "10/17/2017 01:33:27",
      "content": "<p>I feel like this would have been a hell of a lot easier if all the data files were merged and the churns were clearly defined (0/1 or true/false). Especially since I'm doing analysis on a potato.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 232646,
      "author_name": "anwartheravian",
      "author_url": "",
      "post_date": "10/18/2017 02:01:34",
      "content": "<p>There is a very big dataset included in this competition (~30GB uncompressed).</p>\n\n<p>Can I extract the required info from this dataset using distributed computing (outside kaggle), and add that useful info as an additional dataset using the 'Upload a Dataset' option? So that everyone can use that useful information, and competition remains fair to even those who don't have access to distributed computing systems.</p>\n\n<p>I asked this question to Kaggle support, and they replied back with this:</p>\n\n<blockquote>\n  <p>The KKBox churn competition doesn’t allow external data, which this may be considered for this competition. However if you'd like a second opinion, I recommend asking this on the forums and we'll let the WSDM/KKBox team respond!</p>\n</blockquote>\n\n<p>@Wendy Kan, can you please advise us on that?</p>\n\n<p>Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 233006,
      "author_name": "shunjiangxu",
      "author_url": "",
      "post_date": "10/19/2017 01:31:34",
      "content": "<p>Hi, below is a comment I left for a request by @lamthuy to open the training set generation scripts. </p>\n\n<p>'After spending so much time processing the transaction data for the users in the train set, I found out the majority (~90%) of them will be 'out of scope' based on the definition given in the competition. I don't know what's really happening in the test set. Also someone was asking the Log-loss question and how come the loss is not log(2) when submitting all 0.5 as the prediction. Then another question is the 'expiration date' feature and if it's futuristic. In the XGboost model, you can see it's 'the' most important factor for the prediction. It certainly looks like futuristic and should be removed from the data if so. When you have a model with so much difference in the 'validation' and final 'test' result, it points to high probabilities of issues with the data itself, such as they are very different distribution. These questions point to some serious issues of this competition. I hope more people upvote to catch the organizers' attention. If they are serious about it and try to get some good results please do respond to these serious concerns.'</p>\n\n<p>Below are the examples I am talking about for the first comment regarding the majority of the training data is actually out-of-scope for training purpose. As you see, all of them renew their membership right in Feb. This should be 'out-of-scope' by definition. And ~90% of users in the training set do that. With all due respect, if you are serious about the competition, please do address these concerns. Thanks.</p>\n\n<pre><code>                                       msno transaction_date membership_expire_date\n</code></pre>\n\n<p>0  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2016-11-16             2016-12-15\n1  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2016-12-15             2017-01-15\n2  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2017-01-15             2017-02-15\n3  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2017-02-15             2017-03-15\n                                            msno transaction_date membership_expire_date\n18  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-09-30             2016-11-19\n19  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-10-31             2016-12-19\n20  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-11-30             2017-01-19\n21  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-12-31             2017-02-19\n22  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2017-01-31             2017-03-19\n                                            msno transaction_date membership_expire_date\n56  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2016-10-15             2016-11-15\n57  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2016-11-15             2016-12-15\n58  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2016-12-15             2017-01-15\n59  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2017-01-15             2017-02-15\n60  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2017-02-15             2017-03-15</p>",
      "votes": null,
      "replies": [
        {
          "id": 234030,
          "author_name": "whooiif",
          "author_url": "",
          "post_date": "10/22/2017 03:56:04",
          "content": "<p>The result is ln(2), which is correct. I think you have regarded 'log' as log10 rather than loge.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 236658,
          "author_name": "shunjiangxu",
          "author_url": "",
          "post_date": "10/27/2017 19:52:50",
          "content": "<p>This is the original post from the Log-loss question:</p>\n\n<p>\"Is the evaluation metric is actually Log-Loss? One way to check this is to set blindly is_churn = 0.5 for all msno in sample_submission_zero.csv for data submission. According to Log-Loss function, this will result to log(2) = 0.6931472 regardless the actual binary target yi values, either 0 or 1.</p>\n\n<p>I tried to submit this sample_submission_zero.csv with is_churn value is set at 0.5 for all rows and expecting to get a score of 0.6931472. I got a score of 1.70064 instead.\"</p>\n\n<p>ln(2) is 0.6931472 not 1.70064. I did not try to submit all 0.5 myself. But I am not mixing up between ln(2) and log10(2), log10(2) is 0.3010.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 234008,
      "author_name": "whooiif",
      "author_url": "",
      "post_date": "10/22/2017 02:48:38",
      "content": "<p>For those transactions entries whose 'msno' and 'transaction_date' are the same, how can I tell the time order of them? Or which one is the final plan of the user on that day?</p>\n\n<p>Eg.\nxm6fmAfgZx1OYUXaJuHOObD0H2EAtIktv9NYIVlaTf4=\nThe user with 'msno' above have 4 transactions on date 20150311, but the expiration date is respectively 20150403, 20150410, 20150417, 20150424</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 234011,
      "author_name": "whooiif",
      "author_url": "",
      "post_date": "10/22/2017 02:52:09",
      "content": "<p>Why is not 'payment_plan_days' equals to 'membership_expire_date' minus 'transaction_date'? Can you please give a formal definition?</p>\n\n<p>Besides, the 'plan_list_price' and 'actual _amount_paid' is also confusing. What do they mean?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 245892,
      "author_name": "keremt",
      "author_url": "",
      "post_date": "11/20/2017 02:10:05",
      "content": "<p>Hello, Wendy as competition data changed we are now predicting churn of the members that have their latest expiration date within April. But what I've noticed in sample submission v2 is that 4454 members have their expiration date greater than 2017-04-31. Can you clarify the circumstances that determine which members to predict and which not to. Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 252448,
      "author_name": "donlima",
      "author_url": "",
      "post_date": "12/03/2017 00:05:51",
      "content": "<p>Is the presentation at the WSDM conference mandatory to receive the prize?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "222453": "On behalf of KKBox, WSDM planning staff, and Kaggle, I'd like to welcome you to this research competition. This is a series of two competition, the second one will launch in the coming weeks. \n\nPlease ask any questions in this thread and we will answer them here. Good luck!",
    "222454": "Thanks Wendy, Interesting competition!  \nJust for curiosity... is the KKBox benchmark a real thing? it performs very strage.",
    "222505": "KKBOX benchmark is used for validating the competition workflow and scoring mechanism. It is generated without any feature engineering and model selection process. Intuitively, it should only outperforms simple rule based models, e.g. guessing all 0s, 1s or random.",
    "222613": "Hello, \nIs there needed to participate to this one to participate to the second competition ?\nThanks.",
    "222628": "Hi Wendy, both competitions do not award standard ranking points and tiers?",
    "222653": "Hi Skanderbeg,\n\nNo - each competition stands alone, you do not have to participate in both.",
    "222674": "That is correct.",
    "223667": "Hi Wendy,  where I can take part in another task of music recommendation in WSDM CUP 2018?",
    "225000": "Hi Wendy \nI have a question about user ur5l+RJ9n6L7h96PqgBIOfgFxSM95YzhdrA8xS1NvRQ=,\nin train.csv ,this user is churn. but i see the detail transactions ,there will be a discrepany.\ncould you explain this phenomenon,\nthanks!",
    "225016": "Letting people see the detailed transaction data for users in train.csv is intentional.  Recall our target of prediction: A churn user is defined as a user who didn't make a membership renewal 30 days after his/her membership expires in March. I am quite certain that you will not find any transaction data after 2017-02-28. If you want to build a model to make a prediction, you must use history transaction data. The train.csv is a sample data set that tells you whether a user whose membership expired in Feb. made a service renewal within 30 day after the expiration date. Since the data set is imbalanced, we encourage participants to explore the two-year worth of transaction history to discover patterns of churn users.",
    "225051": "oh got it ！thanks Arden",
    "225105": "\"Winners will present their findings at the WSDM conference February 6-8, 2018 in Los Angeles, CA\"\n\nWhat is the package provided for the winners? flight tickets, accomodation...",
    "227183": "Hi @Wendy Kan! Can you please address theese questions to the organizers?\n\nSome clarification regarding the travel grants:\n\nAre master dregree students elegible?\nIs it necessary to have an advisor joining the team in order to be elegible?\nTo be elegible as advisor, is it necessary to have an teacher-student relationship with all the other teamates?\nStudents that have achieved their titles in the second semester of 2017 are elegible?\n\nThank you in advance!",
    "228023": "Please disregard, new to kaggle, and I misinterpreted the data.  Thanks!!\n\nIs there any way we can have the userLog data for the final period?  If the userLog data ended after the final transaction date, it was not included in the given data.  What this means in practice is that in the days to months immediately preceding the prediction of churn (it's variable), critical information is just dropped in a basically arbitrary fashion.\n\nI'm new to kaggle, so I understand if this sort of thing is against the rules or whatever.  I am not looking for future data, just up-to-date data.  My kernel graphs should explain what I mean.  (I hope!)\n\nThanks for putting on this competition!",
    "228170": "Hi,  raddar, \n\nWe provided prizes for winners and travel grants for top 4 student teams that are ranked among top 10.   Flight tickets and accommodations are not included. \nFor teams already are winners are excluded from travel grant receivers.",
    "228224": "Hi, Alosio, \nThanks for your questions. The answers are listed below - \n\nQ: Are master dregree students elegible?\nA: Mater degree students are eligible.\nQ: Is it necessary to have an advisor joining the team in order to be eligible?\nA: No, it is not necessary \nQ: To be eligible as an advisor, is it necessary to have a teacher-student relationship with all the other teammates? \nA: No, it is not necessary, too. '\nQ: Students that have achieved their titles in the second semester of 2017 are eligible?\nA: \"Student participants\" means that the participant has not graduated while the competitions began.",
    "229428": "Thanks a lot!!!",
    "232166": "I feel like this would have been a hell of a lot easier if all the data files were merged and the churns were clearly defined (0/1 or true/false). Especially since I'm doing analysis on a potato.",
    "232646": "There is a very big dataset included in this competition (~30GB uncompressed).\n\nCan I extract the required info from this dataset using distributed computing (outside kaggle), and add that useful info as an additional dataset using the 'Upload a Dataset' option? So that everyone can use that useful information, and competition remains fair to even those who don't have access to distributed computing systems.\n\nI asked this question to Kaggle support, and they replied back with this:\n\n&gt; The KKBox churn competition doesn’t allow external data, which this may be considered for this competition. However if you'd like a second opinion, I recommend asking this on the forums and we'll let the WSDM/KKBox team respond!\n\n@Wendy Kan, can you please advise us on that?\n\nThanks.",
    "233006": "Hi, below is a comment I left for a request by @lamthuy to open the training set generation scripts. \n\n'After spending so much time processing the transaction data for the users in the train set, I found out the majority (~90%) of them will be 'out of scope' based on the definition given in the competition. I don't know what's really happening in the test set. Also someone was asking the Log-loss question and how come the loss is not log(2) when submitting all 0.5 as the prediction. Then another question is the 'expiration date' feature and if it's futuristic. In the XGboost model, you can see it's 'the' most important factor for the prediction. It certainly looks like futuristic and should be removed from the data if so. When you have a model with so much difference in the 'validation' and final 'test' result, it points to high probabilities of issues with the data itself, such as they are very different distribution. These questions point to some serious issues of this competition. I hope more people upvote to catch the organizers' attention. If they are serious about it and try to get some good results please do respond to these serious concerns.'\n\nBelow are the examples I am talking about for the first comment regarding the majority of the training data is actually out-of-scope for training purpose. As you see, all of them renew their membership right in Feb. This should be 'out-of-scope' by definition. And ~90% of users in the training set do that. With all due respect, if you are serious about the competition, please do address these concerns. Thanks.\n\n                                           msno transaction_date membership_expire_date\n0  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2016-11-16             2016-12-15\n1  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2016-12-15             2017-01-15\n2  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2017-01-15             2017-02-15\n3  +++hVY1rZox/33YtvDgmKA2Frg/2qhkz12B9ylCvh8o=       2017-02-15             2017-03-15\n                                            msno transaction_date membership_expire_date\n18  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-09-30             2016-11-19\n19  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-10-31             2016-12-19\n20  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-11-30             2017-01-19\n21  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2016-12-31             2017-02-19\n22  +++l/EXNMLTijfLBa8p2TUVVVp2aFGSuUI/h7mLmthw=       2017-01-31             2017-03-19\n                                            msno transaction_date membership_expire_date\n56  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2016-10-15             2016-11-15\n57  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2016-11-15             2016-12-15\n58  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2016-12-15             2017-01-15\n59  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2017-01-15             2017-02-15\n60  ++/9R3sX37CjxbY/AaGvbwr3QkwElKBCtSvVzhCBDOk=       2017-02-15             2017-03-15",
    "234008": "For those transactions entries whose 'msno' and 'transaction_date' are the same, how can I tell the time order of them? Or which one is the final plan of the user on that day?\n\nEg.\nxm6fmAfgZx1OYUXaJuHOObD0H2EAtIktv9NYIVlaTf4=\nThe user with 'msno' above have 4 transactions on date 20150311, but the expiration date is respectively 20150403, 20150410, 20150417, 20150424",
    "234011": "Why is not 'payment_plan_days' equals to 'membership_expire_date' minus 'transaction_date'? Can you please give a formal definition?\n\nBesides, the 'plan_list_price' and 'actual _amount_paid' is also confusing. What do they mean?",
    "234030": "The result is ln(2), which is correct. I think you have regarded 'log' as log10 rather than loge.",
    "236658": "This is the original post from the Log-loss question:\n\n\"Is the evaluation metric is actually Log-Loss? One way to check this is to set blindly is_churn = 0.5 for all msno in sample_submission_zero.csv for data submission. According to Log-Loss function, this will result to log(2) = 0.6931472 regardless the actual binary target yi values, either 0 or 1.\n\nI tried to submit this sample_submission_zero.csv with is_churn value is set at 0.5 for all rows and expecting to get a score of 0.6931472. I got a score of 1.70064 instead.\"\n\nln(2) is 0.6931472 not 1.70064. I did not try to submit all 0.5 myself. But I am not mixing up between ln(2) and log10(2), log10(2) is 0.3010.",
    "245892": "Hello, Wendy as competition data changed we are now predicting churn of the members that have their latest expiration date within April. But what I've noticed in sample submission v2 is that 4454 members have their expiration date greater than 2017-04-31. Can you clarify the circumstances that determine which members to predict and which not to. Thanks",
    "252448": "Is the presentation at the WSDM conference mandatory to receive the prize?"
  },
  "source": "meta"
}