{
  "id": 44404,
  "title": "Is this competition all about overfitting?",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/44404",
  "author_name": "",
  "post_date": "2017-11-28T07:34:01.732481900Z",
  "votes": 5,
  "comment_count": 27,
  "views": 0,
  "content": "<p>EDIT - Found the solution: my cross validation set should not include any msno used in the dev set. When this condition is met, the cross validation set-up reflects the submission set-up and therefore the cross validation logloss is similar to the leadboard score, i.e. pretty shit (but that's a different issue).  </p>\n\n<p>I get a large logloss difference between validation and leaderboard sets, so I wonder</p>\n\n<p>a) if others have been in a similar situation </p>\n\n<p>b) what I should do now</p>\n\n<p>Using features on the transaction data set, I run xgboost on one of 5 folds and get a logloss of c.0.056 on the validation set (see below). However when I use this model on the test set to make a submission, I get a poor score of c.0.177. I am surprised to see so much overfit. Before the data leak, I got roughly similar scores (dev logloss c.0.06 and leaderboard c.0.20), however we were using February churners to predict March churners. Since we now have Feb and March churners to predict another bunch of March churners, I thought that having more representative train data would help reduce overfit. </p>\n\n<p>I suppose my first question is - does it make sense and have others seen similar gaps between k-fold validation and leadboard submission? \nThen I am left with several options and I would appreciate any guidance / ideas on which option would make most sense for me:</p>\n\n<ul>\n<li><p>build features using the log data set. I think it would be a waste of time. Before the leak, only 1 or 2 log-related features would only slightly improve accuracy, other were useless, so the same might happen now. I am worried I might spend hours to get a tiny bit of validation improvement, translating into a tiny bit of leadboard improvement, say 0.016, which frankly is just as shit as my current 0.177</p></li>\n<li><p>use other xgboost folds to generate other test predictions, then average them to hopefully reduce overfit. EDIT: I have tried that but that did not help.</p></li>\n<li><p>use other models than xgboosst, e.g. rf or neural nets, then average with xgboost test predictions to hopefully reduce overfit</p></li>\n<li><p>Improve the cross validation strategy i.e. set up the dev and cv sets differently to reflect the train-LB set up; for example, for the development set use only msno not seen in the cv set. EDIT: that worked.</p></li>\n<li><p>any other idea?</p></li>\n</ul>\n\n<p><em><strong></strong></em><strong>**<em>*</em>**<em>*</em>**<em></em></strong><em> FOLD N. 1 <strong></strong></em><strong>**<em>*</em>**<em>*</em>****</strong></p>\n\n<p>[0] train-logloss:0.675067  test-logloss:0.675068</p>\n\n<p>Multiple eval metrics have been passed: 'test-logloss' will be used for early stopping.</p>\n\n<p>Will train until test-logloss hasn't improved in 50 rounds.</p>\n\n<p>[50]    train-logloss:0.242682  test-logloss:0.242829</p>\n\n<p>[100]   train-logloss:0.124368  test-logloss:0.124705</p>\n\n<p>[150]   train-logloss:0.084166  test-logloss:0.084569</p>\n\n<p>[200]   train-logloss:0.068575  test-logloss:0.06902</p>\n\n<p>[250]   train-logloss:0.062092  test-logloss:0.062567</p>\n\n<p>[300]   train-logloss:0.058893  test-logloss:0.059417</p>\n\n<p>[350]   train-logloss:0.057189  test-logloss:0.057756</p>\n\n<p>[399]   train-logloss:0.055572  test-logloss:0.056203</p>",
  "messages": [
    {
      "id": "249321",
      "postDate": "11/28/2017 07:34:01",
      "content": "<p>EDIT - Found the solution: my cross validation set should not include any msno used in the dev set. When this condition is met, the cross validation set-up reflects the submission set-up and therefore the cross validation logloss is similar to the leadboard score, i.e. pretty shit (but that's a different issue).  </p>\n\n<p>I get a large logloss difference between validation and leaderboard sets, so I wonder</p>\n\n<p>a) if others have been in a similar situation </p>\n\n<p>b) what I should do now</p>\n\n<p>Using features on the transaction data set, I run xgboost on one of 5 folds and get a logloss of c.0.056 on the validation set (see below). However when I use this model on the test set to make a submission, I get a poor score of c.0.177. I am surprised to see so much overfit. Before the data leak, I got roughly similar scores (dev logloss c.0.06 and leaderboard c.0.20), however we were using February churners to predict March churners. Since we now have Feb and March churners to predict another bunch of March churners, I thought that having more representative train data would help reduce overfit. </p>\n\n<p>I suppose my first question is - does it make sense and have others seen similar gaps between k-fold validation and leadboard submission? \nThen I am left with several options and I would appreciate any guidance / ideas on which option would make most sense for me:</p>\n\n<ul>\n<li><p>build features using the log data set. I think it would be a waste of time. Before the leak, only 1 or 2 log-related features would only slightly improve accuracy, other were useless, so the same might happen now. I am worried I might spend hours to get a tiny bit of validation improvement, translating into a tiny bit of leadboard improvement, say 0.016, which frankly is just as shit as my current 0.177</p></li>\n<li><p>use other xgboost folds to generate other test predictions, then average them to hopefully reduce overfit. EDIT: I have tried that but that did not help.</p></li>\n<li><p>use other models than xgboosst, e.g. rf or neural nets, then average with xgboost test predictions to hopefully reduce overfit</p></li>\n<li><p>Improve the cross validation strategy i.e. set up the dev and cv sets differently to reflect the train-LB set up; for example, for the development set use only msno not seen in the cv set. EDIT: that worked.</p></li>\n<li><p>any other idea?</p></li>\n</ul>\n\n<p><em><strong></strong></em><strong>**<em>*</em>**<em>*</em>**<em></em></strong><em> FOLD N. 1 <strong></strong></em><strong>**<em>*</em>**<em>*</em>****</strong></p>\n\n<p>[0] train-logloss:0.675067  test-logloss:0.675068</p>\n\n<p>Multiple eval metrics have been passed: 'test-logloss' will be used for early stopping.</p>\n\n<p>Will train until test-logloss hasn't improved in 50 rounds.</p>\n\n<p>[50]    train-logloss:0.242682  test-logloss:0.242829</p>\n\n<p>[100]   train-logloss:0.124368  test-logloss:0.124705</p>\n\n<p>[150]   train-logloss:0.084166  test-logloss:0.084569</p>\n\n<p>[200]   train-logloss:0.068575  test-logloss:0.06902</p>\n\n<p>[250]   train-logloss:0.062092  test-logloss:0.062567</p>\n\n<p>[300]   train-logloss:0.058893  test-logloss:0.059417</p>\n\n<p>[350]   train-logloss:0.057189  test-logloss:0.057756</p>\n\n<p>[399]   train-logloss:0.055572  test-logloss:0.056203</p>",
      "rawMarkdown": "EDIT - Found the solution: my cross validation set should not include any msno used in the dev set. When this condition is met, the cross validation set-up reflects the submission set-up and therefore the cross validation logloss is similar to the leadboard score, i.e. pretty shit (but that's a different issue).  \n\nI get a large logloss difference between validation and leaderboard sets, so I wonder\n\na) if others have been in a similar situation \n\nb) what I should do now\n\nUsing features on the transaction data set, I run xgboost on one of 5 folds and get a logloss of c.0.056 on the validation set (see below). However when I use this model on the test set to make a submission, I get a poor score of c.0.177. I am surprised to see so much overfit. Before the data leak, I got roughly similar scores (dev logloss c.0.06 and leaderboard c.0.20), however we were using February churners to predict March churners. Since we now have Feb and March churners to predict another bunch of March churners, I thought that having more representative train data would help reduce overfit. \n\nI suppose my first question is - does it make sense and have others seen similar gaps between k-fold validation and leadboard submission? \nThen I am left with several options and I would appreciate any guidance / ideas on which option would make most sense for me:\n\n- build features using the log data set. I think it would be a waste of time. Before the leak, only 1 or 2 log-related features would only slightly improve accuracy, other were useless, so the same might happen now. I am worried I might spend hours to get a tiny bit of validation improvement, translating into a tiny bit of leadboard improvement, say 0.016, which frankly is just as shit as my current 0.177\n\n- use other xgboost folds to generate other test predictions, then average them to hopefully reduce overfit. EDIT: I have tried that but that did not help.\n\n- use other models than xgboosst, e.g. rf or neural nets, then average with xgboost test predictions to hopefully reduce overfit\n\n- Improve the cross validation strategy i.e. set up the dev and cv sets differently to reflect the train-LB set up; for example, for the development set use only msno not seen in the cv set. EDIT: that worked.\n\n- any other idea?\n\n\n******************* FOLD N. 1 *******************\n\n[0]\ttrain-logloss:0.675067\ttest-logloss:0.675068\n\nMultiple eval metrics have been passed: 'test-logloss' will be used for early stopping.\n\nWill train until test-logloss hasn't improved in 50 rounds.\n\n[50]\ttrain-logloss:0.242682\ttest-logloss:0.242829\n\n[100]\ttrain-logloss:0.124368\ttest-logloss:0.124705\n\n[150]\ttrain-logloss:0.084166\ttest-logloss:0.084569\n\n[200]\ttrain-logloss:0.068575\ttest-logloss:0.06902\n\n[250]\ttrain-logloss:0.062092\ttest-logloss:0.062567\n\n[300]\ttrain-logloss:0.058893\ttest-logloss:0.059417\n\n[350]\ttrain-logloss:0.057189\ttest-logloss:0.057756\n\n[399]\ttrain-logloss:0.055572\ttest-logloss:0.056203",
      "votes": null
    },
    {
      "id": "249325",
      "postDate": "11/28/2017 07:37:26",
      "content": "<p>me too. hh  Just want to know the details in admins generating the test dataset。 </p>",
      "rawMarkdown": "me too. hh  Just want to know the details in admins generating the test dataset。",
      "votes": null
    },
    {
      "id": "249521",
      "postDate": "11/28/2017 16:57:15",
      "content": "<p>my results generated from xgboost also got a large gap between local and LB...</p>",
      "rawMarkdown": "my results generated from xgboost also got a large gap between local and LB...",
      "votes": null
    },
    {
      "id": "251620",
      "postDate": "12/01/2017 13:38:08",
      "content": "<p>XGBoost, average score:0.1245232, 5 folds.  PB: 0.11799</p>",
      "rawMarkdown": "XGBoost, average score:0.1245232, 5 folds.  PB: 0.11799",
      "votes": null
    },
    {
      "id": "251708",
      "postDate": "12/01/2017 15:30:43",
      "content": "<p>Thanks a lot for sharing some actual numbers, that's really helpful. May I ask if you have used simple cross-validation, i.e. simply by shuffling then splitting the entire train set (Feb and Mar months) into 5 parts? </p>\n\n<p>I ask because I have thought about building a cross validation set that only includes users not in the development set. I suspect that my overfit is partly due to having the same customer in both sets, while the test set includes unseen users.</p>\n\n<p>Another question is, do you use the raw expiration_date and membership_expiration_date columns as features? They do have a lot of predicting power (from a xgboost 'feature importance' perspective) but I have a feeling that they may cause PB overfitting</p>",
      "rawMarkdown": "Thanks a lot for sharing some actual numbers, that's really helpful. May I ask if you have used simple cross-validation, i.e. simply by shuffling then splitting the entire train set (Feb and Mar months) into 5 parts? \n\nI ask because I have thought about building a cross validation set that only includes users not in the development set. I suspect that my overfit is partly due to having the same customer in both sets, while the test set includes unseen users.\n\nAnother question is, do you use the raw expiration_date and membership_expiration_date columns as features? They do have a lot of predicting power (from a xgboost 'feature importance' perspective) but I have a feeling that they may cause PB overfitting",
      "votes": null
    },
    {
      "id": "251967",
      "postDate": "12/02/2017 01:39:16",
      "content": "<p>Have the same problem too. I split train data into train/test with 7:3. Then I used train data for 5-fold cv with mean logloss 0.0774. And then the logloss for testing data is 0.0767. But my PB is 0.15971. My guess is there is much difference between train data and test data in PB.</p>",
      "rawMarkdown": "Have the same problem too. I split train data into train/test with 7:3. Then I used train data for 5-fold cv with mean logloss 0.0774. And then the logloss for testing data is 0.0767. But my PB is 0.15971. My guess is there is much difference between train data and test data in PB.",
      "votes": null
    },
    {
      "id": "251985",
      "postDate": "12/02/2017 02:39:11",
      "content": "<p>Hi, I am not sure this comment is relevant to you since I don't know how you generated your training set.  However, based on my understanding of the rules, you can only use data up to Feb 28, 2017 for training.  You can then additionally use the March 2017 data for scoring the submissions.  Your post seems to suggest you are using the March data for training as well.   (Sorry if I have misinterpreted your post.)</p>\n\n<p>If so that would explain what you are seeing.   The submissions are looking for models that can make out-of-time predictions, not just out-of-sample predictions.  </p>",
      "rawMarkdown": "Hi, I am not sure this comment is relevant to you since I don't know how you generated your training set.  However, based on my understanding of the rules, you can only use data up to Feb 28, 2017 for training.  You can then additionally use the March 2017 data for scoring the submissions.  Your post seems to suggest you are using the March data for training as well.   (Sorry if I have misinterpreted your post.)\n\nIf so that would explain what you are seeing.   The submissions are looking for models that can make out-of-time predictions, not just out-of-sample predictions.",
      "votes": null
    },
    {
      "id": "252118",
      "postDate": "12/02/2017 09:29:01",
      "content": "<p>Hi, \nYour comment is 100% relevant. I use March data for training as suggested in Wendy Kan's post (\"Announcement: New test data released\"). Specifically, I use the test set that was previously used to rate submissions on the LB and for which admins have revealed the is_churn column.  Wendy's post suggested that this data was now available to train our models. You say otherwise. This post seems to support your opinion: <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363\">https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363</a></p>\n\n<p>Let's just park the question of what data is allowed because it's not clear and I still can't see how it ties back to the original problem. </p>\n\n<p>You say that what I see can be explained by the fact that I use Feb and Mar data for training, which implies that my model makes out-of-sample predictions. That's completely right. However the submission is for March users - just a different set of msno than those used in my training data. Hence submissions are <strong><em>not</em></strong> looking for models that can me out-of-time predictions - to be clear, they would look for out-of-time predictions <strong><em>only</em></strong> if I restricted my train set to February data.</p>",
      "rawMarkdown": "Hi, \nYour comment is 100% relevant. I use March data for training as suggested in Wendy Kan's post (\"Announcement: New test data released\"). Specifically, I use the test set that was previously used to rate submissions on the LB and for which admins have revealed the is_churn column.  Wendy's post suggested that this data was now available to train our models. You say otherwise. This post seems to support your opinion: https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363\n\nLet's just park the question of what data is allowed because it's not clear and I still can't see how it ties back to the original problem. \n\nYou say that what I see can be explained by the fact that I use Feb and Mar data for training, which implies that my model makes out-of-sample predictions. That's completely right. However the submission is for March users - just a different set of msno than those used in my training data. Hence submissions are ***not*** looking for models that can me out-of-time predictions - to be clear, they would look for out-of-time predictions ***only*** if I restricted my train set to February data.",
      "votes": null
    },
    {
      "id": "252138",
      "postDate": "12/02/2017 10:22:58",
      "content": "<p>Hi, I think the way to think about it is this way.</p>\n\n<p>If you look at the submission msno's, almost all of the membership expiration dates are in April 2017.   Since all of our membership, transactions, and log data end on March 31, 2017, we must train a model that predicts 30 days into the future.  We have no choice because April 2017 data is not available to us.  </p>\n\n<p>With this reasoning, in order to predict March churners, you should only use membership, transactions, and log data up to February 28, 2017.   If you are using March data as well, then your model is not being trained to predict 30 days into the future.  Instead, it is being trained to use data through March to predict March churners.  The model should do very well if it knows what happens in March.</p>\n\n<p>Cheers...</p>",
      "rawMarkdown": "Hi, I think the way to think about it is this way.\n\nIf you look at the submission msno's, almost all of the membership expiration dates are in April 2017.   Since all of our membership, transactions, and log data end on March 31, 2017, we must train a model that predicts 30 days into the future.  We have no choice because April 2017 data is not available to us.  \n\nWith this reasoning, in order to predict March churners, you should only use membership, transactions, and log data up to February 28, 2017.   If you are using March data as well, then your model is not being trained to predict 30 days into the future.  Instead, it is being trained to use data through March to predict March churners.  The model should do very well if it knows what happens in March.\n\nCheers...",
      "votes": null
    },
    {
      "id": "252141",
      "postDate": "12/02/2017 10:45:09",
      "content": "<p>To clarify my last post,</p>\n\n<p>When I say that almost all of the membership expiration dates for the msno's in the submission file are in April 2017, I mean the last entries.   Similarly, almost all of the last entries for the training data msno's have membership expiration dates in March 2017.  This makes sense because, according to the instructions, we are predicting March churn for the training data and April churn for the submission (test) data.</p>\n\n<p>Since we will not have April data to predict the April churners, it seems that we should not use March data to predict the March churners.  Of course, we can always use the March data to train the model.  However, the model will be calibrated to assume that the April data will be available for the April predictions, but that data is not available to us.    </p>\n\n<p>Cheers, ...</p>",
      "rawMarkdown": "To clarify my last post,\n\nWhen I say that almost all of the membership expiration dates for the msno's in the submission file are in April 2017, I mean the last entries.   Similarly, almost all of the last entries for the training data msno's have membership expiration dates in March 2017.  This makes sense because, according to the instructions, we are predicting March churn for the training data and April churn for the submission (test) data.\n\nSince we will not have April data to predict the April churners, it seems that we should not use March data to predict the March churners.  Of course, we can always use the March data to train the model.  However, the model will be calibrated to assume that the April data will be available for the April predictions, but that data is not available to us.    \n\nCheers, ...",
      "votes": null
    },
    {
      "id": "252173",
      "postDate": "12/02/2017 12:01:01",
      "content": "<p>I used StratifiedKFold and no shuffle. I just let the seed of xgboost  change randomly .</p>",
      "rawMarkdown": "I used StratifiedKFold and no shuffle. I just let the seed of xgboost  change randomly .",
      "votes": null
    },
    {
      "id": "252201",
      "postDate": "12/02/2017 13:05:55",
      "content": "<p>Hey we are absolutely on the same page. The data I use to build the features for Feb churners stops on Feb 28. Please note that Feb churners is shorthand for \"users for whom we predict 30 days following their last membership_expiration happening some time in February\". Similarly, the data I use to build the features for March churners has a cut off date on Mar 31.</p>\n\n<p>I have come up with a couple ideas to align my cross validation strategy with the LB submission. One of them is to use only new msno in my cross validation set. I hope that will help me get out of this impasse.</p>",
      "rawMarkdown": "Hey we are absolutely on the same page. The data I use to build the features for Feb churners stops on Feb 28. Please note that Feb churners is shorthand for \"users for whom we predict 30 days following their last membership_expiration happening some time in February\". Similarly, the data I use to build the features for March churners has a cut off date on Mar 31.\n\nI have come up with a couple ideas to align my cross validation strategy with the LB submission. One of them is to use only new msno in my cross validation set. I hope that will help me get out of this impasse.",
      "votes": null
    },
    {
      "id": "252374",
      "postDate": "12/02/2017 20:14:58",
      "content": "<p>Hi Mkffl, thanks for the reply.  After reading many of the discussions, there is one thing that I am not clear about.  Perhaps you can help me understand.</p>\n\n<p>Are the msno's in the train_v2.csv file all March churners?  Or do we have to look at a subset of the train_v2.csv msno's for our training?  </p>\n\n<p>In terms of initial filtering of training data, I only remove data posted after Feb 28, 2017, but I don't remove any msno's from the train_v2.csv file.  For example, some discussions imply that we should remove data with memberships in April because they should be out-of-bounds. and are not legitimate msno's for training.     </p>",
      "rawMarkdown": "Hi Mkffl, thanks for the reply.  After reading many of the discussions, there is one thing that I am not clear about.  Perhaps you can help me understand.\n\nAre the msno's in the train_v2.csv file all March churners?  Or do we have to look at a subset of the train_v2.csv msno's for our training?  \n\nIn terms of initial filtering of training data, I only remove data posted after Feb 28, 2017, but I don't remove any msno's from the train_v2.csv file.  For example, some discussions imply that we should remove data with memberships in April because they should be out-of-bounds. and are not legitimate msno's for training.",
      "votes": null
    },
    {
      "id": "252377",
      "postDate": "12/02/2017 20:17:15",
      "content": "<p>Clarification of my last post.</p>\n\n<p>I didn't mean to say that the train_v2.csv file contains all churners.  What I meant is that they are all customers that either churned in March or are within scope to have churned. </p>",
      "rawMarkdown": "Clarification of my last post.\n\nI didn't mean to say that the train_v2.csv file contains all churners.  What I meant is that they are all customers that either churned in March or are within scope to have churned.",
      "votes": null
    },
    {
      "id": "252793",
      "postDate": "12/03/2017 22:15:55",
      "content": "<p>train_v2.csv includes information for a set of msno about their churn outcome in April i.e. information about their subscription renewal given that their expiration date happens some day in March. It's the same set of msno's as those found in the initial (i.e. before the leak) submission file.</p>\n\n<p>Regarding the cut off date, here's my take: we need to predict churn in April with data available until end of March. To be consistent, the training data must have different cut off dates for train_v2 (March 31) versus train (Feb 28). I have not seen the discussions you mention because it's not top of my worry list to be honest. </p>\n\n<p>Hope that helps</p>",
      "rawMarkdown": "train_v2.csv includes information for a set of msno about their churn outcome in April i.e. information about their subscription renewal given that their expiration date happens some day in March. It's the same set of msno's as those found in the initial (i.e. before the leak) submission file.\n\nRegarding the cut off date, here's my take: we need to predict churn in April with data available until end of March. To be consistent, the training data must have different cut off dates for train_v2 (March 31) versus train (Feb 28). I have not seen the discussions you mention because it's not top of my worry list to be honest. \n\nHope that helps",
      "votes": null
    },
    {
      "id": "252805",
      "postDate": "12/03/2017 22:49:24",
      "content": "<p>Thanks...</p>",
      "rawMarkdown": "Thanks...",
      "votes": null
    },
    {
      "id": "253680",
      "postDate": "12/05/2017 13:08:34",
      "content": "<p>hey I messaged you but no sure it worked. Just wanted to pursue the conversation about cross validation  off this thread: m.kiffel@gmail.com</p>\n\n<p>Cheers</p>",
      "rawMarkdown": "hey I messaged you but no sure it worked. Just wanted to pursue the conversation about cross validation  off this thread: m.kiffel@gmail.com\n\nCheers",
      "votes": null
    },
    {
      "id": "255601",
      "postDate": "12/09/2017 14:25:43",
      "content": "<p>I think feature engineering is the most important part in this competition.</p>\n\n<p>At first my cv is ~0.21, while lb is ~0.16. After working few days, my cv is now ~0.16, while lb is ~0.11.</p>",
      "rawMarkdown": "I think feature engineering is the most important part in this competition.\n\nAt first my cv is ~0.21, while lb is ~0.16. After working few days, my cv is now ~0.16, while lb is ~0.11.",
      "votes": null
    },
    {
      "id": "255671",
      "postDate": "12/09/2017 18:59:59",
      "content": "<p>Hi, I have found that my training and \"test\" scores using the training data are always worse than my score on the lb.  For example, my \"test\" score would be .19 and my leader lb score would be around .14.  It looks like you also get lb results that are much better than what you get using the training data.  Is that right?</p>\n\n<p>I understand why my training and \"test\" scores would be lower than what I get using the submission data.  I don't understand why they are always much higher.  Do you have any ideas?</p>",
      "rawMarkdown": "Hi, I have found that my training and \"test\" scores using the training data are always worse than my score on the lb.  For example, my \"test\" score would be .19 and my leader lb score would be around .14.  It looks like you also get lb results that are much better than what you get using the training data.  Is that right?\n\nI understand why my training and \"test\" scores would be lower than what I get using the submission data.  I don't understand why they are always much higher.  Do you have any ideas?",
      "votes": null
    },
    {
      "id": "255679",
      "postDate": "12/09/2017 19:38:45",
      "content": "<p>Hi, sorry I didn't see this post until now.  Looking at your edit to your post, it looks like you solved your problem.  That is great!</p>",
      "rawMarkdown": "Hi, sorry I didn't see this post until now.  Looking at your edit to your post, it looks like you solved your problem.  That is great!",
      "votes": null
    },
    {
      "id": "255813",
      "postDate": "12/10/2017 07:51:10",
      "content": "<p>I guess maybe there are less churn member in April, so the score on leaderboard is lower.</p>",
      "rawMarkdown": "I guess maybe there are less churn member in April, so the score on leaderboard is lower.",
      "votes": null
    },
    {
      "id": "255829",
      "postDate": "12/10/2017 09:00:02",
      "content": "<p>Thanks for the reply.</p>\n\n<p>The log-loss function is not explicitly a function of the churn rate, since it equally penalizes an incorrect no-churn or an incorrect churn prediction.  In other words, the penalty for predicting 30% for a no-churn is the same as the penalty of predicting 70% for a churner.  In both cases, the penalty is ln(1-.30) = ln(.70) = ~0.36.  So having a higher churn rate does not increase the loss.</p>\n\n<p>However, your answer would make sense if the model is much worse at predicting churners than non-churners, and there are less churners in the submission set.  So, if on average a model predicts 30% for a no-churner and 60% for a churner, then the loss for the no-churner would be ln(.70) = ~0.36, but the loss for the churner would be ln(.60) = ~.51, about 42% higher.</p>\n\n<p>I wonder if that is the case.  </p>\n\n<p>Thanks again!</p>",
      "rawMarkdown": "Thanks for the reply.\n\nThe log-loss function is not explicitly a function of the churn rate, since it equally penalizes an incorrect no-churn or an incorrect churn prediction.  In other words, the penalty for predicting 30% for a no-churn is the same as the penalty of predicting 70% for a churner.  In both cases, the penalty is ln(1-.30) = ln(.70) = ~0.36.  So having a higher churn rate does not increase the loss.\n\nHowever, your answer would make sense if the model is much worse at predicting churners than non-churners, and there are less churners in the submission set.  So, if on average a model predicts 30% for a no-churner and 60% for a churner, then the loss for the no-churner would be ln(.70) = ~0.36, but the loss for the churner would be ln(.60) = ~.51, about 42% higher.\n\nI wonder if that is the case.  \n\nThanks again!",
      "votes": null
    },
    {
      "id": "255831",
      "postDate": "12/10/2017 09:10:45",
      "content": "<p>Yes, or maybe members in April is more rational than March. So the score is lower than before. However m y local cv shows that score is ~0.16 to 0.18</p>",
      "rawMarkdown": "Yes, or maybe members in April is more rational than March. So the score is lower than before. However m y local cv shows that score is ~0.16 to 0.18",
      "votes": null
    },
    {
      "id": "255834",
      "postDate": "12/10/2017 09:27:31",
      "content": "<p>Both of our models do much worse on the development samples (training, validation, test) than on the submission data.  If, as you say, consumers are more rational in April, then our models still should be doing worse on the submission data because the models were trained on irrational March consumers.  </p>\n\n<p>So it seems like the answer has to be that the customers in the submission data happen to be the type of customers that the model happens to be more accurate on. </p>\n\n<p>By the way, you have a great score of .11.  Are you using the original members.csv file or the new members_v3.csv file?  Just wondering...</p>\n\n<p>Great discussion...</p>\n\n<p>Cheers</p>",
      "rawMarkdown": "Both of our models do much worse on the development samples (training, validation, test) than on the submission data.  If, as you say, consumers are more rational in April, then our models still should be doing worse on the submission data because the models were trained on irrational March consumers.  \n\nSo it seems like the answer has to be that the customers in the submission data happen to be the type of customers that the model happens to be more accurate on. \n\nBy the way, you have a great score of .11.  Are you using the original members.csv file or the new members_v3.csv file?  Just wondering...\n\nGreat discussion...\n\nCheers",
      "votes": null
    },
    {
      "id": "255838",
      "postDate": "12/10/2017 09:45:20",
      "content": "<p>I mean 'more rational' is, the outliers in April is less than March. The outliers here means member is thought churn in usual, but is not churn; or member is thought not churn in usual, but is churn.</p>\n\n<p>I didn't have original member.csv file.. only members_v3.csv...</p>",
      "rawMarkdown": "I mean 'more rational' is, the outliers in April is less than March. The outliers here means member is thought churn in usual, but is not churn; or member is thought not churn in usual, but is churn.\n\nI didn't have original member.csv file.. only members_v3.csv...",
      "votes": null
    },
    {
      "id": "255845",
      "postDate": "12/10/2017 10:03:22",
      "content": "<p>OK.....</p>\n\n<p>Thanks for letting me know that you are using members_v3.csv.  I also don't have access to the original file, so I wanted to know whether a ~.11 score is possible without the original members file.  </p>\n\n<p>Cheers....</p>",
      "rawMarkdown": "OK.....\n\nThanks for letting me know that you are using members_v3.csv.  I also don't have access to the original file, so I wanted to know whether a ~.11 score is possible without the original members file.  \n\nCheers....",
      "votes": null
    },
    {
      "id": "256226",
      "postDate": "12/11/2017 14:44:03",
      "content": "<p>Hey Mkffl, could you elaborate on this?</p>\n\n<blockquote>\n  <p>cross validation set should not include any msno used in the dev set</p>\n</blockquote>\n\n<p>train.csv and test.csv have lots of common <em>msno</em>, however, all the rows in train.csv have a unique msno. What's your definition of <em>cross validation set</em> and <em>dev set</em>? Thanks....</p>",
      "rawMarkdown": "Hey Mkffl, could you elaborate on this?\n&gt; cross validation set should not include any msno used in the dev set\n\ntrain.csv and test.csv have lots of common *msno*, however, all the rows in train.csv have a unique msno. What's your definition of *cross validation set* and *dev set*? Thanks....",
      "votes": null
    },
    {
      "id": "256384",
      "postDate": "12/11/2017 21:09:50",
      "content": "<p>Hey,</p>\n\n<p>I use four folds as a development set, and one fold as a validation set. Instead of randomly allocating any data points to either the test or the development sets, as I would do in a classic kfold crossvalidation, I randomly allocate <em>unique msno</em>. The implication is that the model trained on the development set is better at predicting the logloss on unseen msno. </p>\n\n<p>This does not reflect the exercise set up, as the (LB) test set includes lots of msno from the trainv2 set. I appreciate that you would normally seek to replicate the LB set up on your cross validation. So wjy the cross validation set up I use works better than the classic one remains a mystery to me. </p>",
      "rawMarkdown": "Hey,\n \nI use four folds as a development set, and one fold as a validation set. Instead of randomly allocating any data points to either the test or the development sets, as I would do in a classic kfold crossvalidation, I randomly allocate *unique msno*. The implication is that the model trained on the development set is better at predicting the logloss on unseen msno. \n\nThis does not reflect the exercise set up, as the (LB) test set includes lots of msno from the trainv2 set. I appreciate that you would normally seek to replicate the LB set up on your cross validation. So wjy the cross validation set up I use works better than the classic one remains a mystery to me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 249325,
      "author_name": "casscw",
      "author_url": "",
      "post_date": "11/28/2017 07:37:26",
      "content": "<p>me too. hh  Just want to know the details in admins generating the test dataset。 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 249521,
      "author_name": "paullo0106",
      "author_url": "",
      "post_date": "11/28/2017 16:57:15",
      "content": "<p>my results generated from xgboost also got a large gap between local and LB...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 251620,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "12/01/2017 13:38:08",
      "content": "<p>XGBoost, average score:0.1245232, 5 folds.  PB: 0.11799</p>",
      "votes": null,
      "replies": [
        {
          "id": 251708,
          "author_name": "mchlkffl",
          "author_url": "",
          "post_date": "12/01/2017 15:30:43",
          "content": "<p>Thanks a lot for sharing some actual numbers, that's really helpful. May I ask if you have used simple cross-validation, i.e. simply by shuffling then splitting the entire train set (Feb and Mar months) into 5 parts? </p>\n\n<p>I ask because I have thought about building a cross validation set that only includes users not in the development set. I suspect that my overfit is partly due to having the same customer in both sets, while the test set includes unseen users.</p>\n\n<p>Another question is, do you use the raw expiration_date and membership_expiration_date columns as features? They do have a lot of predicting power (from a xgboost 'feature importance' perspective) but I have a feeling that they may cause PB overfitting</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 252173,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "12/02/2017 12:01:01",
          "content": "<p>I used StratifiedKFold and no shuffle. I just let the seed of xgboost  change randomly .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 251967,
      "author_name": "beliefbio",
      "author_url": "",
      "post_date": "12/02/2017 01:39:16",
      "content": "<p>Have the same problem too. I split train data into train/test with 7:3. Then I used train data for 5-fold cv with mean logloss 0.0774. And then the logloss for testing data is 0.0767. But my PB is 0.15971. My guess is there is much difference between train data and test data in PB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 251985,
      "author_name": "",
      "author_url": "",
      "post_date": "12/02/2017 02:39:11",
      "content": "<p>Hi, I am not sure this comment is relevant to you since I don't know how you generated your training set.  However, based on my understanding of the rules, you can only use data up to Feb 28, 2017 for training.  You can then additionally use the March 2017 data for scoring the submissions.  Your post seems to suggest you are using the March data for training as well.   (Sorry if I have misinterpreted your post.)</p>\n\n<p>If so that would explain what you are seeing.   The submissions are looking for models that can make out-of-time predictions, not just out-of-sample predictions.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 252118,
          "author_name": "mchlkffl",
          "author_url": "",
          "post_date": "12/02/2017 09:29:01",
          "content": "<p>Hi, \nYour comment is 100% relevant. I use March data for training as suggested in Wendy Kan's post (\"Announcement: New test data released\"). Specifically, I use the test set that was previously used to rate submissions on the LB and for which admins have revealed the is_churn column.  Wendy's post suggested that this data was now available to train our models. You say otherwise. This post seems to support your opinion: <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363\">https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363</a></p>\n\n<p>Let's just park the question of what data is allowed because it's not clear and I still can't see how it ties back to the original problem. </p>\n\n<p>You say that what I see can be explained by the fact that I use Feb and Mar data for training, which implies that my model makes out-of-sample predictions. That's completely right. However the submission is for March users - just a different set of msno than those used in my training data. Hence submissions are <strong><em>not</em></strong> looking for models that can me out-of-time predictions - to be clear, they would look for out-of-time predictions <strong><em>only</em></strong> if I restricted my train set to February data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 252138,
      "author_name": "",
      "author_url": "",
      "post_date": "12/02/2017 10:22:58",
      "content": "<p>Hi, I think the way to think about it is this way.</p>\n\n<p>If you look at the submission msno's, almost all of the membership expiration dates are in April 2017.   Since all of our membership, transactions, and log data end on March 31, 2017, we must train a model that predicts 30 days into the future.  We have no choice because April 2017 data is not available to us.  </p>\n\n<p>With this reasoning, in order to predict March churners, you should only use membership, transactions, and log data up to February 28, 2017.   If you are using March data as well, then your model is not being trained to predict 30 days into the future.  Instead, it is being trained to use data through March to predict March churners.  The model should do very well if it knows what happens in March.</p>\n\n<p>Cheers...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 252141,
      "author_name": "",
      "author_url": "",
      "post_date": "12/02/2017 10:45:09",
      "content": "<p>To clarify my last post,</p>\n\n<p>When I say that almost all of the membership expiration dates for the msno's in the submission file are in April 2017, I mean the last entries.   Similarly, almost all of the last entries for the training data msno's have membership expiration dates in March 2017.  This makes sense because, according to the instructions, we are predicting March churn for the training data and April churn for the submission (test) data.</p>\n\n<p>Since we will not have April data to predict the April churners, it seems that we should not use March data to predict the March churners.  Of course, we can always use the March data to train the model.  However, the model will be calibrated to assume that the April data will be available for the April predictions, but that data is not available to us.    </p>\n\n<p>Cheers, ...</p>",
      "votes": null,
      "replies": [
        {
          "id": 252201,
          "author_name": "mchlkffl",
          "author_url": "",
          "post_date": "12/02/2017 13:05:55",
          "content": "<p>Hey we are absolutely on the same page. The data I use to build the features for Feb churners stops on Feb 28. Please note that Feb churners is shorthand for \"users for whom we predict 30 days following their last membership_expiration happening some time in February\". Similarly, the data I use to build the features for March churners has a cut off date on Mar 31.</p>\n\n<p>I have come up with a couple ideas to align my cross validation strategy with the LB submission. One of them is to use only new msno in my cross validation set. I hope that will help me get out of this impasse.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 252374,
      "author_name": "",
      "author_url": "",
      "post_date": "12/02/2017 20:14:58",
      "content": "<p>Hi Mkffl, thanks for the reply.  After reading many of the discussions, there is one thing that I am not clear about.  Perhaps you can help me understand.</p>\n\n<p>Are the msno's in the train_v2.csv file all March churners?  Or do we have to look at a subset of the train_v2.csv msno's for our training?  </p>\n\n<p>In terms of initial filtering of training data, I only remove data posted after Feb 28, 2017, but I don't remove any msno's from the train_v2.csv file.  For example, some discussions imply that we should remove data with memberships in April because they should be out-of-bounds. and are not legitimate msno's for training.     </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 252377,
      "author_name": "",
      "author_url": "",
      "post_date": "12/02/2017 20:17:15",
      "content": "<p>Clarification of my last post.</p>\n\n<p>I didn't mean to say that the train_v2.csv file contains all churners.  What I meant is that they are all customers that either churned in March or are within scope to have churned. </p>",
      "votes": null,
      "replies": [
        {
          "id": 252793,
          "author_name": "mchlkffl",
          "author_url": "",
          "post_date": "12/03/2017 22:15:55",
          "content": "<p>train_v2.csv includes information for a set of msno about their churn outcome in April i.e. information about their subscription renewal given that their expiration date happens some day in March. It's the same set of msno's as those found in the initial (i.e. before the leak) submission file.</p>\n\n<p>Regarding the cut off date, here's my take: we need to predict churn in April with data available until end of March. To be consistent, the training data must have different cut off dates for train_v2 (March 31) versus train (Feb 28). I have not seen the discussions you mention because it's not top of my worry list to be honest. </p>\n\n<p>Hope that helps</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 252805,
      "author_name": "",
      "author_url": "",
      "post_date": "12/03/2017 22:49:24",
      "content": "<p>Thanks...</p>",
      "votes": null,
      "replies": [
        {
          "id": 253680,
          "author_name": "mchlkffl",
          "author_url": "",
          "post_date": "12/05/2017 13:08:34",
          "content": "<p>hey I messaged you but no sure it worked. Just wanted to pursue the conversation about cross validation  off this thread: m.kiffel@gmail.com</p>\n\n<p>Cheers</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255679,
          "author_name": "",
          "author_url": "",
          "post_date": "12/09/2017 19:38:45",
          "content": "<p>Hi, sorry I didn't see this post until now.  Looking at your edit to your post, it looks like you solved your problem.  That is great!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 255601,
      "author_name": "infinitewing",
      "author_url": "",
      "post_date": "12/09/2017 14:25:43",
      "content": "<p>I think feature engineering is the most important part in this competition.</p>\n\n<p>At first my cv is ~0.21, while lb is ~0.16. After working few days, my cv is now ~0.16, while lb is ~0.11.</p>",
      "votes": null,
      "replies": [
        {
          "id": 255671,
          "author_name": "",
          "author_url": "",
          "post_date": "12/09/2017 18:59:59",
          "content": "<p>Hi, I have found that my training and \"test\" scores using the training data are always worse than my score on the lb.  For example, my \"test\" score would be .19 and my leader lb score would be around .14.  It looks like you also get lb results that are much better than what you get using the training data.  Is that right?</p>\n\n<p>I understand why my training and \"test\" scores would be lower than what I get using the submission data.  I don't understand why they are always much higher.  Do you have any ideas?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255813,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/10/2017 07:51:10",
          "content": "<p>I guess maybe there are less churn member in April, so the score on leaderboard is lower.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255829,
          "author_name": "",
          "author_url": "",
          "post_date": "12/10/2017 09:00:02",
          "content": "<p>Thanks for the reply.</p>\n\n<p>The log-loss function is not explicitly a function of the churn rate, since it equally penalizes an incorrect no-churn or an incorrect churn prediction.  In other words, the penalty for predicting 30% for a no-churn is the same as the penalty of predicting 70% for a churner.  In both cases, the penalty is ln(1-.30) = ln(.70) = ~0.36.  So having a higher churn rate does not increase the loss.</p>\n\n<p>However, your answer would make sense if the model is much worse at predicting churners than non-churners, and there are less churners in the submission set.  So, if on average a model predicts 30% for a no-churner and 60% for a churner, then the loss for the no-churner would be ln(.70) = ~0.36, but the loss for the churner would be ln(.60) = ~.51, about 42% higher.</p>\n\n<p>I wonder if that is the case.  </p>\n\n<p>Thanks again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255831,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/10/2017 09:10:45",
          "content": "<p>Yes, or maybe members in April is more rational than March. So the score is lower than before. However m y local cv shows that score is ~0.16 to 0.18</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255834,
          "author_name": "",
          "author_url": "",
          "post_date": "12/10/2017 09:27:31",
          "content": "<p>Both of our models do much worse on the development samples (training, validation, test) than on the submission data.  If, as you say, consumers are more rational in April, then our models still should be doing worse on the submission data because the models were trained on irrational March consumers.  </p>\n\n<p>So it seems like the answer has to be that the customers in the submission data happen to be the type of customers that the model happens to be more accurate on. </p>\n\n<p>By the way, you have a great score of .11.  Are you using the original members.csv file or the new members_v3.csv file?  Just wondering...</p>\n\n<p>Great discussion...</p>\n\n<p>Cheers</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255838,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/10/2017 09:45:20",
          "content": "<p>I mean 'more rational' is, the outliers in April is less than March. The outliers here means member is thought churn in usual, but is not churn; or member is thought not churn in usual, but is churn.</p>\n\n<p>I didn't have original member.csv file.. only members_v3.csv...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255845,
          "author_name": "",
          "author_url": "",
          "post_date": "12/10/2017 10:03:22",
          "content": "<p>OK.....</p>\n\n<p>Thanks for letting me know that you are using members_v3.csv.  I also don't have access to the original file, so I wanted to know whether a ~.11 score is possible without the original members file.  </p>\n\n<p>Cheers....</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 256226,
      "author_name": "paullo0106",
      "author_url": "",
      "post_date": "12/11/2017 14:44:03",
      "content": "<p>Hey Mkffl, could you elaborate on this?</p>\n\n<blockquote>\n  <p>cross validation set should not include any msno used in the dev set</p>\n</blockquote>\n\n<p>train.csv and test.csv have lots of common <em>msno</em>, however, all the rows in train.csv have a unique msno. What's your definition of <em>cross validation set</em> and <em>dev set</em>? Thanks....</p>",
      "votes": null,
      "replies": [
        {
          "id": 256384,
          "author_name": "mchlkffl",
          "author_url": "",
          "post_date": "12/11/2017 21:09:50",
          "content": "<p>Hey,</p>\n\n<p>I use four folds as a development set, and one fold as a validation set. Instead of randomly allocating any data points to either the test or the development sets, as I would do in a classic kfold crossvalidation, I randomly allocate <em>unique msno</em>. The implication is that the model trained on the development set is better at predicting the logloss on unseen msno. </p>\n\n<p>This does not reflect the exercise set up, as the (LB) test set includes lots of msno from the trainv2 set. I appreciate that you would normally seek to replicate the LB set up on your cross validation. So wjy the cross validation set up I use works better than the classic one remains a mystery to me. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "249321": "EDIT - Found the solution: my cross validation set should not include any msno used in the dev set. When this condition is met, the cross validation set-up reflects the submission set-up and therefore the cross validation logloss is similar to the leadboard score, i.e. pretty shit (but that's a different issue).  \n\nI get a large logloss difference between validation and leaderboard sets, so I wonder\n\na) if others have been in a similar situation \n\nb) what I should do now\n\nUsing features on the transaction data set, I run xgboost on one of 5 folds and get a logloss of c.0.056 on the validation set (see below). However when I use this model on the test set to make a submission, I get a poor score of c.0.177. I am surprised to see so much overfit. Before the data leak, I got roughly similar scores (dev logloss c.0.06 and leaderboard c.0.20), however we were using February churners to predict March churners. Since we now have Feb and March churners to predict another bunch of March churners, I thought that having more representative train data would help reduce overfit. \n\nI suppose my first question is - does it make sense and have others seen similar gaps between k-fold validation and leadboard submission? \nThen I am left with several options and I would appreciate any guidance / ideas on which option would make most sense for me:\n\n- build features using the log data set. I think it would be a waste of time. Before the leak, only 1 or 2 log-related features would only slightly improve accuracy, other were useless, so the same might happen now. I am worried I might spend hours to get a tiny bit of validation improvement, translating into a tiny bit of leadboard improvement, say 0.016, which frankly is just as shit as my current 0.177\n\n- use other xgboost folds to generate other test predictions, then average them to hopefully reduce overfit. EDIT: I have tried that but that did not help.\n\n- use other models than xgboosst, e.g. rf or neural nets, then average with xgboost test predictions to hopefully reduce overfit\n\n- Improve the cross validation strategy i.e. set up the dev and cv sets differently to reflect the train-LB set up; for example, for the development set use only msno not seen in the cv set. EDIT: that worked.\n\n- any other idea?\n\n\n******************* FOLD N. 1 *******************\n\n[0]\ttrain-logloss:0.675067\ttest-logloss:0.675068\n\nMultiple eval metrics have been passed: 'test-logloss' will be used for early stopping.\n\nWill train until test-logloss hasn't improved in 50 rounds.\n\n[50]\ttrain-logloss:0.242682\ttest-logloss:0.242829\n\n[100]\ttrain-logloss:0.124368\ttest-logloss:0.124705\n\n[150]\ttrain-logloss:0.084166\ttest-logloss:0.084569\n\n[200]\ttrain-logloss:0.068575\ttest-logloss:0.06902\n\n[250]\ttrain-logloss:0.062092\ttest-logloss:0.062567\n\n[300]\ttrain-logloss:0.058893\ttest-logloss:0.059417\n\n[350]\ttrain-logloss:0.057189\ttest-logloss:0.057756\n\n[399]\ttrain-logloss:0.055572\ttest-logloss:0.056203",
    "249325": "me too. hh  Just want to know the details in admins generating the test dataset。",
    "249521": "my results generated from xgboost also got a large gap between local and LB...",
    "251620": "XGBoost, average score:0.1245232, 5 folds.  PB: 0.11799",
    "251708": "Thanks a lot for sharing some actual numbers, that's really helpful. May I ask if you have used simple cross-validation, i.e. simply by shuffling then splitting the entire train set (Feb and Mar months) into 5 parts? \n\nI ask because I have thought about building a cross validation set that only includes users not in the development set. I suspect that my overfit is partly due to having the same customer in both sets, while the test set includes unseen users.\n\nAnother question is, do you use the raw expiration_date and membership_expiration_date columns as features? They do have a lot of predicting power (from a xgboost 'feature importance' perspective) but I have a feeling that they may cause PB overfitting",
    "251967": "Have the same problem too. I split train data into train/test with 7:3. Then I used train data for 5-fold cv with mean logloss 0.0774. And then the logloss for testing data is 0.0767. But my PB is 0.15971. My guess is there is much difference between train data and test data in PB.",
    "251985": "Hi, I am not sure this comment is relevant to you since I don't know how you generated your training set.  However, based on my understanding of the rules, you can only use data up to Feb 28, 2017 for training.  You can then additionally use the March 2017 data for scoring the submissions.  Your post seems to suggest you are using the March data for training as well.   (Sorry if I have misinterpreted your post.)\n\nIf so that would explain what you are seeing.   The submissions are looking for models that can make out-of-time predictions, not just out-of-sample predictions.",
    "252118": "Hi, \nYour comment is 100% relevant. I use March data for training as suggested in Wendy Kan's post (\"Announcement: New test data released\"). Specifically, I use the test set that was previously used to rate submissions on the LB and for which admins have revealed the is_churn column.  Wendy's post suggested that this data was now available to train our models. You say otherwise. This post seems to support your opinion: https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363\n\nLet's just park the question of what data is allowed because it's not clear and I still can't see how it ties back to the original problem. \n\nYou say that what I see can be explained by the fact that I use Feb and Mar data for training, which implies that my model makes out-of-sample predictions. That's completely right. However the submission is for March users - just a different set of msno than those used in my training data. Hence submissions are ***not*** looking for models that can me out-of-time predictions - to be clear, they would look for out-of-time predictions ***only*** if I restricted my train set to February data.",
    "252138": "Hi, I think the way to think about it is this way.\n\nIf you look at the submission msno's, almost all of the membership expiration dates are in April 2017.   Since all of our membership, transactions, and log data end on March 31, 2017, we must train a model that predicts 30 days into the future.  We have no choice because April 2017 data is not available to us.  \n\nWith this reasoning, in order to predict March churners, you should only use membership, transactions, and log data up to February 28, 2017.   If you are using March data as well, then your model is not being trained to predict 30 days into the future.  Instead, it is being trained to use data through March to predict March churners.  The model should do very well if it knows what happens in March.\n\nCheers...",
    "252141": "To clarify my last post,\n\nWhen I say that almost all of the membership expiration dates for the msno's in the submission file are in April 2017, I mean the last entries.   Similarly, almost all of the last entries for the training data msno's have membership expiration dates in March 2017.  This makes sense because, according to the instructions, we are predicting March churn for the training data and April churn for the submission (test) data.\n\nSince we will not have April data to predict the April churners, it seems that we should not use March data to predict the March churners.  Of course, we can always use the March data to train the model.  However, the model will be calibrated to assume that the April data will be available for the April predictions, but that data is not available to us.    \n\nCheers, ...",
    "252173": "I used StratifiedKFold and no shuffle. I just let the seed of xgboost  change randomly .",
    "252201": "Hey we are absolutely on the same page. The data I use to build the features for Feb churners stops on Feb 28. Please note that Feb churners is shorthand for \"users for whom we predict 30 days following their last membership_expiration happening some time in February\". Similarly, the data I use to build the features for March churners has a cut off date on Mar 31.\n\nI have come up with a couple ideas to align my cross validation strategy with the LB submission. One of them is to use only new msno in my cross validation set. I hope that will help me get out of this impasse.",
    "252374": "Hi Mkffl, thanks for the reply.  After reading many of the discussions, there is one thing that I am not clear about.  Perhaps you can help me understand.\n\nAre the msno's in the train_v2.csv file all March churners?  Or do we have to look at a subset of the train_v2.csv msno's for our training?  \n\nIn terms of initial filtering of training data, I only remove data posted after Feb 28, 2017, but I don't remove any msno's from the train_v2.csv file.  For example, some discussions imply that we should remove data with memberships in April because they should be out-of-bounds. and are not legitimate msno's for training.",
    "252377": "Clarification of my last post.\n\nI didn't mean to say that the train_v2.csv file contains all churners.  What I meant is that they are all customers that either churned in March or are within scope to have churned.",
    "252793": "train_v2.csv includes information for a set of msno about their churn outcome in April i.e. information about their subscription renewal given that their expiration date happens some day in March. It's the same set of msno's as those found in the initial (i.e. before the leak) submission file.\n\nRegarding the cut off date, here's my take: we need to predict churn in April with data available until end of March. To be consistent, the training data must have different cut off dates for train_v2 (March 31) versus train (Feb 28). I have not seen the discussions you mention because it's not top of my worry list to be honest. \n\nHope that helps",
    "252805": "Thanks...",
    "253680": "hey I messaged you but no sure it worked. Just wanted to pursue the conversation about cross validation  off this thread: m.kiffel@gmail.com\n\nCheers",
    "255601": "I think feature engineering is the most important part in this competition.\n\nAt first my cv is ~0.21, while lb is ~0.16. After working few days, my cv is now ~0.16, while lb is ~0.11.",
    "255671": "Hi, I have found that my training and \"test\" scores using the training data are always worse than my score on the lb.  For example, my \"test\" score would be .19 and my leader lb score would be around .14.  It looks like you also get lb results that are much better than what you get using the training data.  Is that right?\n\nI understand why my training and \"test\" scores would be lower than what I get using the submission data.  I don't understand why they are always much higher.  Do you have any ideas?",
    "255679": "Hi, sorry I didn't see this post until now.  Looking at your edit to your post, it looks like you solved your problem.  That is great!",
    "255813": "I guess maybe there are less churn member in April, so the score on leaderboard is lower.",
    "255829": "Thanks for the reply.\n\nThe log-loss function is not explicitly a function of the churn rate, since it equally penalizes an incorrect no-churn or an incorrect churn prediction.  In other words, the penalty for predicting 30% for a no-churn is the same as the penalty of predicting 70% for a churner.  In both cases, the penalty is ln(1-.30) = ln(.70) = ~0.36.  So having a higher churn rate does not increase the loss.\n\nHowever, your answer would make sense if the model is much worse at predicting churners than non-churners, and there are less churners in the submission set.  So, if on average a model predicts 30% for a no-churner and 60% for a churner, then the loss for the no-churner would be ln(.70) = ~0.36, but the loss for the churner would be ln(.60) = ~.51, about 42% higher.\n\nI wonder if that is the case.  \n\nThanks again!",
    "255831": "Yes, or maybe members in April is more rational than March. So the score is lower than before. However m y local cv shows that score is ~0.16 to 0.18",
    "255834": "Both of our models do much worse on the development samples (training, validation, test) than on the submission data.  If, as you say, consumers are more rational in April, then our models still should be doing worse on the submission data because the models were trained on irrational March consumers.  \n\nSo it seems like the answer has to be that the customers in the submission data happen to be the type of customers that the model happens to be more accurate on. \n\nBy the way, you have a great score of .11.  Are you using the original members.csv file or the new members_v3.csv file?  Just wondering...\n\nGreat discussion...\n\nCheers",
    "255838": "I mean 'more rational' is, the outliers in April is less than March. The outliers here means member is thought churn in usual, but is not churn; or member is thought not churn in usual, but is churn.\n\nI didn't have original member.csv file.. only members_v3.csv...",
    "255845": "OK.....\n\nThanks for letting me know that you are using members_v3.csv.  I also don't have access to the original file, so I wanted to know whether a ~.11 score is possible without the original members file.  \n\nCheers....",
    "256226": "Hey Mkffl, could you elaborate on this?\n&gt; cross validation set should not include any msno used in the dev set\n\ntrain.csv and test.csv have lots of common *msno*, however, all the rows in train.csv have a unique msno. What's your definition of *cross validation set* and *dev set*? Thanks....",
    "256384": "Hey,\n \nI use four folds as a development set, and one fold as a validation set. Instead of randomly allocating any data points to either the test or the development sets, as I would do in a classic kfold crossvalidation, I randomly allocate *unique msno*. The implication is that the model trained on the development set is better at predicting the logloss on unseen msno. \n\nThis does not reflect the exercise set up, as the (LB) test set includes lots of msno from the trainv2 set. I appreciate that you would normally seek to replicate the LB set up on your cross validation. So wjy the cross validation set up I use works better than the classic one remains a mystery to me."
  },
  "source": "meta"
}