{
  "id": 45926,
  "title": "Features, features, features",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/45926",
  "author_name": "",
  "post_date": "2017-12-18T03:03:02.151234Z",
  "votes": 10,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Considering how rich this dataset was for feature engineering, more so than many past contests I've seen on Kaggle, I am curious how many features did other teams create? And how many made it into your final models?  For me, I had ~250 features total that were created and tried of which ~100 were used in final models.  The bulk of the features came from the user logs, although many came from transactions, a few from members, and some were interactions between them as well as meta features.</p>\n\n<p>This of course doesn't include hyper parameters, base model predictions, etc. as features.  Just discussing strictly features extracted from the data itself.</p>\n\n<p>I continuously saw accuracy increases in both CV and LB every time I went back for another round of feature engineering, so it would be interesting to know at what point the data is finally \"exhausted\", and how many of these would generalize well over time.</p>",
  "messages": [
    {
      "id": "259263",
      "postDate": "12/18/2017 03:03:02",
      "content": "<p>Considering how rich this dataset was for feature engineering, more so than many past contests I've seen on Kaggle, I am curious how many features did other teams create? And how many made it into your final models?  For me, I had ~250 features total that were created and tried of which ~100 were used in final models.  The bulk of the features came from the user logs, although many came from transactions, a few from members, and some were interactions between them as well as meta features.</p>\n\n<p>This of course doesn't include hyper parameters, base model predictions, etc. as features.  Just discussing strictly features extracted from the data itself.</p>\n\n<p>I continuously saw accuracy increases in both CV and LB every time I went back for another round of feature engineering, so it would be interesting to know at what point the data is finally \"exhausted\", and how many of these would generalize well over time.</p>",
      "rawMarkdown": "Considering how rich this dataset was for feature engineering, more so than many past contests I've seen on Kaggle, I am curious how many features did other teams create? And how many made it into your final models?  For me, I had ~250 features total that were created and tried of which ~100 were used in final models.  The bulk of the features came from the user logs, although many came from transactions, a few from members, and some were interactions between them as well as meta features.\n\nThis of course doesn't include hyper parameters, base model predictions, etc. as features.  Just discussing strictly features extracted from the data itself.\n\nI continuously saw accuracy increases in both CV and LB every time I went back for another round of feature engineering, so it would be interesting to know at what point the data is finally \"exhausted\", and how many of these would generalize well over time.",
      "votes": null
    },
    {
      "id": "259289",
      "postDate": "12/18/2017 04:24:31",
      "content": "<p>Congrats for the 1st place! <br>\nI just created 56 features【20 from user_logs，36 from members and transactions】 for the final model【LGB, 5-CV】.   Too few features for me⊙﹏⊙</p>",
      "rawMarkdown": "Congrats for the 1st place!     \nI just created 56 features【20 from user_logs，36 from members and transactions】 for the final model【LGB, 5-CV】.   Too few features for me⊙﹏⊙",
      "votes": null
    },
    {
      "id": "259311",
      "postDate": "12/18/2017 05:22:59",
      "content": "<p>I created 220 features - approx 10 from members,  140 from transactions, and 60 from user logs,  10 from combined. The number looks large, since we used multiple aspects(max, min, avg, first, second, skewness) from each 'groupby' functions.  (e.g., avg, first, skewness value of transaction gap sequence). Our 0.10 loss is mainly achieved by transaction features, and we didn't manage to generate strong features from user logs. </p>\n\n<p>However, I felt tedious and tired while generating features, especially when the leaderboard score is stuck at the certain level. Does anyone tried feature learning or representation learning to tackle this competition?</p>\n\n<p>@Bryan: Congratulations! I'm really curious what kind of features you generated from user logs.</p>",
      "rawMarkdown": "I created 220 features - approx 10 from members,  140 from transactions, and 60 from user logs,  10 from combined. The number looks large, since we used multiple aspects(max, min, avg, first, second, skewness) from each 'groupby' functions.  (e.g., avg, first, skewness value of transaction gap sequence). Our 0.10 loss is mainly achieved by transaction features, and we didn't manage to generate strong features from user logs. \n\nHowever, I felt tedious and tired while generating features, especially when the leaderboard score is stuck at the certain level. Does anyone tried feature learning or representation learning to tackle this competition?\n\n@Bryan: Congratulations! I'm really curious what kind of features you generated from user logs.",
      "votes": null
    },
    {
      "id": "260181",
      "postDate": "12/19/2017 20:13:20",
      "content": "<p>Out of curiosity, did you scale the UL features at all relative to the membership expiration date, or did you stick with grouping by just calendar dates?</p>\n\n<p>For ex., MSNO \"21dh15HZdEDWGomm1AKlnmAvqYbT3qP1zg8cGxZA+0I=\".  For the calendar month of Jan, they have one user login (on 20170131).  So if you take a calendar month count of logins or a sum of secs using the app, then this user appears likely to churn (very low usage relative to user averages for January).  But if you dig deeper, the user just started their subscription on 20170131.  So a single login doesn't have any meaning for this user unless you create the features relative to the # of transaction days in the month.  Calendar month activity is 3.2% (1 login day / 31 potential login days), but relative activity is 100% (1 login day / 1 potential login day).</p>\n\n<p>Hope that makes some sense, I might be doing a bad job of explaining it because I'm in a rush.</p>",
      "rawMarkdown": "Out of curiosity, did you scale the UL features at all relative to the membership expiration date, or did you stick with grouping by just calendar dates?\n\nFor ex., MSNO \"21dh15HZdEDWGomm1AKlnmAvqYbT3qP1zg8cGxZA+0I=\".  For the calendar month of Jan, they have one user login (on 20170131).  So if you take a calendar month count of logins or a sum of secs using the app, then this user appears likely to churn (very low usage relative to user averages for January).  But if you dig deeper, the user just started their subscription on 20170131.  So a single login doesn't have any meaning for this user unless you create the features relative to the # of transaction days in the month.  Calendar month activity is 3.2% (1 login day / 31 potential login days), but relative activity is 100% (1 login day / 1 potential login day).\n\nHope that makes some sense, I might be doing a bad job of explaining it because I'm in a rush.",
      "votes": null
    },
    {
      "id": "260221",
      "postDate": "12/19/2017 22:00:04",
      "content": "<p>Hi Bryan, that makes sense.  I based many of my user log features relative to a user's membership date and / or registration date.  Using the registration date, as you pointed out, takes care of the issue that you raised.</p>\n\n<p>But I also found that measuring any behavior relative to dates (for example, membership expiration dates) was tricky.  For example, there are fewer days in between 3/15 and 2/15 than there are between 4/15 and 3/15.  That can create problems.  For example, on the training set you might find that if a person initiates a transaction 27 days before expiration then the churn rate was 30%.  But that logic cannot be applied to the submission set because the 27 days has a different meaning for the training and submission sets.  For the training set, 27 days from an expiration date in March brings you to the previous expiration date in February (presuming you have a \"30\" day subscription).  For the submission set, 27 days from a expiration date in April brings you to 2 days AFTER the previous expiration date.  That made a difference for some of my features.  So I had to adjust the definition of the features when scoring the submission set.</p>\n\n<p>You also might have noticed that the distribution of the end-of-the-month membership dates in the submission set was skewed versus the training set.  I think that is because membership expiration dates for February 28, captured previous membership expiration dates for January 28, 29, and 30 (presuming you have a \"30\" day subscription).  However, for March, a January 28 expiration date rolled to March 28, a January 29 rolled to March 29, and a January 30 rolled to March 30.  That also made a difference for some of my features that used date math.  </p>\n\n<p>Cheers, </p>",
      "rawMarkdown": "Hi Bryan, that makes sense.  I based many of my user log features relative to a user's membership date and / or registration date.  Using the registration date, as you pointed out, takes care of the issue that you raised.\n\nBut I also found that measuring any behavior relative to dates (for example, membership expiration dates) was tricky.  For example, there are fewer days in between 3/15 and 2/15 than there are between 4/15 and 3/15.  That can create problems.  For example, on the training set you might find that if a person initiates a transaction 27 days before expiration then the churn rate was 30%.  But that logic cannot be applied to the submission set because the 27 days has a different meaning for the training and submission sets.  For the training set, 27 days from an expiration date in March brings you to the previous expiration date in February (presuming you have a \"30\" day subscription).  For the submission set, 27 days from a expiration date in April brings you to 2 days AFTER the previous expiration date.  That made a difference for some of my features.  So I had to adjust the definition of the features when scoring the submission set.\n\nYou also might have noticed that the distribution of the end-of-the-month membership dates in the submission set was skewed versus the training set.  I think that is because membership expiration dates for February 28, captured previous membership expiration dates for January 28, 29, and 30 (presuming you have a \"30\" day subscription).  However, for March, a January 28 expiration date rolled to March 28, a January 29 rolled to March 29, and a January 30 rolled to March 30.  That also made a difference for some of my features that used date math.  \n\nCheers,",
      "votes": null
    },
    {
      "id": "261898",
      "postDate": "12/24/2017 12:28:31",
      "content": "<p>Hi Bryan, do you still plan on sharing your solution?\nIt seems that only <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/46078\">one person has released their solution so far</a> and it's not clear what features were used. So any description of what has helped you rise to 1st place would be interesting to me and surely lots of other people. </p>\n\n<p>A major question mark for me is what is the best score one can achieve using only the provided train data and the transaction data set. What features did you and others from the top 50 leaderboard build off the transaction data? </p>\n\n<p>Reason I ask is that I could not bother to use the scala script to generate more train data, and I noticed that features from the user log data were close to useless so I only used transaction data. </p>\n\n<p>Hence, using only the provided train data and the transaction data I would like to know what the best score one could have achieved? I could not go below 0.14 but I can't help thinking that better feature engineering could have taken me below that, perhaps 0.12, or even 0.11?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi Bryan, do you still plan on sharing your solution?\nIt seems that only [one person has released their solution so far][1] and it's not clear what features were used. So any description of what has helped you rise to 1st place would be interesting to me and surely lots of other people. \n\nA major question mark for me is what is the best score one can achieve using only the provided train data and the transaction data set. What features did you and others from the top 50 leaderboard build off the transaction data? \n\nReason I ask is that I could not bother to use the scala script to generate more train data, and I noticed that features from the user log data were close to useless so I only used transaction data. \n\nHence, using only the provided train data and the transaction data I would like to know what the best score one could have achieved? I could not go below 0.14 but I can't help thinking that better feature engineering could have taken me below that, perhaps 0.12, or even 0.11?\n\nThanks\n\n\n  [1]: https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/46078",
      "votes": null
    },
    {
      "id": "261908",
      "postDate": "12/24/2017 14:09:51",
      "content": "<p>By using train.csv and train_v2.csv, 0.14178 on public LB and 0.14096 on private LB with single lightGBM model. I didn't tune the parameters since I write the labeling code a day after entered the competition. Only two monthes is not enough.</p>",
      "rawMarkdown": "By using train.csv and train_v2.csv, 0.14178 on public LB and 0.14096 on private LB with single lightGBM model. I didn't tune the parameters since I write the labeling code a day after entered the competition. Only two monthes is not enough.",
      "votes": null
    },
    {
      "id": "261909",
      "postDate": "12/24/2017 14:19:56",
      "content": "<p>Got it. It sounds like the only way for me to get below .14 was by labeling more data. Thanks a lot InifiniteWing</p>",
      "rawMarkdown": "Got it. It sounds like the only way for me to get below .14 was by labeling more data. Thanks a lot InifiniteWing",
      "votes": null
    },
    {
      "id": "261913",
      "postDate": "12/24/2017 14:37:43",
      "content": "<p>Thanks, I am also wondering about others' best score by using original training data.</p>",
      "rawMarkdown": "Thanks, I am also wondering about others' best score by using original training data.",
      "votes": null
    },
    {
      "id": "262008",
      "postDate": "12/24/2017 22:04:04",
      "content": "<p>My best score using only the provided training labels was .123 .  I could not run the scala program on my windows machine.  </p>\n\n<p>Mkffl, people who did not use the scala file were at a distinct disadvantage since 60% of the is_churn labels in the provided training files were not consistent with what the scala file generated.  (You can see this if you use the 201703 scala file provided by InfiniteWing.)  Presuming that the scala file generated the submission labels, you were training your model on data that was 60% wrong.  </p>\n\n<p>The user logs were not that useful because a lot of the customers who were identified as churners (using the non-scala labels) were streaming music AFTER their expiration date.  This is because they were mislabeled as churners.  </p>\n\n<p>By the way, based on a recommendation from Bryan, I took my predictions that generated the .123 score and scaled them down so that they predicted a 3.5% churn rate, rather than the 9% churn rate implied by the training data.  My score improved from .123 to about .1165.  That makes sense because the training data implied a significantly higher churn rate than the scala-generated data.</p>\n\n<p>Happy holidays to all!</p>",
      "rawMarkdown": "My best score using only the provided training labels was .123 .  I could not run the scala program on my windows machine.  \n\nMkffl, people who did not use the scala file were at a distinct disadvantage since 60% of the is_churn labels in the provided training files were not consistent with what the scala file generated.  (You can see this if you use the 201703 scala file provided by InfiniteWing.)  Presuming that the scala file generated the submission labels, you were training your model on data that was 60% wrong.  \n\nThe user logs were not that useful because a lot of the customers who were identified as churners (using the non-scala labels) were streaming music AFTER their expiration date.  This is because they were mislabeled as churners.  \n\nBy the way, based on a recommendation from Bryan, I took my predictions that generated the .123 score and scaled them down so that they predicted a 3.5% churn rate, rather than the 9% churn rate implied by the training data.  My score improved from .123 to about .1165.  That makes sense because the training data implied a significantly higher churn rate than the scala-generated data.\n\nHappy holidays to all!",
      "votes": null
    },
    {
      "id": "262115",
      "postDate": "12/25/2017 09:58:33",
      "content": "<p>How did you scale down your predictions to get from 9% to 3.5%? Did you just keep reducing the log-loss scores until you reached 3.5% churn, assuming a 0.5 cut-off (i.e. score &lt; 0..5 &lt;=&gt; non-churn)?</p>\n\n<p>Re the scala file - I knew from Bryan and others' forum posts that the scala script outputs different is_churn labels than the provided training files. But I (wrongly?) assumed that the submission file was generated the same way as the provided training file. If my assumption were true, then I thought I'd be better off ignoring the scala script. </p>\n\n<p>However if the submission file was generated using the scala script, then I agree 100% that generating more training data using the scala script is the best thing to do. I would actually also generate February and March training data and don't use the crappy training data provided by the organiser. Can I ask how many month you extracted to build you entire train set?</p>\n\n<p>Back then it seemed crazy that the organiser would not do anything to reconciliate the training data labels with the scala script labels. But it seems even more crazy that they would use different scripts to generate the training and the submission files. Some have said that IRL data is never clean hence such discrepancies are to be expected, but I argue that what the organiser has done is beyond non-cleanliness, it's a lack of rigour that spawns issues and question marks which could be avoided if they were a little bit more careful.</p>\n\n<p>Thanks for your reply.</p>",
      "rawMarkdown": "How did you scale down your predictions to get from 9% to 3.5%? Did you just keep reducing the log-loss scores until you reached 3.5% churn, assuming a 0.5 cut-off (i.e. score &lt; 0..5 &lt;=&gt; non-churn)?\n\nRe the scala file - I knew from Bryan and others' forum posts that the scala script outputs different is_churn labels than the provided training files. But I (wrongly?) assumed that the submission file was generated the same way as the provided training file. If my assumption were true, then I thought I'd be better off ignoring the scala script. \n\nHowever if the submission file was generated using the scala script, then I agree 100% that generating more training data using the scala script is the best thing to do. I would actually also generate February and March training data and don't use the crappy training data provided by the organiser. Can I ask how many month you extracted to build you entire train set?\n\nBack then it seemed crazy that the organiser would not do anything to reconciliate the training data labels with the scala script labels. But it seems even more crazy that they would use different scripts to generate the training and the submission files. Some have said that IRL data is never clean hence such discrepancies are to be expected, but I argue that what the organiser has done is beyond non-cleanliness, it's a lack of rigour that spawns issues and question marks which could be avoided if they were a little bit more careful.\n\nThanks for your reply.",
      "votes": null
    },
    {
      "id": "262159",
      "postDate": "12/25/2017 16:06:48",
      "content": "<p>Congrats Bryan for the 1st place!\nWe created 150+ features, most of them are from userlogs and transactions.\nWe use all the features in the final model, we don't really know how to select the features -_-.\nBTW, we will share our solution in few days, cause my teammates and I are all busy with the final exams now!</p>",
      "rawMarkdown": "Congrats Bryan for the 1st place!\nWe created 150+ features, most of them are from userlogs and transactions.\nWe use all the features in the final model, we don't really know how to select the features -_-.\nBTW, we will share our solution in few days, cause my teammates and I are all busy with the final exams now!",
      "votes": null
    },
    {
      "id": "262213",
      "postDate": "12/25/2017 20:39:25",
      "content": "<p>There are several ways of scaling down your predictions to generate an overall prediction rate of 3.5%.  I choose the simplest technique.   My prediction rate for my best score had a churn rate of about 6%.  I simply multiplied them by .60 to get an overall prediction rate of 3.6%.  I choose the simplest way because the competition is over and I simply wanted to see what difference it would make.  You should give it a try.</p>\n\n<p>Regarding the scala file, my opinion is that the organizers messed up and provided us with problematic is_churn labels.   I don't think they wanted to admit that they messed up a second time, following the leakage issues (which I read about but had not joined yet).</p>\n\n<p>I don't understand your question regarding how many months I used to build my training set.  My confusion probably means that I never understood how to use the training labels.  I simply assumed that the provided training labels were correct - in other words, they identified all of the March churners.  So, for my training set, I restricted data to before Feb 28 and I removed msno's that did not have an is_churn label.  I did not do anything else to define my training population.  </p>\n\n<p>Like you, I assumed the scala program was a nice-to-have.  I thought it could be used in two ways.  First, it could be used to expand the training msno's because there were a lot of msno's that did not have an is_churn value.  Second, the scala file could be used to generate previous churn behavior of people in my training set and people in the submission file.  In other words, given the definition of churn, it is probably likely that there are msno's that regularly churn and then unchurn.  Using the scala program you can identify these people.  </p>\n\n<p>Regarding your last comment, I believe that the organizers did a poor job in setting up the contest.  They simply provided training labels that were grossly incorrect.  A competitor needed to use the scala file to correct.  The scala file was public so it was fair game.  But, like you, I didn't think it was a requirement. </p>",
      "rawMarkdown": "There are several ways of scaling down your predictions to generate an overall prediction rate of 3.5%.  I choose the simplest technique.   My prediction rate for my best score had a churn rate of about 6%.  I simply multiplied them by .60 to get an overall prediction rate of 3.6%.  I choose the simplest way because the competition is over and I simply wanted to see what difference it would make.  You should give it a try.\n\nRegarding the scala file, my opinion is that the organizers messed up and provided us with problematic is_churn labels.   I don't think they wanted to admit that they messed up a second time, following the leakage issues (which I read about but had not joined yet).\n\nI don't understand your question regarding how many months I used to build my training set.  My confusion probably means that I never understood how to use the training labels.  I simply assumed that the provided training labels were correct - in other words, they identified all of the March churners.  So, for my training set, I restricted data to before Feb 28 and I removed msno's that did not have an is_churn label.  I did not do anything else to define my training population.  \n\nLike you, I assumed the scala program was a nice-to-have.  I thought it could be used in two ways.  First, it could be used to expand the training msno's because there were a lot of msno's that did not have an is_churn value.  Second, the scala file could be used to generate previous churn behavior of people in my training set and people in the submission file.  In other words, given the definition of churn, it is probably likely that there are msno's that regularly churn and then unchurn.  Using the scala program you can identify these people.  \n\nRegarding your last comment, I believe that the organizers did a poor job in setting up the contest.  They simply provided training labels that were grossly incorrect.  A competitor needed to use the scala file to correct.  The scala file was public so it was fair game.  But, like you, I didn't think it was a requirement.",
      "votes": null
    },
    {
      "id": "263017",
      "postDate": "12/28/2017 17:38:59",
      "content": "<p>Hello mkffl,</p>\n\n<p>Sorry for the delayed response, the contest ended right in the middle of Christmas/NY holidays and I've been travelling with family.  </p>\n\n<p>I do still plan on posting an overview of the approach, hopefully today, and I will add it to InfiniteWing's thread. I can definitely tell you that features from the UL data were not useless and had quite a bit of signal, in particular in interaction with the other data (transaction, members, and meta data).   A lot of the UL features ranked high in feature importance in my xgboost model.</p>\n\n<p>-Bryan</p>",
      "rawMarkdown": "Hello mkffl,\n\nSorry for the delayed response, the contest ended right in the middle of Christmas/NY holidays and I've been travelling with family.  \n\nI do still plan on posting an overview of the approach, hopefully today, and I will add it to InfiniteWing's thread. I can definitely tell you that features from the UL data were not useless and had quite a bit of signal, in particular in interaction with the other data (transaction, members, and meta data).   A lot of the UL features ranked high in feature importance in my xgboost model.\n\n-Bryan",
      "votes": null
    },
    {
      "id": "263025",
      "postDate": "12/28/2017 18:06:44",
      "content": "<p>The approach I used for scaling down the April predictions was two-fold:  I used only January data (users exp in Feb) as my training set, which had a much lower churn rate than February training data (users exp in Mar), even in the new scala-generated data sets.   Then I used a hyperparameter of scale_pos_weight=.8 in my xgboost model (the main base model) to scale down the predictions further until the mean churn rate was ~3.6.</p>\n\n<p>I think using scale_pos_weight is more accurate than just linear scaling because, by decreasing the loss of failed predictions, it has the effect of reducing the scale of positive churn predictions that are based on rare data/events, without affecting the scale of predictions that are based on common data/events.  In other words, it makes the predictions less noisy (scales down towards the mean churn rate) on the outlier data.  This intuitively makes sense as we know that the March test data had less churn overall, but users with common features that make them very likely to churn are still very likely to churn, and users with common features that make them very unlikely to churn are still very unlikely to churn.  </p>",
      "rawMarkdown": "The approach I used for scaling down the April predictions was two-fold:  I used only January data (users exp in Feb) as my training set, which had a much lower churn rate than February training data (users exp in Mar), even in the new scala-generated data sets.   Then I used a hyperparameter of scale_pos_weight=.8 in my xgboost model (the main base model) to scale down the predictions further until the mean churn rate was ~3.6.\n\nI think using scale_pos_weight is more accurate than just linear scaling because, by decreasing the loss of failed predictions, it has the effect of reducing the scale of positive churn predictions that are based on rare data/events, without affecting the scale of predictions that are based on common data/events.  In other words, it makes the predictions less noisy (scales down towards the mean churn rate) on the outlier data.  This intuitively makes sense as we know that the March test data had less churn overall, but users with common features that make them very likely to churn are still very likely to churn, and users with common features that make them very unlikely to churn are still very unlikely to churn.",
      "votes": null
    },
    {
      "id": "263461",
      "postDate": "12/30/2017 10:42:33",
      "content": "<p>@ SecondTimeAround, \nWe have broadly used the same training data, however you achieve a 0.018 gain (.141 [my best score] - .123 [your best score]). I wonder why:\n- It could be the user logs, which I have not used\n- Or maybe I have missed a killer feature from the transaction data set</p>\n\n<p>Just to clarify, my question re how many months you used to build your training set referred to getting more labeled data by applying a churn script, e.g. the scala programme, to previous months (i.e. prior to Feb and Mar 2017, which were provided by the organiser). </p>\n\n<p>@Bryan\nApologies for keeping you away from your family and a well-deserved break during the festive season. </p>\n\n<p>I understand that you wanted to align the overall % churn rate of Feb users (exp in Mar) with Jan users (exp in Feb). The reason you give is that the March test data had less churn overall. But, how did you know that? </p>\n\n<p>Thanks a lot!</p>",
      "rawMarkdown": "SecondTimeAround, \nWe have broadly used the same training data, however you achieve a 0.018 gain (.141 [my best score] - .123 [your best score]). I wonder why:\n- It could be the user logs, which I have not used\n- Or maybe I have missed a killer feature from the transaction data set\n\nJust to clarify, my question re how many months you used to build your training set referred to getting more labeled data by applying a churn script, e.g. the scala programme, to previous months (i.e. prior to Feb and Mar 2017, which were provided by the organiser). \n\n@Bryan\nApologies for keeping you away from your family and a well-deserved break during the festive season. \n\nI understand that you wanted to align the overall % churn rate of Feb users (exp in Mar) with Jan users (exp in Feb). The reason you give is that the March test data had less churn overall. But, how did you know that? \n\nThanks a lot!",
      "votes": null
    },
    {
      "id": "263603",
      "postDate": "12/31/2017 01:00:28",
      "content": "<p>@Mkffl, like you I did not find much predictive power using the user log data.  (I am guessing that many people who did not use the scala file to generate the correct churn flags would agree.  After all, if 60% of the churners were incorrectly labeled, then you would not expect 60% of the labeled churners to exhibit churn-like behavior in their user logs.)</p>\n\n<p>I don't think I had any killer features, but here are a few things that I did that might indicate differences in our models.</p>\n\n<p>1) My model contained 25 variables.  </p>\n\n<p>2) I used Naive Bayes estimators to capture interaction effects between the features.  These variables were among the most predictive.</p>\n\n<p>3) I had variables that looked at relationships between expiration dates, lagged expiration dates, transaction dates, and lagged transaction dates.  People who didn't churn tended to have regular behavior.  As an example, a typical non-churner would have one transaction per month and an expiration date one month later.  The spacing between transactions, expiration dates, and transaction and expiration dates were relatively constant.  In contrast, churners tended to have irregular transaction behavior.   You might see multiple transactions in the month leading up to churn, or you might see a sudden change in membership expiration date several days before or after the previous expiration date.  These activities were associated with higher churn rates.  In some cases, churn rates were 90% or higher.  </p>\n\n<p>4) My final model was a combination of the output of a neural network and XGBoost.</p>\n\n<p>Regarding your question to Bryan on knowing March churn rates, I never really understood the terminology used.  I thought that the training data used people who had February expiration dates that churned in March, whereas the submission data used people with March expiration dates who churned in April.  Thus, using that definition, the March churn rate is given by the training data or scala-generated labels.   But if you were asking how to compute the overall churn rate for the submission data, then you can compute it by submitting a constant prediction rate (other than .5) and using the loss-function to solve for it.  This was pointed out to me by InfiniteWing in one of the discussions. </p>\n\n<p>Cheers....</p>",
      "rawMarkdown": "Mkffl, like you I did not find much predictive power using the user log data.  (I am guessing that many people who did not use the scala file to generate the correct churn flags would agree.  After all, if 60% of the churners were incorrectly labeled, then you would not expect 60% of the labeled churners to exhibit churn-like behavior in their user logs.)\n\nI don't think I had any killer features, but here are a few things that I did that might indicate differences in our models.\n\n1) My model contained 25 variables.  \n\n2) I used Naive Bayes estimators to capture interaction effects between the features.  These variables were among the most predictive.\n\n3) I had variables that looked at relationships between expiration dates, lagged expiration dates, transaction dates, and lagged transaction dates.  People who didn't churn tended to have regular behavior.  As an example, a typical non-churner would have one transaction per month and an expiration date one month later.  The spacing between transactions, expiration dates, and transaction and expiration dates were relatively constant.  In contrast, churners tended to have irregular transaction behavior.   You might see multiple transactions in the month leading up to churn, or you might see a sudden change in membership expiration date several days before or after the previous expiration date.  These activities were associated with higher churn rates.  In some cases, churn rates were 90% or higher.  \n\n4) My final model was a combination of the output of a neural network and XGBoost.\n\nRegarding your question to Bryan on knowing March churn rates, I never really understood the terminology used.  I thought that the training data used people who had February expiration dates that churned in March, whereas the submission data used people with March expiration dates who churned in April.  Thus, using that definition, the March churn rate is given by the training data or scala-generated labels.   But if you were asking how to compute the overall churn rate for the submission data, then you can compute it by submitting a constant prediction rate (other than .5) and using the loss-function to solve for it.  This was pointed out to me by InfiniteWing in one of the discussions. \n\nCheers....",
      "votes": null
    },
    {
      "id": "263606",
      "postDate": "12/31/2017 01:06:48",
      "content": "<p>Bryan, I do agree that adjusting the probabilities by a constant factor is not optimal.  But it was the quickest approach and I just wanted to see if it would make a difference.  Cheers...</p>",
      "rawMarkdown": "Bryan, I do agree that adjusting the probabilities by a constant factor is not optimal.  But it was the quickest approach and I just wanted to see if it would make a difference.  Cheers...",
      "votes": null
    },
    {
      "id": "265342",
      "postDate": "01/05/2018 09:13:09",
      "content": "<p>Really useful explanations, thanks a lot. What does 2) relate to? What do you mean by capturing interactions, and how do you use naive bayes? </p>\n\n<p>Re your last point, the confusion is due to the change in data set availability following the leak. Also the terminology used by the organiser is super confusing. With this type of time series / churn data, it's important to be super clear with dates, and they did a poor job. For example in the Data section, you can read things like 'predict user churn in the month of April, 2017.', and you are never really sure what they mean by that. </p>\n\n<p>However, as you said I was asking about how to compute overall churn on the submission test. I will keep the constant prediction rate trick in mind going forward.</p>",
      "rawMarkdown": "Really useful explanations, thanks a lot. What does 2) relate to? What do you mean by capturing interactions, and how do you use naive bayes? \n\nRe your last point, the confusion is due to the change in data set availability following the leak. Also the terminology used by the organiser is super confusing. With this type of time series / churn data, it's important to be super clear with dates, and they did a poor job. For example in the Data section, you can read things like 'predict user churn in the month of April, 2017.', and you are never really sure what they mean by that. \n\nHowever, as you said I was asking about how to compute overall churn on the submission test. I will keep the constant prediction rate trick in mind going forward.",
      "votes": null
    },
    {
      "id": "269333",
      "postDate": "01/16/2018 16:39:43",
      "content": "<p>Dear Mr. Gregory,\nNow I am doing the student research using this dataset. I do not know much about data mining, so I run data mining algorithm such as decision tree, random forest, k-nn and I found that the accuracy of these models are approximately 60% which is a little bit low. Could you please provide me recommendations to improve the accuracy?  Are there any new features that can improve the accuracy?</p>",
      "rawMarkdown": "Dear Mr. Gregory,\nNow I am doing the student research using this dataset. I do not know much about data mining, so I run data mining algorithm such as decision tree, random forest, k-nn and I found that the accuracy of these models are approximately 60% which is a little bit low. Could you please provide me recommendations to improve the accuracy?  Are there any new features that can improve the accuracy?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 259289,
      "author_name": "casscw",
      "author_url": "",
      "post_date": "12/18/2017 04:24:31",
      "content": "<p>Congrats for the 1st place! <br>\nI just created 56 features【20 from user_logs，36 from members and transactions】 for the final model【LGB, 5-CV】.   Too few features for me⊙﹏⊙</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 259311,
      "author_name": "sundong",
      "author_url": "",
      "post_date": "12/18/2017 05:22:59",
      "content": "<p>I created 220 features - approx 10 from members,  140 from transactions, and 60 from user logs,  10 from combined. The number looks large, since we used multiple aspects(max, min, avg, first, second, skewness) from each 'groupby' functions.  (e.g., avg, first, skewness value of transaction gap sequence). Our 0.10 loss is mainly achieved by transaction features, and we didn't manage to generate strong features from user logs. </p>\n\n<p>However, I felt tedious and tired while generating features, especially when the leaderboard score is stuck at the certain level. Does anyone tried feature learning or representation learning to tackle this competition?</p>\n\n<p>@Bryan: Congratulations! I'm really curious what kind of features you generated from user logs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 260181,
          "author_name": "bryangregory",
          "author_url": "",
          "post_date": "12/19/2017 20:13:20",
          "content": "<p>Out of curiosity, did you scale the UL features at all relative to the membership expiration date, or did you stick with grouping by just calendar dates?</p>\n\n<p>For ex., MSNO \"21dh15HZdEDWGomm1AKlnmAvqYbT3qP1zg8cGxZA+0I=\".  For the calendar month of Jan, they have one user login (on 20170131).  So if you take a calendar month count of logins or a sum of secs using the app, then this user appears likely to churn (very low usage relative to user averages for January).  But if you dig deeper, the user just started their subscription on 20170131.  So a single login doesn't have any meaning for this user unless you create the features relative to the # of transaction days in the month.  Calendar month activity is 3.2% (1 login day / 31 potential login days), but relative activity is 100% (1 login day / 1 potential login day).</p>\n\n<p>Hope that makes some sense, I might be doing a bad job of explaining it because I'm in a rush.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260221,
          "author_name": "",
          "author_url": "",
          "post_date": "12/19/2017 22:00:04",
          "content": "<p>Hi Bryan, that makes sense.  I based many of my user log features relative to a user's membership date and / or registration date.  Using the registration date, as you pointed out, takes care of the issue that you raised.</p>\n\n<p>But I also found that measuring any behavior relative to dates (for example, membership expiration dates) was tricky.  For example, there are fewer days in between 3/15 and 2/15 than there are between 4/15 and 3/15.  That can create problems.  For example, on the training set you might find that if a person initiates a transaction 27 days before expiration then the churn rate was 30%.  But that logic cannot be applied to the submission set because the 27 days has a different meaning for the training and submission sets.  For the training set, 27 days from an expiration date in March brings you to the previous expiration date in February (presuming you have a \"30\" day subscription).  For the submission set, 27 days from a expiration date in April brings you to 2 days AFTER the previous expiration date.  That made a difference for some of my features.  So I had to adjust the definition of the features when scoring the submission set.</p>\n\n<p>You also might have noticed that the distribution of the end-of-the-month membership dates in the submission set was skewed versus the training set.  I think that is because membership expiration dates for February 28, captured previous membership expiration dates for January 28, 29, and 30 (presuming you have a \"30\" day subscription).  However, for March, a January 28 expiration date rolled to March 28, a January 29 rolled to March 29, and a January 30 rolled to March 30.  That also made a difference for some of my features that used date math.  </p>\n\n<p>Cheers, </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 261898,
      "author_name": "mchlkffl",
      "author_url": "",
      "post_date": "12/24/2017 12:28:31",
      "content": "<p>Hi Bryan, do you still plan on sharing your solution?\nIt seems that only <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/46078\">one person has released their solution so far</a> and it's not clear what features were used. So any description of what has helped you rise to 1st place would be interesting to me and surely lots of other people. </p>\n\n<p>A major question mark for me is what is the best score one can achieve using only the provided train data and the transaction data set. What features did you and others from the top 50 leaderboard build off the transaction data? </p>\n\n<p>Reason I ask is that I could not bother to use the scala script to generate more train data, and I noticed that features from the user log data were close to useless so I only used transaction data. </p>\n\n<p>Hence, using only the provided train data and the transaction data I would like to know what the best score one could have achieved? I could not go below 0.14 but I can't help thinking that better feature engineering could have taken me below that, perhaps 0.12, or even 0.11?</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 261908,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/24/2017 14:09:51",
          "content": "<p>By using train.csv and train_v2.csv, 0.14178 on public LB and 0.14096 on private LB with single lightGBM model. I didn't tune the parameters since I write the labeling code a day after entered the competition. Only two monthes is not enough.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261909,
          "author_name": "mchlkffl",
          "author_url": "",
          "post_date": "12/24/2017 14:19:56",
          "content": "<p>Got it. It sounds like the only way for me to get below .14 was by labeling more data. Thanks a lot InifiniteWing</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261913,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/24/2017 14:37:43",
          "content": "<p>Thanks, I am also wondering about others' best score by using original training data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 262008,
          "author_name": "",
          "author_url": "",
          "post_date": "12/24/2017 22:04:04",
          "content": "<p>My best score using only the provided training labels was .123 .  I could not run the scala program on my windows machine.  </p>\n\n<p>Mkffl, people who did not use the scala file were at a distinct disadvantage since 60% of the is_churn labels in the provided training files were not consistent with what the scala file generated.  (You can see this if you use the 201703 scala file provided by InfiniteWing.)  Presuming that the scala file generated the submission labels, you were training your model on data that was 60% wrong.  </p>\n\n<p>The user logs were not that useful because a lot of the customers who were identified as churners (using the non-scala labels) were streaming music AFTER their expiration date.  This is because they were mislabeled as churners.  </p>\n\n<p>By the way, based on a recommendation from Bryan, I took my predictions that generated the .123 score and scaled them down so that they predicted a 3.5% churn rate, rather than the 9% churn rate implied by the training data.  My score improved from .123 to about .1165.  That makes sense because the training data implied a significantly higher churn rate than the scala-generated data.</p>\n\n<p>Happy holidays to all!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 262115,
          "author_name": "mchlkffl",
          "author_url": "",
          "post_date": "12/25/2017 09:58:33",
          "content": "<p>How did you scale down your predictions to get from 9% to 3.5%? Did you just keep reducing the log-loss scores until you reached 3.5% churn, assuming a 0.5 cut-off (i.e. score &lt; 0..5 &lt;=&gt; non-churn)?</p>\n\n<p>Re the scala file - I knew from Bryan and others' forum posts that the scala script outputs different is_churn labels than the provided training files. But I (wrongly?) assumed that the submission file was generated the same way as the provided training file. If my assumption were true, then I thought I'd be better off ignoring the scala script. </p>\n\n<p>However if the submission file was generated using the scala script, then I agree 100% that generating more training data using the scala script is the best thing to do. I would actually also generate February and March training data and don't use the crappy training data provided by the organiser. Can I ask how many month you extracted to build you entire train set?</p>\n\n<p>Back then it seemed crazy that the organiser would not do anything to reconciliate the training data labels with the scala script labels. But it seems even more crazy that they would use different scripts to generate the training and the submission files. Some have said that IRL data is never clean hence such discrepancies are to be expected, but I argue that what the organiser has done is beyond non-cleanliness, it's a lack of rigour that spawns issues and question marks which could be avoided if they were a little bit more careful.</p>\n\n<p>Thanks for your reply.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 262213,
          "author_name": "",
          "author_url": "",
          "post_date": "12/25/2017 20:39:25",
          "content": "<p>There are several ways of scaling down your predictions to generate an overall prediction rate of 3.5%.  I choose the simplest technique.   My prediction rate for my best score had a churn rate of about 6%.  I simply multiplied them by .60 to get an overall prediction rate of 3.6%.  I choose the simplest way because the competition is over and I simply wanted to see what difference it would make.  You should give it a try.</p>\n\n<p>Regarding the scala file, my opinion is that the organizers messed up and provided us with problematic is_churn labels.   I don't think they wanted to admit that they messed up a second time, following the leakage issues (which I read about but had not joined yet).</p>\n\n<p>I don't understand your question regarding how many months I used to build my training set.  My confusion probably means that I never understood how to use the training labels.  I simply assumed that the provided training labels were correct - in other words, they identified all of the March churners.  So, for my training set, I restricted data to before Feb 28 and I removed msno's that did not have an is_churn label.  I did not do anything else to define my training population.  </p>\n\n<p>Like you, I assumed the scala program was a nice-to-have.  I thought it could be used in two ways.  First, it could be used to expand the training msno's because there were a lot of msno's that did not have an is_churn value.  Second, the scala file could be used to generate previous churn behavior of people in my training set and people in the submission file.  In other words, given the definition of churn, it is probably likely that there are msno's that regularly churn and then unchurn.  Using the scala program you can identify these people.  </p>\n\n<p>Regarding your last comment, I believe that the organizers did a poor job in setting up the contest.  They simply provided training labels that were grossly incorrect.  A competitor needed to use the scala file to correct.  The scala file was public so it was fair game.  But, like you, I didn't think it was a requirement. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 263017,
          "author_name": "bryangregory",
          "author_url": "",
          "post_date": "12/28/2017 17:38:59",
          "content": "<p>Hello mkffl,</p>\n\n<p>Sorry for the delayed response, the contest ended right in the middle of Christmas/NY holidays and I've been travelling with family.  </p>\n\n<p>I do still plan on posting an overview of the approach, hopefully today, and I will add it to InfiniteWing's thread. I can definitely tell you that features from the UL data were not useless and had quite a bit of signal, in particular in interaction with the other data (transaction, members, and meta data).   A lot of the UL features ranked high in feature importance in my xgboost model.</p>\n\n<p>-Bryan</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 263025,
          "author_name": "bryangregory",
          "author_url": "",
          "post_date": "12/28/2017 18:06:44",
          "content": "<p>The approach I used for scaling down the April predictions was two-fold:  I used only January data (users exp in Feb) as my training set, which had a much lower churn rate than February training data (users exp in Mar), even in the new scala-generated data sets.   Then I used a hyperparameter of scale_pos_weight=.8 in my xgboost model (the main base model) to scale down the predictions further until the mean churn rate was ~3.6.</p>\n\n<p>I think using scale_pos_weight is more accurate than just linear scaling because, by decreasing the loss of failed predictions, it has the effect of reducing the scale of positive churn predictions that are based on rare data/events, without affecting the scale of predictions that are based on common data/events.  In other words, it makes the predictions less noisy (scales down towards the mean churn rate) on the outlier data.  This intuitively makes sense as we know that the March test data had less churn overall, but users with common features that make them very likely to churn are still very likely to churn, and users with common features that make them very unlikely to churn are still very unlikely to churn.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 263461,
          "author_name": "mchlkffl",
          "author_url": "",
          "post_date": "12/30/2017 10:42:33",
          "content": "<p>@ SecondTimeAround, \nWe have broadly used the same training data, however you achieve a 0.018 gain (.141 [my best score] - .123 [your best score]). I wonder why:\n- It could be the user logs, which I have not used\n- Or maybe I have missed a killer feature from the transaction data set</p>\n\n<p>Just to clarify, my question re how many months you used to build your training set referred to getting more labeled data by applying a churn script, e.g. the scala programme, to previous months (i.e. prior to Feb and Mar 2017, which were provided by the organiser). </p>\n\n<p>@Bryan\nApologies for keeping you away from your family and a well-deserved break during the festive season. </p>\n\n<p>I understand that you wanted to align the overall % churn rate of Feb users (exp in Mar) with Jan users (exp in Feb). The reason you give is that the March test data had less churn overall. But, how did you know that? </p>\n\n<p>Thanks a lot!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 263603,
          "author_name": "",
          "author_url": "",
          "post_date": "12/31/2017 01:00:28",
          "content": "<p>@Mkffl, like you I did not find much predictive power using the user log data.  (I am guessing that many people who did not use the scala file to generate the correct churn flags would agree.  After all, if 60% of the churners were incorrectly labeled, then you would not expect 60% of the labeled churners to exhibit churn-like behavior in their user logs.)</p>\n\n<p>I don't think I had any killer features, but here are a few things that I did that might indicate differences in our models.</p>\n\n<p>1) My model contained 25 variables.  </p>\n\n<p>2) I used Naive Bayes estimators to capture interaction effects between the features.  These variables were among the most predictive.</p>\n\n<p>3) I had variables that looked at relationships between expiration dates, lagged expiration dates, transaction dates, and lagged transaction dates.  People who didn't churn tended to have regular behavior.  As an example, a typical non-churner would have one transaction per month and an expiration date one month later.  The spacing between transactions, expiration dates, and transaction and expiration dates were relatively constant.  In contrast, churners tended to have irregular transaction behavior.   You might see multiple transactions in the month leading up to churn, or you might see a sudden change in membership expiration date several days before or after the previous expiration date.  These activities were associated with higher churn rates.  In some cases, churn rates were 90% or higher.  </p>\n\n<p>4) My final model was a combination of the output of a neural network and XGBoost.</p>\n\n<p>Regarding your question to Bryan on knowing March churn rates, I never really understood the terminology used.  I thought that the training data used people who had February expiration dates that churned in March, whereas the submission data used people with March expiration dates who churned in April.  Thus, using that definition, the March churn rate is given by the training data or scala-generated labels.   But if you were asking how to compute the overall churn rate for the submission data, then you can compute it by submitting a constant prediction rate (other than .5) and using the loss-function to solve for it.  This was pointed out to me by InfiniteWing in one of the discussions. </p>\n\n<p>Cheers....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 263606,
          "author_name": "",
          "author_url": "",
          "post_date": "12/31/2017 01:06:48",
          "content": "<p>Bryan, I do agree that adjusting the probabilities by a constant factor is not optimal.  But it was the quickest approach and I just wanted to see if it would make a difference.  Cheers...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 265342,
          "author_name": "mchlkffl",
          "author_url": "",
          "post_date": "01/05/2018 09:13:09",
          "content": "<p>Really useful explanations, thanks a lot. What does 2) relate to? What do you mean by capturing interactions, and how do you use naive bayes? </p>\n\n<p>Re your last point, the confusion is due to the change in data set availability following the leak. Also the terminology used by the organiser is super confusing. With this type of time series / churn data, it's important to be super clear with dates, and they did a poor job. For example in the Data section, you can read things like 'predict user churn in the month of April, 2017.', and you are never really sure what they mean by that. </p>\n\n<p>However, as you said I was asking about how to compute overall churn on the submission test. I will keep the constant prediction rate trick in mind going forward.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 262159,
      "author_name": "shichence",
      "author_url": "",
      "post_date": "12/25/2017 16:06:48",
      "content": "<p>Congrats Bryan for the 1st place!\nWe created 150+ features, most of them are from userlogs and transactions.\nWe use all the features in the final model, we don't really know how to select the features -_-.\nBTW, we will share our solution in few days, cause my teammates and I are all busy with the final exams now!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 269333,
      "author_name": "billiazz",
      "author_url": "",
      "post_date": "01/16/2018 16:39:43",
      "content": "<p>Dear Mr. Gregory,\nNow I am doing the student research using this dataset. I do not know much about data mining, so I run data mining algorithm such as decision tree, random forest, k-nn and I found that the accuracy of these models are approximately 60% which is a little bit low. Could you please provide me recommendations to improve the accuracy?  Are there any new features that can improve the accuracy?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "259263": "Considering how rich this dataset was for feature engineering, more so than many past contests I've seen on Kaggle, I am curious how many features did other teams create? And how many made it into your final models?  For me, I had ~250 features total that were created and tried of which ~100 were used in final models.  The bulk of the features came from the user logs, although many came from transactions, a few from members, and some were interactions between them as well as meta features.\n\nThis of course doesn't include hyper parameters, base model predictions, etc. as features.  Just discussing strictly features extracted from the data itself.\n\nI continuously saw accuracy increases in both CV and LB every time I went back for another round of feature engineering, so it would be interesting to know at what point the data is finally \"exhausted\", and how many of these would generalize well over time.",
    "259289": "Congrats for the 1st place!     \nI just created 56 features【20 from user_logs，36 from members and transactions】 for the final model【LGB, 5-CV】.   Too few features for me⊙﹏⊙",
    "259311": "I created 220 features - approx 10 from members,  140 from transactions, and 60 from user logs,  10 from combined. The number looks large, since we used multiple aspects(max, min, avg, first, second, skewness) from each 'groupby' functions.  (e.g., avg, first, skewness value of transaction gap sequence). Our 0.10 loss is mainly achieved by transaction features, and we didn't manage to generate strong features from user logs. \n\nHowever, I felt tedious and tired while generating features, especially when the leaderboard score is stuck at the certain level. Does anyone tried feature learning or representation learning to tackle this competition?\n\n@Bryan: Congratulations! I'm really curious what kind of features you generated from user logs.",
    "260181": "Out of curiosity, did you scale the UL features at all relative to the membership expiration date, or did you stick with grouping by just calendar dates?\n\nFor ex., MSNO \"21dh15HZdEDWGomm1AKlnmAvqYbT3qP1zg8cGxZA+0I=\".  For the calendar month of Jan, they have one user login (on 20170131).  So if you take a calendar month count of logins or a sum of secs using the app, then this user appears likely to churn (very low usage relative to user averages for January).  But if you dig deeper, the user just started their subscription on 20170131.  So a single login doesn't have any meaning for this user unless you create the features relative to the # of transaction days in the month.  Calendar month activity is 3.2% (1 login day / 31 potential login days), but relative activity is 100% (1 login day / 1 potential login day).\n\nHope that makes some sense, I might be doing a bad job of explaining it because I'm in a rush.",
    "260221": "Hi Bryan, that makes sense.  I based many of my user log features relative to a user's membership date and / or registration date.  Using the registration date, as you pointed out, takes care of the issue that you raised.\n\nBut I also found that measuring any behavior relative to dates (for example, membership expiration dates) was tricky.  For example, there are fewer days in between 3/15 and 2/15 than there are between 4/15 and 3/15.  That can create problems.  For example, on the training set you might find that if a person initiates a transaction 27 days before expiration then the churn rate was 30%.  But that logic cannot be applied to the submission set because the 27 days has a different meaning for the training and submission sets.  For the training set, 27 days from an expiration date in March brings you to the previous expiration date in February (presuming you have a \"30\" day subscription).  For the submission set, 27 days from a expiration date in April brings you to 2 days AFTER the previous expiration date.  That made a difference for some of my features.  So I had to adjust the definition of the features when scoring the submission set.\n\nYou also might have noticed that the distribution of the end-of-the-month membership dates in the submission set was skewed versus the training set.  I think that is because membership expiration dates for February 28, captured previous membership expiration dates for January 28, 29, and 30 (presuming you have a \"30\" day subscription).  However, for March, a January 28 expiration date rolled to March 28, a January 29 rolled to March 29, and a January 30 rolled to March 30.  That also made a difference for some of my features that used date math.  \n\nCheers,",
    "261898": "Hi Bryan, do you still plan on sharing your solution?\nIt seems that only [one person has released their solution so far][1] and it's not clear what features were used. So any description of what has helped you rise to 1st place would be interesting to me and surely lots of other people. \n\nA major question mark for me is what is the best score one can achieve using only the provided train data and the transaction data set. What features did you and others from the top 50 leaderboard build off the transaction data? \n\nReason I ask is that I could not bother to use the scala script to generate more train data, and I noticed that features from the user log data were close to useless so I only used transaction data. \n\nHence, using only the provided train data and the transaction data I would like to know what the best score one could have achieved? I could not go below 0.14 but I can't help thinking that better feature engineering could have taken me below that, perhaps 0.12, or even 0.11?\n\nThanks\n\n\n  [1]: https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/46078",
    "261908": "By using train.csv and train_v2.csv, 0.14178 on public LB and 0.14096 on private LB with single lightGBM model. I didn't tune the parameters since I write the labeling code a day after entered the competition. Only two monthes is not enough.",
    "261909": "Got it. It sounds like the only way for me to get below .14 was by labeling more data. Thanks a lot InifiniteWing",
    "261913": "Thanks, I am also wondering about others' best score by using original training data.",
    "262008": "My best score using only the provided training labels was .123 .  I could not run the scala program on my windows machine.  \n\nMkffl, people who did not use the scala file were at a distinct disadvantage since 60% of the is_churn labels in the provided training files were not consistent with what the scala file generated.  (You can see this if you use the 201703 scala file provided by InfiniteWing.)  Presuming that the scala file generated the submission labels, you were training your model on data that was 60% wrong.  \n\nThe user logs were not that useful because a lot of the customers who were identified as churners (using the non-scala labels) were streaming music AFTER their expiration date.  This is because they were mislabeled as churners.  \n\nBy the way, based on a recommendation from Bryan, I took my predictions that generated the .123 score and scaled them down so that they predicted a 3.5% churn rate, rather than the 9% churn rate implied by the training data.  My score improved from .123 to about .1165.  That makes sense because the training data implied a significantly higher churn rate than the scala-generated data.\n\nHappy holidays to all!",
    "262115": "How did you scale down your predictions to get from 9% to 3.5%? Did you just keep reducing the log-loss scores until you reached 3.5% churn, assuming a 0.5 cut-off (i.e. score &lt; 0..5 &lt;=&gt; non-churn)?\n\nRe the scala file - I knew from Bryan and others' forum posts that the scala script outputs different is_churn labels than the provided training files. But I (wrongly?) assumed that the submission file was generated the same way as the provided training file. If my assumption were true, then I thought I'd be better off ignoring the scala script. \n\nHowever if the submission file was generated using the scala script, then I agree 100% that generating more training data using the scala script is the best thing to do. I would actually also generate February and March training data and don't use the crappy training data provided by the organiser. Can I ask how many month you extracted to build you entire train set?\n\nBack then it seemed crazy that the organiser would not do anything to reconciliate the training data labels with the scala script labels. But it seems even more crazy that they would use different scripts to generate the training and the submission files. Some have said that IRL data is never clean hence such discrepancies are to be expected, but I argue that what the organiser has done is beyond non-cleanliness, it's a lack of rigour that spawns issues and question marks which could be avoided if they were a little bit more careful.\n\nThanks for your reply.",
    "262159": "Congrats Bryan for the 1st place!\nWe created 150+ features, most of them are from userlogs and transactions.\nWe use all the features in the final model, we don't really know how to select the features -_-.\nBTW, we will share our solution in few days, cause my teammates and I are all busy with the final exams now!",
    "262213": "There are several ways of scaling down your predictions to generate an overall prediction rate of 3.5%.  I choose the simplest technique.   My prediction rate for my best score had a churn rate of about 6%.  I simply multiplied them by .60 to get an overall prediction rate of 3.6%.  I choose the simplest way because the competition is over and I simply wanted to see what difference it would make.  You should give it a try.\n\nRegarding the scala file, my opinion is that the organizers messed up and provided us with problematic is_churn labels.   I don't think they wanted to admit that they messed up a second time, following the leakage issues (which I read about but had not joined yet).\n\nI don't understand your question regarding how many months I used to build my training set.  My confusion probably means that I never understood how to use the training labels.  I simply assumed that the provided training labels were correct - in other words, they identified all of the March churners.  So, for my training set, I restricted data to before Feb 28 and I removed msno's that did not have an is_churn label.  I did not do anything else to define my training population.  \n\nLike you, I assumed the scala program was a nice-to-have.  I thought it could be used in two ways.  First, it could be used to expand the training msno's because there were a lot of msno's that did not have an is_churn value.  Second, the scala file could be used to generate previous churn behavior of people in my training set and people in the submission file.  In other words, given the definition of churn, it is probably likely that there are msno's that regularly churn and then unchurn.  Using the scala program you can identify these people.  \n\nRegarding your last comment, I believe that the organizers did a poor job in setting up the contest.  They simply provided training labels that were grossly incorrect.  A competitor needed to use the scala file to correct.  The scala file was public so it was fair game.  But, like you, I didn't think it was a requirement.",
    "263017": "Hello mkffl,\n\nSorry for the delayed response, the contest ended right in the middle of Christmas/NY holidays and I've been travelling with family.  \n\nI do still plan on posting an overview of the approach, hopefully today, and I will add it to InfiniteWing's thread. I can definitely tell you that features from the UL data were not useless and had quite a bit of signal, in particular in interaction with the other data (transaction, members, and meta data).   A lot of the UL features ranked high in feature importance in my xgboost model.\n\n-Bryan",
    "263025": "The approach I used for scaling down the April predictions was two-fold:  I used only January data (users exp in Feb) as my training set, which had a much lower churn rate than February training data (users exp in Mar), even in the new scala-generated data sets.   Then I used a hyperparameter of scale_pos_weight=.8 in my xgboost model (the main base model) to scale down the predictions further until the mean churn rate was ~3.6.\n\nI think using scale_pos_weight is more accurate than just linear scaling because, by decreasing the loss of failed predictions, it has the effect of reducing the scale of positive churn predictions that are based on rare data/events, without affecting the scale of predictions that are based on common data/events.  In other words, it makes the predictions less noisy (scales down towards the mean churn rate) on the outlier data.  This intuitively makes sense as we know that the March test data had less churn overall, but users with common features that make them very likely to churn are still very likely to churn, and users with common features that make them very unlikely to churn are still very unlikely to churn.",
    "263461": "SecondTimeAround, \nWe have broadly used the same training data, however you achieve a 0.018 gain (.141 [my best score] - .123 [your best score]). I wonder why:\n- It could be the user logs, which I have not used\n- Or maybe I have missed a killer feature from the transaction data set\n\nJust to clarify, my question re how many months you used to build your training set referred to getting more labeled data by applying a churn script, e.g. the scala programme, to previous months (i.e. prior to Feb and Mar 2017, which were provided by the organiser). \n\n@Bryan\nApologies for keeping you away from your family and a well-deserved break during the festive season. \n\nI understand that you wanted to align the overall % churn rate of Feb users (exp in Mar) with Jan users (exp in Feb). The reason you give is that the March test data had less churn overall. But, how did you know that? \n\nThanks a lot!",
    "263603": "Mkffl, like you I did not find much predictive power using the user log data.  (I am guessing that many people who did not use the scala file to generate the correct churn flags would agree.  After all, if 60% of the churners were incorrectly labeled, then you would not expect 60% of the labeled churners to exhibit churn-like behavior in their user logs.)\n\nI don't think I had any killer features, but here are a few things that I did that might indicate differences in our models.\n\n1) My model contained 25 variables.  \n\n2) I used Naive Bayes estimators to capture interaction effects between the features.  These variables were among the most predictive.\n\n3) I had variables that looked at relationships between expiration dates, lagged expiration dates, transaction dates, and lagged transaction dates.  People who didn't churn tended to have regular behavior.  As an example, a typical non-churner would have one transaction per month and an expiration date one month later.  The spacing between transactions, expiration dates, and transaction and expiration dates were relatively constant.  In contrast, churners tended to have irregular transaction behavior.   You might see multiple transactions in the month leading up to churn, or you might see a sudden change in membership expiration date several days before or after the previous expiration date.  These activities were associated with higher churn rates.  In some cases, churn rates were 90% or higher.  \n\n4) My final model was a combination of the output of a neural network and XGBoost.\n\nRegarding your question to Bryan on knowing March churn rates, I never really understood the terminology used.  I thought that the training data used people who had February expiration dates that churned in March, whereas the submission data used people with March expiration dates who churned in April.  Thus, using that definition, the March churn rate is given by the training data or scala-generated labels.   But if you were asking how to compute the overall churn rate for the submission data, then you can compute it by submitting a constant prediction rate (other than .5) and using the loss-function to solve for it.  This was pointed out to me by InfiniteWing in one of the discussions. \n\nCheers....",
    "263606": "Bryan, I do agree that adjusting the probabilities by a constant factor is not optimal.  But it was the quickest approach and I just wanted to see if it would make a difference.  Cheers...",
    "265342": "Really useful explanations, thanks a lot. What does 2) relate to? What do you mean by capturing interactions, and how do you use naive bayes? \n\nRe your last point, the confusion is due to the change in data set availability following the leak. Also the terminology used by the organiser is super confusing. With this type of time series / churn data, it's important to be super clear with dates, and they did a poor job. For example in the Data section, you can read things like 'predict user churn in the month of April, 2017.', and you are never really sure what they mean by that. \n\nHowever, as you said I was asking about how to compute overall churn on the submission test. I will keep the constant prediction rate trick in mind going forward.",
    "269333": "Dear Mr. Gregory,\nNow I am doing the student research using this dataset. I do not know much about data mining, so I run data mining algorithm such as decision tree, random forest, k-nn and I found that the accuracy of these models are approximately 60% which is a little bit low. Could you please provide me recommendations to improve the accuracy?  Are there any new features that can improve the accuracy?"
  },
  "source": "meta"
}