{
  "id": 40067,
  "title": "How do I predict churning out with class imbalance using survival analysis?",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/40067",
  "author_name": "",
  "post_date": "2017-09-27T08:03:10.172082600Z",
  "votes": 8,
  "comment_count": 11,
  "views": 0,
  "content": "<p><em>Disclaimer : This is not related directly to this competition. I am working on churning-out model and that's why I need help from you guys.</em></p>\n\n<p>I am working on e-commerce customers data.</p>\n\n<hr>\n\n<p>Initially I just used <strong>classification</strong> algorithms to calculate churn-out. But it <strong>can</strong> predict only whether a customer is going to churn out or not, it <strong>cannot</strong> predict <strong>when</strong> the customer is going to churn out or the probability of surviving on a particular future date.</p>\n\n<hr>\n\n<p>Then I used <strong>survival analysis</strong> to predict churn-out. I made a <strong>cox proportionality hazard model</strong> using <code>coxph</code> function from <code>survival</code> package in R <strong>(I used all data till 1st August</strong>). With that model, I used <code>predictSurvProb</code> function from <code>pec</code> package in R to <strong>calculate probability</strong> of churning out of all non-churned customers (as on 1st August) on <strong>10th August</strong>.</p>\n\n<p>And by using <strong>threshold of 0.5</strong> on probabilities to tell whether a customer has churned out or not, I got the following results -</p>\n\n<p><strong>Prediction Accuracies</strong></p>\n\n<p><strong>84.28 %</strong> - For all customers (who were not churned out on 1st august)</p>\n\n<p><strong>25.72 %</strong> - For customers who actually churned out between 1st to 10th August</p>\n\n<p>Then I checked why I am getting so low accuracy on churned-out customers, I found out that actually there were-</p>\n\n<p>27000 - number of customers who were not churned out as on 1st August</p>\n\n<p>312 - number of members who churned out of the 27000 members above between 1st to 10th August</p>\n\n<p>So definitely there is class imbalance problem. So how do I increase my prediction accuracy on customers who actually churn out?</p>\n\n<p>And if possible how can I use undersampling,oversampling, SMOTE, etc with survival analysis for my problem? <a href=\"https://www.analyticsvidhya.com/blog/2017/03/imbalanced-classification-problem/\">https://www.analyticsvidhya.com/blog/2017/03/imbalanced-classification-problem/</a></p>",
  "messages": [
    {
      "id": "224670",
      "postDate": "09/27/2017 08:03:10",
      "content": "<p><em>Disclaimer : This is not related directly to this competition. I am working on churning-out model and that's why I need help from you guys.</em></p>\n\n<p>I am working on e-commerce customers data.</p>\n\n<hr>\n\n<p>Initially I just used <strong>classification</strong> algorithms to calculate churn-out. But it <strong>can</strong> predict only whether a customer is going to churn out or not, it <strong>cannot</strong> predict <strong>when</strong> the customer is going to churn out or the probability of surviving on a particular future date.</p>\n\n<hr>\n\n<p>Then I used <strong>survival analysis</strong> to predict churn-out. I made a <strong>cox proportionality hazard model</strong> using <code>coxph</code> function from <code>survival</code> package in R <strong>(I used all data till 1st August</strong>). With that model, I used <code>predictSurvProb</code> function from <code>pec</code> package in R to <strong>calculate probability</strong> of churning out of all non-churned customers (as on 1st August) on <strong>10th August</strong>.</p>\n\n<p>And by using <strong>threshold of 0.5</strong> on probabilities to tell whether a customer has churned out or not, I got the following results -</p>\n\n<p><strong>Prediction Accuracies</strong></p>\n\n<p><strong>84.28 %</strong> - For all customers (who were not churned out on 1st august)</p>\n\n<p><strong>25.72 %</strong> - For customers who actually churned out between 1st to 10th August</p>\n\n<p>Then I checked why I am getting so low accuracy on churned-out customers, I found out that actually there were-</p>\n\n<p>27000 - number of customers who were not churned out as on 1st August</p>\n\n<p>312 - number of members who churned out of the 27000 members above between 1st to 10th August</p>\n\n<p>So definitely there is class imbalance problem. So how do I increase my prediction accuracy on customers who actually churn out?</p>\n\n<p>And if possible how can I use undersampling,oversampling, SMOTE, etc with survival analysis for my problem? <a href=\"https://www.analyticsvidhya.com/blog/2017/03/imbalanced-classification-problem/\">https://www.analyticsvidhya.com/blog/2017/03/imbalanced-classification-problem/</a></p>",
      "rawMarkdown": "*Disclaimer : This is not related directly to this competition. I am working on churning-out model and that's why I need help from you guys.*\n\nI am working on e-commerce customers data.\n\n\n----------\n\n\nInitially I just used **classification** algorithms to calculate churn-out. But it **can** predict only whether a customer is going to churn out or not, it **cannot** predict **when** the customer is going to churn out or the probability of surviving on a particular future date.\n\n\n----------\n\n\nThen I used **survival analysis** to predict churn-out. I made a **cox proportionality hazard model** using `coxph` function from `survival` package in R **(I used all data till 1st August**). With that model, I used `predictSurvProb` function from `pec` package in R to **calculate probability** of churning out of all non-churned customers (as on 1st August) on **10th August**.\n\nAnd by using **threshold of 0.5** on probabilities to tell whether a customer has churned out or not, I got the following results -\n\n**Prediction Accuracies**\n\n**84.28 %** - For all customers (who were not churned out on 1st august)\n\n**25.72 %** - For customers who actually churned out between 1st to 10th August\n\nThen I checked why I am getting so low accuracy on churned-out customers, I found out that actually there were-\n\n27000 - number of customers who were not churned out as on 1st August\n\n312 - number of members who churned out of the 27000 members above between 1st to 10th August\n\nSo definitely there is class imbalance problem. So how do I increase my prediction accuracy on customers who actually churn out?\n\nAnd if possible how can I use undersampling,oversampling, SMOTE, etc with survival analysis for my problem? https://www.analyticsvidhya.com/blog/2017/03/imbalanced-classification-problem/",
      "votes": null
    },
    {
      "id": "224742",
      "postDate": "09/27/2017 14:21:03",
      "content": "<p>Hi Manish -</p>\n\n<p>Have you tried a classification algorithm that uses the \"time in\" as a feature? Then, for a given customer, you can vary that feature and see how it responds with increasing time. You should be able to use that to say, e.g., that they user has a certain probability of churn tomorrow, a week from now, etc.</p>",
      "rawMarkdown": "Hi Manish -\n\nHave you tried a classification algorithm that uses the \"time in\" as a feature? Then, for a given customer, you can vary that feature and see how it responds with increasing time. You should be able to use that to say, e.g., that they user has a certain probability of churn tomorrow, a week from now, etc.",
      "votes": null
    },
    {
      "id": "224818",
      "postDate": "09/27/2017 18:20:05",
      "content": "<p>Can you please explain what is that \"time in\" feature you are referring to. And thanks for replying.</p>",
      "rawMarkdown": "Can you please explain what is that \"time in\" feature you are referring to. And thanks for replying.",
      "votes": null
    },
    {
      "id": "224820",
      "postDate": "09/27/2017 18:23:05",
      "content": "<p>Are you referring \"time in\" as the datetime when the customer registers?</p>",
      "rawMarkdown": "Are you referring \"time in\" as the datetime when the customer registers?",
      "votes": null
    },
    {
      "id": "224862",
      "postDate": "09/27/2017 19:54:57",
      "content": "<p>I mean either (a) the time since their registration if they haven't churned, or (b) the amount of time they stayed as a customer if they did churn.</p>",
      "rawMarkdown": "I mean either (a) the time since their registration if they haven't churned, or (b) the amount of time they stayed as a customer if they did churn.",
      "votes": null
    },
    {
      "id": "225007",
      "postDate": "09/28/2017 03:36:44",
      "content": "<p>But there are lot many other features also like total money spent, number of failed transactions, etc. So I can't just increase the \"time in\" to check how the probability of (binary) \"churned_out\" is changing. The other features are also very important and I can't just take their original values while trying to predict for say 60 days later.</p>\n\n<p>So what do u suggest on taking into consideration other important vars also?</p>",
      "rawMarkdown": "But there are lot many other features also like total money spent, number of failed transactions, etc. So I can't just increase the \"time in\" to check how the probability of (binary) \"churned_out\" is changing. The other features are also very important and I can't just take their original values while trying to predict for say 60 days later.\n\nSo what do u suggest on taking into consideration other important vars also?",
      "votes": null
    },
    {
      "id": "225231",
      "postDate": "09/28/2017 15:05:46",
      "content": "<p>What I've done in the past for problems like this is <code>y = churn?</code>, and <code>X = time_in, other_vars, ...</code>. Then I use a Random Forest or XGBoost model to predict the binary outcome of <code>y</code>. Then given a single observation with  fixed <code>other_vars, ...</code>, I vary the <code>time_in</code> and see how <code>y</code> responds for the model, and see how long it will take before that observation has a 50% chance of churning.</p>\n\n<p>If you look at feature importance, it will show you that <code>time_in</code> is a very important predictor, and it will also show you what other variables are important in predicting churn.</p>",
      "rawMarkdown": "What I've done in the past for problems like this is `y = churn?`, and `X = time_in, other_vars, ...`. Then I use a Random Forest or XGBoost model to predict the binary outcome of `y`. Then given a single observation with  fixed `other_vars, ...`, I vary the `time_in` and see how `y` responds for the model, and see how long it will take before that observation has a 50% chance of churning.\n\nIf you look at feature importance, it will show you that `time_in` is a very important predictor, and it will also show you what other variables are important in predicting churn.",
      "votes": null
    },
    {
      "id": "225242",
      "postDate": "09/28/2017 15:15:41",
      "content": "<p>Thank you. I'll build the model you are suggesting and I hope that gives me a good prediction accuracy.</p>",
      "rawMarkdown": "Thank you. I'll build the model you are suggesting and I hope that gives me a good prediction accuracy.",
      "votes": null
    },
    {
      "id": "225475",
      "postDate": "09/29/2017 06:47:03",
      "content": "<p>So I built that <strong>classification</strong> model by training on a dataset with almost 50% customers churned out and 50% not churned out. And as you said, <code>time in</code> was actually the most important feature(way more than even the second most important).</p>\n\n<p>Later for <strong>validation</strong>, I took the data of all my \"not churned-out\" customers as on <strong>1st August 2017</strong>, and tried to predict their \"is_churned\" class on <strong>1st September 2017</strong> (I have the data so I can validate my model). I did that by adding <strong>31</strong> to the \"time in\" feature on the 1st August data and kept all other variables constant.</p>\n\n<p>And I in overall I got <strong>26763/27364</strong> correct prediction. But when I checked the prediction for the customers who actually churned out in that period, I found that I got <strong>0% accuracy</strong> on prediction for them. </p>\n\n<p>So I think the problem is there is class imbalance. In that period out of <strong>27364</strong>, only <strong>587</strong> churned-out. And I was not able to find these <strong>587</strong> churned-out members at all(0% accuracy). </p>\n\n<p>So what do you suggest after that?</p>",
      "rawMarkdown": "So I built that **classification** model by training on a dataset with almost 50% customers churned out and 50% not churned out. And as you said, `time in` was actually the most important feature(way more than even the second most important).\n\nLater for **validation**, I took the data of all my \"not churned-out\" customers as on **1st August 2017**, and tried to predict their \"is_churned\" class on **1st September 2017** (I have the data so I can validate my model). I did that by adding **31** to the \"time in\" feature on the 1st August data and kept all other variables constant.\n\nAnd I in overall I got **26763/27364** correct prediction. But when I checked the prediction for the customers who actually churned out in that period, I found that I got **0% accuracy** on prediction for them. \n\nSo I think the problem is there is class imbalance. In that period out of **27364**, only **587** churned-out. And I was not able to find these **587** churned-out members at all(0% accuracy). \n\nSo what do you suggest after that?",
      "votes": null
    },
    {
      "id": "228718",
      "postDate": "10/07/2017 16:50:18",
      "content": "<p>I think, you can try to find/add more samples for chuned class from past. Ex. churned at August, July, June... etc. </p>",
      "rawMarkdown": "I think, you can try to find/add more samples for chuned class from past. Ex. churned at August, July, June... etc.",
      "votes": null
    },
    {
      "id": "240510",
      "postDate": "11/06/2017 19:24:30",
      "content": "<p>Hi inversion,</p>\n\n<p>Do you mean like calculating the customer tenure?</p>",
      "rawMarkdown": "Hi inversion,\n\nDo you mean like calculating the customer tenure?",
      "votes": null
    },
    {
      "id": "240710",
      "postDate": "11/07/2017 08:03:17",
      "content": "<p>Yes, churning out is the point when the customer cancels his subscription i.e. the end of his lifetime/tenure. But I must remind you that the data is non-contractual in the sense that there is no expiration date of subscription(like netflix). However the customer has to call to the customer care to cancel his subscription.</p>",
      "rawMarkdown": "Yes, churning out is the point when the customer cancels his subscription i.e. the end of his lifetime/tenure. But I must remind you that the data is non-contractual in the sense that there is no expiration date of subscription(like netflix). However the customer has to call to the customer care to cancel his subscription.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 224742,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "09/27/2017 14:21:03",
      "content": "<p>Hi Manish -</p>\n\n<p>Have you tried a classification algorithm that uses the \"time in\" as a feature? Then, for a given customer, you can vary that feature and see how it responds with increasing time. You should be able to use that to say, e.g., that they user has a certain probability of churn tomorrow, a week from now, etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 224818,
          "author_name": "iammangod96",
          "author_url": "",
          "post_date": "09/27/2017 18:20:05",
          "content": "<p>Can you please explain what is that \"time in\" feature you are referring to. And thanks for replying.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224820,
          "author_name": "iammangod96",
          "author_url": "",
          "post_date": "09/27/2017 18:23:05",
          "content": "<p>Are you referring \"time in\" as the datetime when the customer registers?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224862,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "09/27/2017 19:54:57",
          "content": "<p>I mean either (a) the time since their registration if they haven't churned, or (b) the amount of time they stayed as a customer if they did churn.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225007,
          "author_name": "iammangod96",
          "author_url": "",
          "post_date": "09/28/2017 03:36:44",
          "content": "<p>But there are lot many other features also like total money spent, number of failed transactions, etc. So I can't just increase the \"time in\" to check how the probability of (binary) \"churned_out\" is changing. The other features are also very important and I can't just take their original values while trying to predict for say 60 days later.</p>\n\n<p>So what do u suggest on taking into consideration other important vars also?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225231,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "09/28/2017 15:05:46",
          "content": "<p>What I've done in the past for problems like this is <code>y = churn?</code>, and <code>X = time_in, other_vars, ...</code>. Then I use a Random Forest or XGBoost model to predict the binary outcome of <code>y</code>. Then given a single observation with  fixed <code>other_vars, ...</code>, I vary the <code>time_in</code> and see how <code>y</code> responds for the model, and see how long it will take before that observation has a 50% chance of churning.</p>\n\n<p>If you look at feature importance, it will show you that <code>time_in</code> is a very important predictor, and it will also show you what other variables are important in predicting churn.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225242,
          "author_name": "iammangod96",
          "author_url": "",
          "post_date": "09/28/2017 15:15:41",
          "content": "<p>Thank you. I'll build the model you are suggesting and I hope that gives me a good prediction accuracy.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225475,
          "author_name": "iammangod96",
          "author_url": "",
          "post_date": "09/29/2017 06:47:03",
          "content": "<p>So I built that <strong>classification</strong> model by training on a dataset with almost 50% customers churned out and 50% not churned out. And as you said, <code>time in</code> was actually the most important feature(way more than even the second most important).</p>\n\n<p>Later for <strong>validation</strong>, I took the data of all my \"not churned-out\" customers as on <strong>1st August 2017</strong>, and tried to predict their \"is_churned\" class on <strong>1st September 2017</strong> (I have the data so I can validate my model). I did that by adding <strong>31</strong> to the \"time in\" feature on the 1st August data and kept all other variables constant.</p>\n\n<p>And I in overall I got <strong>26763/27364</strong> correct prediction. But when I checked the prediction for the customers who actually churned out in that period, I found that I got <strong>0% accuracy</strong> on prediction for them. </p>\n\n<p>So I think the problem is there is class imbalance. In that period out of <strong>27364</strong>, only <strong>587</strong> churned-out. And I was not able to find these <strong>587</strong> churned-out members at all(0% accuracy). </p>\n\n<p>So what do you suggest after that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228718,
          "author_name": "armamut",
          "author_url": "",
          "post_date": "10/07/2017 16:50:18",
          "content": "<p>I think, you can try to find/add more samples for chuned class from past. Ex. churned at August, July, June... etc. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 240510,
          "author_name": "thunderdome",
          "author_url": "",
          "post_date": "11/06/2017 19:24:30",
          "content": "<p>Hi inversion,</p>\n\n<p>Do you mean like calculating the customer tenure?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 240710,
          "author_name": "iammangod96",
          "author_url": "",
          "post_date": "11/07/2017 08:03:17",
          "content": "<p>Yes, churning out is the point when the customer cancels his subscription i.e. the end of his lifetime/tenure. But I must remind you that the data is non-contractual in the sense that there is no expiration date of subscription(like netflix). However the customer has to call to the customer care to cancel his subscription.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "224670": "*Disclaimer : This is not related directly to this competition. I am working on churning-out model and that's why I need help from you guys.*\n\nI am working on e-commerce customers data.\n\n\n----------\n\n\nInitially I just used **classification** algorithms to calculate churn-out. But it **can** predict only whether a customer is going to churn out or not, it **cannot** predict **when** the customer is going to churn out or the probability of surviving on a particular future date.\n\n\n----------\n\n\nThen I used **survival analysis** to predict churn-out. I made a **cox proportionality hazard model** using `coxph` function from `survival` package in R **(I used all data till 1st August**). With that model, I used `predictSurvProb` function from `pec` package in R to **calculate probability** of churning out of all non-churned customers (as on 1st August) on **10th August**.\n\nAnd by using **threshold of 0.5** on probabilities to tell whether a customer has churned out or not, I got the following results -\n\n**Prediction Accuracies**\n\n**84.28 %** - For all customers (who were not churned out on 1st august)\n\n**25.72 %** - For customers who actually churned out between 1st to 10th August\n\nThen I checked why I am getting so low accuracy on churned-out customers, I found out that actually there were-\n\n27000 - number of customers who were not churned out as on 1st August\n\n312 - number of members who churned out of the 27000 members above between 1st to 10th August\n\nSo definitely there is class imbalance problem. So how do I increase my prediction accuracy on customers who actually churn out?\n\nAnd if possible how can I use undersampling,oversampling, SMOTE, etc with survival analysis for my problem? https://www.analyticsvidhya.com/blog/2017/03/imbalanced-classification-problem/",
    "224742": "Hi Manish -\n\nHave you tried a classification algorithm that uses the \"time in\" as a feature? Then, for a given customer, you can vary that feature and see how it responds with increasing time. You should be able to use that to say, e.g., that they user has a certain probability of churn tomorrow, a week from now, etc.",
    "224818": "Can you please explain what is that \"time in\" feature you are referring to. And thanks for replying.",
    "224820": "Are you referring \"time in\" as the datetime when the customer registers?",
    "224862": "I mean either (a) the time since their registration if they haven't churned, or (b) the amount of time they stayed as a customer if they did churn.",
    "225007": "But there are lot many other features also like total money spent, number of failed transactions, etc. So I can't just increase the \"time in\" to check how the probability of (binary) \"churned_out\" is changing. The other features are also very important and I can't just take their original values while trying to predict for say 60 days later.\n\nSo what do u suggest on taking into consideration other important vars also?",
    "225231": "What I've done in the past for problems like this is `y = churn?`, and `X = time_in, other_vars, ...`. Then I use a Random Forest or XGBoost model to predict the binary outcome of `y`. Then given a single observation with  fixed `other_vars, ...`, I vary the `time_in` and see how `y` responds for the model, and see how long it will take before that observation has a 50% chance of churning.\n\nIf you look at feature importance, it will show you that `time_in` is a very important predictor, and it will also show you what other variables are important in predicting churn.",
    "225242": "Thank you. I'll build the model you are suggesting and I hope that gives me a good prediction accuracy.",
    "225475": "So I built that **classification** model by training on a dataset with almost 50% customers churned out and 50% not churned out. And as you said, `time in` was actually the most important feature(way more than even the second most important).\n\nLater for **validation**, I took the data of all my \"not churned-out\" customers as on **1st August 2017**, and tried to predict their \"is_churned\" class on **1st September 2017** (I have the data so I can validate my model). I did that by adding **31** to the \"time in\" feature on the 1st August data and kept all other variables constant.\n\nAnd I in overall I got **26763/27364** correct prediction. But when I checked the prediction for the customers who actually churned out in that period, I found that I got **0% accuracy** on prediction for them. \n\nSo I think the problem is there is class imbalance. In that period out of **27364**, only **587** churned-out. And I was not able to find these **587** churned-out members at all(0% accuracy). \n\nSo what do you suggest after that?",
    "228718": "I think, you can try to find/add more samples for chuned class from past. Ex. churned at August, July, June... etc.",
    "240510": "Hi inversion,\n\nDo you mean like calculating the customer tenure?",
    "240710": "Yes, churning out is the point when the customer cancels his subscription i.e. the end of his lifetime/tenure. But I must remind you that the data is non-contractual in the sense that there is no expiration date of subscription(like netflix). However the customer has to call to the customer care to cancel his subscription."
  },
  "source": "meta"
}