{
  "id": 40613,
  "title": "Approach on churn prediction",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/40613",
  "author_name": "",
  "post_date": "2017-10-05T10:02:51.353451200Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>it has been awesome that in kaggle we had such a practical case for data learning. But within this competition, i do have some confusion regarding the retention/churn prediction</p>\n\n<p>First one is, we will use the current live user who will mark as 'not churned'  and also the churned user (of course marked as 'churned') to train the data. Then we will use the model to predict on the lived user to see whether / how likely they will be churned in the future. </p>\n\n<p>The problem is, in the training data, we had already told the machine that live customer will not churn,  then i guess almost all the prediction on the live customers will be not churned as well? </p>\n\n<p>anyone tell me if my approach is wrong.</p>",
  "messages": [
    {
      "id": "227861",
      "postDate": "10/05/2017 10:02:51",
      "content": "<p>it has been awesome that in kaggle we had such a practical case for data learning. But within this competition, i do have some confusion regarding the retention/churn prediction</p>\n\n<p>First one is, we will use the current live user who will mark as 'not churned'  and also the churned user (of course marked as 'churned') to train the data. Then we will use the model to predict on the lived user to see whether / how likely they will be churned in the future. </p>\n\n<p>The problem is, in the training data, we had already told the machine that live customer will not churn,  then i guess almost all the prediction on the live customers will be not churned as well? </p>\n\n<p>anyone tell me if my approach is wrong.</p>",
      "rawMarkdown": "it has been awesome that in kaggle we had such a practical case for data learning. But within this competition, i do have some confusion regarding the retention/churn prediction\n\nFirst one is, we will use the current live user who will mark as 'not churned'  and also the churned user (of course marked as 'churned') to train the data. Then we will use the model to predict on the lived user to see whether / how likely they will be churned in the future. \n\nThe problem is, in the training data, we had already told the machine that live customer will not churn,  then i guess almost all the prediction on the live customers will be not churned as well? \n\nanyone tell me if my approach is wrong.",
      "votes": null
    },
    {
      "id": "230665",
      "postDate": "10/12/2017 12:32:07",
      "content": "<p>Depending on the algorithm you are using say logistic regression, the model will try to derive the relation between the IV and DV. You use this model to check the consistency of the training data, as well as the test data. So it won't predict the exact values for training data because it's an algorithm and you use this on several datasets of the same population.</p>",
      "rawMarkdown": "Depending on the algorithm you are using say logistic regression, the model will try to derive the relation between the IV and DV. You use this model to check the consistency of the training data, as well as the test data. So it won't predict the exact values for training data because it's an algorithm and you use this on several datasets of the same population.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 230665,
      "author_name": "adityajanardhan",
      "author_url": "",
      "post_date": "10/12/2017 12:32:07",
      "content": "<p>Depending on the algorithm you are using say logistic regression, the model will try to derive the relation between the IV and DV. You use this model to check the consistency of the training data, as well as the test data. So it won't predict the exact values for training data because it's an algorithm and you use this on several datasets of the same population.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "227861": "it has been awesome that in kaggle we had such a practical case for data learning. But within this competition, i do have some confusion regarding the retention/churn prediction\n\nFirst one is, we will use the current live user who will mark as 'not churned'  and also the churned user (of course marked as 'churned') to train the data. Then we will use the model to predict on the lived user to see whether / how likely they will be churned in the future. \n\nThe problem is, in the training data, we had already told the machine that live customer will not churn,  then i guess almost all the prediction on the live customers will be not churned as well? \n\nanyone tell me if my approach is wrong.",
    "230665": "Depending on the algorithm you are using say logistic regression, the model will try to derive the relation between the IV and DV. You use this model to check the consistency of the training data, as well as the test data. So it won't predict the exact values for training data because it's an algorithm and you use this on several datasets of the same population."
  },
  "source": "meta"
}