{
  "id": 43764,
  "title": "User Only Has Is_Cancel = 1 transaction...",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/43764",
  "author_name": "",
  "post_date": "2017-11-19T07:19:10.632895800Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p><strong>1</strong> For example user with msno '+/w1UrZwyka4C9oNH3+Q8fUf3fD8R3EwWrx57ODIsqk=' has 1 transaction and it's a subscription cancellation. How might this would happen ? </p>\n\n<p><strong>2</strong>  4454 members in the new test data (sample_submission_v2,csv) have expiration date which is greater than the scope we are trying to predict, in fact we should be predicting for customers that have their latest membership expiration data between 2017-04-01 and 2017-04-31. Here is an example member:</p>\n\n<pre><code>transactions[transactions.msno == '+JDM2aOo/iLtUMHtJxDz7+pJ4azJV/UQH85tZPMJzrA='].sort_values('membership_expire_date')\n</code></pre>\n\n<p>You will see that this member had a transaction on 2015-12-15 which expires on 2018-04-09, we shouldn't be predicting for this user. If the problem was to predict whether a user will cancel and/or churn then these cases would be fine. But churn definition is strictly made as \"Will a user have a transaction within 30 days given their membership expires in April\".</p>\n\n<p><strong>Note</strong>: Even though 4454 / 907471 is not a huge proportion. you can simply directly label those users as 0 :)</p>",
  "messages": [
    {
      "id": "245669",
      "postDate": "11/19/2017 07:19:10",
      "content": "<p><strong>1</strong> For example user with msno '+/w1UrZwyka4C9oNH3+Q8fUf3fD8R3EwWrx57ODIsqk=' has 1 transaction and it's a subscription cancellation. How might this would happen ? </p>\n\n<p><strong>2</strong>  4454 members in the new test data (sample_submission_v2,csv) have expiration date which is greater than the scope we are trying to predict, in fact we should be predicting for customers that have their latest membership expiration data between 2017-04-01 and 2017-04-31. Here is an example member:</p>\n\n<pre><code>transactions[transactions.msno == '+JDM2aOo/iLtUMHtJxDz7+pJ4azJV/UQH85tZPMJzrA='].sort_values('membership_expire_date')\n</code></pre>\n\n<p>You will see that this member had a transaction on 2015-12-15 which expires on 2018-04-09, we shouldn't be predicting for this user. If the problem was to predict whether a user will cancel and/or churn then these cases would be fine. But churn definition is strictly made as \"Will a user have a transaction within 30 days given their membership expires in April\".</p>\n\n<p><strong>Note</strong>: Even though 4454 / 907471 is not a huge proportion. you can simply directly label those users as 0 :)</p>",
      "rawMarkdown": "**1** For example user with msno '+/w1UrZwyka4C9oNH3+Q8fUf3fD8R3EwWrx57ODIsqk=' has 1 transaction and it's a subscription cancellation. How might this would happen ? \n\n**2**  4454 members in the new test data (sample_submission_v2,csv) have expiration date which is greater than the scope we are trying to predict, in fact we should be predicting for customers that have their latest membership expiration data between 2017-04-01 and 2017-04-31. Here is an example member:\n\n    transactions[transactions.msno == '+JDM2aOo/iLtUMHtJxDz7+pJ4azJV/UQH85tZPMJzrA='].sort_values('membership_expire_date')\n\nYou will see that this member had a transaction on 2015-12-15 which expires on 2018-04-09, we shouldn't be predicting for this user. If the problem was to predict whether a user will cancel and/or churn then these cases would be fine. But churn definition is strictly made as \"Will a user have a transaction within 30 days given their membership expires in April\".\n\n**Note**: Even though 4454 / 907471 is not a huge proportion. you can simply directly label those users as 0 :)",
      "votes": null
    },
    {
      "id": "246342",
      "postDate": "11/21/2017 01:12:51",
      "content": "<p>would like to see answers to this</p>",
      "rawMarkdown": "would like to see answers to this",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 246342,
      "author_name": "davischumacher",
      "author_url": "",
      "post_date": "11/21/2017 01:12:51",
      "content": "<p>would like to see answers to this</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "245669": "**1** For example user with msno '+/w1UrZwyka4C9oNH3+Q8fUf3fD8R3EwWrx57ODIsqk=' has 1 transaction and it's a subscription cancellation. How might this would happen ? \n\n**2**  4454 members in the new test data (sample_submission_v2,csv) have expiration date which is greater than the scope we are trying to predict, in fact we should be predicting for customers that have their latest membership expiration data between 2017-04-01 and 2017-04-31. Here is an example member:\n\n    transactions[transactions.msno == '+JDM2aOo/iLtUMHtJxDz7+pJ4azJV/UQH85tZPMJzrA='].sort_values('membership_expire_date')\n\nYou will see that this member had a transaction on 2015-12-15 which expires on 2018-04-09, we shouldn't be predicting for this user. If the problem was to predict whether a user will cancel and/or churn then these cases would be fine. But churn definition is strictly made as \"Will a user have a transaction within 30 days given their membership expires in April\".\n\n**Note**: Even though 4454 / 907471 is not a huge proportion. you can simply directly label those users as 0 :)",
    "246342": "would like to see answers to this"
  },
  "source": "meta"
}