{
  "id": 39803,
  "title": "expiration date confusion",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/39803",
  "author_name": "",
  "post_date": "2017-09-21T08:59:53.143698400Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": "223155",
      "postDate": "09/21/2017 08:59:53",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "223865",
      "postDate": "09/23/2017 23:14:22",
      "content": "<p><code>\nmsno is_churn   city    bd  gender  registered_via  registration_init_time  expiration_date <br>\nwaLDQMmcOu2jLDaV1ddDkgCrB/jl6sD66Xzs0Vqax1Y=    1   18.0    36.0    female  9.0 20050406.0  20170907.0\nQA7uiXy8vIbUSPOkCf9RwQ3FsT8jVq2OxDr8zqa7bRQ=    1   10.0    38.0    male    9.0 20050407.0  20170321.0\nfGwBva6hikQmTJzrbz/2Ezjm5Cth5jZUNvXigKK2AFA=    1   11.0    27.0    female  9.0 20051016.0  20170203.0\nmT5V8rEpa+8wuqi6x0DoVd3H5icMKkE9Prt49UlmK+4=    1   13.0    23.0    female  9.0 20051102.0  20170926.0\nXaPhtGLk/5UvvOYHcONTwsnH97P4eGECeq+BARGItRw=    1   3.0 27.0    male    9.0 20051228.0  20170927.0\nGBy8qSz16X5iYWD+3CMxv/Hm6OPSrXBYtmbnlRtknW0=    1   6.0 23.0    female  9.0 20060331.0  20170215.0\nlYLh7TdkWpIoQs3i3o6mIjLH8/IEgMWP9r7OpsLX0Vo=    1   13.0    29.0    female  9.0 20060406.0  20170208.0\nT0FF6lumjKcqEO0O+tUH2ytc+Kb9EkeaLzcVUiTr1aE=    1   11.0    22.0    male    9.0 20060425.0  20170906.0\nNb1ZGEmagQeba5E+nQj8VlQoWl+8SFmLZu+Y8ytIamw=    1   18.0    22.0    female  9.0 20060826.0  20170908.0\nMkuWz0Nq6/Oq5fKqRddWL7oh2SLUSRe3/g+XmAWqW1Q=    1   11.0    30.0    female  9.0 20061123.0  20170324.0\n</code></p>\n\n<p>It is easy to see that there are a plenty (actually most of them) of records in <strong>members.csv joined with train.csv</strong> contains dates like 201709.. 201710.. even for churners. How could that possibly be? According to train dataset description there must be users who churned during March 2017.</p>",
      "rawMarkdown": "```\t\nmsno is_churn\tcity\tbd\tgender\tregistered_via\tregistration_init_time\texpiration_date\t\t\t\t\t\t\t\nwaLDQMmcOu2jLDaV1ddDkgCrB/jl6sD66Xzs0Vqax1Y=\t1\t18.0\t36.0\tfemale\t9.0\t20050406.0\t20170907.0\nQA7uiXy8vIbUSPOkCf9RwQ3FsT8jVq2OxDr8zqa7bRQ=\t1\t10.0\t38.0\tmale\t9.0\t20050407.0\t20170321.0\nfGwBva6hikQmTJzrbz/2Ezjm5Cth5jZUNvXigKK2AFA=\t1\t11.0\t27.0\tfemale\t9.0\t20051016.0\t20170203.0\nmT5V8rEpa+8wuqi6x0DoVd3H5icMKkE9Prt49UlmK+4=\t1\t13.0\t23.0\tfemale\t9.0\t20051102.0\t20170926.0\nXaPhtGLk/5UvvOYHcONTwsnH97P4eGECeq+BARGItRw=\t1\t3.0\t27.0\tmale\t9.0\t20051228.0\t20170927.0\nGBy8qSz16X5iYWD+3CMxv/Hm6OPSrXBYtmbnlRtknW0=\t1\t6.0\t23.0\tfemale\t9.0\t20060331.0\t20170215.0\nlYLh7TdkWpIoQs3i3o6mIjLH8/IEgMWP9r7OpsLX0Vo=\t1\t13.0\t29.0\tfemale\t9.0\t20060406.0\t20170208.0\nT0FF6lumjKcqEO0O+tUH2ytc+Kb9EkeaLzcVUiTr1aE=\t1\t11.0\t22.0\tmale\t9.0\t20060425.0\t20170906.0\nNb1ZGEmagQeba5E+nQj8VlQoWl+8SFmLZu+Y8ytIamw=\t1\t18.0\t22.0\tfemale\t9.0\t20060826.0\t20170908.0\nMkuWz0Nq6/Oq5fKqRddWL7oh2SLUSRe3/g+XmAWqW1Q=\t1\t11.0\t30.0\tfemale\t9.0\t20061123.0\t20170324.0\n```\n\nIt is easy to see that there are a plenty (actually most of them) of records in **members.csv joined with train.csv** contains dates like 201709.. 201710.. even for churners. How could that possibly be? According to train dataset description there must be users who churned during March 2017.",
      "votes": null
    },
    {
      "id": "224079",
      "postDate": "09/25/2017 03:29:45",
      "content": "<p>member.csv is a snapshot of our membership table. It is true that each entry can contains dates in the future. Say if we subscribed a two-year plan from 2017-02-15 to 2019-02-15, then in our member.csv snapshot, say taking on 2017-2-28, you will see that user's expiration date as 2019-02-15. However, if that person made a plan change on 2017-02-28 to cancel the plan on 2017-03-10 and made another subscription on 2017-04-12. We count this user churned because this guy did not make a renewal within 30 days after 2017-03-10. The detail information you find in this segment will not be reflected in member.csv</p>",
      "rawMarkdown": "member.csv is a snapshot of our membership table. It is true that each entry can contains dates in the future. Say if we subscribed a two-year plan from 2017-02-15 to 2019-02-15, then in our member.csv snapshot, say taking on 2017-2-28, you will see that user's expiration date as 2019-02-15. However, if that person made a plan change on 2017-02-28 to cancel the plan on 2017-03-10 and made another subscription on 2017-04-12. We count this user churned because this guy did not make a renewal within 30 days after 2017-03-10. The detail information you find in this segment will not be reflected in member.csv",
      "votes": null
    },
    {
      "id": "224118",
      "postDate": "09/25/2017 07:02:54",
      "content": "<p>the member data contains the newest info of the user, we should use expiration date in transaction dataset, not the member expiration in member dataset, that's my understanding for this problem</p>",
      "rawMarkdown": "the member data contains the newest info of the user, we should use expiration date in transaction dataset, not the member expiration in member dataset, that's my understanding for this problem",
      "votes": null
    },
    {
      "id": "224703",
      "postDate": "09/27/2017 10:59:21",
      "content": "<p>Interesting - that's the opposite of what the data definition claimed/suggested. \nThanks!</p>",
      "rawMarkdown": "Interesting - that's the opposite of what the data definition claimed/suggested. \nThanks!",
      "votes": null
    },
    {
      "id": "229001",
      "postDate": "10/08/2017 14:25:51",
      "content": "<p>So I want to know, the date of snapshot of member.csv is before 2017/02/28?  </p>",
      "rawMarkdown": "So I want to know, the date of snapshot of member.csv is before 2017/02/28?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 223865,
      "author_name": "artlitz",
      "author_url": "",
      "post_date": "09/23/2017 23:14:22",
      "content": "<p><code>\nmsno is_churn   city    bd  gender  registered_via  registration_init_time  expiration_date <br>\nwaLDQMmcOu2jLDaV1ddDkgCrB/jl6sD66Xzs0Vqax1Y=    1   18.0    36.0    female  9.0 20050406.0  20170907.0\nQA7uiXy8vIbUSPOkCf9RwQ3FsT8jVq2OxDr8zqa7bRQ=    1   10.0    38.0    male    9.0 20050407.0  20170321.0\nfGwBva6hikQmTJzrbz/2Ezjm5Cth5jZUNvXigKK2AFA=    1   11.0    27.0    female  9.0 20051016.0  20170203.0\nmT5V8rEpa+8wuqi6x0DoVd3H5icMKkE9Prt49UlmK+4=    1   13.0    23.0    female  9.0 20051102.0  20170926.0\nXaPhtGLk/5UvvOYHcONTwsnH97P4eGECeq+BARGItRw=    1   3.0 27.0    male    9.0 20051228.0  20170927.0\nGBy8qSz16X5iYWD+3CMxv/Hm6OPSrXBYtmbnlRtknW0=    1   6.0 23.0    female  9.0 20060331.0  20170215.0\nlYLh7TdkWpIoQs3i3o6mIjLH8/IEgMWP9r7OpsLX0Vo=    1   13.0    29.0    female  9.0 20060406.0  20170208.0\nT0FF6lumjKcqEO0O+tUH2ytc+Kb9EkeaLzcVUiTr1aE=    1   11.0    22.0    male    9.0 20060425.0  20170906.0\nNb1ZGEmagQeba5E+nQj8VlQoWl+8SFmLZu+Y8ytIamw=    1   18.0    22.0    female  9.0 20060826.0  20170908.0\nMkuWz0Nq6/Oq5fKqRddWL7oh2SLUSRe3/g+XmAWqW1Q=    1   11.0    30.0    female  9.0 20061123.0  20170324.0\n</code></p>\n\n<p>It is easy to see that there are a plenty (actually most of them) of records in <strong>members.csv joined with train.csv</strong> contains dates like 201709.. 201710.. even for churners. How could that possibly be? According to train dataset description there must be users who churned during March 2017.</p>",
      "votes": null,
      "replies": [
        {
          "id": 224079,
          "author_name": "ardenkkbox",
          "author_url": "",
          "post_date": "09/25/2017 03:29:45",
          "content": "<p>member.csv is a snapshot of our membership table. It is true that each entry can contains dates in the future. Say if we subscribed a two-year plan from 2017-02-15 to 2019-02-15, then in our member.csv snapshot, say taking on 2017-2-28, you will see that user's expiration date as 2019-02-15. However, if that person made a plan change on 2017-02-28 to cancel the plan on 2017-03-10 and made another subscription on 2017-04-12. We count this user churned because this guy did not make a renewal within 30 days after 2017-03-10. The detail information you find in this segment will not be reflected in member.csv</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224118,
          "author_name": "franciszmy",
          "author_url": "",
          "post_date": "09/25/2017 07:02:54",
          "content": "<p>the member data contains the newest info of the user, we should use expiration date in transaction dataset, not the member expiration in member dataset, that's my understanding for this problem</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224703,
          "author_name": "danofer",
          "author_url": "",
          "post_date": "09/27/2017 10:59:21",
          "content": "<p>Interesting - that's the opposite of what the data definition claimed/suggested. \nThanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 229001,
          "author_name": "yuanxiaobin",
          "author_url": "",
          "post_date": "10/08/2017 14:25:51",
          "content": "<p>So I want to know, the date of snapshot of member.csv is before 2017/02/28?  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "223155": "",
    "223865": "```\t\nmsno is_churn\tcity\tbd\tgender\tregistered_via\tregistration_init_time\texpiration_date\t\t\t\t\t\t\t\nwaLDQMmcOu2jLDaV1ddDkgCrB/jl6sD66Xzs0Vqax1Y=\t1\t18.0\t36.0\tfemale\t9.0\t20050406.0\t20170907.0\nQA7uiXy8vIbUSPOkCf9RwQ3FsT8jVq2OxDr8zqa7bRQ=\t1\t10.0\t38.0\tmale\t9.0\t20050407.0\t20170321.0\nfGwBva6hikQmTJzrbz/2Ezjm5Cth5jZUNvXigKK2AFA=\t1\t11.0\t27.0\tfemale\t9.0\t20051016.0\t20170203.0\nmT5V8rEpa+8wuqi6x0DoVd3H5icMKkE9Prt49UlmK+4=\t1\t13.0\t23.0\tfemale\t9.0\t20051102.0\t20170926.0\nXaPhtGLk/5UvvOYHcONTwsnH97P4eGECeq+BARGItRw=\t1\t3.0\t27.0\tmale\t9.0\t20051228.0\t20170927.0\nGBy8qSz16X5iYWD+3CMxv/Hm6OPSrXBYtmbnlRtknW0=\t1\t6.0\t23.0\tfemale\t9.0\t20060331.0\t20170215.0\nlYLh7TdkWpIoQs3i3o6mIjLH8/IEgMWP9r7OpsLX0Vo=\t1\t13.0\t29.0\tfemale\t9.0\t20060406.0\t20170208.0\nT0FF6lumjKcqEO0O+tUH2ytc+Kb9EkeaLzcVUiTr1aE=\t1\t11.0\t22.0\tmale\t9.0\t20060425.0\t20170906.0\nNb1ZGEmagQeba5E+nQj8VlQoWl+8SFmLZu+Y8ytIamw=\t1\t18.0\t22.0\tfemale\t9.0\t20060826.0\t20170908.0\nMkuWz0Nq6/Oq5fKqRddWL7oh2SLUSRe3/g+XmAWqW1Q=\t1\t11.0\t30.0\tfemale\t9.0\t20061123.0\t20170324.0\n```\n\nIt is easy to see that there are a plenty (actually most of them) of records in **members.csv joined with train.csv** contains dates like 201709.. 201710.. even for churners. How could that possibly be? According to train dataset description there must be users who churned during March 2017.",
    "224079": "member.csv is a snapshot of our membership table. It is true that each entry can contains dates in the future. Say if we subscribed a two-year plan from 2017-02-15 to 2019-02-15, then in our member.csv snapshot, say taking on 2017-2-28, you will see that user's expiration date as 2019-02-15. However, if that person made a plan change on 2017-02-28 to cancel the plan on 2017-03-10 and made another subscription on 2017-04-12. We count this user churned because this guy did not make a renewal within 30 days after 2017-03-10. The detail information you find in this segment will not be reflected in member.csv",
    "224118": "the member data contains the newest info of the user, we should use expiration date in transaction dataset, not the member expiration in member dataset, that's my understanding for this problem",
    "224703": "Interesting - that's the opposite of what the data definition claimed/suggested. \nThanks!",
    "229001": "So I want to know, the date of snapshot of member.csv is before 2017/02/28?"
  },
  "source": "meta"
}