{
  "id": 43169,
  "title": "Announcement: New test data released",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/43169",
  "author_name": "",
  "post_date": "2017-11-10T17:55:31.779120300Z",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi all, </p>\n\n<p>Over the past week we have released the new test data. The leaderboard has been cleared. </p>\n\n<p>A summary of what is changed:</p>\n\n<ol>\n<li><p>Instead of predicting churn for March, we are now predicting churn for April. That goes into sample_submission_v2.csv</p></li>\n<li><p>Because of that, we released new (March) transactions (transactions_v2.csv) and user logs (user_logs_v2.csv) files. These DO NOT replace the old files since they only contain one month of data. </p></li>\n<li><p>There are some new users in this new test dataset that were previously unseen, so we included some extra historical transactions of those users - hence the dates before 3/1/2017 in transactions_v2.csv</p></li>\n<li><p>We re-generated the members information in members_v2.csv. This file DOES replace the previous version. We have removed the previous expiration date column since it caused controversy. </p></li>\n<li><p>Since there was leakage in the previous test.csv which contains the churn info for March, we decided to release the solution of those and let that become part of the training data. That becomes train_v2.csv. </p></li>\n<li><p>All the new files are now available in Kernels. </p></li>\n<li><p>Since the structure of the data has not changed, and we are only releasing more data, and that there is still more than one month until the end of the competition, we will keep the same timeline as before. </p></li>\n</ol>\n\n<p>Thank you for your understanding and patience as we resolve the issue. </p>\n\n<p>Kaggle/KKBox admins</p>",
  "messages": [
    {
      "id": "242135",
      "postDate": "11/10/2017 17:55:31",
      "content": "<p>Hi all, </p>\n\n<p>Over the past week we have released the new test data. The leaderboard has been cleared. </p>\n\n<p>A summary of what is changed:</p>\n\n<ol>\n<li><p>Instead of predicting churn for March, we are now predicting churn for April. That goes into sample_submission_v2.csv</p></li>\n<li><p>Because of that, we released new (March) transactions (transactions_v2.csv) and user logs (user_logs_v2.csv) files. These DO NOT replace the old files since they only contain one month of data. </p></li>\n<li><p>There are some new users in this new test dataset that were previously unseen, so we included some extra historical transactions of those users - hence the dates before 3/1/2017 in transactions_v2.csv</p></li>\n<li><p>We re-generated the members information in members_v2.csv. This file DOES replace the previous version. We have removed the previous expiration date column since it caused controversy. </p></li>\n<li><p>Since there was leakage in the previous test.csv which contains the churn info for March, we decided to release the solution of those and let that become part of the training data. That becomes train_v2.csv. </p></li>\n<li><p>All the new files are now available in Kernels. </p></li>\n<li><p>Since the structure of the data has not changed, and we are only releasing more data, and that there is still more than one month until the end of the competition, we will keep the same timeline as before. </p></li>\n</ol>\n\n<p>Thank you for your understanding and patience as we resolve the issue. </p>\n\n<p>Kaggle/KKBox admins</p>",
      "rawMarkdown": "Hi all, \n\nOver the past week we have released the new test data. The leaderboard has been cleared. \n\nA summary of what is changed:\n\n1. Instead of predicting churn for March, we are now predicting churn for April. That goes into sample_submission_v2.csv\n\n2. Because of that, we released new (March) transactions (transactions_v2.csv) and user logs (user_logs_v2.csv) files. These DO NOT replace the old files since they only contain one month of data. \n\n3. There are some new users in this new test dataset that were previously unseen, so we included some extra historical transactions of those users - hence the dates before 3/1/2017 in transactions_v2.csv\n\n4. We re-generated the members information in members_v2.csv. This file DOES replace the previous version. We have removed the previous expiration date column since it caused controversy. \n\n5. Since there was leakage in the previous test.csv which contains the churn info for March, we decided to release the solution of those and let that become part of the training data. That becomes train_v2.csv. \n\n6. All the new files are now available in Kernels. \n\n7. Since the structure of the data has not changed, and we are only releasing more data, and that there is still more than one month until the end of the competition, we will keep the same timeline as before. \n\nThank you for your understanding and patience as we resolve the issue. \n\nKaggle/KKBox admins",
      "votes": null
    },
    {
      "id": "242313",
      "postDate": "11/11/2017 07:17:23",
      "content": "<p>Thanks for clarification</p>",
      "rawMarkdown": "Thanks for clarification",
      "votes": null
    },
    {
      "id": "242551",
      "postDate": "11/11/2017 22:02:46",
      "content": "<p>Seems members_v2.csv contains only 795090 entries, while the original member.csv file contains 5116194 entries. Does the new version indeed replace the previous one?</p>",
      "rawMarkdown": "Seems members_v2.csv contains only 795090 entries, while the original member.csv file contains 5116194 entries. Does the new version indeed replace the previous one?",
      "votes": null
    },
    {
      "id": "242630",
      "postDate": "11/12/2017 07:28:15",
      "content": "<p>Based on #3, </p>\n\n<blockquote>\n  <p>There are some new users in this new test dataset that were previously unseen, ... , hence the dates before 3/1/2017 in transactions_v2.csv</p>\n</blockquote>\n\n<p>Should there also be some extra historical user logs (dates before 3/1/2017)?</p>",
      "rawMarkdown": "Based on #3, \n\n&gt; There are some new users in this new test dataset that were previously unseen, ... , hence the dates before 3/1/2017 in transactions_v2.csv\n\nShould there also be some extra historical user logs (dates before 3/1/2017)?",
      "votes": null
    },
    {
      "id": "242885",
      "postDate": "11/12/2017 22:17:17",
      "content": "<p>Thanks for clarifying things Wendy.</p>",
      "rawMarkdown": "Thanks for clarifying things Wendy.",
      "votes": null
    },
    {
      "id": "243080",
      "postDate": "11/13/2017 10:21:26",
      "content": "<p>I have extract the code segment to generate labels for users as \"WSDMChurnLabeller.scala\" in the data section. Hopefully this piece of code can give our participants more insights about how the churn label is applied. The training labels will not be limited to the ones that were provided.</p>",
      "rawMarkdown": "I have extract the code segment to generate labels for users as \"WSDMChurnLabeller.scala\" in the data section. Hopefully this piece of code can give our participants more insights about how the churn label is applied. The training labels will not be limited to the ones that were provided.",
      "votes": null
    },
    {
      "id": "243294",
      "postDate": "11/13/2017 18:33:41",
      "content": "<p>Thanks, this code is very helpful. I had this question before: in transactions file, transaction_date column is at day level granularity, for multiple transactions in one day, the order of them may affect the result of churn. e.g. renew first or cancel first. </p>\n\n<p>Also if I understand correctly, <code>val historyCutoff = \"20170131\"</code> is used to generate train.csv(first training set). Can we think in the way: the snapshot date of generating train.csv(Feb churn or not) is \"20170131\"? Does it mean the snapshot date of generating train_v2(Mar churn or not) is \"20170228\" and snapshot date of generating test(Apr churn or not) target is \"20170331\"? </p>",
      "rawMarkdown": "Thanks, this code is very helpful. I had this question before: in transactions file, transaction_date column is at day level granularity, for multiple transactions in one day, the order of them may affect the result of churn. e.g. renew first or cancel first. \n\nAlso if I understand correctly, ```val historyCutoff = \"20170131\"``` is used to generate train.csv(first training set). Can we think in the way: the snapshot date of generating train.csv(Feb churn or not) is \"20170131\"? Does it mean the snapshot date of generating train_v2(Mar churn or not) is \"20170228\" and snapshot date of generating test(Apr churn or not) target is \"20170331\"?",
      "votes": null
    },
    {
      "id": "243364",
      "postDate": "11/13/2017 22:03:47",
      "content": "<p>Sorry for the mismatch in my description and the data. We have uploaded <code>members_v3</code> to correct the issue. </p>",
      "rawMarkdown": "Sorry for the mismatch in my description and the data. We have uploaded `members_v3` to correct the issue.",
      "votes": null
    },
    {
      "id": "244184",
      "postDate": "11/15/2017 18:58:40",
      "content": "<p>I have run the code, as it is, over the original transaction file, and the results diverge from first train file. Has anyone noticed that?</p>",
      "rawMarkdown": "I have run the code, as it is, over the original transaction file, and the results diverge from first train file. Has anyone noticed that?",
      "votes": null
    },
    {
      "id": "244469",
      "postDate": "11/16/2017 11:22:08",
      "content": "<p>I can't reproduce it either. Not even after including the relevant transactions_v2, and modifying the starting date of the transaction history range to \"20150101\", as suggested in this discussion:\n<a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145\">https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145</a></p>",
      "rawMarkdown": "I can't reproduce it either. Not even after including the relevant transactions_v2, and modifying the starting date of the transaction history range to \"20150101\", as suggested in this discussion:\nhttps://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145",
      "votes": null
    },
    {
      "id": "245723",
      "postDate": "11/19/2017 13:28:34",
      "content": "<p>Hello. I have some troubles during analyzing the data because I don't clearly understand the meaning and all possible choices of these variables.</p>\n\n<p>Could you please explain and clarify more about them?</p>\n\n<p>From table transactions</p>\n\n<ol>\n<li><p>payment method --&gt; please explain the method name of each payment method id</p></li>\n<li><p>payment plan days --&gt; what does this variable mean?</p></li>\n<li><p>plan list price --&gt; what does this variable mean?</p></li>\n<li><p>actual amount paid  --&gt; what does this variable mean?</p></li>\n<li><p>transaction date  --&gt; what does this variable mean?</p></li>\n</ol>\n\n<p>From table members</p>\n\n<ol>\n<li><p>city --&gt; please explain city name of each city number</p></li>\n<li><p>registered via --&gt; please explain registration method name of number</p></li>\n<li><p>registration init time --&gt; what does this variable mean?</p></li>\n</ol>\n\n<p>I would appreciate it if you would answer my question.</p>",
      "rawMarkdown": "Hello. I have some troubles during analyzing the data because I don't clearly understand the meaning and all possible choices of these variables.\n\n\nCould you please explain and clarify more about them?\n\n\nFrom table transactions\n\n1. payment method --&gt; please explain the method name of each payment method id\n\n2. payment plan days --&gt; what does this variable mean?\n\n3. plan list price --&gt; what does this variable mean?\n\n4. actual amount paid  --&gt; what does this variable mean?\n\n5. transaction date  --&gt; what does this variable mean?\n\n\nFrom table members\n\n6. city --&gt; please explain city name of each city number\n\n7. registered via --&gt; please explain registration method name of number\n\n8. registration init time --&gt; what does this variable mean?\n\n\nI would appreciate it if you would answer my question.",
      "votes": null
    },
    {
      "id": "255910",
      "postDate": "12/10/2017 15:06:48",
      "content": "<p>Hi, Wendy</p>\n\n<p>According to <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363\">this post</a>, I am wondering if we can use member.csv (ver. 1), where can we download the file?</p>\n\n<p>For someone entered the competition lately, we are not able to download the original member.csv, while the original member.csv have extra information of expired date.</p>",
      "rawMarkdown": "Hi, Wendy\n\nAccording to [this post](https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363), I am wondering if we can use member.csv (ver. 1), where can we download the file?\n\nFor someone entered the competition lately, we are not able to download the original member.csv, while the original member.csv have extra information of expired date.",
      "votes": null
    },
    {
      "id": "255915",
      "postDate": "12/10/2017 15:12:18",
      "content": "<p>InfiniteWing,</p>\n\n<p>As only the latest versions of the competition files are public, those are the only versions of the competition data that can be used.</p>",
      "rawMarkdown": "InfiniteWing,\n\nAs only the latest versions of the competition files are public, those are the only versions of the competition data that can be used.",
      "votes": null
    },
    {
      "id": "255925",
      "postDate": "12/10/2017 15:21:35",
      "content": "<p>Hi Addison,</p>\n\n<p>Thanks for your reply, I missed that sentence.. sorry. I am curious how to detect people is using the newest version or not. </p>",
      "rawMarkdown": "Hi Addison,\n \nThanks for your reply, I missed that sentence.. sorry. I am curious how to detect people is using the newest version or not.",
      "votes": null
    },
    {
      "id": "258116",
      "postDate": "12/15/2017 14:23:34",
      "content": "<p>According to the data description for train data: <br>\n<strong>The criteria of \"churn\" is no new valid service subscription within 30 days after the current membership expires.</strong> <br>\nWith that definition, how can this user be \"churn\" in March? <br>\n<code>msno = K6fja4+jmoZ5xG6BypqX80Uw/XKpMgrEMdG2edFOxnA=</code> <br>\nThe user's last transaction in Feb was on 2017-02-16 with a membership expiration date of 2017-08-18;  So no matter what their activity is in March, they cannot be a \"churn\" customer in March according to the definition above.\nIs it better to just run a piece of code to generate your own (correct) version of churn status for the set of members in the train_v2 file?  If someone has already done that - could you kindly share that file?   Thanks in advance for answers /guidance.</p>",
      "rawMarkdown": "According to the data description for train data:  \n**The criteria of \"churn\" is no new valid service subscription within 30 days after the current membership expires.**  \nWith that definition, how can this user be \"churn\" in March?  \n```msno = K6fja4+jmoZ5xG6BypqX80Uw/XKpMgrEMdG2edFOxnA= ```  \nThe user's last transaction in Feb was on 2017-02-16 with a membership expiration date of 2017-08-18;  So no matter what their activity is in March, they cannot be a \"churn\" customer in March according to the definition above.\nIs it better to just run a piece of code to generate your own (correct) version of churn status for the set of members in the train_v2 file?  If someone has already done that - could you kindly share that file?   Thanks in advance for answers /guidance.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 242313,
      "author_name": "sundong",
      "author_url": "",
      "post_date": "11/11/2017 07:17:23",
      "content": "<p>Thanks for clarification</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 242551,
      "author_name": "kylexie",
      "author_url": "",
      "post_date": "11/11/2017 22:02:46",
      "content": "<p>Seems members_v2.csv contains only 795090 entries, while the original member.csv file contains 5116194 entries. Does the new version indeed replace the previous one?</p>",
      "votes": null,
      "replies": [
        {
          "id": 243364,
          "author_name": "wendykan",
          "author_url": "",
          "post_date": "11/13/2017 22:03:47",
          "content": "<p>Sorry for the mismatch in my description and the data. We have uploaded <code>members_v3</code> to correct the issue. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 242630,
      "author_name": "soundwaveli00",
      "author_url": "",
      "post_date": "11/12/2017 07:28:15",
      "content": "<p>Based on #3, </p>\n\n<blockquote>\n  <p>There are some new users in this new test dataset that were previously unseen, ... , hence the dates before 3/1/2017 in transactions_v2.csv</p>\n</blockquote>\n\n<p>Should there also be some extra historical user logs (dates before 3/1/2017)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 242885,
      "author_name": "khstangherlin",
      "author_url": "",
      "post_date": "11/12/2017 22:17:17",
      "content": "<p>Thanks for clarifying things Wendy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 243080,
      "author_name": "ardenkkbox",
      "author_url": "",
      "post_date": "11/13/2017 10:21:26",
      "content": "<p>I have extract the code segment to generate labels for users as \"WSDMChurnLabeller.scala\" in the data section. Hopefully this piece of code can give our participants more insights about how the churn label is applied. The training labels will not be limited to the ones that were provided.</p>",
      "votes": null,
      "replies": [
        {
          "id": 243294,
          "author_name": "soundwaveli00",
          "author_url": "",
          "post_date": "11/13/2017 18:33:41",
          "content": "<p>Thanks, this code is very helpful. I had this question before: in transactions file, transaction_date column is at day level granularity, for multiple transactions in one day, the order of them may affect the result of churn. e.g. renew first or cancel first. </p>\n\n<p>Also if I understand correctly, <code>val historyCutoff = \"20170131\"</code> is used to generate train.csv(first training set). Can we think in the way: the snapshot date of generating train.csv(Feb churn or not) is \"20170131\"? Does it mean the snapshot date of generating train_v2(Mar churn or not) is \"20170228\" and snapshot date of generating test(Apr churn or not) target is \"20170331\"? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 244184,
          "author_name": "aloisiodn",
          "author_url": "",
          "post_date": "11/15/2017 18:58:40",
          "content": "<p>I have run the code, as it is, over the original transaction file, and the results diverge from first train file. Has anyone noticed that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 244469,
          "author_name": "brucala",
          "author_url": "",
          "post_date": "11/16/2017 11:22:08",
          "content": "<p>I can't reproduce it either. Not even after including the relevant transactions_v2, and modifying the starting date of the transaction history range to \"20150101\", as suggested in this discussion:\n<a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145\">https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 245723,
      "author_name": "billiazz",
      "author_url": "",
      "post_date": "11/19/2017 13:28:34",
      "content": "<p>Hello. I have some troubles during analyzing the data because I don't clearly understand the meaning and all possible choices of these variables.</p>\n\n<p>Could you please explain and clarify more about them?</p>\n\n<p>From table transactions</p>\n\n<ol>\n<li><p>payment method --&gt; please explain the method name of each payment method id</p></li>\n<li><p>payment plan days --&gt; what does this variable mean?</p></li>\n<li><p>plan list price --&gt; what does this variable mean?</p></li>\n<li><p>actual amount paid  --&gt; what does this variable mean?</p></li>\n<li><p>transaction date  --&gt; what does this variable mean?</p></li>\n</ol>\n\n<p>From table members</p>\n\n<ol>\n<li><p>city --&gt; please explain city name of each city number</p></li>\n<li><p>registered via --&gt; please explain registration method name of number</p></li>\n<li><p>registration init time --&gt; what does this variable mean?</p></li>\n</ol>\n\n<p>I would appreciate it if you would answer my question.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 255910,
      "author_name": "infinitewing",
      "author_url": "",
      "post_date": "12/10/2017 15:06:48",
      "content": "<p>Hi, Wendy</p>\n\n<p>According to <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363\">this post</a>, I am wondering if we can use member.csv (ver. 1), where can we download the file?</p>\n\n<p>For someone entered the competition lately, we are not able to download the original member.csv, while the original member.csv have extra information of expired date.</p>",
      "votes": null,
      "replies": [
        {
          "id": 255915,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "12/10/2017 15:12:18",
          "content": "<p>InfiniteWing,</p>\n\n<p>As only the latest versions of the competition files are public, those are the only versions of the competition data that can be used.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255925,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/10/2017 15:21:35",
          "content": "<p>Hi Addison,</p>\n\n<p>Thanks for your reply, I missed that sentence.. sorry. I am curious how to detect people is using the newest version or not. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258116,
      "author_name": "maruti",
      "author_url": "",
      "post_date": "12/15/2017 14:23:34",
      "content": "<p>According to the data description for train data: <br>\n<strong>The criteria of \"churn\" is no new valid service subscription within 30 days after the current membership expires.</strong> <br>\nWith that definition, how can this user be \"churn\" in March? <br>\n<code>msno = K6fja4+jmoZ5xG6BypqX80Uw/XKpMgrEMdG2edFOxnA=</code> <br>\nThe user's last transaction in Feb was on 2017-02-16 with a membership expiration date of 2017-08-18;  So no matter what their activity is in March, they cannot be a \"churn\" customer in March according to the definition above.\nIs it better to just run a piece of code to generate your own (correct) version of churn status for the set of members in the train_v2 file?  If someone has already done that - could you kindly share that file?   Thanks in advance for answers /guidance.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "242135": "Hi all, \n\nOver the past week we have released the new test data. The leaderboard has been cleared. \n\nA summary of what is changed:\n\n1. Instead of predicting churn for March, we are now predicting churn for April. That goes into sample_submission_v2.csv\n\n2. Because of that, we released new (March) transactions (transactions_v2.csv) and user logs (user_logs_v2.csv) files. These DO NOT replace the old files since they only contain one month of data. \n\n3. There are some new users in this new test dataset that were previously unseen, so we included some extra historical transactions of those users - hence the dates before 3/1/2017 in transactions_v2.csv\n\n4. We re-generated the members information in members_v2.csv. This file DOES replace the previous version. We have removed the previous expiration date column since it caused controversy. \n\n5. Since there was leakage in the previous test.csv which contains the churn info for March, we decided to release the solution of those and let that become part of the training data. That becomes train_v2.csv. \n\n6. All the new files are now available in Kernels. \n\n7. Since the structure of the data has not changed, and we are only releasing more data, and that there is still more than one month until the end of the competition, we will keep the same timeline as before. \n\nThank you for your understanding and patience as we resolve the issue. \n\nKaggle/KKBox admins",
    "242313": "Thanks for clarification",
    "242551": "Seems members_v2.csv contains only 795090 entries, while the original member.csv file contains 5116194 entries. Does the new version indeed replace the previous one?",
    "242630": "Based on #3, \n\n&gt; There are some new users in this new test dataset that were previously unseen, ... , hence the dates before 3/1/2017 in transactions_v2.csv\n\nShould there also be some extra historical user logs (dates before 3/1/2017)?",
    "242885": "Thanks for clarifying things Wendy.",
    "243080": "I have extract the code segment to generate labels for users as \"WSDMChurnLabeller.scala\" in the data section. Hopefully this piece of code can give our participants more insights about how the churn label is applied. The training labels will not be limited to the ones that were provided.",
    "243294": "Thanks, this code is very helpful. I had this question before: in transactions file, transaction_date column is at day level granularity, for multiple transactions in one day, the order of them may affect the result of churn. e.g. renew first or cancel first. \n\nAlso if I understand correctly, ```val historyCutoff = \"20170131\"``` is used to generate train.csv(first training set). Can we think in the way: the snapshot date of generating train.csv(Feb churn or not) is \"20170131\"? Does it mean the snapshot date of generating train_v2(Mar churn or not) is \"20170228\" and snapshot date of generating test(Apr churn or not) target is \"20170331\"?",
    "243364": "Sorry for the mismatch in my description and the data. We have uploaded `members_v3` to correct the issue.",
    "244184": "I have run the code, as it is, over the original transaction file, and the results diverge from first train file. Has anyone noticed that?",
    "244469": "I can't reproduce it either. Not even after including the relevant transactions_v2, and modifying the starting date of the transaction history range to \"20150101\", as suggested in this discussion:\nhttps://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145",
    "245723": "Hello. I have some troubles during analyzing the data because I don't clearly understand the meaning and all possible choices of these variables.\n\n\nCould you please explain and clarify more about them?\n\n\nFrom table transactions\n\n1. payment method --&gt; please explain the method name of each payment method id\n\n2. payment plan days --&gt; what does this variable mean?\n\n3. plan list price --&gt; what does this variable mean?\n\n4. actual amount paid  --&gt; what does this variable mean?\n\n5. transaction date  --&gt; what does this variable mean?\n\n\nFrom table members\n\n6. city --&gt; please explain city name of each city number\n\n7. registered via --&gt; please explain registration method name of number\n\n8. registration init time --&gt; what does this variable mean?\n\n\nI would appreciate it if you would answer my question.",
    "255910": "Hi, Wendy\n\nAccording to [this post](https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363), I am wondering if we can use member.csv (ver. 1), where can we download the file?\n\nFor someone entered the competition lately, we are not able to download the original member.csv, while the original member.csv have extra information of expired date.",
    "255915": "InfiniteWing,\n\nAs only the latest versions of the competition files are public, those are the only versions of the competition data that can be used.",
    "255925": "Hi Addison,\n \nThanks for your reply, I missed that sentence.. sorry. I am curious how to detect people is using the newest version or not.",
    "258116": "According to the data description for train data:  \n**The criteria of \"churn\" is no new valid service subscription within 30 days after the current membership expires.**  \nWith that definition, how can this user be \"churn\" in March?  \n```msno = K6fja4+jmoZ5xG6BypqX80Uw/XKpMgrEMdG2edFOxnA= ```  \nThe user's last transaction in Feb was on 2017-02-16 with a membership expiration date of 2017-08-18;  So no matter what their activity is in March, they cannot be a \"churn\" customer in March according to the definition above.\nIs it better to just run a piece of code to generate your own (correct) version of churn status for the set of members in the train_v2 file?  If someone has already done that - could you kindly share that file?   Thanks in advance for answers /guidance."
  },
  "source": "meta"
}