{
  "id": 43145,
  "title": "something really hard to understand",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/43145",
  "author_name": "",
  "post_date": "2017-11-10T09:48:03.580340800Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>i have read a lot of discussions here,and spent a lot of time try to under stand the data.see the transaction log of  zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=  ,the user was labeled churned twice in train and train_v2,and this user is in the final test set ,but the transaction indicate that this user didn't churn at all,can any one help me to figure this out ?</p>\n\n<p>zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                20161110                20161214 <br>\n zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                20161210                20170114 <br>\nzLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                 20170110                20170214          zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                 20170210                20170314          zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                 20170310                20170414          </p>",
  "messages": [
    {
      "id": "241976",
      "postDate": "11/10/2017 09:48:03",
      "content": "<p>i have read a lot of discussions here,and spent a lot of time try to under stand the data.see the transaction log of  zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=  ,the user was labeled churned twice in train and train_v2,and this user is in the final test set ,but the transaction indicate that this user didn't churn at all,can any one help me to figure this out ?</p>\n\n<p>zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                20161110                20161214 <br>\n zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                20161210                20170114 <br>\nzLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                 20170110                20170214          zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                 20170210                20170314          zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                 20170310                20170414          </p>",
      "rawMarkdown": "i have read a lot of discussions here,and spent a lot of time try to under stand the data.see the transaction log of  zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=  ,the user was labeled churned twice in train and train_v2,and this user is in the final test set ,but the transaction indicate that this user didn't churn at all,can any one help me to figure this out ?\n\n zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                20161110                20161214          \n zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                20161210                20170114         \nzLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                 20170110                20170214          zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                 20170210                20170314          zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                 20170310                20170414",
      "votes": null
    },
    {
      "id": "242010",
      "postDate": "11/10/2017 11:29:03",
      "content": "<p>While this is a very interesting competition and I appreciate the effort of the organizers to provide new data, I believe many cases of non consistent data have been reported without satisfactory explanations. This particular user reported by liulongxiao is an obvious case in which the data seem to contradict the provided definition of churn. Serious predictions cannot be made without a clear and robust definition of churn. Can the organizers please clarify?</p>\n\n<p>More on this user according to both transaction data: </p>\n\n<ul>\n<li>From Jan 2015 till July 2015 he was on an automatic renewal. On 2015-07-18   he canceled the automatic renewal.</li>\n<li>From July 2015 till May 2016 he had sporadic delays in his manual renewals, with a maximum delay of 20 days happening on 2015-09-07. That day he renewed his subscription that had expired 20 days ago on 2015-08-18.</li>\n<li>In May 2015 he went back to automatic renewal without known interruption up to March 2017 (last five transactions shown on liulongxiao comment). </li>\n</ul>\n\n<p>So according to the definition this user has not churned neither on February nor in March 2017, contradicting the provided data. From the 30 days rule of the churn definition he has not churned since 2015.</p>",
      "rawMarkdown": "While this is a very interesting competition and I appreciate the effort of the organizers to provide new data, I believe many cases of non consistent data have been reported without satisfactory explanations. This particular user reported by liulongxiao is an obvious case in which the data seem to contradict the provided definition of churn. Serious predictions cannot be made without a clear and robust definition of churn. Can the organizers please clarify?\n\nMore on this user according to both transaction data: \n\n - From Jan 2015 till July 2015 he was on an automatic renewal. On 2015-07-18\the canceled the automatic renewal.\n - From July 2015 till May 2016 he had sporadic delays in his manual renewals, with a maximum delay of 20 days happening on 2015-09-07. That day he renewed his subscription that had expired 20 days ago on 2015-08-18.\n - In May 2015 he went back to automatic renewal without known interruption up to March 2017 (last five transactions shown on liulongxiao comment). \n\nSo according to the definition this user has not churned neither on February nor in March 2017, contradicting the provided data. From the 30 days rule of the churn definition he has not churned since 2015.",
      "votes": null
    },
    {
      "id": "243076",
      "postDate": "11/13/2017 10:17:09",
      "content": "<p>I have extract the parts of our ETL code for generating the labels for users based on the transaction log. The file is uploaded as \"WSDMChurnLabeller.scala\" in the data section. The code segment should provide more insights into how the labels were applied. Further, we can generate multiple training labels by adjusting the parameters in the code segment.</p>",
      "rawMarkdown": "I have extract the parts of our ETL code for generating the labels for users based on the transaction log. The file is uploaded as \"WSDMChurnLabeller.scala\" in the data section. The code segment should provide more insights into how the labels were applied. Further, we can generate multiple training labels by adjusting the parameters in the code segment.",
      "votes": null
    },
    {
      "id": "243190",
      "postDate": "11/13/2017 14:42:07",
      "content": "<p>I see that as more of a \"real case use\" situation instead than a bad dataset. In real life, data is messy, things never sum up to 100%, and we will never get a perfect algorithm on perfect data predicting perfect humans. Part of the challenge is being able to cope with weird data, we can get solid predictions on this provided dataset by looking at the leaderboard, i find logloss &lt; 0.1 really good for this dataset, which we will never be able to fully predict as there is a lot more in play for a human decision to continue/stop a service than what we have as available data.</p>",
      "rawMarkdown": "I see that as more of a \"real case use\" situation instead than a bad dataset. In real life, data is messy, things never sum up to 100%, and we will never get a perfect algorithm on perfect data predicting perfect humans. Part of the challenge is being able to cope with weird data, we can get solid predictions on this provided dataset by looking at the leaderboard, i find logloss &lt; 0.1 really good for this dataset, which we will never be able to fully predict as there is a lot more in play for a human decision to continue/stop a service than what we have as available data.",
      "votes": null
    },
    {
      "id": "243223",
      "postDate": "11/13/2017 15:33:58",
      "content": "<p>really appreciate that</p>",
      "rawMarkdown": "really appreciate that",
      "votes": null
    },
    {
      "id": "243240",
      "postDate": "11/13/2017 16:09:31",
      "content": "<p>val historyCutoff = \"20170131\"\nval historyData = data.filter(col(\"transaction_date\")&gt;=\"20170101\" and col(\"transaction_date\")&lt;=lit(historyCutoff))\n    val futureData = data.filter(col(\"transaction_date\") &gt; lit(historyCutoff))</p>\n\n#\n\n<p>above is taken from your scrips which says that the train user must have transactions after '20170101', but i have found  many users dont have transactions after 201606 but still contains in the training set,(but they have previous transactions that expire fall in the prediction scope but canceled then),could you please help me with this case ,thks a lot </p>",
      "rawMarkdown": "val historyCutoff = \"20170131\"\nval historyData = data.filter(col(\"transaction_date\")&gt;=\"20170101\" and col(\"transaction_date\")&lt;=lit(historyCutoff))\n    val futureData = data.filter(col(\"transaction_date\") &gt; lit(historyCutoff))\n\n#########################################################\nabove is taken from your scrips which says that the train user must have transactions after '20170101', but i have found  many users dont have transactions after 201606 but still contains in the training set,(but they have previous transactions that expire fall in the prediction scope but canceled then),could you please help me with this case ,thks a lot",
      "votes": null
    },
    {
      "id": "243295",
      "postDate": "11/13/2017 18:38:43",
      "content": "<p>I understand a real life scenario can be messy and subject to uncertain features that escape our control. However here we are given a definition of churn that should be reproducible. Either the algorithm that tags churns does not correspond to what I understand from the description, or I (and few others) am missing something important. Hopefully the provided code (thanks for that) will bring some insight.</p>",
      "rawMarkdown": "I understand a real life scenario can be messy and subject to uncertain features that escape our control. However here we are given a definition of churn that should be reproducible. Either the algorithm that tags churns does not correspond to what I understand from the description, or I (and few others) am missing something important. Hopefully the provided code (thanks for that) will bring some insight.",
      "votes": null
    },
    {
      "id": "243924",
      "postDate": "11/15/2017 06:25:32",
      "content": "<p>The script we provide is  to let people pick the \"window\" at which you observe user churn. Note that the 20170101 in the script is just there so that it is easier to run on my laptop. The actual date we used to generate the transaction log and the test data is \"20150101\".</p>\n\n<p>To generate different test sets, we can change the date values in the script</p>",
      "rawMarkdown": "The script we provide is  to let people pick the \"window\" at which you observe user churn. Note that the 20170101 in the script is just there so that it is easier to run on my laptop. The actual date we used to generate the transaction log and the test data is \"20150101\".\n\nTo generate different test sets, we can change the date values in the script",
      "votes": null
    },
    {
      "id": "253562",
      "postDate": "12/05/2017 07:47:20",
      "content": "<p>Can anyone plz answer me is this user is churn or not after running the scala code?</p>\n\n<p>After reading the scala code, still confused about this example.</p>",
      "rawMarkdown": "Can anyone plz answer me is this user is churn or not after running the scala code?\n\nAfter reading the scala code, still confused about this example.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 242010,
      "author_name": "brucala",
      "author_url": "",
      "post_date": "11/10/2017 11:29:03",
      "content": "<p>While this is a very interesting competition and I appreciate the effort of the organizers to provide new data, I believe many cases of non consistent data have been reported without satisfactory explanations. This particular user reported by liulongxiao is an obvious case in which the data seem to contradict the provided definition of churn. Serious predictions cannot be made without a clear and robust definition of churn. Can the organizers please clarify?</p>\n\n<p>More on this user according to both transaction data: </p>\n\n<ul>\n<li>From Jan 2015 till July 2015 he was on an automatic renewal. On 2015-07-18   he canceled the automatic renewal.</li>\n<li>From July 2015 till May 2016 he had sporadic delays in his manual renewals, with a maximum delay of 20 days happening on 2015-09-07. That day he renewed his subscription that had expired 20 days ago on 2015-08-18.</li>\n<li>In May 2015 he went back to automatic renewal without known interruption up to March 2017 (last five transactions shown on liulongxiao comment). </li>\n</ul>\n\n<p>So according to the definition this user has not churned neither on February nor in March 2017, contradicting the provided data. From the 30 days rule of the churn definition he has not churned since 2015.</p>",
      "votes": null,
      "replies": [
        {
          "id": 243190,
          "author_name": "jeffgrenier",
          "author_url": "",
          "post_date": "11/13/2017 14:42:07",
          "content": "<p>I see that as more of a \"real case use\" situation instead than a bad dataset. In real life, data is messy, things never sum up to 100%, and we will never get a perfect algorithm on perfect data predicting perfect humans. Part of the challenge is being able to cope with weird data, we can get solid predictions on this provided dataset by looking at the leaderboard, i find logloss &lt; 0.1 really good for this dataset, which we will never be able to fully predict as there is a lot more in play for a human decision to continue/stop a service than what we have as available data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 243295,
          "author_name": "brucala",
          "author_url": "",
          "post_date": "11/13/2017 18:38:43",
          "content": "<p>I understand a real life scenario can be messy and subject to uncertain features that escape our control. However here we are given a definition of churn that should be reproducible. Either the algorithm that tags churns does not correspond to what I understand from the description, or I (and few others) am missing something important. Hopefully the provided code (thanks for that) will bring some insight.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 243076,
      "author_name": "ardenkkbox",
      "author_url": "",
      "post_date": "11/13/2017 10:17:09",
      "content": "<p>I have extract the parts of our ETL code for generating the labels for users based on the transaction log. The file is uploaded as \"WSDMChurnLabeller.scala\" in the data section. The code segment should provide more insights into how the labels were applied. Further, we can generate multiple training labels by adjusting the parameters in the code segment.</p>",
      "votes": null,
      "replies": [
        {
          "id": 243223,
          "author_name": "liulongxiao1",
          "author_url": "",
          "post_date": "11/13/2017 15:33:58",
          "content": "<p>really appreciate that</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 243240,
          "author_name": "liulongxiao1",
          "author_url": "",
          "post_date": "11/13/2017 16:09:31",
          "content": "<p>val historyCutoff = \"20170131\"\nval historyData = data.filter(col(\"transaction_date\")&gt;=\"20170101\" and col(\"transaction_date\")&lt;=lit(historyCutoff))\n    val futureData = data.filter(col(\"transaction_date\") &gt; lit(historyCutoff))</p>\n\n#\n\n<p>above is taken from your scrips which says that the train user must have transactions after '20170101', but i have found  many users dont have transactions after 201606 but still contains in the training set,(but they have previous transactions that expire fall in the prediction scope but canceled then),could you please help me with this case ,thks a lot </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 243924,
          "author_name": "ardenkkbox",
          "author_url": "",
          "post_date": "11/15/2017 06:25:32",
          "content": "<p>The script we provide is  to let people pick the \"window\" at which you observe user churn. Note that the 20170101 in the script is just there so that it is easier to run on my laptop. The actual date we used to generate the transaction log and the test data is \"20150101\".</p>\n\n<p>To generate different test sets, we can change the date values in the script</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 253562,
      "author_name": "infinitewing",
      "author_url": "",
      "post_date": "12/05/2017 07:47:20",
      "content": "<p>Can anyone plz answer me is this user is churn or not after running the scala code?</p>\n\n<p>After reading the scala code, still confused about this example.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "241976": "i have read a lot of discussions here,and spent a lot of time try to under stand the data.see the transaction log of  zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=  ,the user was labeled churned twice in train and train_v2,and this user is in the final test set ,but the transaction indicate that this user didn't churn at all,can any one help me to figure this out ?\n\n zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                20161110                20161214          \n zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                20161210                20170114         \nzLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                 20170110                20170214          zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                 20170210                20170314          zLof7igqfvKfIdDOGrs8mmfvQ4xqUnM4nwRU4YgX59I=                 20170310                20170414",
    "242010": "While this is a very interesting competition and I appreciate the effort of the organizers to provide new data, I believe many cases of non consistent data have been reported without satisfactory explanations. This particular user reported by liulongxiao is an obvious case in which the data seem to contradict the provided definition of churn. Serious predictions cannot be made without a clear and robust definition of churn. Can the organizers please clarify?\n\nMore on this user according to both transaction data: \n\n - From Jan 2015 till July 2015 he was on an automatic renewal. On 2015-07-18\the canceled the automatic renewal.\n - From July 2015 till May 2016 he had sporadic delays in his manual renewals, with a maximum delay of 20 days happening on 2015-09-07. That day he renewed his subscription that had expired 20 days ago on 2015-08-18.\n - In May 2015 he went back to automatic renewal without known interruption up to March 2017 (last five transactions shown on liulongxiao comment). \n\nSo according to the definition this user has not churned neither on February nor in March 2017, contradicting the provided data. From the 30 days rule of the churn definition he has not churned since 2015.",
    "243076": "I have extract the parts of our ETL code for generating the labels for users based on the transaction log. The file is uploaded as \"WSDMChurnLabeller.scala\" in the data section. The code segment should provide more insights into how the labels were applied. Further, we can generate multiple training labels by adjusting the parameters in the code segment.",
    "243190": "I see that as more of a \"real case use\" situation instead than a bad dataset. In real life, data is messy, things never sum up to 100%, and we will never get a perfect algorithm on perfect data predicting perfect humans. Part of the challenge is being able to cope with weird data, we can get solid predictions on this provided dataset by looking at the leaderboard, i find logloss &lt; 0.1 really good for this dataset, which we will never be able to fully predict as there is a lot more in play for a human decision to continue/stop a service than what we have as available data.",
    "243223": "really appreciate that",
    "243240": "val historyCutoff = \"20170131\"\nval historyData = data.filter(col(\"transaction_date\")&gt;=\"20170101\" and col(\"transaction_date\")&lt;=lit(historyCutoff))\n    val futureData = data.filter(col(\"transaction_date\") &gt; lit(historyCutoff))\n\n#########################################################\nabove is taken from your scrips which says that the train user must have transactions after '20170101', but i have found  many users dont have transactions after 201606 but still contains in the training set,(but they have previous transactions that expire fall in the prediction scope but canceled then),could you please help me with this case ,thks a lot",
    "243295": "I understand a real life scenario can be messy and subject to uncertain features that escape our control. However here we are given a definition of churn that should be reproducible. Either the algorithm that tags churns does not correspond to what I understand from the description, or I (and few others) am missing something important. Hopefully the provided code (thanks for that) will bring some insight.",
    "243924": "The script we provide is  to let people pick the \"window\" at which you observe user churn. Note that the 20170101 in the script is just there so that it is easier to run on my laptop. The actual date we used to generate the transaction log and the test data is \"20150101\".\n\nTo generate different test sets, we can change the date values in the script",
    "253562": "Can anyone plz answer me is this user is churn or not after running the scala code?\n\nAfter reading the scala code, still confused about this example."
  },
  "source": "meta"
}