{
  "id": 44271,
  "title": "Why is this user is considered as 'is_churn=1' in March?",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/44271",
  "author_name": "Sundong Kim",
  "post_date": "2017-11-26T13:05:44.654000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Why are these users are considered as 'is_churn=1' in March? <br>\n(Please check attached links - churn labels and transaction histories for three users) </p>\n\n<p>Last two examples are what previous Kaggler(@yliu) has been reported.</p>\n\n<p>I have to predict the prediction label, but I can not trust labels in the train data yet.\nThe label of the unveiled test data also seems to have the same problem.</p>\n\n<ul>\n<li>User 1: <a href=\"https://drive.google.com/open?id=1IVf2CWthRHH5CBALAQnQMjQ5cqhKubmG\">https://drive.google.com/open?id=1IVf2CWthRHH5CBALAQnQMjQ5cqhKubmG</a></li>\n<li>User 2: <a href=\"https://drive.google.com/open?id=1dBb81-r66LM8_45eE8A7EJ56TAXkZnia\">https://drive.google.com/open?id=1dBb81-r66LM8_45eE8A7EJ56TAXkZnia</a></li>\n<li>User 3: <a href=\"https://drive.google.com/open?id=1RyvBrSKtNgHa-W0wyCOVgIAcz610VSZ5\">https://drive.google.com/open?id=1RyvBrSKtNgHa-W0wyCOVgIAcz610VSZ5</a></li>\n</ul>\n\n<p>(NOTE: A 'transactions' dataframe is a concatenation of two transactions data given by the competition)</p>",
  "messages": [
    {
      "id": 250080,
      "postDate": "2017-11-29T18:26:52.243Z",
      "content": "<p>I think you can reduce the difference between the extracted files and the official ones if you change the variable <strong>historyCutoff</strong>  to some months  in the past according to Arden Chiu in <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145\">this post</a>.</p>\n\n<p>Even doing that I was not able to exactly match the official train files... </p>",
      "rawMarkdown": "I think you can reduce the difference between the extracted files and the official ones if you change the variable **historyCutoff**  to some months  in the past according to Arden Chiu in [this post][1].\n\nEven doing that I was not able to exactly match the official train files... \n\n  [1]: https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145",
      "votes": 1
    },
    {
      "id": 249983,
      "postDate": "2017-11-29T15:38:55.433Z",
      "content": "<p>The first one you mentioned (Msno=hDW0/cSgIayezNOh5RsbKKskn4WsNUhrx6ZEbInTSQk=) is definitely a mistake of some kind.  According to the scala labeller code, it should not have been included at all in the march test set (train2).  </p>\n\n<p>Scala Explanation:  Using March date cutoffs (historyCutoff=20170228), this msno in the \"historyData\" dataset would have had a \"last_expire\" date of 20170409, which means it is filtered out of the \"predictionCandidates\" dataset which only includes msno with a last_expire date between 20170301 and after 20170331.</p>\n\n<p>Clearly the scala labeller was either ran with incorrect values (improper date cutoffs) to generate our train/test sets, or the code we were given was not the exact code actually used.</p>",
      "rawMarkdown": "The first one you mentioned (Msno=hDW0/cSgIayezNOh5RsbKKskn4WsNUhrx6ZEbInTSQk=) is definitely a mistake of some kind.  According to the scala labeller code, it should not have been included at all in the march test set (train2).  \n\nScala Explanation:  Using March date cutoffs (historyCutoff=20170228), this msno in the \"historyData\" dataset would have had a \"last_expire\" date of 20170409, which means it is filtered out of the \"predictionCandidates\" dataset which only includes msno with a last_expire date between 20170301 and after 20170331.\n\nClearly the scala labeller was either ran with incorrect values (improper date cutoffs) to generate our train/test sets, or the code we were given was not the exact code actually used.",
      "votes": 1,
      "replies": [
        {
          "id": 250810,
          "postDate": "2017-11-30T10:52:24.607Z",
          "content": "<p>This is really helpful for me to understand the Scala labeler which I was not familiar with. \nI wasn't considering the labeler yet, but the key aspect to get the high score seems like doing optimization according to labels from this labeler...! </p>",
          "rawMarkdown": "This is really helpful for me to understand the Scala labeler which I was not familiar with. \nI wasn't considering the labeler yet, but the key aspect to get the high score seems like doing optimization according to labels from this labeler...! "
        }
      ]
    },
    {
      "id": 248549,
      "postDate": "2017-11-26T13:05:44.653Z",
      "content": "<p>Why are these users are considered as 'is_churn=1' in March? <br>\n(Please check attached links - churn labels and transaction histories for three users) </p>\n\n<p>Last two examples are what previous Kaggler(@yliu) has been reported.</p>\n\n<p>I have to predict the prediction label, but I can not trust labels in the train data yet.\nThe label of the unveiled test data also seems to have the same problem.</p>\n\n<ul>\n<li>User 1: <a href=\"https://drive.google.com/open?id=1IVf2CWthRHH5CBALAQnQMjQ5cqhKubmG\">https://drive.google.com/open?id=1IVf2CWthRHH5CBALAQnQMjQ5cqhKubmG</a></li>\n<li>User 2: <a href=\"https://drive.google.com/open?id=1dBb81-r66LM8_45eE8A7EJ56TAXkZnia\">https://drive.google.com/open?id=1dBb81-r66LM8_45eE8A7EJ56TAXkZnia</a></li>\n<li>User 3: <a href=\"https://drive.google.com/open?id=1RyvBrSKtNgHa-W0wyCOVgIAcz610VSZ5\">https://drive.google.com/open?id=1RyvBrSKtNgHa-W0wyCOVgIAcz610VSZ5</a></li>\n</ul>\n\n<p>(NOTE: A 'transactions' dataframe is a concatenation of two transactions data given by the competition)</p>",
      "rawMarkdown": "Why are these users are considered as 'is_churn=1' in March? <br>\n(Please check attached links - churn labels and transaction histories for three users) \n\nLast two examples are what previous Kaggler(@yliu) has been reported.\n\nI have to predict the prediction label, but I can not trust labels in the train data yet.\nThe label of the unveiled test data also seems to have the same problem.\n\n* User 1: https://drive.google.com/open?id=1IVf2CWthRHH5CBALAQnQMjQ5cqhKubmG\n* User 2: https://drive.google.com/open?id=1dBb81-r66LM8_45eE8A7EJ56TAXkZnia\n* User 3: https://drive.google.com/open?id=1RyvBrSKtNgHa-W0wyCOVgIAcz610VSZ5\n\n(NOTE: A 'transactions' dataframe is a concatenation of two transactions data given by the competition)",
      "votes": 1
    },
    {
      "id": 249961,
      "postDate": "2017-11-29T15:02:35.903Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 250080,
      "author_name": "Aloisio Dourado",
      "author_url": "",
      "post_date": "2017-11-29T18:26:52.243000",
      "content": "<p>I think you can reduce the difference between the extracted files and the official ones if you change the variable <strong>historyCutoff</strong>  to some months  in the past according to Arden Chiu in <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145\">this post</a>.</p>\n\n<p>Even doing that I was not able to exactly match the official train files... </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 249983,
      "author_name": "Bryan Gregory",
      "author_url": "",
      "post_date": "2017-11-29T15:38:55.433000",
      "content": "<p>The first one you mentioned (Msno=hDW0/cSgIayezNOh5RsbKKskn4WsNUhrx6ZEbInTSQk=) is definitely a mistake of some kind.  According to the scala labeller code, it should not have been included at all in the march test set (train2).  </p>\n\n<p>Scala Explanation:  Using March date cutoffs (historyCutoff=20170228), this msno in the \"historyData\" dataset would have had a \"last_expire\" date of 20170409, which means it is filtered out of the \"predictionCandidates\" dataset which only includes msno with a last_expire date between 20170301 and after 20170331.</p>\n\n<p>Clearly the scala labeller was either ran with incorrect values (improper date cutoffs) to generate our train/test sets, or the code we were given was not the exact code actually used.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 250810,
          "author_name": "Sundong Kim",
          "author_url": "",
          "post_date": "2017-11-30T10:52:24.607000",
          "content": "<p>This is really helpful for me to understand the Scala labeler which I was not familiar with. \nI wasn't considering the labeler yet, but the key aspect to get the high score seems like doing optimization according to labels from this labeler...! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 249961,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-11-29T15:02:35.903000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "250080": "I think you can reduce the difference between the extracted files and the official ones if you change the variable **historyCutoff**  to some months  in the past according to Arden Chiu in [this post][1].\n\nEven doing that I was not able to exactly match the official train files... \n\n  [1]: https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145",
    "249983": "The first one you mentioned (Msno=hDW0/cSgIayezNOh5RsbKKskn4WsNUhrx6ZEbInTSQk=) is definitely a mistake of some kind.  According to the scala labeller code, it should not have been included at all in the march test set (train2).  \n\nScala Explanation:  Using March date cutoffs (historyCutoff=20170228), this msno in the \"historyData\" dataset would have had a \"last_expire\" date of 20170409, which means it is filtered out of the \"predictionCandidates\" dataset which only includes msno with a last_expire date between 20170301 and after 20170331.\n\nClearly the scala labeller was either ran with incorrect values (improper date cutoffs) to generate our train/test sets, or the code we were given was not the exact code actually used.",
    "248549": "Why are these users are considered as 'is_churn=1' in March? <br>\n(Please check attached links - churn labels and transaction histories for three users) \n\nLast two examples are what previous Kaggler(@yliu) has been reported.\n\nI have to predict the prediction label, but I can not trust labels in the train data yet.\nThe label of the unveiled test data also seems to have the same problem.\n\n* User 1: https://drive.google.com/open?id=1IVf2CWthRHH5CBALAQnQMjQ5cqhKubmG\n* User 2: https://drive.google.com/open?id=1dBb81-r66LM8_45eE8A7EJ56TAXkZnia\n* User 3: https://drive.google.com/open?id=1RyvBrSKtNgHa-W0wyCOVgIAcz610VSZ5\n\n(NOTE: A 'transactions' dataframe is a concatenation of two transactions data given by the competition)",
    "249961": ""
  }
}