{
  "id": 44112,
  "title": "Using WSDMChurnLabeller.scala",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/44112",
  "author_name": "",
  "post_date": "2017-11-23T16:25:45.039895400Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello everybody,</p>\n\n<p>I am trying to to use the scala code to generate new labeled data. At this stage I am just trying the replicate the labels using the hard-coded dates :\nval historyCutoff = \"20170131\"</p>\n\n<p>I then compared with the labels provided in train.csv.</p>\n\n<p>The number of labeled customers in train.csv is : 992931\nThe number of labeled customers in the csv generated from WSDMChurnLabeller.scala is : 929460</p>\n\n<p>There are also some disagreements between the 2 sets of customers and some labels (~5%) in disagreement between the 2 datasets.</p>\n\n<p>Am I using the script correctly / comparing with the right dataset or should I modify the dates to match the labels from train.csv ?</p>\n\n<p>Thanks for your help,</p>\n\n<p>Cheers,</p>\n\n<p>Bertrand</p>",
  "messages": [
    {
      "id": "247693",
      "postDate": "11/23/2017 16:25:45",
      "content": "<p>Hello everybody,</p>\n\n<p>I am trying to to use the scala code to generate new labeled data. At this stage I am just trying the replicate the labels using the hard-coded dates :\nval historyCutoff = \"20170131\"</p>\n\n<p>I then compared with the labels provided in train.csv.</p>\n\n<p>The number of labeled customers in train.csv is : 992931\nThe number of labeled customers in the csv generated from WSDMChurnLabeller.scala is : 929460</p>\n\n<p>There are also some disagreements between the 2 sets of customers and some labels (~5%) in disagreement between the 2 datasets.</p>\n\n<p>Am I using the script correctly / comparing with the right dataset or should I modify the dates to match the labels from train.csv ?</p>\n\n<p>Thanks for your help,</p>\n\n<p>Cheers,</p>\n\n<p>Bertrand</p>",
      "rawMarkdown": "Hello everybody,\n\nI am trying to to use the scala code to generate new labeled data. At this stage I am just trying the replicate the labels using the hard-coded dates :\nval historyCutoff = \"20170131\"\n\nI then compared with the labels provided in train.csv.\n\nThe number of labeled customers in train.csv is : 992931\nThe number of labeled customers in the csv generated from WSDMChurnLabeller.scala is : 929460\n\nThere are also some disagreements between the 2 sets of customers and some labels (~5%) in disagreement between the 2 datasets.\n\nAm I using the script correctly / comparing with the right dataset or should I modify the dates to match the labels from train.csv ?\n\nThanks for your help,\n\nCheers,\n\nBertrand",
      "votes": null
    },
    {
      "id": "247735",
      "postDate": "11/23/2017 18:09:42",
      "content": "<p>Hi Bertrand! Take a look at this: <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145\">https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145</a>.</p>\n\n<p>Even changing dates and using the full transactions dataset, I was not able to replicate train files...</p>",
      "rawMarkdown": "Hi Bertrand! Take a look at this: [https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145][1].\n\nEven changing dates and using the full transactions dataset, I was not able to replicate train files...\n\n  [1]: https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145",
      "votes": null
    },
    {
      "id": "247980",
      "postDate": "11/24/2017 13:27:28",
      "content": "<p>Thank you Aloisio. Following the link, the answers are not very satisfying. I do not see the point of providing a script to create more training examples if we cannot replicate the given training dataset. </p>",
      "rawMarkdown": "Thank you Aloisio. Following the link, the answers are not very satisfying. I do not see the point of providing a script to create more training examples if we cannot replicate the given training dataset.",
      "votes": null
    },
    {
      "id": "251716",
      "postDate": "12/01/2017 15:43:30",
      "content": "<p>Hi Aloisio! Can I ask how many records do you get when running with \"20150101\" and cutoff=\"20170301\"? Or do you use other values?</p>",
      "rawMarkdown": "Hi Aloisio! Can I ask how many records do you get when running with \"20150101\" and cutoff=\"20170301\"? Or do you use other values?",
      "votes": null
    },
    {
      "id": "251969",
      "postDate": "12/02/2017 01:51:03",
      "content": "<p>Sorry! I did not run with this setup and I think I can't tell what setup I am using without exposing too much my strategy.\nBut I can tell you that even using the recommendend settings I did not match the exact number of records of the official train files... </p>",
      "rawMarkdown": "Sorry! I did not run with this setup and I think I can't tell what setup I am using without exposing too much my strategy.\nBut I can tell you that even using the recommendend settings I did not match the exact number of records of the official train files...",
      "votes": null
    },
    {
      "id": "258763",
      "postDate": "12/17/2017 01:09:41",
      "content": "<p>Javad, this might be a little late, but I get 867530 records. Hope that helps!</p>",
      "rawMarkdown": "Javad, this might be a little late, but I get 867530 records. Hope that helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 247735,
      "author_name": "aloisiodn",
      "author_url": "",
      "post_date": "11/23/2017 18:09:42",
      "content": "<p>Hi Bertrand! Take a look at this: <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145\">https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145</a>.</p>\n\n<p>Even changing dates and using the full transactions dataset, I was not able to replicate train files...</p>",
      "votes": null,
      "replies": [
        {
          "id": 251716,
          "author_name": "jvdnouri",
          "author_url": "",
          "post_date": "12/01/2017 15:43:30",
          "content": "<p>Hi Aloisio! Can I ask how many records do you get when running with \"20150101\" and cutoff=\"20170301\"? Or do you use other values?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 251969,
          "author_name": "aloisiodn",
          "author_url": "",
          "post_date": "12/02/2017 01:51:03",
          "content": "<p>Sorry! I did not run with this setup and I think I can't tell what setup I am using without exposing too much my strategy.\nBut I can tell you that even using the recommendend settings I did not match the exact number of records of the official train files... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258763,
          "author_name": "akshaysubr",
          "author_url": "",
          "post_date": "12/17/2017 01:09:41",
          "content": "<p>Javad, this might be a little late, but I get 867530 records. Hope that helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 247980,
      "author_name": "bertrandb",
      "author_url": "",
      "post_date": "11/24/2017 13:27:28",
      "content": "<p>Thank you Aloisio. Following the link, the answers are not very satisfying. I do not see the point of providing a script to create more training examples if we cannot replicate the given training dataset. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "247693": "Hello everybody,\n\nI am trying to to use the scala code to generate new labeled data. At this stage I am just trying the replicate the labels using the hard-coded dates :\nval historyCutoff = \"20170131\"\n\nI then compared with the labels provided in train.csv.\n\nThe number of labeled customers in train.csv is : 992931\nThe number of labeled customers in the csv generated from WSDMChurnLabeller.scala is : 929460\n\nThere are also some disagreements between the 2 sets of customers and some labels (~5%) in disagreement between the 2 datasets.\n\nAm I using the script correctly / comparing with the right dataset or should I modify the dates to match the labels from train.csv ?\n\nThanks for your help,\n\nCheers,\n\nBertrand",
    "247735": "Hi Bertrand! Take a look at this: [https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145][1].\n\nEven changing dates and using the full transactions dataset, I was not able to replicate train files...\n\n  [1]: https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43145",
    "247980": "Thank you Aloisio. Following the link, the answers are not very satisfying. I do not see the point of providing a script to create more training examples if we cannot replicate the given training dataset.",
    "251716": "Hi Aloisio! Can I ask how many records do you get when running with \"20150101\" and cutoff=\"20170301\"? Or do you use other values?",
    "251969": "Sorry! I did not run with this setup and I think I can't tell what setup I am using without exposing too much my strategy.\nBut I can tell you that even using the recommendend settings I did not match the exact number of records of the official train files...",
    "258763": "Javad, this might be a little late, but I get 867530 records. Hope that helps!"
  },
  "source": "meta"
}