{
  "id": 45615,
  "title": "Different label pre-processing method,  which to trust?",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/45615",
  "author_name": "",
  "post_date": "2017-12-13T14:42:27.097618200Z",
  "votes": null,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Currently I run the official scala code, and I got ~0.116 on LB. (~0.146 CV).<br>\nHowever, I can get ~0.108 on LB (~0.166 CV) with my own labeling method. I use same features for both result, only the training label is different.</p>\n\n<p>In addition, both the label result of two label methods are different from train.csv, train_v2.csv.</p>\n\n<p><br><br>\nAnyone faced the same issue? I am wondering is top teams use the official scala code or not.</p>",
  "messages": [
    {
      "id": "257156",
      "postDate": "12/13/2017 14:42:27",
      "content": "<p>Currently I run the official scala code, and I got ~0.116 on LB. (~0.146 CV).<br>\nHowever, I can get ~0.108 on LB (~0.166 CV) with my own labeling method. I use same features for both result, only the training label is different.</p>\n\n<p>In addition, both the label result of two label methods are different from train.csv, train_v2.csv.</p>\n\n<p><br><br>\nAnyone faced the same issue? I am wondering is top teams use the official scala code or not.</p>",
      "rawMarkdown": "Currently I run the official scala code, and I got ~0.116 on LB. (~0.146 CV).<br>\nHowever, I can get ~0.108 on LB (~0.166 CV) with my own labeling method. I use same features for both result, only the training label is different.\n\nIn addition, both the label result of two label methods are different from train.csv, train_v2.csv.\n\n<br><br>\nAnyone faced the same issue? I am wondering is top teams use the official scala code or not.",
      "votes": null
    },
    {
      "id": "257248",
      "postDate": "12/13/2017 19:43:17",
      "content": "<p>I prefer to use the official labeller because, teoretically, it is used to generate the labels used in public and private LB. But if you are getting a better result with your own code, I think you should go ahead!</p>\n\n<p>The competition end is coming soon. Good luck to us all!!!</p>",
      "rawMarkdown": "I prefer to use the official labeller because, teoretically, it is used to generate the labels used in public and private LB. But if you are getting a better result with your own code, I think you should go ahead!\n\nThe competition end is coming soon. Good luck to us all!!!",
      "votes": null
    },
    {
      "id": "257350",
      "postDate": "12/14/2017 04:40:51",
      "content": "<p>Hi, do you find a big difference between training your model on the train.csv and train_v2.csv files versus generating the churn flag using the scala file? </p>",
      "rawMarkdown": "Hi, do you find a big difference between training your model on the train.csv and train_v2.csv files versus generating the churn flag using the scala file?",
      "votes": null
    },
    {
      "id": "257475",
      "postDate": "12/14/2017 10:14:33",
      "content": "<p>Good questions, I did not use train.csv and train_v2.csv to train my model since the second day after I entered the competition..</p>",
      "rawMarkdown": "Good questions, I did not use train.csv and train_v2.csv to train my model since the second day after I entered the competition..",
      "votes": null
    },
    {
      "id": "257480",
      "postDate": "12/14/2017 10:18:22",
      "content": "<p>Good luck tooo!</p>\n\n<p>Thanks for your reply, I also believe CV score more, but the risk is that the CV score is highly based on the labeling method...</p>",
      "rawMarkdown": "Good luck tooo!\n\nThanks for your reply, I also believe CV score more, but the risk is that the CV score is highly based on the labeling method...",
      "votes": null
    },
    {
      "id": "257484",
      "postDate": "12/14/2017 10:32:00",
      "content": "<p>Thanks for the reply.  It seems that the train.csv and train_v2.csv files are not right.  There are people who are identified as churners but are still streaming music in March, after their expiration. </p>\n\n<p>I looked at the scala file and it looks like it can only be run in linux because I see references to apache and spark.  Do you know if it can be run on a Windows based machine?  </p>\n\n<p>I wonder if everyone on top of the LB is using the scala file to construct the labels......</p>",
      "rawMarkdown": "Thanks for the reply.  It seems that the train.csv and train_v2.csv files are not right.  There are people who are identified as churners but are still streaming music in March, after their expiration. \n\nI looked at the scala file and it looks like it can only be run in linux because I see references to apache and spark.  Do you know if it can be run on a Windows based machine?  \n\nI wonder if everyone on top of the LB is using the scala file to construct the labels......",
      "votes": null
    },
    {
      "id": "257521",
      "postDate": "12/14/2017 11:58:24",
      "content": "<p>Use vmware, or virtualbox to install linux. Spark is not recommended to install on windows. Or you can rent a VM from cloud computing vendors such as GCP, Amazon.</p>",
      "rawMarkdown": "Use vmware, or virtualbox to install linux. Spark is not recommended to install on windows. Or you can rent a VM from cloud computing vendors such as GCP, Amazon.",
      "votes": null
    },
    {
      "id": "257748",
      "postDate": "12/14/2017 22:11:35",
      "content": "<p>Thanks.. </p>",
      "rawMarkdown": "Thanks..",
      "votes": null
    },
    {
      "id": "260167",
      "postDate": "12/19/2017 18:55:46",
      "content": "<p>I had issues attempting to run scala on my windows development machine as well, and moved to a cloud Linux VM.  Everything went much smoother after that.  I used free Azure credits.</p>",
      "rawMarkdown": "I had issues attempting to run scala on my windows development machine as well, and moved to a cloud Linux VM.  Everything went much smoother after that.  I used free Azure credits.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 257248,
      "author_name": "aloisiodn",
      "author_url": "",
      "post_date": "12/13/2017 19:43:17",
      "content": "<p>I prefer to use the official labeller because, teoretically, it is used to generate the labels used in public and private LB. But if you are getting a better result with your own code, I think you should go ahead!</p>\n\n<p>The competition end is coming soon. Good luck to us all!!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 257480,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/14/2017 10:18:22",
          "content": "<p>Good luck tooo!</p>\n\n<p>Thanks for your reply, I also believe CV score more, but the risk is that the CV score is highly based on the labeling method...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 257350,
      "author_name": "",
      "author_url": "",
      "post_date": "12/14/2017 04:40:51",
      "content": "<p>Hi, do you find a big difference between training your model on the train.csv and train_v2.csv files versus generating the churn flag using the scala file? </p>",
      "votes": null,
      "replies": [
        {
          "id": 257475,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/14/2017 10:14:33",
          "content": "<p>Good questions, I did not use train.csv and train_v2.csv to train my model since the second day after I entered the competition..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 257484,
          "author_name": "",
          "author_url": "",
          "post_date": "12/14/2017 10:32:00",
          "content": "<p>Thanks for the reply.  It seems that the train.csv and train_v2.csv files are not right.  There are people who are identified as churners but are still streaming music in March, after their expiration. </p>\n\n<p>I looked at the scala file and it looks like it can only be run in linux because I see references to apache and spark.  Do you know if it can be run on a Windows based machine?  </p>\n\n<p>I wonder if everyone on top of the LB is using the scala file to construct the labels......</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 257521,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/14/2017 11:58:24",
          "content": "<p>Use vmware, or virtualbox to install linux. Spark is not recommended to install on windows. Or you can rent a VM from cloud computing vendors such as GCP, Amazon.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 257748,
          "author_name": "",
          "author_url": "",
          "post_date": "12/14/2017 22:11:35",
          "content": "<p>Thanks.. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260167,
          "author_name": "bryangregory",
          "author_url": "",
          "post_date": "12/19/2017 18:55:46",
          "content": "<p>I had issues attempting to run scala on my windows development machine as well, and moved to a cloud Linux VM.  Everything went much smoother after that.  I used free Azure credits.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "257156": "Currently I run the official scala code, and I got ~0.116 on LB. (~0.146 CV).<br>\nHowever, I can get ~0.108 on LB (~0.166 CV) with my own labeling method. I use same features for both result, only the training label is different.\n\nIn addition, both the label result of two label methods are different from train.csv, train_v2.csv.\n\n<br><br>\nAnyone faced the same issue? I am wondering is top teams use the official scala code or not.",
    "257248": "I prefer to use the official labeller because, teoretically, it is used to generate the labels used in public and private LB. But if you are getting a better result with your own code, I think you should go ahead!\n\nThe competition end is coming soon. Good luck to us all!!!",
    "257350": "Hi, do you find a big difference between training your model on the train.csv and train_v2.csv files versus generating the churn flag using the scala file?",
    "257475": "Good questions, I did not use train.csv and train_v2.csv to train my model since the second day after I entered the competition..",
    "257480": "Good luck tooo!\n\nThanks for your reply, I also believe CV score more, but the risk is that the CV score is highly based on the labeling method...",
    "257484": "Thanks for the reply.  It seems that the train.csv and train_v2.csv files are not right.  There are people who are identified as churners but are still streaming music in March, after their expiration. \n\nI looked at the scala file and it looks like it can only be run in linux because I see references to apache and spark.  Do you know if it can be run on a Windows based machine?  \n\nI wonder if everyone on top of the LB is using the scala file to construct the labels......",
    "257521": "Use vmware, or virtualbox to install linux. Spark is not recommended to install on windows. Or you can rent a VM from cloud computing vendors such as GCP, Amazon.",
    "257748": "Thanks..",
    "260167": "I had issues attempting to run scala on my windows development machine as well, and moved to a cloud Linux VM.  Everything went much smoother after that.  I used free Azure credits."
  },
  "source": "meta"
}