{
  "id": 45932,
  "title": "About training data",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/45932",
  "author_name": "",
  "post_date": "2017-12-18T04:24:48.676147800Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi, congrats to the winners, and thanks everyone.</p>\n\n<p>This is a interesting competition and makes me to learn many new things and restudy old things. However, did anyone find that there's a huge gap between LB scores when you use the official training data and the training data you make with local scala code?</p>\n\n<p>In my case, I use training label 201703 which generated by scala code to get 0.13443 private LB score, but 0.15708 with train_v2.csv. The demo code is now <a href=\"https://www.kaggle.com/infinitewing/lightgbm-using-scala-label-201703\">posted here</a>.</p>\n\n<p>The scala generated label is now <a href=\"https://www.kaggle.com/infinitewing/kkbox-churn-scala-label\">available here</a></p>",
  "messages": [
    {
      "id": "259290",
      "postDate": "12/18/2017 04:24:48",
      "content": "<p>Hi, congrats to the winners, and thanks everyone.</p>\n\n<p>This is a interesting competition and makes me to learn many new things and restudy old things. However, did anyone find that there's a huge gap between LB scores when you use the official training data and the training data you make with local scala code?</p>\n\n<p>In my case, I use training label 201703 which generated by scala code to get 0.13443 private LB score, but 0.15708 with train_v2.csv. The demo code is now <a href=\"https://www.kaggle.com/infinitewing/lightgbm-using-scala-label-201703\">posted here</a>.</p>\n\n<p>The scala generated label is now <a href=\"https://www.kaggle.com/infinitewing/kkbox-churn-scala-label\">available here</a></p>",
      "rawMarkdown": "Hi, congrats to the winners, and thanks everyone.\n\nThis is a interesting competition and makes me to learn many new things and restudy old things. However, did anyone find that there's a huge gap between LB scores when you use the official training data and the training data you make with local scala code?\n\nIn my case, I use training label 201703 which generated by scala code to get 0.13443 private LB score, but 0.15708 with train_v2.csv. The demo code is now [posted here](https://www.kaggle.com/infinitewing/lightgbm-using-scala-label-201703).\n\nThe scala generated label is now [available here](https://www.kaggle.com/infinitewing/kkbox-churn-scala-label)",
      "votes": null
    },
    {
      "id": "259302",
      "postDate": "12/18/2017 05:08:25",
      "content": "<p>We generated monthly 'is_churn' labels by given scala code, and the public LB score improved from 0.113 to 0.102(equivalent to private LB score 0.10).</p>",
      "rawMarkdown": "We generated monthly 'is_churn' labels by given scala code, and the public LB score improved from 0.113 to 0.102(equivalent to private LB score 0.10).",
      "votes": null
    },
    {
      "id": "259307",
      "postDate": "12/18/2017 05:14:30",
      "content": "<p>That's interesting, my model can also get ~0.113 by using scala label. However by using the label generated by myself, I can get ~0.0108. </p>\n\n<p>So I am also wondering how did the test label been generated.</p>\n\n<p>By the way, can you check did <a href=\"https://www.kaggle.com/infinitewing/kkbox-churn-scala-label\">my scala label result</a> is same as you generated?</p>",
      "rawMarkdown": "That's interesting, my model can also get ~0.113 by using scala label. However by using the label generated by myself, I can get ~0.0108. \n\nSo I am also wondering how did the test label been generated.\n\nBy the way, can you check did [my scala label result](https://www.kaggle.com/infinitewing/kkbox-churn-scala-label) is same as you generated?",
      "votes": null
    },
    {
      "id": "259480",
      "postDate": "12/18/2017 13:39:48",
      "content": "<p>Yes I checked, and it is exactly same. (Below code)</p>\n\n<p>2017-03 scala label is unreliable, since we do not have 2017-04 transaction data, which is needed to calculate the gap between the 2017-03 membership expire date and 2017-04 transaction date. So for 2017-03, we substitute existing train_v2.csv instead of using scala label. By the way, what is your rule of generating labels? How is it different from scala code?</p>\n\n<hr>\n\n<pre><code>##  is_churn_x: label generated by me, is_churn_y: label generated by you\n&lt;2017-02 data&gt;\n(train2['is_churn_x']-train2['is_churn_y']).value_counts()     \n0    879537\ndtype: int64\n\n&lt;2017-03 data&gt;\n(train2['is_churn_x']-train2['is_churn_y']).value_counts()\n0    886500\ndtype: int64\n</code></pre>",
      "rawMarkdown": "Yes I checked, and it is exactly same. (Below code)\n\n2017-03 scala label is unreliable, since we do not have 2017-04 transaction data, which is needed to calculate the gap between the 2017-03 membership expire date and 2017-04 transaction date. So for 2017-03, we substitute existing train_v2.csv instead of using scala label. By the way, what is your rule of generating labels? How is it different from scala code?\n\n-----------------------\n    ##  is_churn_x: label generated by me, is_churn_y: label generated by you\n    &lt;2017-02 data&gt;\n    (train2['is_churn_x']-train2['is_churn_y']).value_counts()     \n    0    879537\n    dtype: int64\n    \n    &lt;2017-03 data&gt;\n    (train2['is_churn_x']-train2['is_churn_y']).value_counts()\n    0    886500\n    dtype: int64",
      "votes": null
    },
    {
      "id": "259856",
      "postDate": "12/19/2017 04:33:07",
      "content": "<p>InfiniteWing.  Thanks for posting your labels.  I really appreciate it.   </p>\n\n<p>I don't understand why the scala labels (based on your 201703 file) that you produced are so different than what is in the training_v2.csv file.  Do you know why?  I started a discussion topic to understand why.</p>\n\n<p>Thanks again for posting the file.....</p>",
      "rawMarkdown": "InfiniteWing.  Thanks for posting your labels.  I really appreciate it.   \n\nI don't understand why the scala labels (based on your 201703 file) that you produced are so different than what is in the training_v2.csv file.  Do you know why?  I started a discussion topic to understand why.\n\nThanks again for posting the file.....",
      "votes": null
    },
    {
      "id": "259971",
      "postDate": "12/19/2017 10:27:01",
      "content": "<p>I have no idea, so I am wondering how did the test label been generated.</p>",
      "rawMarkdown": "I have no idea, so I am wondering how did the test label been generated.",
      "votes": null
    },
    {
      "id": "363883",
      "postDate": "07/30/2018 08:48:08",
      "content": "<p>Thanks, this is very useful to understanding just what went on here. A shame the competition hosts never updated the data with the correct labels (or answered this, presumably due to embarrassment after the previous target leak issues).</p>",
      "rawMarkdown": "Thanks, this is very useful to understanding just what went on here. A shame the competition hosts never updated the data with the correct labels (or answered this, presumably due to embarrassment after the previous target leak issues).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 259302,
      "author_name": "sundong",
      "author_url": "",
      "post_date": "12/18/2017 05:08:25",
      "content": "<p>We generated monthly 'is_churn' labels by given scala code, and the public LB score improved from 0.113 to 0.102(equivalent to private LB score 0.10).</p>",
      "votes": null,
      "replies": [
        {
          "id": 259307,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/18/2017 05:14:30",
          "content": "<p>That's interesting, my model can also get ~0.113 by using scala label. However by using the label generated by myself, I can get ~0.0108. </p>\n\n<p>So I am also wondering how did the test label been generated.</p>\n\n<p>By the way, can you check did <a href=\"https://www.kaggle.com/infinitewing/kkbox-churn-scala-label\">my scala label result</a> is same as you generated?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259480,
          "author_name": "sundong",
          "author_url": "",
          "post_date": "12/18/2017 13:39:48",
          "content": "<p>Yes I checked, and it is exactly same. (Below code)</p>\n\n<p>2017-03 scala label is unreliable, since we do not have 2017-04 transaction data, which is needed to calculate the gap between the 2017-03 membership expire date and 2017-04 transaction date. So for 2017-03, we substitute existing train_v2.csv instead of using scala label. By the way, what is your rule of generating labels? How is it different from scala code?</p>\n\n<hr>\n\n<pre><code>##  is_churn_x: label generated by me, is_churn_y: label generated by you\n&lt;2017-02 data&gt;\n(train2['is_churn_x']-train2['is_churn_y']).value_counts()     \n0    879537\ndtype: int64\n\n&lt;2017-03 data&gt;\n(train2['is_churn_x']-train2['is_churn_y']).value_counts()\n0    886500\ndtype: int64\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 259856,
      "author_name": "",
      "author_url": "",
      "post_date": "12/19/2017 04:33:07",
      "content": "<p>InfiniteWing.  Thanks for posting your labels.  I really appreciate it.   </p>\n\n<p>I don't understand why the scala labels (based on your 201703 file) that you produced are so different than what is in the training_v2.csv file.  Do you know why?  I started a discussion topic to understand why.</p>\n\n<p>Thanks again for posting the file.....</p>",
      "votes": null,
      "replies": [
        {
          "id": 259971,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/19/2017 10:27:01",
          "content": "<p>I have no idea, so I am wondering how did the test label been generated.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 363883,
      "author_name": "danofer",
      "author_url": "",
      "post_date": "07/30/2018 08:48:08",
      "content": "<p>Thanks, this is very useful to understanding just what went on here. A shame the competition hosts never updated the data with the correct labels (or answered this, presumably due to embarrassment after the previous target leak issues).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "259290": "Hi, congrats to the winners, and thanks everyone.\n\nThis is a interesting competition and makes me to learn many new things and restudy old things. However, did anyone find that there's a huge gap between LB scores when you use the official training data and the training data you make with local scala code?\n\nIn my case, I use training label 201703 which generated by scala code to get 0.13443 private LB score, but 0.15708 with train_v2.csv. The demo code is now [posted here](https://www.kaggle.com/infinitewing/lightgbm-using-scala-label-201703).\n\nThe scala generated label is now [available here](https://www.kaggle.com/infinitewing/kkbox-churn-scala-label)",
    "259302": "We generated monthly 'is_churn' labels by given scala code, and the public LB score improved from 0.113 to 0.102(equivalent to private LB score 0.10).",
    "259307": "That's interesting, my model can also get ~0.113 by using scala label. However by using the label generated by myself, I can get ~0.0108. \n\nSo I am also wondering how did the test label been generated.\n\nBy the way, can you check did [my scala label result](https://www.kaggle.com/infinitewing/kkbox-churn-scala-label) is same as you generated?",
    "259480": "Yes I checked, and it is exactly same. (Below code)\n\n2017-03 scala label is unreliable, since we do not have 2017-04 transaction data, which is needed to calculate the gap between the 2017-03 membership expire date and 2017-04 transaction date. So for 2017-03, we substitute existing train_v2.csv instead of using scala label. By the way, what is your rule of generating labels? How is it different from scala code?\n\n-----------------------\n    ##  is_churn_x: label generated by me, is_churn_y: label generated by you\n    &lt;2017-02 data&gt;\n    (train2['is_churn_x']-train2['is_churn_y']).value_counts()     \n    0    879537\n    dtype: int64\n    \n    &lt;2017-03 data&gt;\n    (train2['is_churn_x']-train2['is_churn_y']).value_counts()\n    0    886500\n    dtype: int64",
    "259856": "InfiniteWing.  Thanks for posting your labels.  I really appreciate it.   \n\nI don't understand why the scala labels (based on your 201703 file) that you produced are so different than what is in the training_v2.csv file.  Do you know why?  I started a discussion topic to understand why.\n\nThanks again for posting the file.....",
    "259971": "I have no idea, so I am wondering how did the test label been generated.",
    "363883": "Thanks, this is very useful to understanding just what went on here. A shame the competition hosts never updated the data with the correct labels (or answered this, presumably due to embarrassment after the previous target leak issues)."
  },
  "source": "meta"
}