{
  "id": 56250,
  "title": "11th place features",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/writeups/tkm2261-11th-place-features",
  "author_name": "",
  "post_date": "2018-05-16T07:27:12.650Z",
  "votes": 98,
  "comment_count": 35,
  "views": 0,
  "content": "<p>Congrats winners!</p>\n\n<p>I would like to share my features briefly before I forget it. </p>\n\n<h3>Features</h3>\n\n<p>I used BigQuery on the almost all feature engineering parts. <br>\n <a href=\"https://gist.github.com/tkm2261/1b3c3c37753e55ed2914577c0f96d222\">https://gist.github.com/tkm2261/1b3c3c37753e55ed2914577c0f96d222</a></p>\n\n<p>It takes only about 20 minutes. very fast!</p>\n\n<p>There may not be surprising features if you have seen the shared kernels. I tried to find good features by brute force way.</p>\n\n<h3>Machine</h3>\n\n<p>CPU: 16 cores, MEM: 100GB on GCP</p>\n\n<h3>Data</h3>\n\n<ul>\n<li>training data: day 7 and day 8 (13,188,695,398 rows)</li>\n<li>validation data: day 9 (53,016,937 rows)</li>\n<li>test data: day 10 (18,790,469 rows)</li>\n</ul>\n\n<p>I used validation data to determin the boosting round. Then, I used train+valid data for training.</p>\n\n<h3>Training</h3>\n\n<p>Simply, I used LightGBM and ensembled with different seeds.</p>\n\n<p>The single best model is 0.9823 on the public LB.</p>\n\n<p>The parameter is here: </p>\n\n<pre><code>{'colsample_bytree': 0.6, 'learning_rate': 0.1, 'max_bin': 1023, 'max_depth': -1, 'metric': ['binary_logloss', 'auc'], 'min_child_weight': 30, 'min_split_gain': 0.0001, 'num_leaves': 127, 'objective': 'binary', 'reg_alpha': 0, 'scale_pos_weight': 1, 'seed': 1142, 'subsample': 0.9, 'subsample_freq': 1, 'verbose': -1}\n</code></pre>\n\n<p>source code is here: <a href=\"https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/train_lgb.py\">https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/train_lgb.py</a></p>\n\n<p>I tried to earn my tuition for MS CS program starting next fall. But, there was no free lunch!</p>",
  "messages": [
    {
      "id": "324921",
      "postDate": "05/08/2018 01:40:19",
      "content": "<p>Congrats winners!</p>\n\n<p>I would like to share my features briefly before I forget it. </p>\n\n<h3>Features</h3>\n\n<p>I used BigQuery on the almost all feature engineering parts. <br>\n <a href=\"https://gist.github.com/tkm2261/1b3c3c37753e55ed2914577c0f96d222\">https://gist.github.com/tkm2261/1b3c3c37753e55ed2914577c0f96d222</a></p>\n\n<p>It takes only about 20 minutes. very fast!</p>\n\n<p>There may not be surprising features if you have seen the shared kernels. I tried to find good features by brute force way.</p>\n\n<h3>Machine</h3>\n\n<p>CPU: 16 cores, MEM: 100GB on GCP</p>\n\n<h3>Data</h3>\n\n<ul>\n<li>training data: day 7 and day 8 (13,188,695,398 rows)</li>\n<li>validation data: day 9 (53,016,937 rows)</li>\n<li>test data: day 10 (18,790,469 rows)</li>\n</ul>\n\n<p>I used validation data to determin the boosting round. Then, I used train+valid data for training.</p>\n\n<h3>Training</h3>\n\n<p>Simply, I used LightGBM and ensembled with different seeds.</p>\n\n<p>The single best model is 0.9823 on the public LB.</p>\n\n<p>The parameter is here: </p>\n\n<pre><code>{'colsample_bytree': 0.6, 'learning_rate': 0.1, 'max_bin': 1023, 'max_depth': -1, 'metric': ['binary_logloss', 'auc'], 'min_child_weight': 30, 'min_split_gain': 0.0001, 'num_leaves': 127, 'objective': 'binary', 'reg_alpha': 0, 'scale_pos_weight': 1, 'seed': 1142, 'subsample': 0.9, 'subsample_freq': 1, 'verbose': -1}\n</code></pre>\n\n<p>source code is here: <a href=\"https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/train_lgb.py\">https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/train_lgb.py</a></p>\n\n<p>I tried to earn my tuition for MS CS program starting next fall. But, there was no free lunch!</p>",
      "rawMarkdown": "Congrats winners!\n\nI would like to share my features briefly before I forget it. \n\n### Features\n\nI used BigQuery on the almost all feature engineering parts.  \n https://gist.github.com/tkm2261/1b3c3c37753e55ed2914577c0f96d222\n\nIt takes only about 20 minutes. very fast!\n\nThere may not be surprising features if you have seen the shared kernels. I tried to find good features by brute force way.\n\n### Machine\n\nCPU: 16 cores, MEM: 100GB on GCP\n\n### Data\n\n* training data: day 7 and day 8 (13,188,695,398 rows)\n* validation data: day 9 (53,016,937 rows)\n* test data: day 10 (18,790,469 rows)\n\nI used validation data to determin the boosting round. Then, I used train+valid data for training.\n\n### Training\n\nSimply, I used LightGBM and ensembled with different seeds.\n\nThe single best model is 0.9823 on the public LB.\n\nThe parameter is here: \n\n    {'colsample_bytree': 0.6, 'learning_rate': 0.1, 'max_bin': 1023, 'max_depth': -1, 'metric': ['binary_logloss', 'auc'], 'min_child_weight': 30, 'min_split_gain': 0.0001, 'num_leaves': 127, 'objective': 'binary', 'reg_alpha': 0, 'scale_pos_weight': 1, 'seed': 1142, 'subsample': 0.9, 'subsample_freq': 1, 'verbose': -1}\n\nsource code is here: https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/train_lgb.py\n\nI tried to earn my tuition for MS CS program starting next fall. But, there was no free lunch!",
      "votes": null
    },
    {
      "id": "324925",
      "postDate": "05/08/2018 01:42:23",
      "content": "<p>Good game and well played!</p>",
      "rawMarkdown": "Good game and well played!",
      "votes": null
    },
    {
      "id": "324929",
      "postDate": "05/08/2018 01:44:21",
      "content": "<p>Congrats and thanks for sharing~</p>",
      "rawMarkdown": "Congrats and thanks for sharing~",
      "votes": null
    },
    {
      "id": "324936",
      "postDate": "05/08/2018 01:54:36",
      "content": "<p>Cool! Thanks for sharing! I hope you do well in your MS program!</p>",
      "rawMarkdown": "Cool! Thanks for sharing! I hope you do well in your MS program!",
      "votes": null
    },
    {
      "id": "324943",
      "postDate": "05/08/2018 02:01:50",
      "content": "<p>Congrats and thanks for sharing. so first step is make train csv file to sql dataset.</p>",
      "rawMarkdown": "Congrats and thanks for sharing. so first step is make train csv file to sql dataset.",
      "votes": null
    },
    {
      "id": "324949",
      "postDate": "05/08/2018 02:06:15",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "324957",
      "postDate": "05/08/2018 02:15:32",
      "content": "<p>Congrats and thanks for sharing! <br>\nAnd may I ask here <br>\n<code>\nI used validation data to determin the boosting round. Then, I used train+valid data for training.\n</code> <br>\nThat means you use the same parameters and rounds you get in validation to retrain the model without using validation in retraining, is that right? <br>\nTHANKS.</p>",
      "rawMarkdown": "Congrats and thanks for sharing!   \nAnd may I ask here    \n```\nI used validation data to determin the boosting round. Then, I used train+valid data for training.\n```  \nThat means you use the same parameters and rounds you get in validation to retrain the model without using validation in retraining, is that right?  \nTHANKS.",
      "votes": null
    },
    {
      "id": "324962",
      "postDate": "05/08/2018 02:17:55",
      "content": "<p>'scale_pos_weight': 1 ? I have not seen anyone using such a small number.\nCongratulation on your accomplishment!</p>",
      "rawMarkdown": "'scale_pos_weight': 1 ? I have not seen anyone using such a small number.\nCongratulation on your accomplishment!",
      "votes": null
    },
    {
      "id": "324966",
      "postDate": "05/08/2018 02:22:18",
      "content": "<p>I mean the validation data is used like fold-out data.</p>",
      "rawMarkdown": "I mean the validation data is used like fold-out data.",
      "votes": null
    },
    {
      "id": "324975",
      "postDate": "05/08/2018 02:29:33",
      "content": "<p>I handle the imbalance by min_sample_weight, num_leaves, and other parameters. Since the eval metric is AUC, the scale is not important.</p>",
      "rawMarkdown": "I handle the imbalance by min_sample_weight, num_leaves, and other parameters. Since the eval metric is AUC, the scale is not important.",
      "votes": null
    },
    {
      "id": "324978",
      "postDate": "05/08/2018 02:31:31",
      "content": "<p>Thank you for sharing! I'll try BigQuery</p>",
      "rawMarkdown": "Thank you for sharing! I'll try BigQuery",
      "votes": null
    },
    {
      "id": "325006",
      "postDate": "05/08/2018 03:16:02",
      "content": "<p>Nice Work!</p>",
      "rawMarkdown": "Nice Work!",
      "votes": null
    },
    {
      "id": "325018",
      "postDate": "05/08/2018 03:23:41",
      "content": "<p>just wondering about whether using validation(or early_stopping) in retraining or not. thx</p>",
      "rawMarkdown": "just wondering about whether using validation(or early_stopping) in retraining or not. thx",
      "votes": null
    },
    {
      "id": "325025",
      "postDate": "05/08/2018 03:30:07",
      "content": "<p>haha, i am going for MS CS , in the coming fall ! how did you come up with the hyperparameters </p>",
      "rawMarkdown": "haha, i am going for MS CS , in the coming fall ! how did you come up with the hyperparameters",
      "votes": null
    },
    {
      "id": "325028",
      "postDate": "05/08/2018 03:33:03",
      "content": "<p>One quick question: do you refit on the whole set after get the best number of trees from the validation?</p>",
      "rawMarkdown": "One quick question: do you refit on the whole set after get the best number of trees from the validation?",
      "votes": null
    },
    {
      "id": "325034",
      "postDate": "05/08/2018 03:42:59",
      "content": "<p>Wow that's so fast to build all those features, thanks for sharing &amp; congrats! I never knew it can be so much faster using BiqQuery, I only use VM instances for this competittion.</p>",
      "rawMarkdown": "Wow that's so fast to build all those features, thanks for sharing &amp; congrats! I never knew it can be so much faster using BiqQuery, I only use VM instances for this competittion.",
      "votes": null
    },
    {
      "id": "325061",
      "postDate": "05/08/2018 04:38:00",
      "content": "<p>Brilliant, \nThanks for sharing. </p>",
      "rawMarkdown": "Brilliant, \nThanks for sharing.",
      "votes": null
    },
    {
      "id": "325090",
      "postDate": "05/08/2018 05:30:14",
      "content": "<p>great solution! most suitable for production</p>",
      "rawMarkdown": "great solution! most suitable for production",
      "votes": null
    },
    {
      "id": "325094",
      "postDate": "05/08/2018 05:44:19",
      "content": "<p>Congrats! Vote for the \"next_click\" part, you've done way more than I did. Thanks for sharing.</p>",
      "rawMarkdown": "Congrats! Vote for the \"next_click\" part, you've done way more than I did. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "325097",
      "postDate": "05/08/2018 05:48:38",
      "content": "<p>Could you tell a bit about tuning of the hyper-parameters, if you did any?</p>",
      "rawMarkdown": "Could you tell a bit about tuning of the hyper-parameters, if you did any?",
      "votes": null
    },
    {
      "id": "325099",
      "postDate": "05/08/2018 05:54:12",
      "content": "<p>yes, I refit.</p>",
      "rawMarkdown": "yes, I refit.",
      "votes": null
    },
    {
      "id": "325100",
      "postDate": "05/08/2018 05:54:57",
      "content": "<p>I used the validation data in retraining.</p>",
      "rawMarkdown": "I used the validation data in retraining.",
      "votes": null
    },
    {
      "id": "325123",
      "postDate": "05/08/2018 06:30:33",
      "content": "<p>Congrats. <br> Did you skip day 8 from your training set?</p>",
      "rawMarkdown": "Congrats. <br> Did you skip day 8 from your training set?",
      "votes": null
    },
    {
      "id": "325124",
      "postDate": "05/08/2018 06:32:10",
      "content": "<p>sorry, day 8 is included in my training data.</p>",
      "rawMarkdown": "sorry, day 8 is included in my training data.",
      "votes": null
    },
    {
      "id": "325137",
      "postDate": "05/08/2018 06:51:37",
      "content": "<p>May I ask if you made features for each day separately? Like below</p>\n\n<pre><code>day7 = make_features(day7)\n...\n...\nday9 = make_features(day9)\n\nX_train = pd.concat([day7,day8])\nX_val = day9[hours[4,11]]\n</code></pre>\n\n<p>or</p>\n\n<pre><code>all_days = make_features(all_days)\nX_train = all_days[days[7,8]]\nX_val = all_days[day[9]]\n</code></pre>",
      "rawMarkdown": "May I ask if you made features for each day separately? Like below\n\n    day7 = make_features(day7)\n    ...\n    ...\n    day9 = make_features(day9)\n    \n    X_train = pd.concat([day7,day8])\n    X_val = day9[hours[4,11]]\n\nor\n\n    all_days = make_features(all_days)\n    X_train = all_days[days[7,8]]\n    X_val = all_days[day[9]]",
      "votes": null
    },
    {
      "id": "325150",
      "postDate": "05/08/2018 07:12:47",
      "content": "<p>the latter is correct. plz see my sql gist for more detail.</p>",
      "rawMarkdown": "the latter is correct. plz see my sql gist for more detail.",
      "votes": null
    },
    {
      "id": "325190",
      "postDate": "05/08/2018 07:55:43",
      "content": "<p>Something still confuses me, so I try to make it clear. <br>\n1.You use train data and valid data to train models-&gt;find the best parameters and rounds. <br>\n2.You use the parameters and rounds in 1 and the whole dataset to retrain a model, in this process you won't split the whole dataset for validation, and just train the number of rounds in 1 without using early stopping. <br>\nAm I get the right thing here?</p>",
      "rawMarkdown": "Something still confuses me, so I try to make it clear.  \n1.You use train data and valid data to train models-&gt;find the best parameters and rounds.  \n2.You use the parameters and rounds in 1 and the whole dataset to retrain a model, in this process you won't split the whole dataset for validation, and just train the number of rounds in 1 without using early stopping.  \nAm I get the right thing here?",
      "votes": null
    },
    {
      "id": "325223",
      "postDate": "05/08/2018 08:36:41",
      "content": "<p>maybe right. if you want to learn more, please see my code: <a href=\"https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/train_lgb.py\">https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/train_lgb.py</a></p>",
      "rawMarkdown": "maybe right. if you want to learn more, please see my code: https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/train_lgb.py",
      "votes": null
    },
    {
      "id": "325245",
      "postDate": "05/08/2018 08:56:16",
      "content": "<p>Cool, I think I got what I want in your code, many thanks.</p>",
      "rawMarkdown": "Cool, I think I got what I want in your code, many thanks.",
      "votes": null
    },
    {
      "id": "325248",
      "postDate": "05/08/2018 08:59:36",
      "content": "<p>Well done, you have been the lead solo competitor till the last few days if I remember well, chasing you during the competition helped me a lot!</p>",
      "rawMarkdown": "Well done, you have been the lead solo competitor till the last few days if I remember well, chasing you during the competition helped me a lot!",
      "votes": null
    },
    {
      "id": "325285",
      "postDate": "05/08/2018 09:22:47",
      "content": "<p>Well done tkm2261!\nAfter reading all post, I realized I completely missed that it was valuable to split data before and after day 9 and 10. \nAnd thanks for sharing the code! tkm2261, I am surprised changing the seeds plays a role. I will try in the next competition. And good luck for your MS.</p>",
      "rawMarkdown": "Well done tkm2261!\nAfter reading all post, I realized I completely missed that it was valuable to split data before and after day 9 and 10. \nAnd thanks for sharing the code! tkm2261, I am surprised changing the seeds plays a role. I will try in the next competition. And good luck for your MS.",
      "votes": null
    },
    {
      "id": "329250",
      "postDate": "05/16/2018 05:03:39",
      "content": "<p>Thanks for sharing. 2 questions regarding to your solution:</p>\n\n<ol>\n<li>All data before day 9 should have 131886953 rows and all day 9 data should have 53016937 rows. In your post looks like they have ~ 2 million rows less, did you filter out some data? Duplication? </li>\n<li>In your load data code <a href=\"https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/load_data.py#L43\">https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/load_data.py#L43</a> what does \"prev\" data set refer to? Looks like it is not part of your queries. </li>\n</ol>\n\n<p>Thanks</p>",
      "rawMarkdown": "Thanks for sharing. 2 questions regarding to your solution:\n\n1. All data before day 9 should have 131886953 rows and all day 9 data should have 53016937 rows. In your post looks like they have ~ 2 million rows less, did you filter out some data? Duplication? \n2. In your load data code https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/load_data.py#L43 what does \"prev\" data set refer to? Looks like it is not part of your queries. \n\nThanks",
      "votes": null
    },
    {
      "id": "329305",
      "postDate": "05/16/2018 07:35:48",
      "content": "<p>Thanks for questioning! <br>\n1. You are right. I copied these numbers from the log when I deleted duplicated rows by mistake. And, deleting duplicated rows did not improve my score. <br>\n2.  At first, I used only day 8 data as train data. Then, I added day 7 data that I called \"prev data\". So, in my query, prev and train are merged.</p>",
      "rawMarkdown": "Thanks for questioning!  \n1. You are right. I copied these numbers from the log when I deleted duplicated rows by mistake. And, deleting duplicated rows did not improve my score.  \n2.  At first, I used only day 8 data as train data. Then, I added day 7 data that I called \"prev data\". So, in my query, prev and train are merged.",
      "votes": null
    },
    {
      "id": "329307",
      "postDate": "05/16/2018 07:42:17",
      "content": "<p>Gotcha! Thanks for answering these questions. I'm on the half way of reproducing your best single model. </p>\n\n<p>I realized you excluded several features from BigQuery results(<a href=\"https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/load_data.py#L10\">https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/load_data.py#L10</a>), how did you do the feature selection? I'm also curious how you find the best param of lightGBM, from your code, looks like you did some manual tuning? </p>",
      "rawMarkdown": "Gotcha! Thanks for answering these questions. I'm on the half way of reproducing your best single model. \n\nI realized you excluded several features from BigQuery results(https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/load_data.py#L10), how did you do the feature selection? I'm also curious how you find the best param of lightGBM, from your code, looks like you did some manual tuning?",
      "votes": null
    },
    {
      "id": "329316",
      "postDate": "05/16/2018 08:11:07",
      "content": "<ol>\n<li>I left top about 50 features ordered by the feature importance of lightgbm when I made new features. I sometimes dropped features in python code.</li>\n<li>For parameter tuning, I used day 8 data as train data and day 9 data as validation data. Since training only  day 8 data is not relatively heavy, I could try about 20 parameters sets. But, I did not devote much effort to the parameter tuning.</li>\n</ol>",
      "rawMarkdown": "1. I left top about 50 features ordered by the feature importance of lightgbm when I made new features. I sometimes dropped features in python code.\n2. For parameter tuning, I used day 8 data as train data and day 9 data as validation data. Since training only  day 8 data is not relatively heavy, I could try about 20 parameters sets. But, I did not devote much effort to the parameter tuning.",
      "votes": null
    },
    {
      "id": "329530",
      "postDate": "05/16/2018 16:45:02",
      "content": "<p>Thanks for sharing. Really helpful.</p>",
      "rawMarkdown": "Thanks for sharing. Really helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 324925,
      "author_name": "wythhh",
      "author_url": "",
      "post_date": "05/08/2018 01:42:23",
      "content": "<p>Good game and well played!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324929,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "05/08/2018 01:44:21",
      "content": "<p>Congrats and thanks for sharing~</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324936,
      "author_name": "kenkoooo",
      "author_url": "",
      "post_date": "05/08/2018 01:54:36",
      "content": "<p>Cool! Thanks for sharing! I hope you do well in your MS program!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324943,
      "author_name": "liuhdsgoal",
      "author_url": "",
      "post_date": "05/08/2018 02:01:50",
      "content": "<p>Congrats and thanks for sharing. so first step is make train csv file to sql dataset.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324949,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "05/08/2018 02:06:15",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324957,
      "author_name": "creamiracle",
      "author_url": "",
      "post_date": "05/08/2018 02:15:32",
      "content": "<p>Congrats and thanks for sharing! <br>\nAnd may I ask here <br>\n<code>\nI used validation data to determin the boosting round. Then, I used train+valid data for training.\n</code> <br>\nThat means you use the same parameters and rounds you get in validation to retrain the model without using validation in retraining, is that right? <br>\nTHANKS.</p>",
      "votes": null,
      "replies": [
        {
          "id": 324966,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "05/08/2018 02:22:18",
          "content": "<p>I mean the validation data is used like fold-out data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325018,
          "author_name": "creamiracle",
          "author_url": "",
          "post_date": "05/08/2018 03:23:41",
          "content": "<p>just wondering about whether using validation(or early_stopping) in retraining or not. thx</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325100,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "05/08/2018 05:54:57",
          "content": "<p>I used the validation data in retraining.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325190,
          "author_name": "creamiracle",
          "author_url": "",
          "post_date": "05/08/2018 07:55:43",
          "content": "<p>Something still confuses me, so I try to make it clear. <br>\n1.You use train data and valid data to train models-&gt;find the best parameters and rounds. <br>\n2.You use the parameters and rounds in 1 and the whole dataset to retrain a model, in this process you won't split the whole dataset for validation, and just train the number of rounds in 1 without using early stopping. <br>\nAm I get the right thing here?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325223,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "05/08/2018 08:36:41",
          "content": "<p>maybe right. if you want to learn more, please see my code: <a href=\"https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/train_lgb.py\">https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/train_lgb.py</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325245,
          "author_name": "creamiracle",
          "author_url": "",
          "post_date": "05/08/2018 08:56:16",
          "content": "<p>Cool, I think I got what I want in your code, many thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324962,
      "author_name": "tapioca",
      "author_url": "",
      "post_date": "05/08/2018 02:17:55",
      "content": "<p>'scale_pos_weight': 1 ? I have not seen anyone using such a small number.\nCongratulation on your accomplishment!</p>",
      "votes": null,
      "replies": [
        {
          "id": 324975,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "05/08/2018 02:29:33",
          "content": "<p>I handle the imbalance by min_sample_weight, num_leaves, and other parameters. Since the eval metric is AUC, the scale is not important.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324978,
      "author_name": "takpme",
      "author_url": "",
      "post_date": "05/08/2018 02:31:31",
      "content": "<p>Thank you for sharing! I'll try BigQuery</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325006,
      "author_name": "fufufukakaka",
      "author_url": "",
      "post_date": "05/08/2018 03:16:02",
      "content": "<p>Nice Work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325025,
      "author_name": "nick7hill",
      "author_url": "",
      "post_date": "05/08/2018 03:30:07",
      "content": "<p>haha, i am going for MS CS , in the coming fall ! how did you come up with the hyperparameters </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325028,
      "author_name": "chengju",
      "author_url": "",
      "post_date": "05/08/2018 03:33:03",
      "content": "<p>One quick question: do you refit on the whole set after get the best number of trees from the validation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 325099,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "05/08/2018 05:54:12",
          "content": "<p>yes, I refit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325034,
      "author_name": "sdarmawan",
      "author_url": "",
      "post_date": "05/08/2018 03:42:59",
      "content": "<p>Wow that's so fast to build all those features, thanks for sharing &amp; congrats! I never knew it can be so much faster using BiqQuery, I only use VM instances for this competittion.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325061,
      "author_name": "a45632",
      "author_url": "",
      "post_date": "05/08/2018 04:38:00",
      "content": "<p>Brilliant, \nThanks for sharing. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325090,
      "author_name": "izmaylov",
      "author_url": "",
      "post_date": "05/08/2018 05:30:14",
      "content": "<p>great solution! most suitable for production</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325094,
      "author_name": "niuddd",
      "author_url": "",
      "post_date": "05/08/2018 05:44:19",
      "content": "<p>Congrats! Vote for the \"next_click\" part, you've done way more than I did. Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325097,
      "author_name": "burhanusm",
      "author_url": "",
      "post_date": "05/08/2018 05:48:38",
      "content": "<p>Could you tell a bit about tuning of the hyper-parameters, if you did any?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325123,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "05/08/2018 06:30:33",
      "content": "<p>Congrats. <br> Did you skip day 8 from your training set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 325124,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "05/08/2018 06:32:10",
          "content": "<p>sorry, day 8 is included in my training data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325137,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "05/08/2018 06:51:37",
          "content": "<p>May I ask if you made features for each day separately? Like below</p>\n\n<pre><code>day7 = make_features(day7)\n...\n...\nday9 = make_features(day9)\n\nX_train = pd.concat([day7,day8])\nX_val = day9[hours[4,11]]\n</code></pre>\n\n<p>or</p>\n\n<pre><code>all_days = make_features(all_days)\nX_train = all_days[days[7,8]]\nX_val = all_days[day[9]]\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325150,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "05/08/2018 07:12:47",
          "content": "<p>the latter is correct. plz see my sql gist for more detail.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325248,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/08/2018 08:59:36",
      "content": "<p>Well done, you have been the lead solo competitor till the last few days if I remember well, chasing you during the competition helped me a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325285,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "05/08/2018 09:22:47",
      "content": "<p>Well done tkm2261!\nAfter reading all post, I realized I completely missed that it was valuable to split data before and after day 9 and 10. \nAnd thanks for sharing the code! tkm2261, I am surprised changing the seeds plays a role. I will try in the next competition. And good luck for your MS.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 329250,
      "author_name": "soundwaveli00",
      "author_url": "",
      "post_date": "05/16/2018 05:03:39",
      "content": "<p>Thanks for sharing. 2 questions regarding to your solution:</p>\n\n<ol>\n<li>All data before day 9 should have 131886953 rows and all day 9 data should have 53016937 rows. In your post looks like they have ~ 2 million rows less, did you filter out some data? Duplication? </li>\n<li>In your load data code <a href=\"https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/load_data.py#L43\">https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/load_data.py#L43</a> what does \"prev\" data set refer to? Looks like it is not part of your queries. </li>\n</ol>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 329305,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "05/16/2018 07:35:48",
          "content": "<p>Thanks for questioning! <br>\n1. You are right. I copied these numbers from the log when I deleted duplicated rows by mistake. And, deleting duplicated rows did not improve my score. <br>\n2.  At first, I used only day 8 data as train data. Then, I added day 7 data that I called \"prev data\". So, in my query, prev and train are merged.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 329307,
          "author_name": "soundwaveli00",
          "author_url": "",
          "post_date": "05/16/2018 07:42:17",
          "content": "<p>Gotcha! Thanks for answering these questions. I'm on the half way of reproducing your best single model. </p>\n\n<p>I realized you excluded several features from BigQuery results(<a href=\"https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/load_data.py#L10\">https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/load_data.py#L10</a>), how did you do the feature selection? I'm also curious how you find the best param of lightGBM, from your code, looks like you did some manual tuning? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 329316,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "05/16/2018 08:11:07",
          "content": "<ol>\n<li>I left top about 50 features ordered by the feature importance of lightgbm when I made new features. I sometimes dropped features in python code.</li>\n<li>For parameter tuning, I used day 8 data as train data and day 9 data as validation data. Since training only  day 8 data is not relatively heavy, I could try about 20 parameters sets. But, I did not devote much effort to the parameter tuning.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 329530,
          "author_name": "soundwaveli00",
          "author_url": "",
          "post_date": "05/16/2018 16:45:02",
          "content": "<p>Thanks for sharing. Really helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "324921": "Congrats winners!\n\nI would like to share my features briefly before I forget it. \n\n### Features\n\nI used BigQuery on the almost all feature engineering parts.  \n https://gist.github.com/tkm2261/1b3c3c37753e55ed2914577c0f96d222\n\nIt takes only about 20 minutes. very fast!\n\nThere may not be surprising features if you have seen the shared kernels. I tried to find good features by brute force way.\n\n### Machine\n\nCPU: 16 cores, MEM: 100GB on GCP\n\n### Data\n\n* training data: day 7 and day 8 (13,188,695,398 rows)\n* validation data: day 9 (53,016,937 rows)\n* test data: day 10 (18,790,469 rows)\n\nI used validation data to determin the boosting round. Then, I used train+valid data for training.\n\n### Training\n\nSimply, I used LightGBM and ensembled with different seeds.\n\nThe single best model is 0.9823 on the public LB.\n\nThe parameter is here: \n\n    {'colsample_bytree': 0.6, 'learning_rate': 0.1, 'max_bin': 1023, 'max_depth': -1, 'metric': ['binary_logloss', 'auc'], 'min_child_weight': 30, 'min_split_gain': 0.0001, 'num_leaves': 127, 'objective': 'binary', 'reg_alpha': 0, 'scale_pos_weight': 1, 'seed': 1142, 'subsample': 0.9, 'subsample_freq': 1, 'verbose': -1}\n\nsource code is here: https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/train_lgb.py\n\nI tried to earn my tuition for MS CS program starting next fall. But, there was no free lunch!",
    "324925": "Good game and well played!",
    "324929": "Congrats and thanks for sharing~",
    "324936": "Cool! Thanks for sharing! I hope you do well in your MS program!",
    "324943": "Congrats and thanks for sharing. so first step is make train csv file to sql dataset.",
    "324949": "Thanks for sharing.",
    "324957": "Congrats and thanks for sharing!   \nAnd may I ask here    \n```\nI used validation data to determin the boosting round. Then, I used train+valid data for training.\n```  \nThat means you use the same parameters and rounds you get in validation to retrain the model without using validation in retraining, is that right?  \nTHANKS.",
    "324962": "'scale_pos_weight': 1 ? I have not seen anyone using such a small number.\nCongratulation on your accomplishment!",
    "324966": "I mean the validation data is used like fold-out data.",
    "324975": "I handle the imbalance by min_sample_weight, num_leaves, and other parameters. Since the eval metric is AUC, the scale is not important.",
    "324978": "Thank you for sharing! I'll try BigQuery",
    "325006": "Nice Work!",
    "325018": "just wondering about whether using validation(or early_stopping) in retraining or not. thx",
    "325025": "haha, i am going for MS CS , in the coming fall ! how did you come up with the hyperparameters",
    "325028": "One quick question: do you refit on the whole set after get the best number of trees from the validation?",
    "325034": "Wow that's so fast to build all those features, thanks for sharing &amp; congrats! I never knew it can be so much faster using BiqQuery, I only use VM instances for this competittion.",
    "325061": "Brilliant, \nThanks for sharing.",
    "325090": "great solution! most suitable for production",
    "325094": "Congrats! Vote for the \"next_click\" part, you've done way more than I did. Thanks for sharing.",
    "325097": "Could you tell a bit about tuning of the hyper-parameters, if you did any?",
    "325099": "yes, I refit.",
    "325100": "I used the validation data in retraining.",
    "325123": "Congrats. <br> Did you skip day 8 from your training set?",
    "325124": "sorry, day 8 is included in my training data.",
    "325137": "May I ask if you made features for each day separately? Like below\n\n    day7 = make_features(day7)\n    ...\n    ...\n    day9 = make_features(day9)\n    \n    X_train = pd.concat([day7,day8])\n    X_val = day9[hours[4,11]]\n\nor\n\n    all_days = make_features(all_days)\n    X_train = all_days[days[7,8]]\n    X_val = all_days[day[9]]",
    "325150": "the latter is correct. plz see my sql gist for more detail.",
    "325190": "Something still confuses me, so I try to make it clear.  \n1.You use train data and valid data to train models-&gt;find the best parameters and rounds.  \n2.You use the parameters and rounds in 1 and the whole dataset to retrain a model, in this process you won't split the whole dataset for validation, and just train the number of rounds in 1 without using early stopping.  \nAm I get the right thing here?",
    "325223": "maybe right. if you want to learn more, please see my code: https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/train_lgb.py",
    "325245": "Cool, I think I got what I want in your code, many thanks.",
    "325248": "Well done, you have been the lead solo competitor till the last few days if I remember well, chasing you during the competition helped me a lot!",
    "325285": "Well done tkm2261!\nAfter reading all post, I realized I completely missed that it was valuable to split data before and after day 9 and 10. \nAnd thanks for sharing the code! tkm2261, I am surprised changing the seeds plays a role. I will try in the next competition. And good luck for your MS.",
    "329250": "Thanks for sharing. 2 questions regarding to your solution:\n\n1. All data before day 9 should have 131886953 rows and all day 9 data should have 53016937 rows. In your post looks like they have ~ 2 million rows less, did you filter out some data? Duplication? \n2. In your load data code https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/load_data.py#L43 what does \"prev\" data set refer to? Looks like it is not part of your queries. \n\nThanks",
    "329305": "Thanks for questioning!  \n1. You are right. I copied these numbers from the log when I deleted duplicated rows by mistake. And, deleting duplicated rows did not improve my score.  \n2.  At first, I used only day 8 data as train data. Then, I added day 7 data that I called \"prev data\". So, in my query, prev and train are merged.",
    "329307": "Gotcha! Thanks for answering these questions. I'm on the half way of reproducing your best single model. \n\nI realized you excluded several features from BigQuery results(https://github.com/tkm2261/kaggle_talkingdata/blob/master/protos/load_data.py#L10), how did you do the feature selection? I'm also curious how you find the best param of lightGBM, from your code, looks like you did some manual tuning?",
    "329316": "1. I left top about 50 features ordered by the feature importance of lightgbm when I made new features. I sometimes dropped features in python code.\n2. For parameter tuning, I used day 8 data as train data and day 9 data as validation data. Since training only  day 8 data is not relatively heavy, I could try about 20 parameters sets. But, I did not devote much effort to the parameter tuning.",
    "329530": "Thanks for sharing. Really helpful."
  },
  "source": "meta"
}