{
  "id": 56406,
  "title": "5th place solution (0.9836 with pure feature engineering)",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/writeups/mmdp-5th-place-solution-0-9836-with-pure-feature-e",
  "author_name": "",
  "post_date": "2021-02-15T10:12:22.987Z",
  "votes": 57,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Thank you my briliant team, all teams who competed with us, and all people in the competition. <br>\nand Congrats Danijel, who became grandmaster in this competition! <br><br>\nIt's my first Kaggle competition and I'm really happy to get 5th in this competition. <br><br>\nMy model is based on simple LightGBM, with lotta feature engineering. I made 2 models, one with 561 features, the other with 695 features. but increasing the number of features from 561 to 695 gave very small score up, so I'll explain only about 561 features model. <br><br>\nI posted the '9836_features.txt', which shows the feature name and the importance(num_split). <br></p>\n<h1>Features (only effective ones)</h1>\n<p>actually I made more than 30 types of features. Here I show the 6 types which are effective.</p>\n<h2>1. time_diff_k_for, time_diff_k_back</h2>\n<p>time_diff_1_for means the click time difference with next click, <br><br>\ntime_diff_1_back means the click time difference with previous click. <br><br>\ntime_diff_2_for means the click time difference with next next click. <br><br>\nwhen k = 1, 2, it is very effective.<br></p>\n<h2>2. nunique_counts_ratio, 3. top_counts_ratio, 4. top10_hoge</h2>\n<p>For example, if specific ip has 100 samples and the distribution of the device is <br><br>\n2(30 samples), 3(25 samples), 5(20 samples), 7(15 samples), 8(10 samples), <br><br>\nnunique_counts is 5, nunique_counts_ratio is 0.02, top_counts is 30, top_counts_ratio is 0.3, <br><br>\ntop2_counts is 25, top2_counts_ratio is 0.25, top3_counts is 15, top3_counts_ratio is 0.15. <br><br>\ntop_device is 2, top2_device is 3, top3_device is 5, top4_device is 7, top5_device is 8. <br></p>\n<h2>5. tdf_top_counts_ratio</h2>\n<p>sometimes, there are ip or (ip, os, device) that has always similar timedeltas. for example, if specific ip's click log is like this, <br><br>\n00:00:00 <br><br>\n00:00:05 <br><br>\n00:00:10 <br><br>\n00:00:15 <br><br>\nthey always have 5 sec timedeltas. <br><br>\nso, I got top_counts, top2_counts, top3_counts from timedeltas.<br></p>\n<h2>6. lagged_dl_counts_ratio</h2>\n<p>lagged target encoding. <br><br>\nin day1, this features are missing. <br><br>\nfor day2, I used target mean in day1. <br><br>\nfor day3, I used target mean in day2. <br><br>\nfor day4, I used target mean in day3.<br></p>\n<h1>Adding Features &amp; Feature selection</h1>\n<p>I tried many patterns and select features using feature_importances, which were really time consuming. I made more than 7000 features, but most of the experiments failed.<br></p>\n<h1>Ensemble</h1>\n<p>As <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56319\" target=\"_blank\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56319</a> says, Michael's blend NN improved the score and finally got 0.9840 :)</p>\n<h1>Comments</h1>\n<p>I tried very hard to deal with duplicate problem, but finally couldn't find the leak. <br><br>\nI was very sad to see this. We could have won this competition if we had noticed it… <br><br>\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268\" target=\"_blank\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268</a> <br><br>\nAnyway I really enjoyed my first kaggle competition, I'll continue kaggle and want to win the next competition. <br></p>\n<h3>Thanks all, see you again!</h3>",
  "messages": [
    {
      "id": "326225",
      "postDate": "05/09/2018 13:15:27",
      "content": "<p>Thank you my briliant team, all teams who competed with us, and all people in the competition. <br>\nand Congrats Danijel, who became grandmaster in this competition! <br><br>\nIt's my first Kaggle competition and I'm really happy to get 5th in this competition. <br><br>\nMy model is based on simple LightGBM, with lotta feature engineering. I made 2 models, one with 561 features, the other with 695 features. but increasing the number of features from 561 to 695 gave very small score up, so I'll explain only about 561 features model. <br><br>\nI posted the '9836_features.txt', which shows the feature name and the importance(num_split). <br></p>\n<h1>Features (only effective ones)</h1>\n<p>actually I made more than 30 types of features. Here I show the 6 types which are effective.</p>\n<h2>1. time_diff_k_for, time_diff_k_back</h2>\n<p>time_diff_1_for means the click time difference with next click, <br><br>\ntime_diff_1_back means the click time difference with previous click. <br><br>\ntime_diff_2_for means the click time difference with next next click. <br><br>\nwhen k = 1, 2, it is very effective.<br></p>\n<h2>2. nunique_counts_ratio, 3. top_counts_ratio, 4. top10_hoge</h2>\n<p>For example, if specific ip has 100 samples and the distribution of the device is <br><br>\n2(30 samples), 3(25 samples), 5(20 samples), 7(15 samples), 8(10 samples), <br><br>\nnunique_counts is 5, nunique_counts_ratio is 0.02, top_counts is 30, top_counts_ratio is 0.3, <br><br>\ntop2_counts is 25, top2_counts_ratio is 0.25, top3_counts is 15, top3_counts_ratio is 0.15. <br><br>\ntop_device is 2, top2_device is 3, top3_device is 5, top4_device is 7, top5_device is 8. <br></p>\n<h2>5. tdf_top_counts_ratio</h2>\n<p>sometimes, there are ip or (ip, os, device) that has always similar timedeltas. for example, if specific ip's click log is like this, <br><br>\n00:00:00 <br><br>\n00:00:05 <br><br>\n00:00:10 <br><br>\n00:00:15 <br><br>\nthey always have 5 sec timedeltas. <br><br>\nso, I got top_counts, top2_counts, top3_counts from timedeltas.<br></p>\n<h2>6. lagged_dl_counts_ratio</h2>\n<p>lagged target encoding. <br><br>\nin day1, this features are missing. <br><br>\nfor day2, I used target mean in day1. <br><br>\nfor day3, I used target mean in day2. <br><br>\nfor day4, I used target mean in day3.<br></p>\n<h1>Adding Features &amp; Feature selection</h1>\n<p>I tried many patterns and select features using feature_importances, which were really time consuming. I made more than 7000 features, but most of the experiments failed.<br></p>\n<h1>Ensemble</h1>\n<p>As <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56319\" target=\"_blank\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56319</a> says, Michael's blend NN improved the score and finally got 0.9840 :)</p>\n<h1>Comments</h1>\n<p>I tried very hard to deal with duplicate problem, but finally couldn't find the leak. <br><br>\nI was very sad to see this. We could have won this competition if we had noticed it… <br><br>\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268\" target=\"_blank\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268</a> <br><br>\nAnyway I really enjoyed my first kaggle competition, I'll continue kaggle and want to win the next competition. <br></p>\n<h3>Thanks all, see you again!</h3>",
      "rawMarkdown": "Thank you my briliant team, all teams who competed with us, and all people in the competition. \nand Congrats Danijel, who became grandmaster in this competition! <br>\nIt's my first Kaggle competition and I'm really happy to get 5th in this competition. <br>\nMy model is based on simple LightGBM, with lotta feature engineering. I made 2 models, one with 561 features, the other with 695 features. but increasing the number of features from 561 to 695 gave very small score up, so I'll explain only about 561 features model. <br>\nI posted the '9836\\_features.txt', which shows the feature name and the importance(num\\_split). <br>\n\n# Features (only effective ones)\nactually I made more than 30 types of features. Here I show the 6 types which are effective.\n\n## 1. time\\_diff\\_k\\_for, time\\_diff\\_k\\_back\ntime\\_diff\\_1\\_for means the click time difference with next click, <br>\ntime\\_diff\\_1\\_back means the click time difference with previous click. <br>\ntime\\_diff\\_2\\_for means the click time difference with next next click. <br>\nwhen k = 1, 2, it is very effective.<br>\n\n## 2. nunique\\_counts\\_ratio, 3. top\\_counts\\_ratio, 4. top10\\_hoge\nFor example, if specific ip has 100 samples and the distribution of the device is <br>\n2(30 samples), 3(25 samples), 5(20 samples), 7(15 samples), 8(10 samples), <br>\nnunique\\_counts is 5, nunique\\_counts\\_ratio is 0.02, top\\_counts is 30, top\\_counts\\_ratio is 0.3, <br>\ntop2\\_counts is 25, top2\\_counts\\_ratio is 0.25, top3\\_counts is 15, top3\\_counts\\_ratio is 0.15. <br>\ntop\\_device is 2, top2\\_device is 3, top3\\_device is 5, top4\\_device is 7, top5\\_device is 8. <br>\n\n## 5. tdf\\_top\\_counts\\_ratio\nsometimes, there are ip or (ip, os, device) that has always similar timedeltas. for example, if specific ip's click log is like this, <br>\n00:00:00 <br>\n00:00:05 <br>\n00:00:10 <br>\n00:00:15 <br>\nthey always have 5 sec timedeltas. <br>\nso, I got top\\_counts, top2\\_counts, top3\\_counts from timedeltas.<br>\n\n## 6. lagged\\_dl\\_counts\\_ratio\nlagged target encoding. <br>\nin day1, this features are missing. <br>\nfor day2, I used target mean in day1. <br>\nfor day3, I used target mean in day2. <br>\nfor day4, I used target mean in day3.<br>\n\n# Adding Features &amp; Feature selection\nI tried many patterns and select features using feature\\_importances, which were really time consuming. I made more than 7000 features, but most of the experiments failed.<br>\n\n#Ensemble\nAs https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56319 says, Michael's blend NN improved the score and finally got 0.9840 :)\n\n#Comments\nI tried very hard to deal with duplicate problem, but finally couldn't find the leak. <br>\nI was very sad to see this. We could have won this competition if we had noticed it... <br>\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268 <br>\nAnyway I really enjoyed my first kaggle competition, I'll continue kaggle and want to win the next competition. <br>\n<h3>Thanks all, see you again!</h3>",
      "votes": null
    },
    {
      "id": "326231",
      "postDate": "05/09/2018 13:20:15",
      "content": "<p>If we were to rank people on how hard they tried to win this competition, I am sure you will be ranked #1 <br>\nYou worked day and night. I never saw you going to bed.... <br>\nReally great teammate, hope you will win the next competition!</p>",
      "rawMarkdown": "If we were to rank people on how hard they tried to win this competition, I am sure you will be ranked #1 <br>\nYou worked day and night. I never saw you going to bed.... <br>\nReally great teammate, hope you will win the next competition!",
      "votes": null
    },
    {
      "id": "326239",
      "postDate": "05/09/2018 13:24:16",
      "content": "<p>Thanks pocket! I'm very happy to team-up with you, I learned a lot from you :) I really hope you become grandmaster :)</p>",
      "rawMarkdown": "Thanks pocket! I'm very happy to team-up with you, I learned a lot from you :) I really hope you become grandmaster :)",
      "votes": null
    },
    {
      "id": "326248",
      "postDate": "05/09/2018 13:33:59",
      "content": "<p>wonderful. can you explain what is the meaning of \"tdf_top_counts_ratio\"??</p>",
      "rawMarkdown": "wonderful. can you explain what is the meaning of \"tdf_top_counts_ratio\"??",
      "votes": null
    },
    {
      "id": "326250",
      "postDate": "05/09/2018 13:37:04",
      "content": "<p>Experiments on over 7000 features!!! Impressive! I assume you have a very strong and stable pipeline to work with these amount of features, right?</p>",
      "rawMarkdown": "Experiments on over 7000 features!!! Impressive! I assume you have a very strong and stable pipeline to work with these amount of features, right?",
      "votes": null
    },
    {
      "id": "326251",
      "postDate": "05/09/2018 13:37:16",
      "content": "<p>for example, if specific ip's click log is like this, <br>\n00:00:00 <br>\n00:00:05 <br>\n00:00:10 <br>\n00:00:15 <br>\n00:00:23 <br>\n00:00:30 <br>\n00:00:35 <br>\nthey have 5 sec timedeltas 4 times. so, tdf_top_counts_ratio can be caluclated by this. <br>\n4/(7-1) = 0.67 <br>\nThis is tdf_top_counts_ratio. </p>",
      "rawMarkdown": "for example, if specific ip's click log is like this, <br>\n00:00:00 <br>\n00:00:05 <br>\n00:00:10 <br>\n00:00:15 <br>\n00:00:23 <br>\n00:00:30 <br>\n00:00:35 <br>\nthey have 5 sec timedeltas 4 times. so, tdf_top_counts_ratio can be caluclated by this. <br>\n4/(7-1) = 0.67 <br>\nThis is tdf_top_counts_ratio.",
      "votes": null
    },
    {
      "id": "326254",
      "postDate": "05/09/2018 13:40:21",
      "content": "<p>Thank you for sharing your solution. Just curious, </p>\n\n<blockquote>\n  <p>I just tried all patterns</p>\n</blockquote>\n\n<p>means all combination of features? If we have 10 features here, you tried 10C1 + 10C2 + 10C3 ... 10C10? So you made 1022 models? Stepwise worked for me.</p>",
      "rawMarkdown": "Thank you for sharing your solution. Just curious, \n\n&gt; I just tried all patterns\n\nmeans all combination of features? If we have 10 features here, you tried 10C1 + 10C2 + 10C3 ... 10C10? So you made 1022 models? Stepwise worked for me.",
      "votes": null
    },
    {
      "id": "326256",
      "postDate": "05/09/2018 13:41:20",
      "content": "<p>I can only confirm what <a href=\"https://www.kaggle.com/pocketsuteado\">pocket</a> said. Incredible how much time and energy you dedicated to this competition. You literally tried all the features. You'd have deserved to win. Too bad about the leak, I'm also really upset about this. But still, without you we wouldn't have finished #5. Thank you so much for that! Looking forward to meet you in another competition. Hopefully in the same a team because it will be really hard to beat you.</p>",
      "rawMarkdown": "I can only confirm what [pocket][1] said. Incredible how much time and energy you dedicated to this competition. You literally tried all the features. You'd have deserved to win. Too bad about the leak, I'm also really upset about this. But still, without you we wouldn't have finished #5. Thank you so much for that! Looking forward to meet you in another competition. Hopefully in the same a team because it will be really hard to beat you.\n\n\n  [1]: https://www.kaggle.com/pocketsuteado",
      "votes": null
    },
    {
      "id": "326257",
      "postDate": "05/09/2018 13:42:06",
      "content": "<p>I used n1-megamem-96 instance, which have 1.4TB RAM. It's not so expensive because I used preemptive instance.</p>",
      "rawMarkdown": "I used n1-megamem-96 instance, which have 1.4TB RAM. It's not so expensive because I used preemptive instance.",
      "votes": null
    },
    {
      "id": "326258",
      "postDate": "05/09/2018 13:43:16",
      "content": "<p>Thanks for sharing and congrats on the result.  You have some interesting features I did not try.</p>",
      "rawMarkdown": "Thanks for sharing and congrats on the result.  You have some interesting features I did not try.",
      "votes": null
    },
    {
      "id": "326259",
      "postDate": "05/09/2018 13:44:12",
      "content": "<p>I get! 4 mean：the 5 timedelta have 4 times. 7 - 1 mean all deltas count...</p>",
      "rawMarkdown": "I get! 4 mean：the 5 timedelta have 4 times. 7 - 1 mean all deltas count...",
      "votes": null
    },
    {
      "id": "326261",
      "postDate": "05/09/2018 13:45:44",
      "content": "<p>Thanks for reply! I mean the feature engineering and feature selection pipeline (framework/code). If you can share something about this.</p>",
      "rawMarkdown": "Thanks for reply! I mean the feature engineering and feature selection pipeline (framework/code). If you can share something about this.",
      "votes": null
    },
    {
      "id": "326262",
      "postDate": "05/09/2018 13:46:25",
      "content": "<p>No, It's impossible :( It means, For example, time_diff_k_for has 2^5 -1 = 31 patterns because there are (ip, os, device, app, channel). and nunique_counts_ratio has 180 patterns. I added all of them and select by feature_importances. I didn't do accurate feature selection, did very rough one.</p>",
      "rawMarkdown": "No, It's impossible :( It means, For example, time_diff_k_for has 2^5 -1 = 31 patterns because there are (ip, os, device, app, channel). and nunique_counts_ratio has 180 patterns. I added all of them and select by feature_importances. I didn't do accurate feature selection, did very rough one.",
      "votes": null
    },
    {
      "id": "326263",
      "postDate": "05/09/2018 13:47:37",
      "content": "<p>what is the means of \"back_10_mean_level_1_ip_os_device\"?? </p>",
      "rawMarkdown": "what is the means of \"back_10_mean_level_1_ip_os_device\"??",
      "votes": null
    },
    {
      "id": "326265",
      "postDate": "05/09/2018 13:47:59",
      "content": "<p>I did nothing special, just used .npy file to store features and load it and concat them, which took much time. but I didn't want to use DataFrame to avoid memory explodes...</p>",
      "rawMarkdown": "I did nothing special, just used .npy file to store features and load it and concat them, which took much time. but I didn't want to use DataFrame to avoid memory explodes...",
      "votes": null
    },
    {
      "id": "326267",
      "postDate": "05/09/2018 13:50:15",
      "content": "<p>Ah I completely forgot this. this is the rolling mean of timedeltas of 10 samples before the click. </p>",
      "rawMarkdown": "Ah I completely forgot this. this is the rolling mean of timedeltas of 10 samples before the click.",
      "votes": null
    },
    {
      "id": "326269",
      "postDate": "05/09/2018 13:52:39",
      "content": "<p>Thanks Danijel, now you are grandmaster. :) I think I can't improve my model to 0.9836 without the discussion with you. I'm really looking forward to see you in another competition, and hopefully I want to see you in the real world :)</p>",
      "rawMarkdown": "Thanks Danijel, now you are grandmaster. :) I think I can't improve my model to 0.9836 without the discussion with you. I'm really looking forward to see you in another competition, and hopefully I want to see you in the real world :)",
      "votes": null
    },
    {
      "id": "326271",
      "postDate": "05/09/2018 13:54:02",
      "content": "<p>thanks</p>",
      "rawMarkdown": "thanks",
      "votes": null
    },
    {
      "id": "326273",
      "postDate": "05/09/2018 13:55:25",
      "content": "<p>Thanks CPMP, your information in the discussion/kernel were really valuable and I was always impressed with you :)</p>",
      "rawMarkdown": "Thanks CPMP, your information in the discussion/kernel were really valuable and I was always impressed with you :)",
      "votes": null
    },
    {
      "id": "326284",
      "postDate": "05/09/2018 14:05:46",
      "content": "<p>Thank you mamasingks for these feedbacks and hints. For noob like me it is highly valuable to read feedbacks from top kagglers. All the best to you</p>",
      "rawMarkdown": "Thank you mamasingks for these feedbacks and hints. For noob like me it is highly valuable to read feedbacks from top kagglers. All the best to you",
      "votes": null
    },
    {
      "id": "326285",
      "postDate": "05/09/2018 14:06:58",
      "content": "<p>Agree with you that CPMP comments and posts are very valuable. </p>",
      "rawMarkdown": "Agree with you that CPMP comments and posts are very valuable.",
      "votes": null
    },
    {
      "id": "326556",
      "postDate": "05/09/2018 22:55:47",
      "content": "<p>I believe some of features make your AUC worse. We can not use feature importance for selecting features at these kind of competition. Because datetime is different.</p>",
      "rawMarkdown": "I believe some of features make your AUC worse. We can not use feature importance for selecting features at these kind of competition. Because datetime is different.",
      "votes": null
    },
    {
      "id": "326627",
      "postDate": "05/10/2018 04:38:44",
      "content": "<p>I used feature importances to get the effective subset of the specific features group. For example, there are 180 patterns of \"nunique_counts_ratio\", but I didn't use all of them because if I use all of them training time becomes very long but the gain is small. <br>\nIn contrast, I didn't use feature_importance but used valid score to check if the specific features group is effective or not. For example, I made \"lagged_ca_diff_mean\", which means the average of difference between click_time and attributed_time in the previous day. Actually some of them have high feature_importance but the valid score doesn't improve, so I threw away all of them. I think it's very important to use both valid score and feature importances when selecting features.</p>",
      "rawMarkdown": "I used feature importances to get the effective subset of the specific features group. For example, there are 180 patterns of \"nunique_counts_ratio\", but I didn't use all of them because if I use all of them training time becomes very long but the gain is small. <br>\nIn contrast, I didn't use feature_importance but used valid score to check if the specific features group is effective or not. For example, I made \"lagged_ca_diff_mean\", which means the average of difference between click_time and attributed_time in the previous day. Actually some of them have high feature_importance but the valid score doesn't improve, so I threw away all of them. I think it's very important to use both valid score and feature importances when selecting features.",
      "votes": null
    },
    {
      "id": "326696",
      "postDate": "05/10/2018 07:30:46",
      "content": "<p>Fair enough! I could understand your solution!</p>",
      "rawMarkdown": "Fair enough! I could understand your solution!",
      "votes": null
    },
    {
      "id": "326951",
      "postDate": "05/10/2018 15:03:20",
      "content": "<p>Hi mamasingkgs, how did you choose the depth or num_leaves in GBM, since you used over 500 features? In most of the public kernels(including mine), they set depth as 3-5 using only about 30 features.</p>",
      "rawMarkdown": "Hi mamasingkgs, how did you choose the depth or num_leaves in GBM, since you used over 500 features? In most of the public kernels(including mine), they set depth as 3-5 using only about 30 features.",
      "votes": null
    },
    {
      "id": "326956",
      "postDate": "05/10/2018 15:10:04",
      "content": "<p>This is my parameter.\nparams = {\n    'boosting_type': 'gbdt',\n    'objective':     'binary',\n    'tree_learner':  'serial',\n    'metric':        'binary_logloss',\n    'learning_rate': 0.05,\n    'max_bin':       255,\n    'num_leaves':    63,\n    'max_depth':     -1, \n    'min_data_in_leaf': 1000, \n    'feature_fraction': 0.7, \n    'bagging_freq':     1,\n    'bagging_fraction': 0.7,\n    'lambda_l1':       1, \n    'lambda_l2':       1, \n    'sigmoid':         1,\n    'verbosity':       0, \n}</p>",
      "rawMarkdown": "This is my parameter.\nparams = {\n    'boosting_type': 'gbdt',\n    'objective':     'binary',\n    'tree_learner':  'serial',\n    'metric':        'binary_logloss',\n    'learning_rate': 0.05,\n    'max_bin':       255,\n    'num_leaves':    63,\n    'max_depth':     -1, \n    'min_data_in_leaf': 1000, \n    'feature_fraction': 0.7, \n    'bagging_freq':     1,\n    'bagging_fraction': 0.7,\n    'lambda_l1':       1, \n    'lambda_l2':       1, \n    'sigmoid':         1,\n    'verbosity':       0, \n}",
      "votes": null
    },
    {
      "id": "326964",
      "postDate": "05/10/2018 15:20:15",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!",
      "votes": null
    },
    {
      "id": "327741",
      "postDate": "05/12/2018 10:28:39",
      "content": "<p>I'm confused by lagged_dl_counts_ratio. <code>in day1, this features are missing.</code> If they are missing, is it good for training? </p>",
      "rawMarkdown": "I'm confused by lagged_dl_counts_ratio. `in day1, this features are missing.` If they are missing, is it good for training?",
      "votes": null
    },
    {
      "id": "327828",
      "postDate": "05/12/2018 15:17:18",
      "content": "<p>It doesn't give bad effect for me. If there is no sample or no positive sample in previous day, I set the value 0. so, there are no NaN samples in day2, day3, day4 and no problem happened. </p>",
      "rawMarkdown": "It doesn't give bad effect for me. If there is no sample or no positive sample in previous day, I set the value 0. so, there are no NaN samples in day2, day3, day4 and no problem happened.",
      "votes": null
    },
    {
      "id": "329920",
      "postDate": "05/17/2018 15:03:24",
      "content": "<p>What a amazing feature engineering!</p>",
      "rawMarkdown": "What a amazing feature engineering!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 326231,
      "author_name": "pocketsuteado",
      "author_url": "",
      "post_date": "05/09/2018 13:20:15",
      "content": "<p>If we were to rank people on how hard they tried to win this competition, I am sure you will be ranked #1 <br>\nYou worked day and night. I never saw you going to bed.... <br>\nReally great teammate, hope you will win the next competition!</p>",
      "votes": null,
      "replies": [
        {
          "id": 326239,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/09/2018 13:24:16",
          "content": "<p>Thanks pocket! I'm very happy to team-up with you, I learned a lot from you :) I really hope you become grandmaster :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326248,
      "author_name": "huayupeng",
      "author_url": "",
      "post_date": "05/09/2018 13:33:59",
      "content": "<p>wonderful. can you explain what is the meaning of \"tdf_top_counts_ratio\"??</p>",
      "votes": null,
      "replies": [
        {
          "id": 326251,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/09/2018 13:37:16",
          "content": "<p>for example, if specific ip's click log is like this, <br>\n00:00:00 <br>\n00:00:05 <br>\n00:00:10 <br>\n00:00:15 <br>\n00:00:23 <br>\n00:00:30 <br>\n00:00:35 <br>\nthey have 5 sec timedeltas 4 times. so, tdf_top_counts_ratio can be caluclated by this. <br>\n4/(7-1) = 0.67 <br>\nThis is tdf_top_counts_ratio. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326259,
          "author_name": "huayupeng",
          "author_url": "",
          "post_date": "05/09/2018 13:44:12",
          "content": "<p>I get! 4 mean：the 5 timedelta have 4 times. 7 - 1 mean all deltas count...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326250,
      "author_name": "cczaixian",
      "author_url": "",
      "post_date": "05/09/2018 13:37:04",
      "content": "<p>Experiments on over 7000 features!!! Impressive! I assume you have a very strong and stable pipeline to work with these amount of features, right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 326257,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/09/2018 13:42:06",
          "content": "<p>I used n1-megamem-96 instance, which have 1.4TB RAM. It's not so expensive because I used preemptive instance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326261,
          "author_name": "cczaixian",
          "author_url": "",
          "post_date": "05/09/2018 13:45:44",
          "content": "<p>Thanks for reply! I mean the feature engineering and feature selection pipeline (framework/code). If you can share something about this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326265,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/09/2018 13:47:59",
          "content": "<p>I did nothing special, just used .npy file to store features and load it and concat them, which took much time. but I didn't want to use DataFrame to avoid memory explodes...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326254,
      "author_name": "onodera",
      "author_url": "",
      "post_date": "05/09/2018 13:40:21",
      "content": "<p>Thank you for sharing your solution. Just curious, </p>\n\n<blockquote>\n  <p>I just tried all patterns</p>\n</blockquote>\n\n<p>means all combination of features? If we have 10 features here, you tried 10C1 + 10C2 + 10C3 ... 10C10? So you made 1022 models? Stepwise worked for me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 326262,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/09/2018 13:46:25",
          "content": "<p>No, It's impossible :( It means, For example, time_diff_k_for has 2^5 -1 = 31 patterns because there are (ip, os, device, app, channel). and nunique_counts_ratio has 180 patterns. I added all of them and select by feature_importances. I didn't do accurate feature selection, did very rough one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326556,
          "author_name": "onodera",
          "author_url": "",
          "post_date": "05/09/2018 22:55:47",
          "content": "<p>I believe some of features make your AUC worse. We can not use feature importance for selecting features at these kind of competition. Because datetime is different.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326627,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/10/2018 04:38:44",
          "content": "<p>I used feature importances to get the effective subset of the specific features group. For example, there are 180 patterns of \"nunique_counts_ratio\", but I didn't use all of them because if I use all of them training time becomes very long but the gain is small. <br>\nIn contrast, I didn't use feature_importance but used valid score to check if the specific features group is effective or not. For example, I made \"lagged_ca_diff_mean\", which means the average of difference between click_time and attributed_time in the previous day. Actually some of them have high feature_importance but the valid score doesn't improve, so I threw away all of them. I think it's very important to use both valid score and feature importances when selecting features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326696,
          "author_name": "onodera",
          "author_url": "",
          "post_date": "05/10/2018 07:30:46",
          "content": "<p>Fair enough! I could understand your solution!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326256,
      "author_name": "danijelk",
      "author_url": "",
      "post_date": "05/09/2018 13:41:20",
      "content": "<p>I can only confirm what <a href=\"https://www.kaggle.com/pocketsuteado\">pocket</a> said. Incredible how much time and energy you dedicated to this competition. You literally tried all the features. You'd have deserved to win. Too bad about the leak, I'm also really upset about this. But still, without you we wouldn't have finished #5. Thank you so much for that! Looking forward to meet you in another competition. Hopefully in the same a team because it will be really hard to beat you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 326269,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/09/2018 13:52:39",
          "content": "<p>Thanks Danijel, now you are grandmaster. :) I think I can't improve my model to 0.9836 without the discussion with you. I'm really looking forward to see you in another competition, and hopefully I want to see you in the real world :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326258,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/09/2018 13:43:16",
      "content": "<p>Thanks for sharing and congrats on the result.  You have some interesting features I did not try.</p>",
      "votes": null,
      "replies": [
        {
          "id": 326273,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/09/2018 13:55:25",
          "content": "<p>Thanks CPMP, your information in the discussion/kernel were really valuable and I was always impressed with you :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326285,
          "author_name": "ericbenhamou",
          "author_url": "",
          "post_date": "05/09/2018 14:06:58",
          "content": "<p>Agree with you that CPMP comments and posts are very valuable. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326263,
      "author_name": "huayupeng",
      "author_url": "",
      "post_date": "05/09/2018 13:47:37",
      "content": "<p>what is the means of \"back_10_mean_level_1_ip_os_device\"?? </p>",
      "votes": null,
      "replies": [
        {
          "id": 326267,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/09/2018 13:50:15",
          "content": "<p>Ah I completely forgot this. this is the rolling mean of timedeltas of 10 samples before the click. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326271,
      "author_name": "xiaohuihui",
      "author_url": "",
      "post_date": "05/09/2018 13:54:02",
      "content": "<p>thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326284,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "05/09/2018 14:05:46",
      "content": "<p>Thank you mamasingks for these feedbacks and hints. For noob like me it is highly valuable to read feedbacks from top kagglers. All the best to you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326951,
      "author_name": "fangao",
      "author_url": "",
      "post_date": "05/10/2018 15:03:20",
      "content": "<p>Hi mamasingkgs, how did you choose the depth or num_leaves in GBM, since you used over 500 features? In most of the public kernels(including mine), they set depth as 3-5 using only about 30 features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 326956,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/10/2018 15:10:04",
          "content": "<p>This is my parameter.\nparams = {\n    'boosting_type': 'gbdt',\n    'objective':     'binary',\n    'tree_learner':  'serial',\n    'metric':        'binary_logloss',\n    'learning_rate': 0.05,\n    'max_bin':       255,\n    'num_leaves':    63,\n    'max_depth':     -1, \n    'min_data_in_leaf': 1000, \n    'feature_fraction': 0.7, \n    'bagging_freq':     1,\n    'bagging_fraction': 0.7,\n    'lambda_l1':       1, \n    'lambda_l2':       1, \n    'sigmoid':         1,\n    'verbosity':       0, \n}</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326964,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/10/2018 15:20:15",
          "content": "<p>Thanks for sharing!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327741,
      "author_name": "shenchengguang",
      "author_url": "",
      "post_date": "05/12/2018 10:28:39",
      "content": "<p>I'm confused by lagged_dl_counts_ratio. <code>in day1, this features are missing.</code> If they are missing, is it good for training? </p>",
      "votes": null,
      "replies": [
        {
          "id": 327828,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/12/2018 15:17:18",
          "content": "<p>It doesn't give bad effect for me. If there is no sample or no positive sample in previous day, I set the value 0. so, there are no NaN samples in day2, day3, day4 and no problem happened. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 329920,
      "author_name": "shawnyxiao",
      "author_url": "",
      "post_date": "05/17/2018 15:03:24",
      "content": "<p>What a amazing feature engineering!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "326225": "Thank you my briliant team, all teams who competed with us, and all people in the competition. \nand Congrats Danijel, who became grandmaster in this competition! <br>\nIt's my first Kaggle competition and I'm really happy to get 5th in this competition. <br>\nMy model is based on simple LightGBM, with lotta feature engineering. I made 2 models, one with 561 features, the other with 695 features. but increasing the number of features from 561 to 695 gave very small score up, so I'll explain only about 561 features model. <br>\nI posted the '9836\\_features.txt', which shows the feature name and the importance(num\\_split). <br>\n\n# Features (only effective ones)\nactually I made more than 30 types of features. Here I show the 6 types which are effective.\n\n## 1. time\\_diff\\_k\\_for, time\\_diff\\_k\\_back\ntime\\_diff\\_1\\_for means the click time difference with next click, <br>\ntime\\_diff\\_1\\_back means the click time difference with previous click. <br>\ntime\\_diff\\_2\\_for means the click time difference with next next click. <br>\nwhen k = 1, 2, it is very effective.<br>\n\n## 2. nunique\\_counts\\_ratio, 3. top\\_counts\\_ratio, 4. top10\\_hoge\nFor example, if specific ip has 100 samples and the distribution of the device is <br>\n2(30 samples), 3(25 samples), 5(20 samples), 7(15 samples), 8(10 samples), <br>\nnunique\\_counts is 5, nunique\\_counts\\_ratio is 0.02, top\\_counts is 30, top\\_counts\\_ratio is 0.3, <br>\ntop2\\_counts is 25, top2\\_counts\\_ratio is 0.25, top3\\_counts is 15, top3\\_counts\\_ratio is 0.15. <br>\ntop\\_device is 2, top2\\_device is 3, top3\\_device is 5, top4\\_device is 7, top5\\_device is 8. <br>\n\n## 5. tdf\\_top\\_counts\\_ratio\nsometimes, there are ip or (ip, os, device) that has always similar timedeltas. for example, if specific ip's click log is like this, <br>\n00:00:00 <br>\n00:00:05 <br>\n00:00:10 <br>\n00:00:15 <br>\nthey always have 5 sec timedeltas. <br>\nso, I got top\\_counts, top2\\_counts, top3\\_counts from timedeltas.<br>\n\n## 6. lagged\\_dl\\_counts\\_ratio\nlagged target encoding. <br>\nin day1, this features are missing. <br>\nfor day2, I used target mean in day1. <br>\nfor day3, I used target mean in day2. <br>\nfor day4, I used target mean in day3.<br>\n\n# Adding Features &amp; Feature selection\nI tried many patterns and select features using feature\\_importances, which were really time consuming. I made more than 7000 features, but most of the experiments failed.<br>\n\n#Ensemble\nAs https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56319 says, Michael's blend NN improved the score and finally got 0.9840 :)\n\n#Comments\nI tried very hard to deal with duplicate problem, but finally couldn't find the leak. <br>\nI was very sad to see this. We could have won this competition if we had noticed it... <br>\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268 <br>\nAnyway I really enjoyed my first kaggle competition, I'll continue kaggle and want to win the next competition. <br>\n<h3>Thanks all, see you again!</h3>",
    "326231": "If we were to rank people on how hard they tried to win this competition, I am sure you will be ranked #1 <br>\nYou worked day and night. I never saw you going to bed.... <br>\nReally great teammate, hope you will win the next competition!",
    "326239": "Thanks pocket! I'm very happy to team-up with you, I learned a lot from you :) I really hope you become grandmaster :)",
    "326248": "wonderful. can you explain what is the meaning of \"tdf_top_counts_ratio\"??",
    "326250": "Experiments on over 7000 features!!! Impressive! I assume you have a very strong and stable pipeline to work with these amount of features, right?",
    "326251": "for example, if specific ip's click log is like this, <br>\n00:00:00 <br>\n00:00:05 <br>\n00:00:10 <br>\n00:00:15 <br>\n00:00:23 <br>\n00:00:30 <br>\n00:00:35 <br>\nthey have 5 sec timedeltas 4 times. so, tdf_top_counts_ratio can be caluclated by this. <br>\n4/(7-1) = 0.67 <br>\nThis is tdf_top_counts_ratio.",
    "326254": "Thank you for sharing your solution. Just curious, \n\n&gt; I just tried all patterns\n\nmeans all combination of features? If we have 10 features here, you tried 10C1 + 10C2 + 10C3 ... 10C10? So you made 1022 models? Stepwise worked for me.",
    "326256": "I can only confirm what [pocket][1] said. Incredible how much time and energy you dedicated to this competition. You literally tried all the features. You'd have deserved to win. Too bad about the leak, I'm also really upset about this. But still, without you we wouldn't have finished #5. Thank you so much for that! Looking forward to meet you in another competition. Hopefully in the same a team because it will be really hard to beat you.\n\n\n  [1]: https://www.kaggle.com/pocketsuteado",
    "326257": "I used n1-megamem-96 instance, which have 1.4TB RAM. It's not so expensive because I used preemptive instance.",
    "326258": "Thanks for sharing and congrats on the result.  You have some interesting features I did not try.",
    "326259": "I get! 4 mean：the 5 timedelta have 4 times. 7 - 1 mean all deltas count...",
    "326261": "Thanks for reply! I mean the feature engineering and feature selection pipeline (framework/code). If you can share something about this.",
    "326262": "No, It's impossible :( It means, For example, time_diff_k_for has 2^5 -1 = 31 patterns because there are (ip, os, device, app, channel). and nunique_counts_ratio has 180 patterns. I added all of them and select by feature_importances. I didn't do accurate feature selection, did very rough one.",
    "326263": "what is the means of \"back_10_mean_level_1_ip_os_device\"??",
    "326265": "I did nothing special, just used .npy file to store features and load it and concat them, which took much time. but I didn't want to use DataFrame to avoid memory explodes...",
    "326267": "Ah I completely forgot this. this is the rolling mean of timedeltas of 10 samples before the click.",
    "326269": "Thanks Danijel, now you are grandmaster. :) I think I can't improve my model to 0.9836 without the discussion with you. I'm really looking forward to see you in another competition, and hopefully I want to see you in the real world :)",
    "326271": "thanks",
    "326273": "Thanks CPMP, your information in the discussion/kernel were really valuable and I was always impressed with you :)",
    "326284": "Thank you mamasingks for these feedbacks and hints. For noob like me it is highly valuable to read feedbacks from top kagglers. All the best to you",
    "326285": "Agree with you that CPMP comments and posts are very valuable.",
    "326556": "I believe some of features make your AUC worse. We can not use feature importance for selecting features at these kind of competition. Because datetime is different.",
    "326627": "I used feature importances to get the effective subset of the specific features group. For example, there are 180 patterns of \"nunique_counts_ratio\", but I didn't use all of them because if I use all of them training time becomes very long but the gain is small. <br>\nIn contrast, I didn't use feature_importance but used valid score to check if the specific features group is effective or not. For example, I made \"lagged_ca_diff_mean\", which means the average of difference between click_time and attributed_time in the previous day. Actually some of them have high feature_importance but the valid score doesn't improve, so I threw away all of them. I think it's very important to use both valid score and feature importances when selecting features.",
    "326696": "Fair enough! I could understand your solution!",
    "326951": "Hi mamasingkgs, how did you choose the depth or num_leaves in GBM, since you used over 500 features? In most of the public kernels(including mine), they set depth as 3-5 using only about 30 features.",
    "326956": "This is my parameter.\nparams = {\n    'boosting_type': 'gbdt',\n    'objective':     'binary',\n    'tree_learner':  'serial',\n    'metric':        'binary_logloss',\n    'learning_rate': 0.05,\n    'max_bin':       255,\n    'num_leaves':    63,\n    'max_depth':     -1, \n    'min_data_in_leaf': 1000, \n    'feature_fraction': 0.7, \n    'bagging_freq':     1,\n    'bagging_fraction': 0.7,\n    'lambda_l1':       1, \n    'lambda_l2':       1, \n    'sigmoid':         1,\n    'verbosity':       0, \n}",
    "326964": "Thanks for sharing!!",
    "327741": "I'm confused by lagged_dl_counts_ratio. `in day1, this features are missing.` If they are missing, is it good for training?",
    "327828": "It doesn't give bad effect for me. If there is no sample or no positive sample in previous day, I set the value 0. so, there are no NaN samples in day2, day3, day4 and no problem happened.",
    "329920": "What a amazing feature engineering!"
  },
  "source": "meta"
}