{
  "id": 56777,
  "title": "My first kaggle competition journey and what I have learned from winning teams",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56777",
  "author_name": "",
  "post_date": "2018-05-14T22:49:19.153051700Z",
  "votes": 13,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Big thanks to sponsor <em>TalkingData</em> and <em>Kaggle</em> for providing such an interesting competition. Congradulations to those top teams and appreciations of kernal contributions from @Pranav Pandya and <a href=\"/anttip\">@anttip</a>. </p>\n\n<h2>Some thoughts</h2>\n\n<p>This is my first Kaggle competition and I can't tell you how much fun for me being a part of it. I had a fulltime job and I knew that I can only commit my weekend free time to the competition. As a newby on kaggle, I did not anticipate a good LB score at all before going into the competition. Just one week right before the final submission deadline, I was so pumped up that I got myself solo ranked <strong>top 3%</strong> in <em>public LB</em>. However, that didn't last long, and my final submission is ranked <strong>top 15%</strong> in <em>private LB</em>, which I think it is a reasonable rank for me. Overall, I think this is one of the most competitive competition and it's very hard to get in <strong>top 5%</strong> without a team.</p>\n\n<p>My results are shown below, I won't share too much about my strategy because it's not a winning strategy anyway and most of my stuff is taken from public kernels. However, I will share what I have learned and what makes a winning strategy. For those who are curious about my strategy, you are more than welcome to come to <a href=\"https://github.com/KevinLiao159/TalkingData\">my git repo</a> and check it out.</p>\n\n<h2>My Model and LB score (AUC-ROC)</h2>\n\n<p>model definition can be found in <a href=\"https://github.com/KevinLiao159/TalkingData/blob/master/scripts/train_lightgbm-v3.py\">scripts/train_lightgbm-v3.py</a></p>\n\n<p>feature engineering can be found in <a href=\"https://github.com/KevinLiao159/TalkingData/blob/master/scripts/feature_eng-v3.py\">scripts/feature_eng-v3.py</a></p>\n\n<ul>\n<li><p><strong>model</strong> LGBM with 42 (36 numerical, 6 categorical) features.</p>\n\n<ul><li>public score: 0.9806721</li>\n<li>private score: 0.9811112</li>\n<li>final rank: 586th (<em>top15%</em>)</li></ul></li>\n</ul>\n\n<h2>What I have learned from kaggle competition winners?</h2>\n\n<p><strong>We have to understand the game before wasting time.</strong></p>\n\n<p>In this compeition, the data set is huge but we only have six features. This means that 1). we need to spend a lot of time in feature engineering 2). feature engineering and model validation cycle would take long time (because data is huge). Unless we have a good team, time resource allocation is crucial in this particular competition. A suggested time table will be like following <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\">6th place solution</a>:</p>\n\n<ul>\n<li>80% feature engineering</li>\n<li>10% making local validation as fast as possible</li>\n<li>5% hyper parameter tuning</li>\n<li>5% ensembling</li>\n</ul>\n\n<p><strong>Establishing a high speed research cycle is the key to win</strong></p>\n\n<p>This competition is about training model in past historical data and predicting future fraudulant clicks (which is a big-time imbalanced classification). For imbalanced future classification problem, using tradititonal five-fold cross-validation may not be a good strategy (or you have to be really careful about sampling ratio, the timing and future information leakage).</p>\n\n<ol>\n<li><p>Basic strategy: a good practice research framework for this kind would be like following <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\">6th place solution</a>:</p>\n\n<ul><li><p>Understanding that training data starts from day 7 and ends at day 9. Testing data is day 10, in hours of 4, 5, 9, 10, 13, 14.</p></li>\n<li><p>Introducing a insample hold-out set bright line. So we can enforce a bright line between day 8 and day 9 for insample research cycle.</p></li>\n<li><p>Training on day &lt;= 8, and validating on both day 9 - hour 4 (mirror public LP), and day-9, hours 5, 9, 10, 13, 14 (mirror private LP).</p></li>\n<li><p>For out-of-sample (public LB score) iteration, we retrain on all data using 1.2 times the number of trees found by early stopping in insample validation</p></li></ul></li>\n<li><p>Advanced strategy: a fast run-time and light weight memory usage iteration would be <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56475\">1st place solution</a>:</p>\n\n<ul><li><p>Understanding that there are 99.85% of negative examples in the data and dropping out tons of negative example DOES NOT deteriorate out-of-sample performance.</p></li>\n<li><p>Using negative down-sampling strategy, which means that we use all positive examples (i.e., is_attributed == 1) and down-sampled negative examples on model training. We down-sampled negative examples such that their size becomes equal to the number of positive ones. It discards about 99.8% of negative examples.</p></li>\n<li><p>Using sample bagging technique, which means we bag five predictors trained on five sampled datasets created from different random seeds.</p></li>\n<li><p>This technique allows us to use hundreds of features while keeping LGB training time less than 30 minutes.</p></li>\n<li><p>Or use <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105\">memory trick</a> in numpy </p></li></ul></li>\n<li><p>Good principle: keep your insample hold-out score align with LB score:</p>\n\n<ul><li><p>Do not rely solely on either pulic LB score or insample hold-out score. If you do that, you will end up overfitting to one of them eventually </p></li>\n<li><p>Discard features that increase the gap between pulic LB score and insample hold-out score even though it increases your insample hold-out score</p></li></ul></li>\n</ol>\n\n<p><strong>Feature engineering is the winning secret sauce</strong></p>\n\n<p>We have five original categorical features and one timestamp feature in the data set. Unless you have some crazy NN models with proper data preprocessing (<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262\">3rd place solution</a>), you definitely need some magic features to separate youself from the crowd. If you have no idea about how to engineer some new features, please see this <a href=\"https://github.com/h2oai/h2o-meetups/blob/master/2017_11_29_Feature_Engineering/Feature%20Engineering.pdf\">good feature engineering guidance</a>.</p>\n\n<p>Here are some general ideas taken from top winners:</p>\n\n<ul>\n<li><p>dropping original worse-than-noise features [<em>ip</em>, maybe <em>device</em>]</p></li>\n<li><p>encode timestamp into day and hour</p></li>\n<li><p>user concepts: ip, device, os triplets</p></li>\n<li><p>(require brute-force) aggregates on various feature groups (click series-based feature sets (i.e., each feature set consists of 31 (=(2^5) - 1) features)) <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56475\">1st place solution</a></p>\n\n<ul><li>count features, unique count features, cumcount features</li>\n<li>time delta with previous value, delta with next value</li>\n<li>mean and variance with respect to hour</li>\n<li>standard target encoding</li>\n<li><a href=\"https://github.com/h2oai/h2o-meetups/blob/master/2017_11_29_Feature_Engineering/Feature%20Engineering.pdf\">Weights of Evidence target encoding</a></li></ul></li>\n<li><p>ratios features</p>\n\n<ul><li>number of clicks per ip, app to number of click per app</li>\n<li>nunique_counts_ratio</li>\n<li>top_counts_ratio</li></ul></li>\n<li><p>magic additions:</p>\n\n<ul><li><p>feature extraction (topic models): categorical feature embedding by using LDA/NMF/LSA <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56475\">1st place solution</a></p></li>\n<li><p>matrix factorization: truncated svd from sklearn and FM-like embedding</p></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268\">data leakage in test set</a> </p></li>\n</ul>\n\n<p><strong>Appropriate models for categorical features with large data</strong></p>\n\n<ul>\n<li><p>LightGBM is crowned over XGBoost in this competition in terms of memory usage and run-time optimization</p></li>\n<li><p>Do NOT spend too much time on hyper-param tuning (not too much juice from hyper-params)</p></li>\n<li><p>Some NN models for me to learn </p>\n\n<ul><li><p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56328\">2nd place solution</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262\">3rd place solution</a></p></li>\n<li><p><a href=\"https://github.com/CuteChibiko/TalkingData/blob/master/model.png\">4th place solution</a></p></li>\n<li><p><a href=\"https://www.ibm.com/developerworks/community/blogs/jfp/entry/Implementing_Libfm_in_Keras?lang=en_us\">libFM in Keras</a></p></li></ul></li>\n</ul>\n\n<p><strong>Extra slight boost from ensembling</strong></p>\n\n<ul>\n<li><p>most people ensemble their predictions based on LB score</p></li>\n<li><p>good practice in blending - average the logit of the predictions (aka raw predictions)</p></li>\n<li><p><a href=\"https://github.com/kaz-Anova/StackNet#restacking-mode\">restacking</a> barely helps in this competition</p></li>\n</ul>\n\n<h2>Some baseline benchmark from my observations (this is meant for ranking roughly estimates)</h2>\n\n<ol>\n<li><p>To be in <em>top 30%</em>, use solely LightGBM and trained (without too much tuning) it on some good features from public kernels </p></li>\n<li><p>To be in <em>top 20%</em>, use solely LightGBM and trained (with some proper tuning) it on at least top 40 features from public kernels (must include time delta, count, unique count types aggregates with various feature groups)</p></li>\n<li><p>To be in <em>top 10%</em>, must have beast machine (I am talking about at least 128G RAM) and train models with minimum of 100 proven-to-be-useful features or use NN models based on 20+ aggregate level features </p></li>\n<li><p>To be in <em>top 5%</em>, all above + feature extraction (categorical feature embedding) or FM-like algos</p></li>\n<li><p>To be in <em>top 1%</em>, this is really hard. Not sure how to do it.</p></li>\n</ol>\n\n<h2>Reference</h2>\n\n<p>[1]<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52374\">IP address encoding issues</a></p>\n\n<p>[2]<a href=\"https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683\">EDA by @Pranav Pandya</a></p>\n\n<p>[3]<a href=\"https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9752/code\">FM_FTRL by </a><a href=\"/anttip\">@anttip</a></p>\n\n<p>[4]<a href=\"http://quinonero.net/Publications/predicting-clicks-facebook.pdf\">Practical Lessons from Predicting Clicks on Ads at Facebook</a></p>\n\n<p>[5]<a href=\"https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf\">Ad Click Prediction: a View from the Trenches</a></p>\n\n<p>[6]<a href=\"https://github.com/KevinLiao159/TalkingData\">My TalkingData git repo</a></p>",
  "messages": [
    {
      "id": "328687",
      "postDate": "05/14/2018 22:49:19",
      "content": "<p>Big thanks to sponsor <em>TalkingData</em> and <em>Kaggle</em> for providing such an interesting competition. Congradulations to those top teams and appreciations of kernal contributions from @Pranav Pandya and <a href=\"/anttip\">@anttip</a>. </p>\n\n<h2>Some thoughts</h2>\n\n<p>This is my first Kaggle competition and I can't tell you how much fun for me being a part of it. I had a fulltime job and I knew that I can only commit my weekend free time to the competition. As a newby on kaggle, I did not anticipate a good LB score at all before going into the competition. Just one week right before the final submission deadline, I was so pumped up that I got myself solo ranked <strong>top 3%</strong> in <em>public LB</em>. However, that didn't last long, and my final submission is ranked <strong>top 15%</strong> in <em>private LB</em>, which I think it is a reasonable rank for me. Overall, I think this is one of the most competitive competition and it's very hard to get in <strong>top 5%</strong> without a team.</p>\n\n<p>My results are shown below, I won't share too much about my strategy because it's not a winning strategy anyway and most of my stuff is taken from public kernels. However, I will share what I have learned and what makes a winning strategy. For those who are curious about my strategy, you are more than welcome to come to <a href=\"https://github.com/KevinLiao159/TalkingData\">my git repo</a> and check it out.</p>\n\n<h2>My Model and LB score (AUC-ROC)</h2>\n\n<p>model definition can be found in <a href=\"https://github.com/KevinLiao159/TalkingData/blob/master/scripts/train_lightgbm-v3.py\">scripts/train_lightgbm-v3.py</a></p>\n\n<p>feature engineering can be found in <a href=\"https://github.com/KevinLiao159/TalkingData/blob/master/scripts/feature_eng-v3.py\">scripts/feature_eng-v3.py</a></p>\n\n<ul>\n<li><p><strong>model</strong> LGBM with 42 (36 numerical, 6 categorical) features.</p>\n\n<ul><li>public score: 0.9806721</li>\n<li>private score: 0.9811112</li>\n<li>final rank: 586th (<em>top15%</em>)</li></ul></li>\n</ul>\n\n<h2>What I have learned from kaggle competition winners?</h2>\n\n<p><strong>We have to understand the game before wasting time.</strong></p>\n\n<p>In this compeition, the data set is huge but we only have six features. This means that 1). we need to spend a lot of time in feature engineering 2). feature engineering and model validation cycle would take long time (because data is huge). Unless we have a good team, time resource allocation is crucial in this particular competition. A suggested time table will be like following <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\">6th place solution</a>:</p>\n\n<ul>\n<li>80% feature engineering</li>\n<li>10% making local validation as fast as possible</li>\n<li>5% hyper parameter tuning</li>\n<li>5% ensembling</li>\n</ul>\n\n<p><strong>Establishing a high speed research cycle is the key to win</strong></p>\n\n<p>This competition is about training model in past historical data and predicting future fraudulant clicks (which is a big-time imbalanced classification). For imbalanced future classification problem, using tradititonal five-fold cross-validation may not be a good strategy (or you have to be really careful about sampling ratio, the timing and future information leakage).</p>\n\n<ol>\n<li><p>Basic strategy: a good practice research framework for this kind would be like following <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\">6th place solution</a>:</p>\n\n<ul><li><p>Understanding that training data starts from day 7 and ends at day 9. Testing data is day 10, in hours of 4, 5, 9, 10, 13, 14.</p></li>\n<li><p>Introducing a insample hold-out set bright line. So we can enforce a bright line between day 8 and day 9 for insample research cycle.</p></li>\n<li><p>Training on day &lt;= 8, and validating on both day 9 - hour 4 (mirror public LP), and day-9, hours 5, 9, 10, 13, 14 (mirror private LP).</p></li>\n<li><p>For out-of-sample (public LB score) iteration, we retrain on all data using 1.2 times the number of trees found by early stopping in insample validation</p></li></ul></li>\n<li><p>Advanced strategy: a fast run-time and light weight memory usage iteration would be <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56475\">1st place solution</a>:</p>\n\n<ul><li><p>Understanding that there are 99.85% of negative examples in the data and dropping out tons of negative example DOES NOT deteriorate out-of-sample performance.</p></li>\n<li><p>Using negative down-sampling strategy, which means that we use all positive examples (i.e., is_attributed == 1) and down-sampled negative examples on model training. We down-sampled negative examples such that their size becomes equal to the number of positive ones. It discards about 99.8% of negative examples.</p></li>\n<li><p>Using sample bagging technique, which means we bag five predictors trained on five sampled datasets created from different random seeds.</p></li>\n<li><p>This technique allows us to use hundreds of features while keeping LGB training time less than 30 minutes.</p></li>\n<li><p>Or use <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105\">memory trick</a> in numpy </p></li></ul></li>\n<li><p>Good principle: keep your insample hold-out score align with LB score:</p>\n\n<ul><li><p>Do not rely solely on either pulic LB score or insample hold-out score. If you do that, you will end up overfitting to one of them eventually </p></li>\n<li><p>Discard features that increase the gap between pulic LB score and insample hold-out score even though it increases your insample hold-out score</p></li></ul></li>\n</ol>\n\n<p><strong>Feature engineering is the winning secret sauce</strong></p>\n\n<p>We have five original categorical features and one timestamp feature in the data set. Unless you have some crazy NN models with proper data preprocessing (<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262\">3rd place solution</a>), you definitely need some magic features to separate youself from the crowd. If you have no idea about how to engineer some new features, please see this <a href=\"https://github.com/h2oai/h2o-meetups/blob/master/2017_11_29_Feature_Engineering/Feature%20Engineering.pdf\">good feature engineering guidance</a>.</p>\n\n<p>Here are some general ideas taken from top winners:</p>\n\n<ul>\n<li><p>dropping original worse-than-noise features [<em>ip</em>, maybe <em>device</em>]</p></li>\n<li><p>encode timestamp into day and hour</p></li>\n<li><p>user concepts: ip, device, os triplets</p></li>\n<li><p>(require brute-force) aggregates on various feature groups (click series-based feature sets (i.e., each feature set consists of 31 (=(2^5) - 1) features)) <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56475\">1st place solution</a></p>\n\n<ul><li>count features, unique count features, cumcount features</li>\n<li>time delta with previous value, delta with next value</li>\n<li>mean and variance with respect to hour</li>\n<li>standard target encoding</li>\n<li><a href=\"https://github.com/h2oai/h2o-meetups/blob/master/2017_11_29_Feature_Engineering/Feature%20Engineering.pdf\">Weights of Evidence target encoding</a></li></ul></li>\n<li><p>ratios features</p>\n\n<ul><li>number of clicks per ip, app to number of click per app</li>\n<li>nunique_counts_ratio</li>\n<li>top_counts_ratio</li></ul></li>\n<li><p>magic additions:</p>\n\n<ul><li><p>feature extraction (topic models): categorical feature embedding by using LDA/NMF/LSA <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56475\">1st place solution</a></p></li>\n<li><p>matrix factorization: truncated svd from sklearn and FM-like embedding</p></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268\">data leakage in test set</a> </p></li>\n</ul>\n\n<p><strong>Appropriate models for categorical features with large data</strong></p>\n\n<ul>\n<li><p>LightGBM is crowned over XGBoost in this competition in terms of memory usage and run-time optimization</p></li>\n<li><p>Do NOT spend too much time on hyper-param tuning (not too much juice from hyper-params)</p></li>\n<li><p>Some NN models for me to learn </p>\n\n<ul><li><p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56328\">2nd place solution</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262\">3rd place solution</a></p></li>\n<li><p><a href=\"https://github.com/CuteChibiko/TalkingData/blob/master/model.png\">4th place solution</a></p></li>\n<li><p><a href=\"https://www.ibm.com/developerworks/community/blogs/jfp/entry/Implementing_Libfm_in_Keras?lang=en_us\">libFM in Keras</a></p></li></ul></li>\n</ul>\n\n<p><strong>Extra slight boost from ensembling</strong></p>\n\n<ul>\n<li><p>most people ensemble their predictions based on LB score</p></li>\n<li><p>good practice in blending - average the logit of the predictions (aka raw predictions)</p></li>\n<li><p><a href=\"https://github.com/kaz-Anova/StackNet#restacking-mode\">restacking</a> barely helps in this competition</p></li>\n</ul>\n\n<h2>Some baseline benchmark from my observations (this is meant for ranking roughly estimates)</h2>\n\n<ol>\n<li><p>To be in <em>top 30%</em>, use solely LightGBM and trained (without too much tuning) it on some good features from public kernels </p></li>\n<li><p>To be in <em>top 20%</em>, use solely LightGBM and trained (with some proper tuning) it on at least top 40 features from public kernels (must include time delta, count, unique count types aggregates with various feature groups)</p></li>\n<li><p>To be in <em>top 10%</em>, must have beast machine (I am talking about at least 128G RAM) and train models with minimum of 100 proven-to-be-useful features or use NN models based on 20+ aggregate level features </p></li>\n<li><p>To be in <em>top 5%</em>, all above + feature extraction (categorical feature embedding) or FM-like algos</p></li>\n<li><p>To be in <em>top 1%</em>, this is really hard. Not sure how to do it.</p></li>\n</ol>\n\n<h2>Reference</h2>\n\n<p>[1]<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52374\">IP address encoding issues</a></p>\n\n<p>[2]<a href=\"https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683\">EDA by @Pranav Pandya</a></p>\n\n<p>[3]<a href=\"https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9752/code\">FM_FTRL by </a><a href=\"/anttip\">@anttip</a></p>\n\n<p>[4]<a href=\"http://quinonero.net/Publications/predicting-clicks-facebook.pdf\">Practical Lessons from Predicting Clicks on Ads at Facebook</a></p>\n\n<p>[5]<a href=\"https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf\">Ad Click Prediction: a View from the Trenches</a></p>\n\n<p>[6]<a href=\"https://github.com/KevinLiao159/TalkingData\">My TalkingData git repo</a></p>",
      "rawMarkdown": "Big thanks to sponsor *TalkingData* and *Kaggle* for providing such an interesting competition. Congradulations to those top teams and appreciations of kernal contributions from @Pranav Pandya and @anttip. \n\n## Some thoughts\nThis is my first Kaggle competition and I can't tell you how much fun for me being a part of it. I had a fulltime job and I knew that I can only commit my weekend free time to the competition. As a newby on kaggle, I did not anticipate a good LB score at all before going into the competition. Just one week right before the final submission deadline, I was so pumped up that I got myself solo ranked **top 3%** in *public LB*. However, that didn't last long, and my final submission is ranked **top 15%** in *private LB*, which I think it is a reasonable rank for me. Overall, I think this is one of the most competitive competition and it's very hard to get in **top 5%** without a team.\n\nMy results are shown below, I won't share too much about my strategy because it's not a winning strategy anyway and most of my stuff is taken from public kernels. However, I will share what I have learned and what makes a winning strategy. For those who are curious about my strategy, you are more than welcome to come to [my git repo](https://github.com/KevinLiao159/TalkingData) and check it out.\n\n## My Model and LB score (AUC-ROC)\nmodel definition can be found in [scripts/train_lightgbm-v3.py](https://github.com/KevinLiao159/TalkingData/blob/master/scripts/train_lightgbm-v3.py)\n\nfeature engineering can be found in [scripts/feature_eng-v3.py](https://github.com/KevinLiao159/TalkingData/blob/master/scripts/feature_eng-v3.py)\n\n  - **model** LGBM with 42 (36 numerical, 6 categorical) features.\n\n* public score: 0.9806721\n* private score: 0.9811112\n* final rank: 586th (*top15%*)\n\n## What I have learned from kaggle competition winners?\n\n__We have to understand the game before wasting time.__\n\nIn this compeition, the data set is huge but we only have six features. This means that 1). we need to spend a lot of time in feature engineering 2). feature engineering and model validation cycle would take long time (because data is huge). Unless we have a good team, time resource allocation is crucial in this particular competition. A suggested time table will be like following [6th place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283):\n\n* 80% feature engineering\n* 10% making local validation as fast as possible\n* 5% hyper parameter tuning\n* 5% ensembling\n\n__Establishing a high speed research cycle is the key to win__\n\nThis competition is about training model in past historical data and predicting future fraudulant clicks (which is a big-time imbalanced classification). For imbalanced future classification problem, using tradititonal five-fold cross-validation may not be a good strategy (or you have to be really careful about sampling ratio, the timing and future information leakage).\n\n1. Basic strategy: a good practice research framework for this kind would be like following [6th place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283):\n\n    * Understanding that training data starts from day 7 and ends at day 9. Testing data is day 10, in hours of 4, 5, 9, 10, 13, 14.\n\n    * Introducing a insample hold-out set bright line. So we can enforce a bright line between day 8 and day 9 for insample research cycle.\n\n    * Training on day &lt;= 8, and validating on both day 9 - hour 4 (mirror public LP), and day-9, hours 5, 9, 10, 13, 14 (mirror private LP).\n\n    * For out-of-sample (public LB score) iteration, we retrain on all data using 1.2 times the number of trees found by early stopping in insample validation\n\n\n2. Advanced strategy: a fast run-time and light weight memory usage iteration would be [1st place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56475):\n\n    * Understanding that there are 99.85% of negative examples in the data and dropping out tons of negative example DOES NOT deteriorate out-of-sample performance.\n\n    * Using negative down-sampling strategy, which means that we use all positive examples (i.e., is_attributed == 1) and down-sampled negative examples on model training. We down-sampled negative examples such that their size becomes equal to the number of positive ones. It discards about 99.8% of negative examples.\n\n    * Using sample bagging technique, which means we bag five predictors trained on five sampled datasets created from different random seeds.\n\n    * This technique allows us to use hundreds of features while keeping LGB training time less than 30 minutes.\n\n    * Or use [memory trick](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105) in numpy \n\n\n3. Good principle: keep your insample hold-out score align with LB score:\n\n    * Do not rely solely on either pulic LB score or insample hold-out score. If you do that, you will end up overfitting to one of them eventually \n\n    * Discard features that increase the gap between pulic LB score and insample hold-out score even though it increases your insample hold-out score\n\n\n__Feature engineering is the winning secret sauce__\n\nWe have five original categorical features and one timestamp feature in the data set. Unless you have some crazy NN models with proper data preprocessing ([3rd place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262)), you definitely need some magic features to separate youself from the crowd. If you have no idea about how to engineer some new features, please see this [good feature engineering guidance](https://github.com/h2oai/h2o-meetups/blob/master/2017_11_29_Feature_Engineering/Feature%20Engineering.pdf).\n\nHere are some general ideas taken from top winners:\n\n* dropping original worse-than-noise features [*ip*, maybe *device*]\n\n* encode timestamp into day and hour\n\n* user concepts: ip, device, os triplets\n\n* (require brute-force) aggregates on various feature groups (click series-based feature sets (i.e., each feature set consists of 31 (=(2^5) - 1) features)) [1st place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56475)\n    * count features, unique count features, cumcount features\n    * time delta with previous value, delta with next value\n    * mean and variance with respect to hour\n    * standard target encoding\n    * [Weights of Evidence target encoding](https://github.com/h2oai/h2o-meetups/blob/master/2017_11_29_Feature_Engineering/Feature%20Engineering.pdf)\n\n* ratios features\n    * number of clicks per ip, app to number of click per app\n    * nunique_counts_ratio\n    * top_counts_ratio\n\n* magic additions:\n    * feature extraction (topic models): categorical feature embedding by using LDA/NMF/LSA [1st place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56475)\n\n    * matrix factorization: truncated svd from sklearn and FM-like embedding\n\n* [data leakage in test set](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268) \n\n\n__Appropriate models for categorical features with large data__\n\n* LightGBM is crowned over XGBoost in this competition in terms of memory usage and run-time optimization\n\n* Do NOT spend too much time on hyper-param tuning (not too much juice from hyper-params)\n\n* Some NN models for me to learn \n    * [2nd place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56328)\n\n    * [3rd place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262)\n\n    * [4th place solution](https://github.com/CuteChibiko/TalkingData/blob/master/model.png)\n\n    * [libFM in Keras](https://www.ibm.com/developerworks/community/blogs/jfp/entry/Implementing_Libfm_in_Keras?lang=en_us)\n\n\n__Extra slight boost from ensembling__\n\n* most people ensemble their predictions based on LB score\n\n* good practice in blending - average the logit of the predictions (aka raw predictions)\n\n* [restacking](https://github.com/kaz-Anova/StackNet#restacking-mode) barely helps in this competition\n\n\n## Some baseline benchmark from my observations (this is meant for ranking roughly estimates)\n\n1. To be in *top 30%*, use solely LightGBM and trained (without too much tuning) it on some good features from public kernels \n\n2. To be in *top 20%*, use solely LightGBM and trained (with some proper tuning) it on at least top 40 features from public kernels (must include time delta, count, unique count types aggregates with various feature groups)\n\n3. To be in *top 10%*, must have beast machine (I am talking about at least 128G RAM) and train models with minimum of 100 proven-to-be-useful features or use NN models based on 20+ aggregate level features \n\n4. To be in *top 5%*, all above + feature extraction (categorical feature embedding) or FM-like algos\n\n5. To be in *top 1%*, this is really hard. Not sure how to do it.\n\n\n## Reference\n\n[1][IP address encoding issues](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52374)\n\n[2][EDA by @Pranav Pandya](https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683)\n\n[3][FM_FTRL by @anttip](https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9752/code)\n\n[4][Practical Lessons from Predicting Clicks on Ads at Facebook](http://quinonero.net/Publications/predicting-clicks-facebook.pdf)\n\n[5][Ad Click Prediction: a View from the Trenches](https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf)\n\n[6][My TalkingData git repo](https://github.com/KevinLiao159/TalkingData)",
      "votes": null
    },
    {
      "id": "328821",
      "postDate": "05/15/2018 06:24:54",
      "content": "<p>Good. A good summary. Thank you </p>",
      "rawMarkdown": "Good. A good summary. Thank you",
      "votes": null
    },
    {
      "id": "328848",
      "postDate": "05/15/2018 07:20:39",
      "content": "<p>Interesting summary.</p>\n\n<blockquote>\n  <p>To be in top 10%, must have beast machine (I am talking about at least 128G RAM) and train models with minimum of 100 proven-to-be-useful features or use NN models based on 20+ aggregate level features </p>\n</blockquote>\n\n<p>I disagree with the above.  I was 6th with a 64 GB machine and 48 features only.</p>",
      "rawMarkdown": "Interesting summary.\n\n&gt; To be in top 10%, must have beast machine (I am talking about at least 128G RAM) and train models with minimum of 100 proven-to-be-useful features or use NN models based on 20+ aggregate level features \n\nI disagree with the above.  I was 6th with a 64 GB machine and 48 features only.",
      "votes": null
    },
    {
      "id": "328997",
      "postDate": "05/15/2018 13:50:45",
      "content": "<p>Thanks for this!\nThere is a lot of chalanges on those competitions, is good to see a summary like this.</p>",
      "rawMarkdown": "Thanks for this!\nThere is a lot of chalanges on those competitions, is good to see a summary like this.",
      "votes": null
    },
    {
      "id": "329070",
      "postDate": "05/15/2018 16:57:15",
      "content": "<p>Haha! You are right. 128G is overstated. But it has to be better than a personal laptop with only 16G RAM. It would be hard to imagine a person with slow iteration speed can find a good strategy to be in top 10%</p>",
      "rawMarkdown": "Haha! You are right. 128G is overstated. But it has to be better than a personal laptop with only 16G RAM. It would be hard to imagine a person with slow iteration speed can find a good strategy to be in top 10%",
      "votes": null
    },
    {
      "id": "329919",
      "postDate": "05/17/2018 15:02:46",
      "content": "<p>Thank you for the summary. </p>",
      "rawMarkdown": "Thank you for the summary.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 328821,
      "author_name": "huaguo",
      "author_url": "",
      "post_date": "05/15/2018 06:24:54",
      "content": "<p>Good. A good summary. Thank you </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 328848,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/15/2018 07:20:39",
      "content": "<p>Interesting summary.</p>\n\n<blockquote>\n  <p>To be in top 10%, must have beast machine (I am talking about at least 128G RAM) and train models with minimum of 100 proven-to-be-useful features or use NN models based on 20+ aggregate level features </p>\n</blockquote>\n\n<p>I disagree with the above.  I was 6th with a 64 GB machine and 48 features only.</p>",
      "votes": null,
      "replies": [
        {
          "id": 329070,
          "author_name": "lwk723",
          "author_url": "",
          "post_date": "05/15/2018 16:57:15",
          "content": "<p>Haha! You are right. 128G is overstated. But it has to be better than a personal laptop with only 16G RAM. It would be hard to imagine a person with slow iteration speed can find a good strategy to be in top 10%</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 328997,
      "author_name": "gledsonpicharski",
      "author_url": "",
      "post_date": "05/15/2018 13:50:45",
      "content": "<p>Thanks for this!\nThere is a lot of chalanges on those competitions, is good to see a summary like this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 329919,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "05/17/2018 15:02:46",
      "content": "<p>Thank you for the summary. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "328687": "Big thanks to sponsor *TalkingData* and *Kaggle* for providing such an interesting competition. Congradulations to those top teams and appreciations of kernal contributions from @Pranav Pandya and @anttip. \n\n## Some thoughts\nThis is my first Kaggle competition and I can't tell you how much fun for me being a part of it. I had a fulltime job and I knew that I can only commit my weekend free time to the competition. As a newby on kaggle, I did not anticipate a good LB score at all before going into the competition. Just one week right before the final submission deadline, I was so pumped up that I got myself solo ranked **top 3%** in *public LB*. However, that didn't last long, and my final submission is ranked **top 15%** in *private LB*, which I think it is a reasonable rank for me. Overall, I think this is one of the most competitive competition and it's very hard to get in **top 5%** without a team.\n\nMy results are shown below, I won't share too much about my strategy because it's not a winning strategy anyway and most of my stuff is taken from public kernels. However, I will share what I have learned and what makes a winning strategy. For those who are curious about my strategy, you are more than welcome to come to [my git repo](https://github.com/KevinLiao159/TalkingData) and check it out.\n\n## My Model and LB score (AUC-ROC)\nmodel definition can be found in [scripts/train_lightgbm-v3.py](https://github.com/KevinLiao159/TalkingData/blob/master/scripts/train_lightgbm-v3.py)\n\nfeature engineering can be found in [scripts/feature_eng-v3.py](https://github.com/KevinLiao159/TalkingData/blob/master/scripts/feature_eng-v3.py)\n\n  - **model** LGBM with 42 (36 numerical, 6 categorical) features.\n\n* public score: 0.9806721\n* private score: 0.9811112\n* final rank: 586th (*top15%*)\n\n## What I have learned from kaggle competition winners?\n\n__We have to understand the game before wasting time.__\n\nIn this compeition, the data set is huge but we only have six features. This means that 1). we need to spend a lot of time in feature engineering 2). feature engineering and model validation cycle would take long time (because data is huge). Unless we have a good team, time resource allocation is crucial in this particular competition. A suggested time table will be like following [6th place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283):\n\n* 80% feature engineering\n* 10% making local validation as fast as possible\n* 5% hyper parameter tuning\n* 5% ensembling\n\n__Establishing a high speed research cycle is the key to win__\n\nThis competition is about training model in past historical data and predicting future fraudulant clicks (which is a big-time imbalanced classification). For imbalanced future classification problem, using tradititonal five-fold cross-validation may not be a good strategy (or you have to be really careful about sampling ratio, the timing and future information leakage).\n\n1. Basic strategy: a good practice research framework for this kind would be like following [6th place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283):\n\n    * Understanding that training data starts from day 7 and ends at day 9. Testing data is day 10, in hours of 4, 5, 9, 10, 13, 14.\n\n    * Introducing a insample hold-out set bright line. So we can enforce a bright line between day 8 and day 9 for insample research cycle.\n\n    * Training on day &lt;= 8, and validating on both day 9 - hour 4 (mirror public LP), and day-9, hours 5, 9, 10, 13, 14 (mirror private LP).\n\n    * For out-of-sample (public LB score) iteration, we retrain on all data using 1.2 times the number of trees found by early stopping in insample validation\n\n\n2. Advanced strategy: a fast run-time and light weight memory usage iteration would be [1st place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56475):\n\n    * Understanding that there are 99.85% of negative examples in the data and dropping out tons of negative example DOES NOT deteriorate out-of-sample performance.\n\n    * Using negative down-sampling strategy, which means that we use all positive examples (i.e., is_attributed == 1) and down-sampled negative examples on model training. We down-sampled negative examples such that their size becomes equal to the number of positive ones. It discards about 99.8% of negative examples.\n\n    * Using sample bagging technique, which means we bag five predictors trained on five sampled datasets created from different random seeds.\n\n    * This technique allows us to use hundreds of features while keeping LGB training time less than 30 minutes.\n\n    * Or use [memory trick](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105) in numpy \n\n\n3. Good principle: keep your insample hold-out score align with LB score:\n\n    * Do not rely solely on either pulic LB score or insample hold-out score. If you do that, you will end up overfitting to one of them eventually \n\n    * Discard features that increase the gap between pulic LB score and insample hold-out score even though it increases your insample hold-out score\n\n\n__Feature engineering is the winning secret sauce__\n\nWe have five original categorical features and one timestamp feature in the data set. Unless you have some crazy NN models with proper data preprocessing ([3rd place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262)), you definitely need some magic features to separate youself from the crowd. If you have no idea about how to engineer some new features, please see this [good feature engineering guidance](https://github.com/h2oai/h2o-meetups/blob/master/2017_11_29_Feature_Engineering/Feature%20Engineering.pdf).\n\nHere are some general ideas taken from top winners:\n\n* dropping original worse-than-noise features [*ip*, maybe *device*]\n\n* encode timestamp into day and hour\n\n* user concepts: ip, device, os triplets\n\n* (require brute-force) aggregates on various feature groups (click series-based feature sets (i.e., each feature set consists of 31 (=(2^5) - 1) features)) [1st place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56475)\n    * count features, unique count features, cumcount features\n    * time delta with previous value, delta with next value\n    * mean and variance with respect to hour\n    * standard target encoding\n    * [Weights of Evidence target encoding](https://github.com/h2oai/h2o-meetups/blob/master/2017_11_29_Feature_Engineering/Feature%20Engineering.pdf)\n\n* ratios features\n    * number of clicks per ip, app to number of click per app\n    * nunique_counts_ratio\n    * top_counts_ratio\n\n* magic additions:\n    * feature extraction (topic models): categorical feature embedding by using LDA/NMF/LSA [1st place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56475)\n\n    * matrix factorization: truncated svd from sklearn and FM-like embedding\n\n* [data leakage in test set](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268) \n\n\n__Appropriate models for categorical features with large data__\n\n* LightGBM is crowned over XGBoost in this competition in terms of memory usage and run-time optimization\n\n* Do NOT spend too much time on hyper-param tuning (not too much juice from hyper-params)\n\n* Some NN models for me to learn \n    * [2nd place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56328)\n\n    * [3rd place solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262)\n\n    * [4th place solution](https://github.com/CuteChibiko/TalkingData/blob/master/model.png)\n\n    * [libFM in Keras](https://www.ibm.com/developerworks/community/blogs/jfp/entry/Implementing_Libfm_in_Keras?lang=en_us)\n\n\n__Extra slight boost from ensembling__\n\n* most people ensemble their predictions based on LB score\n\n* good practice in blending - average the logit of the predictions (aka raw predictions)\n\n* [restacking](https://github.com/kaz-Anova/StackNet#restacking-mode) barely helps in this competition\n\n\n## Some baseline benchmark from my observations (this is meant for ranking roughly estimates)\n\n1. To be in *top 30%*, use solely LightGBM and trained (without too much tuning) it on some good features from public kernels \n\n2. To be in *top 20%*, use solely LightGBM and trained (with some proper tuning) it on at least top 40 features from public kernels (must include time delta, count, unique count types aggregates with various feature groups)\n\n3. To be in *top 10%*, must have beast machine (I am talking about at least 128G RAM) and train models with minimum of 100 proven-to-be-useful features or use NN models based on 20+ aggregate level features \n\n4. To be in *top 5%*, all above + feature extraction (categorical feature embedding) or FM-like algos\n\n5. To be in *top 1%*, this is really hard. Not sure how to do it.\n\n\n## Reference\n\n[1][IP address encoding issues](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52374)\n\n[2][EDA by @Pranav Pandya](https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683)\n\n[3][FM_FTRL by @anttip](https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9752/code)\n\n[4][Practical Lessons from Predicting Clicks on Ads at Facebook](http://quinonero.net/Publications/predicting-clicks-facebook.pdf)\n\n[5][Ad Click Prediction: a View from the Trenches](https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf)\n\n[6][My TalkingData git repo](https://github.com/KevinLiao159/TalkingData)",
    "328821": "Good. A good summary. Thank you",
    "328848": "Interesting summary.\n\n&gt; To be in top 10%, must have beast machine (I am talking about at least 128G RAM) and train models with minimum of 100 proven-to-be-useful features or use NN models based on 20+ aggregate level features \n\nI disagree with the above.  I was 6th with a 64 GB machine and 48 features only.",
    "328997": "Thanks for this!\nThere is a lot of chalanges on those competitions, is good to see a summary like this.",
    "329070": "Haha! You are right. 128G is overstated. But it has to be better than a personal laptop with only 16G RAM. It would be hard to imagine a person with slow iteration speed can find a good strategy to be in top 10%",
    "329919": "Thank you for the summary."
  },
  "source": "meta"
}