{
  "id": 56429,
  "title": "Solution #50 Story - One Week Hackathon ",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/writeups/diveintotd-solution-50-story-one-week-hackathon",
  "author_name": "",
  "post_date": "2018-05-10T08:30:20.720Z",
  "votes": 27,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First of all, let me say up-front that I don't particularly think we have much to boast about this result - but it was an effort of very nice team effort in a very nice competition that stretches our RAM management techniques</p>\n\n<p>We did this competition as an ongoing effort to work with friends in <a href=\"https://diveintocode.jp/ai_curriculum\">DiveIntoCode</a> to provide data science mentorship to their students - so we spend most of our effort to help students to understand the <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54765\">business context of click fraud</a>, as well as creating <a href=\"https://www.kaggle.com/ogrellier/basic-predictions-using-target-encoding-on-app/code\">friendly kernels</a> to ease our students into handling the rather large dataset. So it wasn't until the last week we decided to make a stronger effort on the LB, and we kind of end-up doing this in a hackathon manner in the last 3-4 days. Our best submissions were submitted and selected within the last 20 minutes before the close of play, so that was entirely bonkers and fun way to keep kaggle addictive :)</p>\n\n<p>So here we go, here is a description of our approach:</p>\n\n<p><strong>Information Sharing</strong>\nWe use slack for communication, google drive for sharing feature/data, and then a google spreadsheet to manage various aspect of information. For us, this way of working turned out to be very important - we started so late and needed a way to gather \"intelligence\" that is disclosed in via kernels and discussion. We specifically have a dedicated list of bookmarks shared among us, and also key EDA results (both from various kernels but also our own effort) captured in a shared google slide</p>\n\n<p><strong>Modelling Approach</strong>\nSince we started so late, we had very little choice but to start with the most secure methods - that means LightGBM, and midway through we discovered that extracting features using the test supplement gave a much better score.  We also ruled out traditional K-fold CV setting in the belief that it would take too much time for us to iterate - well, we now know there are much more clever methods for faster iteration as shown by <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56328\">PPP's solution</a> by <a href=\"/panfeiyang\">@panfeiyang</a>.</p>\n\n<p>The one thing I would recommend people to check out is the<code>boosting=dart</code> option from LightGBM. As <a href=\"/kazanova\">@kazanova</a> kindly shared in his <a href=\"https://www.kaggle.com/c/instacart-market-basket-analysis/discussion/38100\">Instartcart solution</a> Dart tree can out-perform typical GBM. What I have since discovered by experience is that it tends to outperform GBM when the dataset is relatively large. My intuition is that this allows the dropout mechanism to fully explore its potential.  I had tried to use dart tee on the much smaller dataset from Porto Seguro without success at all. </p>\n\n<p>Nevertheless, tuning Dart tree can be quite tricky, some of the typical logic on GBM doesn't apply for instance decreasing learning rate doesn't necessarily lead to better accuracy because it interacts with drop-out related parameters.  Anyway, we started with the parameter shared by <a href=\"/kazanova\">@kazanova</a> and started out tuning using the last 25millions rows in training data as local validation. Here is one of several sets of parameters that gave us good model:</p>\n\n<pre><code>num_leaves': 22,\nmin_sum_hessian_in_leaf':20,\nmax_depth': 7,\nlearning_rate': 0.2,    \nnum_threads': NUM_CORES,\nfeature_fraction':0.7,\nbagging_fraction':0.85,\nmax_drop':6,\ndrop_rate':0.01,\nmin_data_in_leaf':10,\nbagging_freq': 1,\nscale_pos_weight':200,\n</code></pre>\n\n<p>I don't have deep knowledge about how dart trees work, so I can only speak by experience. I shall now go and have a good read at <a href=\"https://arxiv.org/abs/1505.01866\">this paper</a>, folks with better knowledge please also share your experience and we would be much appreciated!</p>\n\n<p><strong>Feature Engineering</strong>\nnot much to share here, we were mainly scavaging features from shared kernels, but <a href=\"/ogrellier\">@ogrellier</a> did implement a much neater way to compute \"time delta\" features. In addition to channel, app, os, and hour which were the strongest, we used different variations of next clicks and previous clicks - we generated additional features like \"the click after next click\" and some of them played an important role in our models. </p>\n\n<p>Other features are your usual count and unique_count features.  We also have a <code>click_rate</code> feature which measures the frequent were the clicks from specific IP during the day which helped a bit   By the last day of the competition, we had 34 features,  we could have generated more but time was running out.</p>\n\n<p>A work on feature selection: one by one feature selection was too long a process for us, so throughout the week I have been throwing 5-6 features by batch, and keep them if they increase the local validation score which thankfully corresponded well with LB.  I personally trusted a lot on how LightGBM prioritises features, and I operate in the following principles: </p>\n\n<ul>\n<li>whenever a batch of features improve local validation score, keep them</li>\n<li>whenever a batch of feature decrease score - I will remove the most import features among this batch according to LightGBM, and re-test</li>\n<li>whenever features were not used by LightGBM at all i.e. 0 importance score I will drop them. </li>\n</ul>\n\n<p>Not sure this is the best approach, but in general it worked in this instance, and our ultimate constraint was that we couldn't manufacture enough feature fast enough within a limited amount of time - ideally we would like to have more than 50+ features which I believe will improve (or overfit) our models more.  </p>\n\n<p><strong>Post-processing</strong>\nSo again this little secret was kindly shared by <a href=\"/plantsgo\">@plantsgo</a>, and it gave a 0.0002 boost to our models, we weren't so sure if this will overfit private LB, so we submit one with post-processing and one without.</p>\n\n<p><strong>The End Game</strong>\nWe spent a large part of the last two day trying to get FM working - in the hope that we could have an alternative model to merge with LightGBM - we weren't familiar with FM algorithms and it took a lot of time without getting useful results.  In highlight we could have better invested this time in NN - we took a choice and it didn't work out so next time we will approach with new knowledge. I personally enjoyed learning a lot from <a href=\"/anttip\">@anttip</a>'s wonderful FM kernel so no harm was done there. </p>\n\n<p>What caused a bit of panic was that we had a bug while trying to reduce the RAM requirement of our data, and it caused invalid values in our predictions - we spent hours fixing that while testing our LightGBM model with various parameters. Lucky the bug was fixed, and we manage to submit our best solutions within the last 20 minutes of the competition!  The final solution was a weight average of two LightGBM model, so again pretty boring stuff :)</p>\n\n<p><strong>Take Away</strong></p>\n\n<ul>\n<li>Dart tree is good, in fact very good for large dataset - our parameters for dart tree was faster and more accurate than the gbm equivalent </li>\n<li>Trust NN - I really should have considered using NN as an alternative model - especially since I notice the drop-out mechanism from dart tree was working well for this dataset. </li>\n<li>START EARLIER IF YOU WANT GOLD: well, at least this applies to us\naverage folks. I think we could have generated a lot better solution\neven if one more week - but then it might be next intense and less\nfun. ^_^</li>\n<li>[update] Downsampling when the data is large and imbalanced - now looking at both the 1st and 2nd solutions, it is clear that downsampling on negative observation were key steps for top teams - it didn't work out for me in the past, but I guess this depends on dataset. I will make sure I try this next time! </li>\n</ul>\n\n<p>By the way, we are the biggest overfitter in top 50 (i.e. biggest drop between public and private LB), so perhaps this is exactly what you shouldn't be doing if you want to get a gold medal - LoL </p>",
  "messages": [
    {
      "id": "326403",
      "postDate": "05/09/2018 16:33:46",
      "content": "<p>First of all, let me say up-front that I don't particularly think we have much to boast about this result - but it was an effort of very nice team effort in a very nice competition that stretches our RAM management techniques</p>\n\n<p>We did this competition as an ongoing effort to work with friends in <a href=\"https://diveintocode.jp/ai_curriculum\">DiveIntoCode</a> to provide data science mentorship to their students - so we spend most of our effort to help students to understand the <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54765\">business context of click fraud</a>, as well as creating <a href=\"https://www.kaggle.com/ogrellier/basic-predictions-using-target-encoding-on-app/code\">friendly kernels</a> to ease our students into handling the rather large dataset. So it wasn't until the last week we decided to make a stronger effort on the LB, and we kind of end-up doing this in a hackathon manner in the last 3-4 days. Our best submissions were submitted and selected within the last 20 minutes before the close of play, so that was entirely bonkers and fun way to keep kaggle addictive :)</p>\n\n<p>So here we go, here is a description of our approach:</p>\n\n<p><strong>Information Sharing</strong>\nWe use slack for communication, google drive for sharing feature/data, and then a google spreadsheet to manage various aspect of information. For us, this way of working turned out to be very important - we started so late and needed a way to gather \"intelligence\" that is disclosed in via kernels and discussion. We specifically have a dedicated list of bookmarks shared among us, and also key EDA results (both from various kernels but also our own effort) captured in a shared google slide</p>\n\n<p><strong>Modelling Approach</strong>\nSince we started so late, we had very little choice but to start with the most secure methods - that means LightGBM, and midway through we discovered that extracting features using the test supplement gave a much better score.  We also ruled out traditional K-fold CV setting in the belief that it would take too much time for us to iterate - well, we now know there are much more clever methods for faster iteration as shown by <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56328\">PPP's solution</a> by <a href=\"/panfeiyang\">@panfeiyang</a>.</p>\n\n<p>The one thing I would recommend people to check out is the<code>boosting=dart</code> option from LightGBM. As <a href=\"/kazanova\">@kazanova</a> kindly shared in his <a href=\"https://www.kaggle.com/c/instacart-market-basket-analysis/discussion/38100\">Instartcart solution</a> Dart tree can out-perform typical GBM. What I have since discovered by experience is that it tends to outperform GBM when the dataset is relatively large. My intuition is that this allows the dropout mechanism to fully explore its potential.  I had tried to use dart tee on the much smaller dataset from Porto Seguro without success at all. </p>\n\n<p>Nevertheless, tuning Dart tree can be quite tricky, some of the typical logic on GBM doesn't apply for instance decreasing learning rate doesn't necessarily lead to better accuracy because it interacts with drop-out related parameters.  Anyway, we started with the parameter shared by <a href=\"/kazanova\">@kazanova</a> and started out tuning using the last 25millions rows in training data as local validation. Here is one of several sets of parameters that gave us good model:</p>\n\n<pre><code>num_leaves': 22,\nmin_sum_hessian_in_leaf':20,\nmax_depth': 7,\nlearning_rate': 0.2,    \nnum_threads': NUM_CORES,\nfeature_fraction':0.7,\nbagging_fraction':0.85,\nmax_drop':6,\ndrop_rate':0.01,\nmin_data_in_leaf':10,\nbagging_freq': 1,\nscale_pos_weight':200,\n</code></pre>\n\n<p>I don't have deep knowledge about how dart trees work, so I can only speak by experience. I shall now go and have a good read at <a href=\"https://arxiv.org/abs/1505.01866\">this paper</a>, folks with better knowledge please also share your experience and we would be much appreciated!</p>\n\n<p><strong>Feature Engineering</strong>\nnot much to share here, we were mainly scavaging features from shared kernels, but <a href=\"/ogrellier\">@ogrellier</a> did implement a much neater way to compute \"time delta\" features. In addition to channel, app, os, and hour which were the strongest, we used different variations of next clicks and previous clicks - we generated additional features like \"the click after next click\" and some of them played an important role in our models. </p>\n\n<p>Other features are your usual count and unique_count features.  We also have a <code>click_rate</code> feature which measures the frequent were the clicks from specific IP during the day which helped a bit   By the last day of the competition, we had 34 features,  we could have generated more but time was running out.</p>\n\n<p>A work on feature selection: one by one feature selection was too long a process for us, so throughout the week I have been throwing 5-6 features by batch, and keep them if they increase the local validation score which thankfully corresponded well with LB.  I personally trusted a lot on how LightGBM prioritises features, and I operate in the following principles: </p>\n\n<ul>\n<li>whenever a batch of features improve local validation score, keep them</li>\n<li>whenever a batch of feature decrease score - I will remove the most import features among this batch according to LightGBM, and re-test</li>\n<li>whenever features were not used by LightGBM at all i.e. 0 importance score I will drop them. </li>\n</ul>\n\n<p>Not sure this is the best approach, but in general it worked in this instance, and our ultimate constraint was that we couldn't manufacture enough feature fast enough within a limited amount of time - ideally we would like to have more than 50+ features which I believe will improve (or overfit) our models more.  </p>\n\n<p><strong>Post-processing</strong>\nSo again this little secret was kindly shared by <a href=\"/plantsgo\">@plantsgo</a>, and it gave a 0.0002 boost to our models, we weren't so sure if this will overfit private LB, so we submit one with post-processing and one without.</p>\n\n<p><strong>The End Game</strong>\nWe spent a large part of the last two day trying to get FM working - in the hope that we could have an alternative model to merge with LightGBM - we weren't familiar with FM algorithms and it took a lot of time without getting useful results.  In highlight we could have better invested this time in NN - we took a choice and it didn't work out so next time we will approach with new knowledge. I personally enjoyed learning a lot from <a href=\"/anttip\">@anttip</a>'s wonderful FM kernel so no harm was done there. </p>\n\n<p>What caused a bit of panic was that we had a bug while trying to reduce the RAM requirement of our data, and it caused invalid values in our predictions - we spent hours fixing that while testing our LightGBM model with various parameters. Lucky the bug was fixed, and we manage to submit our best solutions within the last 20 minutes of the competition!  The final solution was a weight average of two LightGBM model, so again pretty boring stuff :)</p>\n\n<p><strong>Take Away</strong></p>\n\n<ul>\n<li>Dart tree is good, in fact very good for large dataset - our parameters for dart tree was faster and more accurate than the gbm equivalent </li>\n<li>Trust NN - I really should have considered using NN as an alternative model - especially since I notice the drop-out mechanism from dart tree was working well for this dataset. </li>\n<li>START EARLIER IF YOU WANT GOLD: well, at least this applies to us\naverage folks. I think we could have generated a lot better solution\neven if one more week - but then it might be next intense and less\nfun. ^_^</li>\n<li>[update] Downsampling when the data is large and imbalanced - now looking at both the 1st and 2nd solutions, it is clear that downsampling on negative observation were key steps for top teams - it didn't work out for me in the past, but I guess this depends on dataset. I will make sure I try this next time! </li>\n</ul>\n\n<p>By the way, we are the biggest overfitter in top 50 (i.e. biggest drop between public and private LB), so perhaps this is exactly what you shouldn't be doing if you want to get a gold medal - LoL </p>",
      "rawMarkdown": "First of all, let me say up-front that I don't particularly think we have much to boast about this result - but it was an effort of very nice team effort in a very nice competition that stretches our RAM management techniques\n\nWe did this competition as an ongoing effort to work with friends in [DiveIntoCode][1] to provide data science mentorship to their students - so we spend most of our effort to help students to understand the [business context of click fraud][2], as well as creating [friendly kernels][3] to ease our students into handling the rather large dataset. So it wasn't until the last week we decided to make a stronger effort on the LB, and we kind of end-up doing this in a hackathon manner in the last 3-4 days. Our best submissions were submitted and selected within the last 20 minutes before the close of play, so that was entirely bonkers and fun way to keep kaggle addictive :)\n\nSo here we go, here is a description of our approach:\n\n**Information Sharing**\nWe use slack for communication, google drive for sharing feature/data, and then a google spreadsheet to manage various aspect of information. For us, this way of working turned out to be very important - we started so late and needed a way to gather \"intelligence\" that is disclosed in via kernels and discussion. We specifically have a dedicated list of bookmarks shared among us, and also key EDA results (both from various kernels but also our own effort) captured in a shared google slide\n\n**Modelling Approach**\nSince we started so late, we had very little choice but to start with the most secure methods - that means LightGBM, and midway through we discovered that extracting features using the test supplement gave a much better score.  We also ruled out traditional K-fold CV setting in the belief that it would take too much time for us to iterate - well, we now know there are much more clever methods for faster iteration as shown by [PPP's solution][4] by @panfeiyang.\n\nThe one thing I would recommend people to check out is the`boosting=dart` option from LightGBM. As @kazanova kindly shared in his [Instartcart solution][5] Dart tree can out-perform typical GBM. What I have since discovered by experience is that it tends to outperform GBM when the dataset is relatively large. My intuition is that this allows the dropout mechanism to fully explore its potential.  I had tried to use dart tee on the much smaller dataset from Porto Seguro without success at all. \n\nNevertheless, tuning Dart tree can be quite tricky, some of the typical logic on GBM doesn't apply for instance decreasing learning rate doesn't necessarily lead to better accuracy because it interacts with drop-out related parameters.  Anyway, we started with the parameter shared by @kazanova and started out tuning using the last 25millions rows in training data as local validation. Here is one of several sets of parameters that gave us good model:\n\n    num_leaves': 22,\n    min_sum_hessian_in_leaf':20,\n    max_depth': 7,\n    learning_rate': 0.2,    \n    num_threads': NUM_CORES,\n    feature_fraction':0.7,\n    bagging_fraction':0.85,\n    max_drop':6,\n    drop_rate':0.01,\n    min_data_in_leaf':10,\n    bagging_freq': 1,\n    scale_pos_weight':200,\n\n\nI don't have deep knowledge about how dart trees work, so I can only speak by experience. I shall now go and have a good read at [this paper][6], folks with better knowledge please also share your experience and we would be much appreciated!\n\n**Feature Engineering**\nnot much to share here, we were mainly scavaging features from shared kernels, but @ogrellier did implement a much neater way to compute \"time delta\" features. In addition to channel, app, os, and hour which were the strongest, we used different variations of next clicks and previous clicks - we generated additional features like \"the click after next click\" and some of them played an important role in our models. \n\nOther features are your usual count and unique_count features.  We also have a `click_rate` feature which measures the frequent were the clicks from specific IP during the day which helped a bit   By the last day of the competition, we had 34 features,  we could have generated more but time was running out.\n\nA work on feature selection: one by one feature selection was too long a process for us, so throughout the week I have been throwing 5-6 features by batch, and keep them if they increase the local validation score which thankfully corresponded well with LB.  I personally trusted a lot on how LightGBM prioritises features, and I operate in the following principles: \n\n - whenever a batch of features improve local validation score, keep them\n - whenever a batch of feature decrease score - I will remove the most import features among this batch according to LightGBM, and re-test\n - whenever features were not used by LightGBM at all i.e. 0 importance score I will drop them. \n\nNot sure this is the best approach, but in general it worked in this instance, and our ultimate constraint was that we couldn't manufacture enough feature fast enough within a limited amount of time - ideally we would like to have more than 50+ features which I believe will improve (or overfit) our models more.  \n\n**Post-processing**\nSo again this little secret was kindly shared by @plantsgo, and it gave a 0.0002 boost to our models, we weren't so sure if this will overfit private LB, so we submit one with post-processing and one without.\n\n **The End Game**\nWe spent a large part of the last two day trying to get FM working - in the hope that we could have an alternative model to merge with LightGBM - we weren't familiar with FM algorithms and it took a lot of time without getting useful results.  In highlight we could have better invested this time in NN - we took a choice and it didn't work out so next time we will approach with new knowledge. I personally enjoyed learning a lot from @anttip's wonderful FM kernel so no harm was done there. \n\nWhat caused a bit of panic was that we had a bug while trying to reduce the RAM requirement of our data, and it caused invalid values in our predictions - we spent hours fixing that while testing our LightGBM model with various parameters. Lucky the bug was fixed, and we manage to submit our best solutions within the last 20 minutes of the competition!  The final solution was a weight average of two LightGBM model, so again pretty boring stuff :)\n\n**Take Away**\n\n - Dart tree is good, in fact very good for large dataset - our parameters for dart tree was faster and more accurate than the gbm equivalent \n - Trust NN - I really should have considered using NN as an alternative model - especially since I notice the drop-out mechanism from dart tree was working well for this dataset. \n - START EARLIER IF YOU WANT GOLD: well, at least this applies to us\n   average folks. I think we could have generated a lot better solution\n   even if one more week - but then it might be next intense and less\n   fun. ^_^\n - [update] Downsampling when the data is large and imbalanced - now looking at both the 1st and 2nd solutions, it is clear that downsampling on negative observation were key steps for top teams - it didn't work out for me in the past, but I guess this depends on dataset. I will make sure I try this next time! \n\nBy the way, we are the biggest overfitter in top 50 (i.e. biggest drop between public and private LB), so perhaps this is exactly what you shouldn't be doing if you want to get a gold medal - LoL \n\n\n  [1]: https://diveintocode.jp/ai_curriculum\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54765\n  [3]: https://www.kaggle.com/ogrellier/basic-predictions-using-target-encoding-on-app/code\n  [4]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56328\n  [5]: https://www.kaggle.com/c/instacart-market-basket-analysis/discussion/38100\n  [6]: https://arxiv.org/abs/1505.01866",
      "votes": null
    },
    {
      "id": "326415",
      "postDate": "05/09/2018 17:01:06",
      "content": "<p>Thanks for sharing my friend, and congrats on the good result.</p>",
      "rawMarkdown": "Thanks for sharing my friend, and congrats on the good result.",
      "votes": null
    },
    {
      "id": "326436",
      "postDate": "05/09/2018 17:44:49",
      "content": "<p>For what I've seen this far I can tell you this guy rocks ! Thanks Yifan for your outstanding work. </p>\n\n<p>Zhiqiang has been amazing too. A very special mention to DiveIntoCode Students who put lots of energy and effort in this competition. Very nice team spirit here.  </p>\n\n<p>Like someone I loved to watch on TV would say: I love when a plan comes together ;-) (Hope I'm not H.M. \"Howling Mad\" Murdock though)</p>",
      "rawMarkdown": "For what I've seen this far I can tell you this guy rocks ! Thanks Yifan for your outstanding work. \n\nZhiqiang has been amazing too. A very special mention to DiveIntoCode Students who put lots of energy and effort in this competition. Very nice team spirit here.  \n\nLike someone I loved to watch on TV would say: I love when a plan comes together ;-) (Hope I'm not H.M. \"Howling Mad\" Murdock though)",
      "votes": null
    },
    {
      "id": "326456",
      "postDate": "05/09/2018 18:54:22",
      "content": "<p>Thanks Yifan for sharing and congrats to all your team for such great achievement in short timeline.</p>\n\n<p>BTW I  realise now I am not the only one still using only Dart mode since Kazanova shared his list of params...:)</p>\n\n<p>It outperformed gbdt on Mercari (even though our largest part was deep learning and FM_FTRL), so I keep using it :)</p>",
      "rawMarkdown": "Thanks Yifan for sharing and congrats to all your team for such great achievement in short timeline.\n\nBTW I  realise now I am not the only one still using only Dart mode since Kazanova shared his list of params...:)\n\nIt outperformed gbdt on Mercari (even though our largest part was deep learning and FM_FTRL), so I keep using it :)",
      "votes": null
    },
    {
      "id": "326493",
      "postDate": "05/09/2018 20:14:43",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "326546",
      "postDate": "05/09/2018 22:38:22",
      "content": "<p><a href=\"/serigne\">@serigne</a> haha, good to know some else are using dart tree out there! do you have any good tips at tuning? there isn't much shared experience out there so I am more or less just followed what worked before, and try to follow whatever work in a given dataset. </p>",
      "rawMarkdown": "serigne haha, good to know some else are using dart tree out there! do you have any good tips at tuning? there isn't much shared experience out there so I am more or less just followed what worked before, and try to follow whatever work in a given dataset.",
      "votes": null
    },
    {
      "id": "326827",
      "postDate": "05/10/2018 11:44:40",
      "content": "<p>Thanks! And awesome! Will try dart next time.</p>",
      "rawMarkdown": "Thanks! And awesome! Will try dart next time.",
      "votes": null
    },
    {
      "id": "327081",
      "postDate": "05/10/2018 19:42:43",
      "content": "<blockquote>\n  <p><strong>Yifan Xie wrote</strong></p>\n  \n  <p>do you have any good tips at tuning?</p>\n</blockquote>\n\n<p>Sadly No !</p>\n\n<p>I just start with guess of params and tune them manually. </p>\n\n<p>I use sometimes <a href=\"https://github.com/fmfn/BayesianOptimization\">Bayesian Optimization</a>  but just when the dataset is small. </p>",
      "rawMarkdown": "&gt; **Yifan Xie wrote**\n\n&gt;  do you have any good tips at tuning?\n \nSadly No !\n\nI just start with guess of params and tune them manually. \n\nI use sometimes [Bayesian Optimization][1]  but just when the dataset is small. \n\n\n  [1]: https://github.com/fmfn/BayesianOptimization",
      "votes": null
    },
    {
      "id": "327104",
      "postDate": "05/10/2018 20:22:56",
      "content": "<p>lol, looks like we are on the same boat there :)</p>",
      "rawMarkdown": "lol, looks like we are on the same boat there :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 326415,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/09/2018 17:01:06",
      "content": "<p>Thanks for sharing my friend, and congrats on the good result.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326436,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "05/09/2018 17:44:49",
      "content": "<p>For what I've seen this far I can tell you this guy rocks ! Thanks Yifan for your outstanding work. </p>\n\n<p>Zhiqiang has been amazing too. A very special mention to DiveIntoCode Students who put lots of energy and effort in this competition. Very nice team spirit here.  </p>\n\n<p>Like someone I loved to watch on TV would say: I love when a plan comes together ;-) (Hope I'm not H.M. \"Howling Mad\" Murdock though)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326456,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "05/09/2018 18:54:22",
      "content": "<p>Thanks Yifan for sharing and congrats to all your team for such great achievement in short timeline.</p>\n\n<p>BTW I  realise now I am not the only one still using only Dart mode since Kazanova shared his list of params...:)</p>\n\n<p>It outperformed gbdt on Mercari (even though our largest part was deep learning and FM_FTRL), so I keep using it :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 326546,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "05/09/2018 22:38:22",
          "content": "<p><a href=\"/serigne\">@serigne</a> haha, good to know some else are using dart tree out there! do you have any good tips at tuning? there isn't much shared experience out there so I am more or less just followed what worked before, and try to follow whatever work in a given dataset. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327081,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/10/2018 19:42:43",
          "content": "<blockquote>\n  <p><strong>Yifan Xie wrote</strong></p>\n  \n  <p>do you have any good tips at tuning?</p>\n</blockquote>\n\n<p>Sadly No !</p>\n\n<p>I just start with guess of params and tune them manually. </p>\n\n<p>I use sometimes <a href=\"https://github.com/fmfn/BayesianOptimization\">Bayesian Optimization</a>  but just when the dataset is small. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327104,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "05/10/2018 20:22:56",
          "content": "<p>lol, looks like we are on the same boat there :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326493,
      "author_name": "wang6609",
      "author_url": "",
      "post_date": "05/09/2018 20:14:43",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326827,
      "author_name": "rayarrow",
      "author_url": "",
      "post_date": "05/10/2018 11:44:40",
      "content": "<p>Thanks! And awesome! Will try dart next time.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "326403": "First of all, let me say up-front that I don't particularly think we have much to boast about this result - but it was an effort of very nice team effort in a very nice competition that stretches our RAM management techniques\n\nWe did this competition as an ongoing effort to work with friends in [DiveIntoCode][1] to provide data science mentorship to their students - so we spend most of our effort to help students to understand the [business context of click fraud][2], as well as creating [friendly kernels][3] to ease our students into handling the rather large dataset. So it wasn't until the last week we decided to make a stronger effort on the LB, and we kind of end-up doing this in a hackathon manner in the last 3-4 days. Our best submissions were submitted and selected within the last 20 minutes before the close of play, so that was entirely bonkers and fun way to keep kaggle addictive :)\n\nSo here we go, here is a description of our approach:\n\n**Information Sharing**\nWe use slack for communication, google drive for sharing feature/data, and then a google spreadsheet to manage various aspect of information. For us, this way of working turned out to be very important - we started so late and needed a way to gather \"intelligence\" that is disclosed in via kernels and discussion. We specifically have a dedicated list of bookmarks shared among us, and also key EDA results (both from various kernels but also our own effort) captured in a shared google slide\n\n**Modelling Approach**\nSince we started so late, we had very little choice but to start with the most secure methods - that means LightGBM, and midway through we discovered that extracting features using the test supplement gave a much better score.  We also ruled out traditional K-fold CV setting in the belief that it would take too much time for us to iterate - well, we now know there are much more clever methods for faster iteration as shown by [PPP's solution][4] by @panfeiyang.\n\nThe one thing I would recommend people to check out is the`boosting=dart` option from LightGBM. As @kazanova kindly shared in his [Instartcart solution][5] Dart tree can out-perform typical GBM. What I have since discovered by experience is that it tends to outperform GBM when the dataset is relatively large. My intuition is that this allows the dropout mechanism to fully explore its potential.  I had tried to use dart tee on the much smaller dataset from Porto Seguro without success at all. \n\nNevertheless, tuning Dart tree can be quite tricky, some of the typical logic on GBM doesn't apply for instance decreasing learning rate doesn't necessarily lead to better accuracy because it interacts with drop-out related parameters.  Anyway, we started with the parameter shared by @kazanova and started out tuning using the last 25millions rows in training data as local validation. Here is one of several sets of parameters that gave us good model:\n\n    num_leaves': 22,\n    min_sum_hessian_in_leaf':20,\n    max_depth': 7,\n    learning_rate': 0.2,    \n    num_threads': NUM_CORES,\n    feature_fraction':0.7,\n    bagging_fraction':0.85,\n    max_drop':6,\n    drop_rate':0.01,\n    min_data_in_leaf':10,\n    bagging_freq': 1,\n    scale_pos_weight':200,\n\n\nI don't have deep knowledge about how dart trees work, so I can only speak by experience. I shall now go and have a good read at [this paper][6], folks with better knowledge please also share your experience and we would be much appreciated!\n\n**Feature Engineering**\nnot much to share here, we were mainly scavaging features from shared kernels, but @ogrellier did implement a much neater way to compute \"time delta\" features. In addition to channel, app, os, and hour which were the strongest, we used different variations of next clicks and previous clicks - we generated additional features like \"the click after next click\" and some of them played an important role in our models. \n\nOther features are your usual count and unique_count features.  We also have a `click_rate` feature which measures the frequent were the clicks from specific IP during the day which helped a bit   By the last day of the competition, we had 34 features,  we could have generated more but time was running out.\n\nA work on feature selection: one by one feature selection was too long a process for us, so throughout the week I have been throwing 5-6 features by batch, and keep them if they increase the local validation score which thankfully corresponded well with LB.  I personally trusted a lot on how LightGBM prioritises features, and I operate in the following principles: \n\n - whenever a batch of features improve local validation score, keep them\n - whenever a batch of feature decrease score - I will remove the most import features among this batch according to LightGBM, and re-test\n - whenever features were not used by LightGBM at all i.e. 0 importance score I will drop them. \n\nNot sure this is the best approach, but in general it worked in this instance, and our ultimate constraint was that we couldn't manufacture enough feature fast enough within a limited amount of time - ideally we would like to have more than 50+ features which I believe will improve (or overfit) our models more.  \n\n**Post-processing**\nSo again this little secret was kindly shared by @plantsgo, and it gave a 0.0002 boost to our models, we weren't so sure if this will overfit private LB, so we submit one with post-processing and one without.\n\n **The End Game**\nWe spent a large part of the last two day trying to get FM working - in the hope that we could have an alternative model to merge with LightGBM - we weren't familiar with FM algorithms and it took a lot of time without getting useful results.  In highlight we could have better invested this time in NN - we took a choice and it didn't work out so next time we will approach with new knowledge. I personally enjoyed learning a lot from @anttip's wonderful FM kernel so no harm was done there. \n\nWhat caused a bit of panic was that we had a bug while trying to reduce the RAM requirement of our data, and it caused invalid values in our predictions - we spent hours fixing that while testing our LightGBM model with various parameters. Lucky the bug was fixed, and we manage to submit our best solutions within the last 20 minutes of the competition!  The final solution was a weight average of two LightGBM model, so again pretty boring stuff :)\n\n**Take Away**\n\n - Dart tree is good, in fact very good for large dataset - our parameters for dart tree was faster and more accurate than the gbm equivalent \n - Trust NN - I really should have considered using NN as an alternative model - especially since I notice the drop-out mechanism from dart tree was working well for this dataset. \n - START EARLIER IF YOU WANT GOLD: well, at least this applies to us\n   average folks. I think we could have generated a lot better solution\n   even if one more week - but then it might be next intense and less\n   fun. ^_^\n - [update] Downsampling when the data is large and imbalanced - now looking at both the 1st and 2nd solutions, it is clear that downsampling on negative observation were key steps for top teams - it didn't work out for me in the past, but I guess this depends on dataset. I will make sure I try this next time! \n\nBy the way, we are the biggest overfitter in top 50 (i.e. biggest drop between public and private LB), so perhaps this is exactly what you shouldn't be doing if you want to get a gold medal - LoL \n\n\n  [1]: https://diveintocode.jp/ai_curriculum\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54765\n  [3]: https://www.kaggle.com/ogrellier/basic-predictions-using-target-encoding-on-app/code\n  [4]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56328\n  [5]: https://www.kaggle.com/c/instacart-market-basket-analysis/discussion/38100\n  [6]: https://arxiv.org/abs/1505.01866",
    "326415": "Thanks for sharing my friend, and congrats on the good result.",
    "326436": "For what I've seen this far I can tell you this guy rocks ! Thanks Yifan for your outstanding work. \n\nZhiqiang has been amazing too. A very special mention to DiveIntoCode Students who put lots of energy and effort in this competition. Very nice team spirit here.  \n\nLike someone I loved to watch on TV would say: I love when a plan comes together ;-) (Hope I'm not H.M. \"Howling Mad\" Murdock though)",
    "326456": "Thanks Yifan for sharing and congrats to all your team for such great achievement in short timeline.\n\nBTW I  realise now I am not the only one still using only Dart mode since Kazanova shared his list of params...:)\n\nIt outperformed gbdt on Mercari (even though our largest part was deep learning and FM_FTRL), so I keep using it :)",
    "326493": "Thanks for sharing!",
    "326546": "serigne haha, good to know some else are using dart tree out there! do you have any good tips at tuning? there isn't much shared experience out there so I am more or less just followed what worked before, and try to follow whatever work in a given dataset.",
    "326827": "Thanks! And awesome! Will try dart next time.",
    "327081": "&gt; **Yifan Xie wrote**\n\n&gt;  do you have any good tips at tuning?\n \nSadly No !\n\nI just start with guess of params and tune them manually. \n\nI use sometimes [Bayesian Optimization][1]  but just when the dataset is small. \n\n\n  [1]: https://github.com/fmfn/BayesianOptimization",
    "327104": "lol, looks like we are on the same boat there :)"
  },
  "source": "meta"
}