{
  "id": 56325,
  "title": "8'th place solution. [ods.ai] blenders in game )",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/writeups/ods-ai-blenders-8-th-place-solution-ods-ai-blender",
  "author_name": "",
  "post_date": "2018-05-08T15:49:30.201343500Z",
  "votes": 45,
  "comment_count": 7,
  "views": 0,
  "content": "<p>At first, as usual, I want to say Thank you to all kagglers who share their approaches and thoughts during this and other competitions. Despite some annoying events (unfortunately such events can happen in any competition) kaggle platform is still the best place for practitioner Data Scientist who want to lift up their skills and go deeper into this DS Rabbit's hole )</p>\n\n<p>Our Team was built last week just before the end of competition (before merging deadline), I was at 80th place (at that moment), <a href=\"/johnpateha\">@johnpateha</a> just decided to take part in this game that day, <a href=\"/ppleskov\">@ppleskov</a>, <a href=\"/maksimovka\">@maksimovka</a>, <a href=\"/yaroshevskiy\">@yaroshevskiy</a> and <a href=\"/ddanevskyi\">@ddanevskyi</a> occupied ~100th places on the LB. I was ready to get a bronze (in case of luck) and finish my attempts to climb up, but one morning...</p>\n\n<p>\"Knock, Knock, Kruegger\" ) say <a href=\"/johnpateha\">@johnpateha</a>. \"We have a good chance to catch a gold here, just trust me, wake up and go <strong>\"вджобывать\"</strong> (to do a really hard work). \"The metric is fine, there is a lot of data, and in the worst case we don't lose anything\".</p>\n\n<p>Hmm. Why not? We made the team, several hours later combined our team with <a href=\"/ppleskov\">@ppleskov</a> &amp; Co, and start <strong>\"вджобывать\"</strong>.</p>\n\n<p><strong>Here I just describe my part of solution, my colleagues will add their own approaches later.</strong></p>\n\n<p>This competition is really hard due to size of the dataset. We need a lot of RAM to train our models especially if we have a lot of features. So this is why my solution before merging has been based on day 9 for training, day 8 for target encoding calculation - and 10% of train (shuffled) for validation. It was enough to climb up to 9805 score, but I feel that for this pipeline it is a ceiling.</p>\n\n<p><strong>Dataset</strong>:  </p>\n\n<p>In the final solution I use day 7 for target encoding, day 8+9 for training and last 2.5M rows from train as holdout. We decided to just blend our final solutions and not to use stacking on that level. (and of course, we had to comply with our team name! :)</p>\n\n<p>I spent a lot of hours and made a lot of attempts to fit my dataset to memory, I used all tricks I knew before and found there on the forum - but my machine (with 64Gb RAM) time after time told me \"Out of Memory\", \"Out of memory\"...</p>\n\n<p>The final approach that I found - is to use numpy memmap as data storage, I describe this method here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105</a></p>\n\n<p><strong>Features</strong>:</p>\n\n<p>I have several group of features, most of them you can find in the public kernels, nothing special:  </p>\n\n<ul>\n<li>Count by several groups</li>\n<li>NextClicks</li>\n<li>TargetEncoding over groups</li>\n<li>Statistics (mean/var)</li>\n<li>and so on...</li>\n</ul>\n\n<p>After merging I added to the dataset some features from my colleagues' solution, like duplicate orders, and so on. No \"Killer features\", but overall score became more stable.</p>\n\n<p><strong>Feature Engineering</strong></p>\n\n<p>At first I tried to use my favorite method of feature selection (random shuffling - like Boruta) but without success, so I returned back to old-fashioned style - greedy forward selection by small groups (3-5 attrs at a time). If my val score was raising I added that group to the dataset, in other case - gave up the whole group.</p>\n\n<p>After adding a group I tried to \"cut tail\" of features based on feature important but in very \"conservative\" manner.</p>\n\n<p>The final solution contains 74 features.</p>\n\n<p>Finally I selected three big groups of features in addition to the full set and built models for all my best parameters on these sets.</p>\n\n<p><strong>Models and parameters tuning</strong></p>\n\n<p>I used lgb as my base model, tried to build FM-FTRL/XGB, but with no overall improvement. I also was unlucky in CatBoost - it refused to run on my data, I don't know, probably it took offense ).</p>\n\n<p>I tried to use some parameters I found in the public kernels, but the best one I got - I got from Bayasian optimization process. I really a big fan of this method so I suggest trying it if you haven't tried it before.</p>\n\n<p>I use this implementation and it gives me very nice results in all competitions I use it.</p>\n\n<p><a href=\"https://github.com/fmfn/BayesianOptimization/blob/master/README.md\">https://github.com/fmfn/BayesianOptimization/blob/master/README.md</a></p>\n\n<p><strong>Diversity in dataset</strong></p>\n\n<p>Initially I built my models on last 75M rows (according to memory constraints), than on last 100M rows (thanks to my teammates for one more power computer) and finally (when I implemented dark numpy magic :) ) - on the whole dataset.</p>\n\n<p><strong>Ensembling</strong></p>\n\n<p>I used method that dropped me down to ~1500 place in the Toxic competition (I didn't prepare it well in Toxic), but in this competition it gave a huge improvement. The method is Scipy optimize with softmax restriction over weights. I tried to use power ensembing, caruana's hillclimbing, just geometric averaging - but scipy (in case you have holdout prediction for all of your model) gave the best score in all environment I tested it.</p>\n\n<p>My final solution was blending with scipy weights over 7 best models (from ~30 overall).</p>\n\n<p>The oof score of my best has .99235 -&gt; .9823x on public LB. Combining it with models from <a href=\"/johnpateha\">@johnpateha</a> and other teammates we had the final score that lifted us up to 8'th place )</p>\n\n<p><strong>Final words</strong></p>\n\n<p>Thank you to all my teammates, especially <a href=\"/johnpateha\">@johnpateha</a> who pushed me to Gold medal ) We did a very nice cooperative work inside our team, got the results in short time, so my dear colleagues - you are the best! )</p>\n\n<p>It was my first team competition (I played solo before) and I really appreciate the results of working in a team.</p>\n\n<p>Happy kaggling! <br>\n(C) Kruegger</p>\n\n<p>P.S. Some Intrigue - <a href=\"/johnpateha\">@johnpateha</a> did huge investigation over data, one of his advices helped to lift score of my model up to +0.0001. Advice was \"just put 0 to this 4(four)! rows in the dataset.\" Dark magic? ) But let him to describe his solution himself.</p>",
  "messages": [
    {
      "id": "325589",
      "postDate": "05/08/2018 15:49:30",
      "content": "<p>At first, as usual, I want to say Thank you to all kagglers who share their approaches and thoughts during this and other competitions. Despite some annoying events (unfortunately such events can happen in any competition) kaggle platform is still the best place for practitioner Data Scientist who want to lift up their skills and go deeper into this DS Rabbit's hole )</p>\n\n<p>Our Team was built last week just before the end of competition (before merging deadline), I was at 80th place (at that moment), <a href=\"/johnpateha\">@johnpateha</a> just decided to take part in this game that day, <a href=\"/ppleskov\">@ppleskov</a>, <a href=\"/maksimovka\">@maksimovka</a>, <a href=\"/yaroshevskiy\">@yaroshevskiy</a> and <a href=\"/ddanevskyi\">@ddanevskyi</a> occupied ~100th places on the LB. I was ready to get a bronze (in case of luck) and finish my attempts to climb up, but one morning...</p>\n\n<p>\"Knock, Knock, Kruegger\" ) say <a href=\"/johnpateha\">@johnpateha</a>. \"We have a good chance to catch a gold here, just trust me, wake up and go <strong>\"вджобывать\"</strong> (to do a really hard work). \"The metric is fine, there is a lot of data, and in the worst case we don't lose anything\".</p>\n\n<p>Hmm. Why not? We made the team, several hours later combined our team with <a href=\"/ppleskov\">@ppleskov</a> &amp; Co, and start <strong>\"вджобывать\"</strong>.</p>\n\n<p><strong>Here I just describe my part of solution, my colleagues will add their own approaches later.</strong></p>\n\n<p>This competition is really hard due to size of the dataset. We need a lot of RAM to train our models especially if we have a lot of features. So this is why my solution before merging has been based on day 9 for training, day 8 for target encoding calculation - and 10% of train (shuffled) for validation. It was enough to climb up to 9805 score, but I feel that for this pipeline it is a ceiling.</p>\n\n<p><strong>Dataset</strong>:  </p>\n\n<p>In the final solution I use day 7 for target encoding, day 8+9 for training and last 2.5M rows from train as holdout. We decided to just blend our final solutions and not to use stacking on that level. (and of course, we had to comply with our team name! :)</p>\n\n<p>I spent a lot of hours and made a lot of attempts to fit my dataset to memory, I used all tricks I knew before and found there on the forum - but my machine (with 64Gb RAM) time after time told me \"Out of Memory\", \"Out of memory\"...</p>\n\n<p>The final approach that I found - is to use numpy memmap as data storage, I describe this method here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105</a></p>\n\n<p><strong>Features</strong>:</p>\n\n<p>I have several group of features, most of them you can find in the public kernels, nothing special:  </p>\n\n<ul>\n<li>Count by several groups</li>\n<li>NextClicks</li>\n<li>TargetEncoding over groups</li>\n<li>Statistics (mean/var)</li>\n<li>and so on...</li>\n</ul>\n\n<p>After merging I added to the dataset some features from my colleagues' solution, like duplicate orders, and so on. No \"Killer features\", but overall score became more stable.</p>\n\n<p><strong>Feature Engineering</strong></p>\n\n<p>At first I tried to use my favorite method of feature selection (random shuffling - like Boruta) but without success, so I returned back to old-fashioned style - greedy forward selection by small groups (3-5 attrs at a time). If my val score was raising I added that group to the dataset, in other case - gave up the whole group.</p>\n\n<p>After adding a group I tried to \"cut tail\" of features based on feature important but in very \"conservative\" manner.</p>\n\n<p>The final solution contains 74 features.</p>\n\n<p>Finally I selected three big groups of features in addition to the full set and built models for all my best parameters on these sets.</p>\n\n<p><strong>Models and parameters tuning</strong></p>\n\n<p>I used lgb as my base model, tried to build FM-FTRL/XGB, but with no overall improvement. I also was unlucky in CatBoost - it refused to run on my data, I don't know, probably it took offense ).</p>\n\n<p>I tried to use some parameters I found in the public kernels, but the best one I got - I got from Bayasian optimization process. I really a big fan of this method so I suggest trying it if you haven't tried it before.</p>\n\n<p>I use this implementation and it gives me very nice results in all competitions I use it.</p>\n\n<p><a href=\"https://github.com/fmfn/BayesianOptimization/blob/master/README.md\">https://github.com/fmfn/BayesianOptimization/blob/master/README.md</a></p>\n\n<p><strong>Diversity in dataset</strong></p>\n\n<p>Initially I built my models on last 75M rows (according to memory constraints), than on last 100M rows (thanks to my teammates for one more power computer) and finally (when I implemented dark numpy magic :) ) - on the whole dataset.</p>\n\n<p><strong>Ensembling</strong></p>\n\n<p>I used method that dropped me down to ~1500 place in the Toxic competition (I didn't prepare it well in Toxic), but in this competition it gave a huge improvement. The method is Scipy optimize with softmax restriction over weights. I tried to use power ensembing, caruana's hillclimbing, just geometric averaging - but scipy (in case you have holdout prediction for all of your model) gave the best score in all environment I tested it.</p>\n\n<p>My final solution was blending with scipy weights over 7 best models (from ~30 overall).</p>\n\n<p>The oof score of my best has .99235 -&gt; .9823x on public LB. Combining it with models from <a href=\"/johnpateha\">@johnpateha</a> and other teammates we had the final score that lifted us up to 8'th place )</p>\n\n<p><strong>Final words</strong></p>\n\n<p>Thank you to all my teammates, especially <a href=\"/johnpateha\">@johnpateha</a> who pushed me to Gold medal ) We did a very nice cooperative work inside our team, got the results in short time, so my dear colleagues - you are the best! )</p>\n\n<p>It was my first team competition (I played solo before) and I really appreciate the results of working in a team.</p>\n\n<p>Happy kaggling! <br>\n(C) Kruegger</p>\n\n<p>P.S. Some Intrigue - <a href=\"/johnpateha\">@johnpateha</a> did huge investigation over data, one of his advices helped to lift score of my model up to +0.0001. Advice was \"just put 0 to this 4(four)! rows in the dataset.\" Dark magic? ) But let him to describe his solution himself.</p>",
      "rawMarkdown": "At first, as usual, I want to say Thank you to all kagglers who share their approaches and thoughts during this and other competitions. Despite some annoying events (unfortunately such events can happen in any competition) kaggle platform is still the best place for practitioner Data Scientist who want to lift up their skills and go deeper into this DS Rabbit's hole )\n\nOur Team was built last week just before the end of competition (before merging deadline), I was at 80th place (at that moment), @johnpateha just decided to take part in this game that day, @ppleskov, @maksimovka, @yaroshevskiy and @ddanevskyi occupied ~100th places on the LB. I was ready to get a bronze (in case of luck) and finish my attempts to climb up, but one morning...\n\n\"Knock, Knock, Kruegger\" ) say @johnpateha. \"We have a good chance to catch a gold here, just trust me, wake up and go **\"вджобывать\"** (to do a really hard work). \"The metric is fine, there is a lot of data, and in the worst case we don't lose anything\".\n\nHmm. Why not? We made the team, several hours later combined our team with @ppleskov &amp; Co, and start **\"вджобывать\"**.\n\n**Here I just describe my part of solution, my colleagues will add their own approaches later.**\n\nThis competition is really hard due to size of the dataset. We need a lot of RAM to train our models especially if we have a lot of features. So this is why my solution before merging has been based on day 9 for training, day 8 for target encoding calculation - and 10% of train (shuffled) for validation. It was enough to climb up to 9805 score, but I feel that for this pipeline it is a ceiling.\n\n**Dataset**:  \n\nIn the final solution I use day 7 for target encoding, day 8+9 for training and last 2.5M rows from train as holdout. We decided to just blend our final solutions and not to use stacking on that level. (and of course, we had to comply with our team name! :)\n\nI spent a lot of hours and made a lot of attempts to fit my dataset to memory, I used all tricks I knew before and found there on the forum - but my machine (with 64Gb RAM) time after time told me \"Out of Memory\", \"Out of memory\"...\n\nThe final approach that I found - is to use numpy memmap as data storage, I describe this method here:\n\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105\n\n**Features**:\n\nI have several group of features, most of them you can find in the public kernels, nothing special:  \n\n* Count by several groups\n* NextClicks\n* TargetEncoding over groups\n* Statistics (mean/var)\n* and so on...\n\nAfter merging I added to the dataset some features from my colleagues' solution, like duplicate orders, and so on. No \"Killer features\", but overall score became more stable.\n\n**Feature Engineering**\n\nAt first I tried to use my favorite method of feature selection (random shuffling - like Boruta) but without success, so I returned back to old-fashioned style - greedy forward selection by small groups (3-5 attrs at a time). If my val score was raising I added that group to the dataset, in other case - gave up the whole group.\n\nAfter adding a group I tried to \"cut tail\" of features based on feature important but in very \"conservative\" manner.\n\nThe final solution contains 74 features.\n\nFinally I selected three big groups of features in addition to the full set and built models for all my best parameters on these sets.\n\n\n**Models and parameters tuning**\n\nI used lgb as my base model, tried to build FM-FTRL/XGB, but with no overall improvement. I also was unlucky in CatBoost - it refused to run on my data, I don't know, probably it took offense ).\n\nI tried to use some parameters I found in the public kernels, but the best one I got - I got from Bayasian optimization process. I really a big fan of this method so I suggest trying it if you haven't tried it before.\n\nI use this implementation and it gives me very nice results in all competitions I use it.\n\nhttps://github.com/fmfn/BayesianOptimization/blob/master/README.md\n\n\n**Diversity in dataset**\n\nInitially I built my models on last 75M rows (according to memory constraints), than on last 100M rows (thanks to my teammates for one more power computer) and finally (when I implemented dark numpy magic :) ) - on the whole dataset.\n\n**Ensembling**\n\nI used method that dropped me down to ~1500 place in the Toxic competition (I didn't prepare it well in Toxic), but in this competition it gave a huge improvement. The method is Scipy optimize with softmax restriction over weights. I tried to use power ensembing, caruana's hillclimbing, just geometric averaging - but scipy (in case you have holdout prediction for all of your model) gave the best score in all environment I tested it.\n\nMy final solution was blending with scipy weights over 7 best models (from ~30 overall).\n\nThe oof score of my best has .99235 -&gt; .9823x on public LB. Combining it with models from @johnpateha and other teammates we had the final score that lifted us up to 8'th place )\n\n**Final words**\n\nThank you to all my teammates, especially @johnpateha who pushed me to Gold medal ) We did a very nice cooperative work inside our team, got the results in short time, so my dear colleagues - you are the best! )\n\nIt was my first team competition (I played solo before) and I really appreciate the results of working in a team.\n\nHappy kaggling!  \n(C) Kruegger\n\nP.S. Some Intrigue - @johnpateha did huge investigation over data, one of his advices helped to lift score of my model up to +0.0001. Advice was \"just put 0 to this 4(four)! rows in the dataset.\" Dark magic? ) But let him to describe his solution himself.",
      "votes": null
    },
    {
      "id": "325621",
      "postDate": "05/08/2018 16:30:11",
      "content": "<p>Thanks @Kruegger for sharing your solution.</p>",
      "rawMarkdown": "Thanks @Kruegger for sharing your solution.",
      "votes": null
    },
    {
      "id": "326000",
      "postDate": "05/09/2018 06:46:12",
      "content": "<p>@Kruegger Thank you for the post, could you share your solution?</p>",
      "rawMarkdown": "Kruegger Thank you for the post, could you share your solution?",
      "votes": null
    },
    {
      "id": "326046",
      "postDate": "05/09/2018 07:29:21",
      "content": "<p>Thanks! Good discovery about memmap.</p>",
      "rawMarkdown": "Thanks! Good discovery about memmap.",
      "votes": null
    },
    {
      "id": "326057",
      "postDate": "05/09/2018 07:55:03",
      "content": "<p>Thanks for sharing and congrats on the result!  Your numpy dark magic came in too late for me, but i'll use it for sure.</p>",
      "rawMarkdown": "Thanks for sharing and congrats on the result!  Your numpy dark magic came in too late for me, but i'll use it for sure.",
      "votes": null
    },
    {
      "id": "326128",
      "postDate": "05/09/2018 10:28:12",
      "content": "<p>First of all, I want to say thanks to @Kruegger. He let me a chance to jump into the competition at the last day before merge deadline with only some ideas without real results at that moment. Also I want to thank all my teammates for good team work.</p>\n\n<p>Within one week I created lgb model with 41 features with public/private scores 0.98193/0.98259 which was a part of our blend.</p>\n\n<p>Some features were taken from models of my teammates, some were created by my own but nothing impressive. Mostly many of counts and unique.</p>\n\n<p>For feature selection and tuning of parameters I used 3-fold CV. I shifted time to local China time and put each day to separate fold.  I thought this approach should be very close to the competition's task, especially for work with categories which appeared for one day only. </p>\n\n<p>For speed up I used only 3 hours for feature check (hours were selected according to test hours). I chose the features which showed improve at each fold (day). May be little bit conservative for Kaggle, but very reliable.  It helped me to use only few submissions during the week and save more attempts for my teammates. </p>\n\n<p>Training of 1 fold lasts 10 to 20 minutes (depends on number of features). </p>\n\n<p>For final test prediction I used same CV approach with full train data and averaged 3 predictions.  </p>\n\n<p>May be one interesting thing. When I started to analyze the data, I tried to separate Apple and Android devices and found some anomaly - it was a computer-based click-farm with strong pattern each day (at 00:00 by local time ) one new OS was appeared (33,607,748,866). For masking they used some downloads (very minor for android devices and near 10% for Apple devices (apple has higher conversion for normal cases). As it is impossible to use one os whith apple and android devices I supposed that it could be computer based farm which used fake idetnification and I changed all the fake categories to one. This farm also used some real os-device categories for masking but not too much. I have no idea how strong it improved my results because I made it before I created model. But my advice to clear positive target for very few masking rows  (os %in% c(33,57,607,748,866) &amp; device %in% c(1,2,3,1525,3032,3543,3866) help my teammates to improve their results too.</p>\n\n<p>full farm description:</p>\n\n<pre><code> DT[os %in% c(33,57,607,748,866) &amp;amp; device %in% c(1,2,3,1525,3032,3543,3866), \":=\"(os=999, device=9999, os_type=9)] # pure fake\n    DT[os %in% c(67) &amp;amp; device %in% c(3208,3866,3543,3, 3032), \":=\"(os=999, device=9999, os_type=9)] # pure fake\n    DT[os %in% c(33,57,607,748,866) &amp;amp; device %in% c(0,2694,2691,59,3208,3858), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n    DT[os %in% c(223, 0) &amp;amp; device %in% c(3543), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n    DT[os %in% c(24) &amp;amp; device %in% c(3866), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n    DT[os %in% c(59) &amp;amp; device %in% c(57), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n</code></pre>",
      "rawMarkdown": "First of all, I want to say thanks to @Kruegger. He let me a chance to jump into the competition at the last day before merge deadline with only some ideas without real results at that moment. Also I want to thank all my teammates for good team work.\n\nWithin one week I created lgb model with 41 features with public/private scores 0.98193/0.98259 which was a part of our blend.\n\nSome features were taken from models of my teammates, some were created by my own but nothing impressive. Mostly many of counts and unique.\n\nFor feature selection and tuning of parameters I used 3-fold CV. I shifted time to local China time and put each day to separate fold.  I thought this approach should be very close to the competition's task, especially for work with categories which appeared for one day only. \n\nFor speed up I used only 3 hours for feature check (hours were selected according to test hours). I chose the features which showed improve at each fold (day). May be little bit conservative for Kaggle, but very reliable.  It helped me to use only few submissions during the week and save more attempts for my teammates. \n\nTraining of 1 fold lasts 10 to 20 minutes (depends on number of features). \n\nFor final test prediction I used same CV approach with full train data and averaged 3 predictions.  \n\n\nMay be one interesting thing. When I started to analyze the data, I tried to separate Apple and Android devices and found some anomaly - it was a computer-based click-farm with strong pattern each day (at 00:00 by local time ) one new OS was appeared (33,607,748,866). For masking they used some downloads (very minor for android devices and near 10% for Apple devices (apple has higher conversion for normal cases). As it is impossible to use one os whith apple and android devices I supposed that it could be computer based farm which used fake idetnification and I changed all the fake categories to one. This farm also used some real os-device categories for masking but not too much. I have no idea how strong it improved my results because I made it before I created model. But my advice to clear positive target for very few masking rows  (os %in% c(33,57,607,748,866) &amp; device %in% c(1,2,3,1525,3032,3543,3866) help my teammates to improve their results too.\n\n\nfull farm description:\n\n     DT[os %in% c(33,57,607,748,866) &amp; device %in% c(1,2,3,1525,3032,3543,3866), \":=\"(os=999, device=9999, os_type=9)] # pure fake\n        DT[os %in% c(67) &amp; device %in% c(3208,3866,3543,3, 3032), \":=\"(os=999, device=9999, os_type=9)] # pure fake\n        DT[os %in% c(33,57,607,748,866) &amp; device %in% c(0,2694,2691,59,3208,3858), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n        DT[os %in% c(223, 0) &amp; device %in% c(3543), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n        DT[os %in% c(24) &amp; device %in% c(3866), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n        DT[os %in% c(59) &amp; device %in% c(57), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads",
      "votes": null
    },
    {
      "id": "326143",
      "postDate": "05/09/2018 10:50:01",
      "content": "<p>Thanks for sharing.  </p>\n\n<p>&gt; I shifted time to local China time and put each day to separate fold.</p>\n\n<p>I wish I had done this since the start, thought of it too late.  Having oof predictions on all data was key for stacking I think.</p>",
      "rawMarkdown": "Thanks for sharing.  \n\n&gt; I shifted time to local China time and put each day to separate fold.\n\nI wish I had done this since the start, thought of it too late.  Having oof predictions on all data was key for stacking I think.",
      "votes": null
    },
    {
      "id": "326246",
      "postDate": "05/09/2018 13:31:44",
      "content": "<p>Thanks for sharing and well done on the score!</p>",
      "rawMarkdown": "Thanks for sharing and well done on the score!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 325621,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "05/08/2018 16:30:11",
      "content": "<p>Thanks @Kruegger for sharing your solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326000,
      "author_name": "insaff",
      "author_url": "",
      "post_date": "05/09/2018 06:46:12",
      "content": "<p>@Kruegger Thank you for the post, could you share your solution?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326046,
      "author_name": "rayarrow",
      "author_url": "",
      "post_date": "05/09/2018 07:29:21",
      "content": "<p>Thanks! Good discovery about memmap.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326057,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/09/2018 07:55:03",
      "content": "<p>Thanks for sharing and congrats on the result!  Your numpy dark magic came in too late for me, but i'll use it for sure.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326128,
      "author_name": "johnpateha",
      "author_url": "",
      "post_date": "05/09/2018 10:28:12",
      "content": "<p>First of all, I want to say thanks to @Kruegger. He let me a chance to jump into the competition at the last day before merge deadline with only some ideas without real results at that moment. Also I want to thank all my teammates for good team work.</p>\n\n<p>Within one week I created lgb model with 41 features with public/private scores 0.98193/0.98259 which was a part of our blend.</p>\n\n<p>Some features were taken from models of my teammates, some were created by my own but nothing impressive. Mostly many of counts and unique.</p>\n\n<p>For feature selection and tuning of parameters I used 3-fold CV. I shifted time to local China time and put each day to separate fold.  I thought this approach should be very close to the competition's task, especially for work with categories which appeared for one day only. </p>\n\n<p>For speed up I used only 3 hours for feature check (hours were selected according to test hours). I chose the features which showed improve at each fold (day). May be little bit conservative for Kaggle, but very reliable.  It helped me to use only few submissions during the week and save more attempts for my teammates. </p>\n\n<p>Training of 1 fold lasts 10 to 20 minutes (depends on number of features). </p>\n\n<p>For final test prediction I used same CV approach with full train data and averaged 3 predictions.  </p>\n\n<p>May be one interesting thing. When I started to analyze the data, I tried to separate Apple and Android devices and found some anomaly - it was a computer-based click-farm with strong pattern each day (at 00:00 by local time ) one new OS was appeared (33,607,748,866). For masking they used some downloads (very minor for android devices and near 10% for Apple devices (apple has higher conversion for normal cases). As it is impossible to use one os whith apple and android devices I supposed that it could be computer based farm which used fake idetnification and I changed all the fake categories to one. This farm also used some real os-device categories for masking but not too much. I have no idea how strong it improved my results because I made it before I created model. But my advice to clear positive target for very few masking rows  (os %in% c(33,57,607,748,866) &amp; device %in% c(1,2,3,1525,3032,3543,3866) help my teammates to improve their results too.</p>\n\n<p>full farm description:</p>\n\n<pre><code> DT[os %in% c(33,57,607,748,866) &amp;amp; device %in% c(1,2,3,1525,3032,3543,3866), \":=\"(os=999, device=9999, os_type=9)] # pure fake\n    DT[os %in% c(67) &amp;amp; device %in% c(3208,3866,3543,3, 3032), \":=\"(os=999, device=9999, os_type=9)] # pure fake\n    DT[os %in% c(33,57,607,748,866) &amp;amp; device %in% c(0,2694,2691,59,3208,3858), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n    DT[os %in% c(223, 0) &amp;amp; device %in% c(3543), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n    DT[os %in% c(24) &amp;amp; device %in% c(3866), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n    DT[os %in% c(59) &amp;amp; device %in% c(57), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 326143,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/09/2018 10:50:01",
          "content": "<p>Thanks for sharing.  </p>\n\n<p>&gt; I shifted time to local China time and put each day to separate fold.</p>\n\n<p>I wish I had done this since the start, thought of it too late.  Having oof predictions on all data was key for stacking I think.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326246,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "05/09/2018 13:31:44",
      "content": "<p>Thanks for sharing and well done on the score!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "325589": "At first, as usual, I want to say Thank you to all kagglers who share their approaches and thoughts during this and other competitions. Despite some annoying events (unfortunately such events can happen in any competition) kaggle platform is still the best place for practitioner Data Scientist who want to lift up their skills and go deeper into this DS Rabbit's hole )\n\nOur Team was built last week just before the end of competition (before merging deadline), I was at 80th place (at that moment), @johnpateha just decided to take part in this game that day, @ppleskov, @maksimovka, @yaroshevskiy and @ddanevskyi occupied ~100th places on the LB. I was ready to get a bronze (in case of luck) and finish my attempts to climb up, but one morning...\n\n\"Knock, Knock, Kruegger\" ) say @johnpateha. \"We have a good chance to catch a gold here, just trust me, wake up and go **\"вджобывать\"** (to do a really hard work). \"The metric is fine, there is a lot of data, and in the worst case we don't lose anything\".\n\nHmm. Why not? We made the team, several hours later combined our team with @ppleskov &amp; Co, and start **\"вджобывать\"**.\n\n**Here I just describe my part of solution, my colleagues will add their own approaches later.**\n\nThis competition is really hard due to size of the dataset. We need a lot of RAM to train our models especially if we have a lot of features. So this is why my solution before merging has been based on day 9 for training, day 8 for target encoding calculation - and 10% of train (shuffled) for validation. It was enough to climb up to 9805 score, but I feel that for this pipeline it is a ceiling.\n\n**Dataset**:  \n\nIn the final solution I use day 7 for target encoding, day 8+9 for training and last 2.5M rows from train as holdout. We decided to just blend our final solutions and not to use stacking on that level. (and of course, we had to comply with our team name! :)\n\nI spent a lot of hours and made a lot of attempts to fit my dataset to memory, I used all tricks I knew before and found there on the forum - but my machine (with 64Gb RAM) time after time told me \"Out of Memory\", \"Out of memory\"...\n\nThe final approach that I found - is to use numpy memmap as data storage, I describe this method here:\n\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105\n\n**Features**:\n\nI have several group of features, most of them you can find in the public kernels, nothing special:  \n\n* Count by several groups\n* NextClicks\n* TargetEncoding over groups\n* Statistics (mean/var)\n* and so on...\n\nAfter merging I added to the dataset some features from my colleagues' solution, like duplicate orders, and so on. No \"Killer features\", but overall score became more stable.\n\n**Feature Engineering**\n\nAt first I tried to use my favorite method of feature selection (random shuffling - like Boruta) but without success, so I returned back to old-fashioned style - greedy forward selection by small groups (3-5 attrs at a time). If my val score was raising I added that group to the dataset, in other case - gave up the whole group.\n\nAfter adding a group I tried to \"cut tail\" of features based on feature important but in very \"conservative\" manner.\n\nThe final solution contains 74 features.\n\nFinally I selected three big groups of features in addition to the full set and built models for all my best parameters on these sets.\n\n\n**Models and parameters tuning**\n\nI used lgb as my base model, tried to build FM-FTRL/XGB, but with no overall improvement. I also was unlucky in CatBoost - it refused to run on my data, I don't know, probably it took offense ).\n\nI tried to use some parameters I found in the public kernels, but the best one I got - I got from Bayasian optimization process. I really a big fan of this method so I suggest trying it if you haven't tried it before.\n\nI use this implementation and it gives me very nice results in all competitions I use it.\n\nhttps://github.com/fmfn/BayesianOptimization/blob/master/README.md\n\n\n**Diversity in dataset**\n\nInitially I built my models on last 75M rows (according to memory constraints), than on last 100M rows (thanks to my teammates for one more power computer) and finally (when I implemented dark numpy magic :) ) - on the whole dataset.\n\n**Ensembling**\n\nI used method that dropped me down to ~1500 place in the Toxic competition (I didn't prepare it well in Toxic), but in this competition it gave a huge improvement. The method is Scipy optimize with softmax restriction over weights. I tried to use power ensembing, caruana's hillclimbing, just geometric averaging - but scipy (in case you have holdout prediction for all of your model) gave the best score in all environment I tested it.\n\nMy final solution was blending with scipy weights over 7 best models (from ~30 overall).\n\nThe oof score of my best has .99235 -&gt; .9823x on public LB. Combining it with models from @johnpateha and other teammates we had the final score that lifted us up to 8'th place )\n\n**Final words**\n\nThank you to all my teammates, especially @johnpateha who pushed me to Gold medal ) We did a very nice cooperative work inside our team, got the results in short time, so my dear colleagues - you are the best! )\n\nIt was my first team competition (I played solo before) and I really appreciate the results of working in a team.\n\nHappy kaggling!  \n(C) Kruegger\n\nP.S. Some Intrigue - @johnpateha did huge investigation over data, one of his advices helped to lift score of my model up to +0.0001. Advice was \"just put 0 to this 4(four)! rows in the dataset.\" Dark magic? ) But let him to describe his solution himself.",
    "325621": "Thanks @Kruegger for sharing your solution.",
    "326000": "Kruegger Thank you for the post, could you share your solution?",
    "326046": "Thanks! Good discovery about memmap.",
    "326057": "Thanks for sharing and congrats on the result!  Your numpy dark magic came in too late for me, but i'll use it for sure.",
    "326128": "First of all, I want to say thanks to @Kruegger. He let me a chance to jump into the competition at the last day before merge deadline with only some ideas without real results at that moment. Also I want to thank all my teammates for good team work.\n\nWithin one week I created lgb model with 41 features with public/private scores 0.98193/0.98259 which was a part of our blend.\n\nSome features were taken from models of my teammates, some were created by my own but nothing impressive. Mostly many of counts and unique.\n\nFor feature selection and tuning of parameters I used 3-fold CV. I shifted time to local China time and put each day to separate fold.  I thought this approach should be very close to the competition's task, especially for work with categories which appeared for one day only. \n\nFor speed up I used only 3 hours for feature check (hours were selected according to test hours). I chose the features which showed improve at each fold (day). May be little bit conservative for Kaggle, but very reliable.  It helped me to use only few submissions during the week and save more attempts for my teammates. \n\nTraining of 1 fold lasts 10 to 20 minutes (depends on number of features). \n\nFor final test prediction I used same CV approach with full train data and averaged 3 predictions.  \n\n\nMay be one interesting thing. When I started to analyze the data, I tried to separate Apple and Android devices and found some anomaly - it was a computer-based click-farm with strong pattern each day (at 00:00 by local time ) one new OS was appeared (33,607,748,866). For masking they used some downloads (very minor for android devices and near 10% for Apple devices (apple has higher conversion for normal cases). As it is impossible to use one os whith apple and android devices I supposed that it could be computer based farm which used fake idetnification and I changed all the fake categories to one. This farm also used some real os-device categories for masking but not too much. I have no idea how strong it improved my results because I made it before I created model. But my advice to clear positive target for very few masking rows  (os %in% c(33,57,607,748,866) &amp; device %in% c(1,2,3,1525,3032,3543,3866) help my teammates to improve their results too.\n\n\nfull farm description:\n\n     DT[os %in% c(33,57,607,748,866) &amp; device %in% c(1,2,3,1525,3032,3543,3866), \":=\"(os=999, device=9999, os_type=9)] # pure fake\n        DT[os %in% c(67) &amp; device %in% c(3208,3866,3543,3, 3032), \":=\"(os=999, device=9999, os_type=9)] # pure fake\n        DT[os %in% c(33,57,607,748,866) &amp; device %in% c(0,2694,2691,59,3208,3858), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n        DT[os %in% c(223, 0) &amp; device %in% c(3543), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n        DT[os %in% c(24) &amp; device %in% c(3866), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads\n        DT[os %in% c(59) &amp; device %in% c(57), \":=\"(os=998, device=9999, os_type=9)] # fake with downloads",
    "326143": "Thanks for sharing.  \n\n&gt; I shifted time to local China time and put each day to separate fold.\n\nI wish I had done this since the start, thought of it too late.  Having oof predictions on all data was key for stacking I think.",
    "326246": "Thanks for sharing and well done on the score!"
  },
  "source": "meta"
}