{
  "id": 56243,
  "title": "4th place (brief) tips",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56243",
  "author_name": "Μαριος Μιχαηλιδης KazAnova",
  "post_date": "2018-05-08T00:47:18.452000",
  "votes": 136,
  "comment_count": 46,
  "views": 0,
  "content": "<p>I would like to start via praising my teamates for an excellent effort and congratulate the winners for  an intense last-days' race. </p>\n\n<p>Also thank you to all the people who shared code and ideas - they made it a great competition (with the exception of the latest high-scoring kernels) . Special thanks to <a href=\"https://www.kaggle.com/anttip\">anttip</a> for his overall contribution with <a href=\"https://github.com/anttttti/Wordbatch\">wordbatch</a> and kernels in general .</p>\n\n<p>My favourite positions in a kaggle  competition are 1st and 4th. 1st you get most points/money. 4th, you dont get as many points, but you dont have to reproduce your solution :)</p>\n\n<p>Our validation schema was as follows: </p>\n\n<p>We used days 7,8 for training and we were making predictions for 9th day (hours [4,14])\nFor test predictions, we were training on all 7,8,9 and making predictions for the test day (10th)</p>\n\n<p>Things that work for us apart from what it is already in forums/public kernels:</p>\n\n<p>1) <a href=\"https://github.com/kaz-Anova/StackNet#restacking-mode\">Restacking</a> - When we first started doing Stacking, we could barely get 1,2 points our of it (like from 0.9818 to 0.9820). After adding ALL the features used in our standard modelling to the predictions of the ninth day - we got another +3 boost (to 0.9823). We ended up having around 50 models - mostly lightgbms, but also nns , FMs and some linear models. NNs were on par with LGB models. </p>\n\n<p>2) We got another +4 from creating WoE (<a href=\"https://github.com/h2oai/h2o-meetups/blob/master/2017_11_29_Feature_Engineering/Feature%20Engineering.pdf\">Weights of Evidence</a>) features for many combinations of all variables (ip,app,device,os)  . This is very strange , because we tried standard target encoding and <strong>it did not work</strong>. This is very strange, because the ordering of likelihood features and woe should be the same/similar (only the range in woe is more condensed) . We are not sure why this happens, maybe it has to do with the binning of Lightgbm?</p>\n\n<p>3)  We got a boost of +4 via sorting ties of time in the same groups of ip,app,device,os, making certain the is_attributed==1 <strong>is always last</strong> <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677\">as it was pointed out in the forums</a>. </p>\n\n<p>Each one of the team members had different features - mine were more related with time-series. Like counts of previous (and next ) days, hours, minutes and seconds of ips,apps,device (and their combinations) . </p>\n\n<p>Other than that features that measure time between next/previous clicks for various sortings (like ip and app or ip,app,device and os)  were also important. </p>\n\n<p>The most important feature (importance-wise)  was by far the app (treated as categorical). You could see that certain apps had very different (HIGH) probabilities of <strong>is_attributed</strong>  - I called them for fun tr-APPS!</p>",
  "messages": [
    {
      "id": 324861,
      "postDate": "2018-05-08T00:47:18.453Z",
      "content": "<p>I would like to start via praising my teamates for an excellent effort and congratulate the winners for  an intense last-days' race. </p>\n\n<p>Also thank you to all the people who shared code and ideas - they made it a great competition (with the exception of the latest high-scoring kernels) . Special thanks to <a href=\"https://www.kaggle.com/anttip\">anttip</a> for his overall contribution with <a href=\"https://github.com/anttttti/Wordbatch\">wordbatch</a> and kernels in general .</p>\n\n<p>My favourite positions in a kaggle  competition are 1st and 4th. 1st you get most points/money. 4th, you dont get as many points, but you dont have to reproduce your solution :)</p>\n\n<p>Our validation schema was as follows: </p>\n\n<p>We used days 7,8 for training and we were making predictions for 9th day (hours [4,14])\nFor test predictions, we were training on all 7,8,9 and making predictions for the test day (10th)</p>\n\n<p>Things that work for us apart from what it is already in forums/public kernels:</p>\n\n<p>1) <a href=\"https://github.com/kaz-Anova/StackNet#restacking-mode\">Restacking</a> - When we first started doing Stacking, we could barely get 1,2 points our of it (like from 0.9818 to 0.9820). After adding ALL the features used in our standard modelling to the predictions of the ninth day - we got another +3 boost (to 0.9823). We ended up having around 50 models - mostly lightgbms, but also nns , FMs and some linear models. NNs were on par with LGB models. </p>\n\n<p>2) We got another +4 from creating WoE (<a href=\"https://github.com/h2oai/h2o-meetups/blob/master/2017_11_29_Feature_Engineering/Feature%20Engineering.pdf\">Weights of Evidence</a>) features for many combinations of all variables (ip,app,device,os)  . This is very strange , because we tried standard target encoding and <strong>it did not work</strong>. This is very strange, because the ordering of likelihood features and woe should be the same/similar (only the range in woe is more condensed) . We are not sure why this happens, maybe it has to do with the binning of Lightgbm?</p>\n\n<p>3)  We got a boost of +4 via sorting ties of time in the same groups of ip,app,device,os, making certain the is_attributed==1 <strong>is always last</strong> <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677\">as it was pointed out in the forums</a>. </p>\n\n<p>Each one of the team members had different features - mine were more related with time-series. Like counts of previous (and next ) days, hours, minutes and seconds of ips,apps,device (and their combinations) . </p>\n\n<p>Other than that features that measure time between next/previous clicks for various sortings (like ip and app or ip,app,device and os)  were also important. </p>\n\n<p>The most important feature (importance-wise)  was by far the app (treated as categorical). You could see that certain apps had very different (HIGH) probabilities of <strong>is_attributed</strong>  - I called them for fun tr-APPS!</p>",
      "rawMarkdown": "I would like to start via praising my teamates for an excellent effort and congratulate the winners for  an intense last-days' race. \n\nAlso thank you to all the people who shared code and ideas - they made it a great competition (with the exception of the latest high-scoring kernels) . Special thanks to [anttip][1] for his overall contribution with [wordbatch][2] and kernels in general .\n\nMy favourite positions in a kaggle  competition are 1st and 4th. 1st you get most points/money. 4th, you dont get as many points, but you dont have to reproduce your solution :)\n\nOur validation schema was as follows: \n\nWe used days 7,8 for training and we were making predictions for 9th day (hours [4,14])\nFor test predictions, we were training on all 7,8,9 and making predictions for the test day (10th)\n\nThings that work for us apart from what it is already in forums/public kernels:\n\n1) [Restacking][3] - When we first started doing Stacking, we could barely get 1,2 points our of it (like from 0.9818 to 0.9820). After adding ALL the features used in our standard modelling to the predictions of the ninth day - we got another +3 boost (to 0.9823). We ended up having around 50 models - mostly lightgbms, but also nns , FMs and some linear models. NNs were on par with LGB models. \n\n2) We got another +4 from creating WoE ([Weights of Evidence][4]) features for many combinations of all variables (ip,app,device,os)  . This is very strange , because we tried standard target encoding and **it did not work**. This is very strange, because the ordering of likelihood features and woe should be the same/similar (only the range in woe is more condensed) . We are not sure why this happens, maybe it has to do with the binning of Lightgbm?\n\n3)  We got a boost of +4 via sorting ties of time in the same groups of ip,app,device,os, making certain the is_attributed==1 **is always last** [as it was pointed out in the forums][5]. \n\nEach one of the team members had different features - mine were more related with time-series. Like counts of previous (and next ) days, hours, minutes and seconds of ips,apps,device (and their combinations) . \n\nOther than that features that measure time between next/previous clicks for various sortings (like ip and app or ip,app,device and os)  were also important. \n\nThe most important feature (importance-wise)  was by far the app (treated as categorical). You could see that certain apps had very different (HIGH) probabilities of **is_attributed**  - I called them for fun tr-APPS!\n\n  [1]: https://www.kaggle.com/anttip\n  [2]: https://github.com/anttttti/Wordbatch\n  [3]: https://github.com/kaz-Anova/StackNet#restacking-mode\n  [4]: https://github.com/h2oai/h2o-meetups/blob/master/2017_11_29_Feature_Engineering/Feature%20Engineering.pdf\n  [5]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677",
      "votes": 136
    },
    {
      "id": 324869,
      "postDate": "2018-05-08T00:57:37.810Z",
      "content": "<p>Thanks for sharing and I have one question. For your no.3 tips, you sort ties of time in the same group to make sure is_attributed = 1 is the last, but I think it is not doable in test data because we dont know which one has is_attributed = 1 for test data, right? Thank you again look forward to receiving your reply.</p>",
      "rawMarkdown": "Thanks for sharing and I have one question. For your no.3 tips, you sort ties of time in the same group to make sure is_attributed = 1 is the last, but I think it is not doable in test data because we dont know which one has is_attributed = 1 for test data, right? Thank you again look forward to receiving your reply.",
      "votes": 3,
      "replies": [
        {
          "id": 325329,
          "postDate": "2018-05-08T10:09:06.040Z",
          "content": "<p>Please read here : <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268</a></p>",
          "rawMarkdown": "Please read here : https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268"
        }
      ]
    },
    {
      "id": 325102,
      "postDate": "2018-05-08T05:57:01.737Z",
      "content": "<p>@KazAnova, thanks for sharing and congrats for your 4th place. Your thoughts on WoE vs target encoding with LightGBM are very interesting indeed.</p>",
      "rawMarkdown": "@KazAnova, thanks for sharing and congrats for your 4th place. Your thoughts on WoE vs target encoding with LightGBM are very interesting indeed.",
      "votes": 4,
      "replies": [
        {
          "id": 325344,
          "postDate": "2018-05-08T10:17:34.913Z",
          "content": "<p>Yeah - I will investigate more too - I still think that was very strange</p>",
          "rawMarkdown": "Yeah - I will investigate more too - I still think that was very strange"
        }
      ]
    },
    {
      "id": 325716,
      "postDate": "2018-05-08T19:25:11.997Z",
      "content": "<p>Congrats! Thanks for sharing your experience!</p>",
      "rawMarkdown": "Congrats! Thanks for sharing your experience!",
      "votes": 1
    },
    {
      "id": 325636,
      "postDate": "2018-05-08T16:49:05.287Z",
      "content": "<p>Congrats. <br> May I ask if you did feature engineering for each day separately or you concat train + test_supplement to do feature engineering? </p>",
      "rawMarkdown": "Congrats. <br> May I ask if you did feature engineering for each day separately or you concat train + test_supplement to do feature engineering? ",
      "votes": 1
    },
    {
      "id": 325309,
      "postDate": "2018-05-08T09:41:12.870Z",
      "content": "<p>Congrats @KazAnova and thanks for sharing. </p>",
      "rawMarkdown": "Congrats @KazAnova and thanks for sharing. ",
      "votes": 1
    },
    {
      "id": 325297,
      "postDate": "2018-05-08T09:31:12.817Z",
      "content": "<p>Thanks for this nice infos! As you I learn in this competition wordbatch and I will definitely have a more thorough look at it now that the competition is over. This lib (wordbatch) seems very promising</p>",
      "rawMarkdown": "Thanks for this nice infos! As you I learn in this competition wordbatch and I will definitely have a more thorough look at it now that the competition is over. This lib (wordbatch) seems very promising",
      "votes": 1
    },
    {
      "id": 325283,
      "postDate": "2018-05-08T09:20:59.107Z",
      "content": "<p>Congratz Marios and team ... especially on coming 4th so you don't need to reproduce the solution... Was that on purpose ? haha  :) \nInteresting to see FMs in here, we tried with them a lot and they did not help - but did not add to a stack. </p>",
      "rawMarkdown": "Congratz Marios and team ... especially on coming 4th so you don't need to reproduce the solution... Was that on purpose ? haha  :) \nInteresting to see FMs in here, we tried with them a lot and they did not help - but did not add to a stack. ",
      "votes": 1,
      "replies": [
        {
          "id": 325341,
          "postDate": "2018-05-08T10:16:15.877Z",
          "content": "<p>haha - not exactly on purpose, but I certainly did not mind slipping down at the end :). We definitely got  something from FMs and some Vowpal wabbit models trained only on the categorical features (so 5 variables as input) exploring all possible interactions . These accounted for more than 8% of the total gain (as in lightgbm) of the last stack. </p>",
          "rawMarkdown": "haha - not exactly on purpose, but I certainly did not mind slipping down at the end :). We definitely got  something from FMs and some Vowpal wabbit models trained only on the categorical features (so 5 variables as input) exploring all possible interactions . These accounted for more than 8% of the total gain (as in lightgbm) of the last stack. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 325060,
      "postDate": "2018-05-08T04:37:35.593Z",
      "content": "<p>Congratulations,I should learn carefully tomorrow!</p>",
      "rawMarkdown": "Congratulations,I should learn carefully tomorrow!",
      "votes": 1,
      "replies": [
        {
          "id": 325342,
          "postDate": "2018-05-08T10:17:04.480Z",
          "content": "<p>Likewise - thank you for sharing and congrats for the 3rd place!</p>",
          "rawMarkdown": "Likewise - thank you for sharing and congrats for the 3rd place!"
        }
      ]
    },
    {
      "id": 325051,
      "postDate": "2018-05-08T04:10:54.197Z",
      "content": "<blockquote>\n  <p>3) We got a boost of +4 via sorting ties of time in the same groups of ip,app,device,os, making certain the is_attributed==1 is always last</p>\n</blockquote>\n\n<p>Thanks for the quick and nice sharing. One more question here: For my model, I trusted the original ordering of clicks to work on the time_till_next_click feature, and that gave me reasonably good results. But as you've mentioned, have you reordered clicks in groupby (ip, app, device, os) for time-delta features in training data?</p>",
      "rawMarkdown": "&gt; 3) We got a boost of +4 via sorting ties of time in the same groups of ip,app,device,os, making certain the is_attributed==1 is always last\n\nThanks for the quick and nice sharing. One more question here: For my model, I trusted the original ordering of clicks to work on the time_till_next_click feature, and that gave me reasonably good results. But as you've mentioned, have you reordered clicks in groupby (ip, app, device, os) for time-delta features in training data?",
      "votes": 1,
      "replies": [
        {
          "id": 325323,
          "postDate": "2018-05-08T10:05:33.427Z",
          "content": "<p>please have a read here - I did : <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268</a></p>",
          "rawMarkdown": "please have a read here - I did : https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268"
        }
      ]
    },
    {
      "id": 325019,
      "postDate": "2018-05-08T03:24:17.523Z",
      "content": "<p>Thanks for sharing KazAnova and congrats! I tried target encoding features and they didn't work too in LightGBM.</p>\n\n<p>Your participation and other top Kagglers' are the main reasons why I still want to participate in Kaggle, because you guys show that hard &amp; smart works always pay off at the end, not just by blending other public kernels/csvs !</p>",
      "rawMarkdown": "Thanks for sharing KazAnova and congrats! I tried target encoding features and they didn't work too in LightGBM.\n\nYour participation and other top Kagglers' are the main reasons why I still want to participate in Kaggle, because you guys show that hard &amp; smart works always pay off at the end, not just by blending other public kernels/csvs !",
      "votes": 1
    },
    {
      "id": 324917,
      "postDate": "2018-05-08T01:38:44.573Z",
      "content": "<p>thank you for sharing. This is the moment, that people share top solutions, I waited for two month.</p>",
      "rawMarkdown": "thank you for sharing. This is the moment, that people share top solutions, I waited for two month.",
      "votes": 1,
      "replies": [
        {
          "id": 325347,
          "postDate": "2018-05-08T10:19:31.953Z",
          "content": "<p>yeah, me too :)</p>",
          "rawMarkdown": "yeah, me too :)"
        }
      ]
    },
    {
      "id": 324885,
      "postDate": "2018-05-08T01:12:12.653Z",
      "content": "<p>Thanks for sharing! Could you share the hardware used for training + validation of your team's models? How much RAM did it take?</p>",
      "rawMarkdown": "Thanks for sharing! Could you share the hardware used for training + validation of your team's models? How much RAM did it take?",
      "votes": 1,
      "replies": [
        {
          "id": 325330,
          "postDate": "2018-05-08T10:10:33.410Z",
          "content": "<p>We used a lot of machinery :) . We used a few 256 GB (linux)  servers with 40 cores and we also had one with  512 GB and 64 cores. </p>",
          "rawMarkdown": "We used a lot of machinery :) . We used a few 256 GB (linux)  servers with 40 cores and we also had one with  512 GB and 64 cores. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 324884,
      "postDate": "2018-05-08T01:12:06.560Z",
      "content": "<p>Thanks for the quick share and the github links, good stuff to learn :)</p>",
      "rawMarkdown": "Thanks for the quick share and the github links, good stuff to learn :)",
      "votes": 1
    },
    {
      "id": 324878,
      "postDate": "2018-05-08T01:08:09.623Z",
      "content": "<p>Спасибо</p>",
      "rawMarkdown": "Спасибо",
      "votes": 1,
      "replies": [
        {
          "id": 325345,
          "postDate": "2018-05-08T10:19:13.913Z",
          "content": "<p>Также спасибо</p>",
          "rawMarkdown": "Также спасибо"
        }
      ]
    },
    {
      "id": 324867,
      "postDate": "2018-05-08T00:56:05.607Z",
      "content": "<p>Thanks for sharing so quickly, and congrats on the result!</p>",
      "rawMarkdown": "Thanks for sharing so quickly, and congrats on the result!"
    },
    {
      "id": 365184,
      "postDate": "2018-08-02T04:17:51.207Z",
      "content": "<p>Sorry this seems late. May I ask when you tried target encoding / WOE, are the time series factor taken into account or did you simply just encode w.r.t the target. I.e we ignore the time component since encoding throughout will result in past / future data might mixed. </p>",
      "rawMarkdown": "Sorry this seems late. May I ask when you tried target encoding / WOE, are the time series factor taken into account or did you simply just encode w.r.t the target. I.e we ignore the time component since encoding throughout will result in past / future data might mixed. "
    },
    {
      "id": 328910,
      "postDate": "2018-05-15T10:56:40.827Z",
      "content": "<p>Nice solution, I didn't notice about duplicate problem at all. btw, I'm curious about how much the score will become better if we combine the best submission. could you upload your best submission if possible? 1st &amp; 5th &amp; 6th has already uploaded. <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423</a></p>",
      "rawMarkdown": "Nice solution, I didn't notice about duplicate problem at all. btw, I'm curious about how much the score will become better if we combine the best submission. could you upload your best submission if possible? 1st &amp; 5th &amp; 6th has already uploaded. https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423"
    },
    {
      "id": 327747,
      "postDate": "2018-05-12T11:05:04.480Z",
      "content": "<p>Congrats! How to deal with \"counts of previous  days\" if it is first day?</p>",
      "rawMarkdown": "Congrats! How to deal with \"counts of previous  days\" if it is first day?"
    },
    {
      "id": 326038,
      "postDate": "2018-05-09T07:19:04.090Z",
      "content": "<p>Thank you for your sharing. I don't understand the last word of Restacking part, what does \"NNs were on par with LGB models\" mean ? Could you explain more ?</p>",
      "rawMarkdown": "Thank you for your sharing. I don't understand the last word of Restacking part, what does \"NNs were on par with LGB models\" mean ? Could you explain more ?",
      "replies": [
        {
          "id": 326554,
          "postDate": "2018-05-09T22:51:22.573Z",
          "content": "<p>He means their neural net models were performing as well as the lgbm ones</p>",
          "rawMarkdown": "He means their neural net models were performing as well as the lgbm ones"
        },
        {
          "id": 326563,
          "postDate": "2018-05-09T23:26:00.317Z",
          "content": "<p>Got it. Thank you.</p>",
          "rawMarkdown": "Got it. Thank you."
        }
      ]
    },
    {
      "id": 325155,
      "postDate": "2018-05-08T07:16:48.887Z",
      "content": "<p>Congratulations! Really nice job ;)</p>",
      "rawMarkdown": "Congratulations! Really nice job ;)"
    },
    {
      "id": 325078,
      "postDate": "2018-05-08T05:14:28.050Z",
      "content": "<p>hello, I have read the ppt in your tip 2, and I have a question. The metric '<strong>Information value</strong>'  show predictiveness of original feature or WOE?</p>",
      "rawMarkdown": "hello, I have read the ppt in your tip 2, and I have a question. The metric '**Information value**'  show predictiveness of original feature or WOE?",
      "replies": [
        {
          "id": 325335,
          "postDate": "2018-05-08T10:12:43.660Z",
          "content": "<p>Yes -it can be used as a measure to gauge the \"predictiveness\" of a feature.  You can see another exampl here : <a href=\"http://ucanalytics.com/blogs/information-value-and-weight-of-evidencebanking-case/\">http://ucanalytics.com/blogs/information-value-and-weight-of-evidencebanking-case/</a></p>",
          "rawMarkdown": "Yes -it can be used as a measure to gauge the \"predictiveness\" of a feature.  You can see another exampl here : http://ucanalytics.com/blogs/information-value-and-weight-of-evidencebanking-case/"
        },
        {
          "id": 325389,
          "postDate": "2018-05-08T11:09:16.953Z",
          "content": "<p>Thank you very much,your tip is so useful</p>",
          "rawMarkdown": "Thank you very much,your tip is so useful"
        }
      ]
    },
    {
      "id": 325069,
      "postDate": "2018-05-08T04:55:05.907Z",
      "content": "<p>Thanks for sharing. Lots of useful tips.</p>",
      "rawMarkdown": "Thanks for sharing. Lots of useful tips."
    },
    {
      "id": 324953,
      "postDate": "2018-05-08T02:13:06.190Z",
      "content": "<p>Thank you for sharing and congratulation!</p>",
      "rawMarkdown": "Thank you for sharing and congratulation!"
    },
    {
      "id": 324931,
      "postDate": "2018-05-08T01:47:49.950Z",
      "content": "<p>Thanks for sharing. May I ask one question, if you don't mind? How do you deal with tr-APPS?</p>",
      "rawMarkdown": "Thanks for sharing. May I ask one question, if you don't mind? How do you deal with tr-APPS?",
      "replies": [
        {
          "id": 325397,
          "postDate": "2018-05-08T11:30:24.930Z",
          "content": "<p>Many app-focused features and made certain app was in restacking too (as categorical) </p>",
          "rawMarkdown": "Many app-focused features and made certain app was in restacking too (as categorical) "
        }
      ]
    },
    {
      "id": 324903,
      "postDate": "2018-05-08T01:27:03.010Z",
      "content": "<p>Thanks for sharing and congratulations on a strong finish.</p>",
      "rawMarkdown": "Thanks for sharing and congratulations on a strong finish."
    },
    {
      "id": 325704,
      "postDate": "2018-05-08T18:49:50.883Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 325314,
      "postDate": "2018-05-08T09:46:49.400Z",
      "content": "<p>as usual learn a lot from you.. thanks</p>",
      "rawMarkdown": "as usual learn a lot from you.. thanks",
      "votes": 1
    },
    {
      "id": 325198,
      "postDate": "2018-05-08T08:03:45.357Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 324922,
      "postDate": "2018-05-08T01:40:28.360Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 324880,
      "postDate": "2018-05-08T01:09:48.957Z",
      "content": "<p>Thanks for sharing. Good Job.</p>",
      "rawMarkdown": "Thanks for sharing. Good Job.",
      "votes": 1
    },
    {
      "id": 324866,
      "postDate": "2018-05-08T00:55:50.637Z",
      "content": "<p>Thanks for sharing and congrats!</p>",
      "rawMarkdown": "Thanks for sharing and congrats!",
      "votes": 1
    },
    {
      "id": 324864,
      "postDate": "2018-05-08T00:51:35.753Z",
      "content": "<p>thanks for your sharing </p>",
      "rawMarkdown": "thanks for your sharing ",
      "votes": 1
    },
    {
      "id": 325055,
      "postDate": "2018-05-08T04:26:22.360Z",
      "content": "<p>Thanks for sharing and congratulation.</p>",
      "rawMarkdown": "Thanks for sharing and congratulation."
    }
  ],
  "comments": [
    {
      "id": 324869,
      "author_name": "Snorlax",
      "author_url": "",
      "post_date": "2018-05-08T00:57:37.810000",
      "content": "<p>Thanks for sharing and I have one question. For your no.3 tips, you sort ties of time in the same group to make sure is_attributed = 1 is the last, but I think it is not doable in test data because we dont know which one has is_attributed = 1 for test data, right? Thank you again look forward to receiving your reply.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 325329,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-05-08T10:09:06.040000",
          "content": "<p>Please read here : <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325102,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2018-05-08T05:57:01.737000",
      "content": "<p>@KazAnova, thanks for sharing and congrats for your 4th place. Your thoughts on WoE vs target encoding with LightGBM are very interesting indeed.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 325344,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-05-08T10:17:34.913000",
          "content": "<p>Yeah - I will investigate more too - I still think that was very strange</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325716,
      "author_name": "TaoL",
      "author_url": "",
      "post_date": "2018-05-08T19:25:11.997000",
      "content": "<p>Congrats! Thanks for sharing your experience!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325636,
      "author_name": "Sohaib Omar",
      "author_url": "",
      "post_date": "2018-05-08T16:49:05.287000",
      "content": "<p>Congrats. <br> May I ask if you did feature engineering for each day separately or you concat train + test_supplement to do feature engineering? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325309,
      "author_name": "Pranav Pandya",
      "author_url": "",
      "post_date": "2018-05-08T09:41:12.870000",
      "content": "<p>Congrats @KazAnova and thanks for sharing. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325297,
      "author_name": "Eric",
      "author_url": "",
      "post_date": "2018-05-08T09:31:12.817000",
      "content": "<p>Thanks for this nice infos! As you I learn in this competition wordbatch and I will definitely have a more thorough look at it now that the competition is over. This lib (wordbatch) seems very promising</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325283,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2018-05-08T09:20:59.107000",
      "content": "<p>Congratz Marios and team ... especially on coming 4th so you don't need to reproduce the solution... Was that on purpose ? haha  :) \nInteresting to see FMs in here, we tried with them a lot and they did not help - but did not add to a stack. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 325341,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-05-08T10:16:15.877000",
          "content": "<p>haha - not exactly on purpose, but I certainly did not mind slipping down at the end :). We definitely got  something from FMs and some Vowpal wabbit models trained only on the categorical features (so 5 variables as input) exploring all possible interactions . These accounted for more than 8% of the total gain (as in lightgbm) of the last stack. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 325060,
      "author_name": "bestfitting",
      "author_url": "",
      "post_date": "2018-05-08T04:37:35.593000",
      "content": "<p>Congratulations,I should learn carefully tomorrow!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325342,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-05-08T10:17:04.480000",
          "content": "<p>Likewise - thank you for sharing and congrats for the 3rd place!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325051,
      "author_name": "Endi Niu",
      "author_url": "",
      "post_date": "2018-05-08T04:10:54.197000",
      "content": "<blockquote>\n  <p>3) We got a boost of +4 via sorting ties of time in the same groups of ip,app,device,os, making certain the is_attributed==1 is always last</p>\n</blockquote>\n\n<p>Thanks for the quick and nice sharing. One more question here: For my model, I trusted the original ordering of clicks to work on the time_till_next_click feature, and that gave me reasonably good results. But as you've mentioned, have you reordered clicks in groupby (ip, app, device, os) for time-delta features in training data?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325323,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-05-08T10:05:33.427000",
          "content": "<p>please have a read here - I did : <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325019,
      "author_name": "SDA",
      "author_url": "",
      "post_date": "2018-05-08T03:24:17.523000",
      "content": "<p>Thanks for sharing KazAnova and congrats! I tried target encoding features and they didn't work too in LightGBM.</p>\n\n<p>Your participation and other top Kagglers' are the main reasons why I still want to participate in Kaggle, because you guys show that hard &amp; smart works always pay off at the end, not just by blending other public kernels/csvs !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 324917,
      "author_name": "Johnny Liu",
      "author_url": "",
      "post_date": "2018-05-08T01:38:44.573000",
      "content": "<p>thank you for sharing. This is the moment, that people share top solutions, I waited for two month.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325347,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-05-08T10:19:31.953000",
          "content": "<p>yeah, me too :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 324885,
      "author_name": "astraldawn",
      "author_url": "",
      "post_date": "2018-05-08T01:12:12.653000",
      "content": "<p>Thanks for sharing! Could you share the hardware used for training + validation of your team's models? How much RAM did it take?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325330,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-05-08T10:10:33.410000",
          "content": "<p>We used a lot of machinery :) . We used a few 256 GB (linux)  servers with 40 cores and we also had one with  512 GB and 64 cores. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 324884,
      "author_name": "KenMa",
      "author_url": "",
      "post_date": "2018-05-08T01:12:06.560000",
      "content": "<p>Thanks for the quick share and the github links, good stuff to learn :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 324878,
      "author_name": "MengYe",
      "author_url": "",
      "post_date": "2018-05-08T01:08:09.623000",
      "content": "<p>Спасибо</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325345,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-05-08T10:19:13.913000",
          "content": "<p>Также спасибо</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 324867,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-05-08T00:56:05.607000",
      "content": "<p>Thanks for sharing so quickly, and congrats on the result!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 365184,
      "author_name": "germayne",
      "author_url": "",
      "post_date": "2018-08-02T04:17:51.207000",
      "content": "<p>Sorry this seems late. May I ask when you tried target encoding / WOE, are the time series factor taken into account or did you simply just encode w.r.t the target. I.e we ignore the time component since encoding throughout will result in past / future data might mixed. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 328910,
      "author_name": "mamas",
      "author_url": "",
      "post_date": "2018-05-15T10:56:40.827000",
      "content": "<p>Nice solution, I didn't notice about duplicate problem at all. btw, I'm curious about how much the score will become better if we combine the best submission. could you upload your best submission if possible? 1st &amp; 5th &amp; 6th has already uploaded. <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 327747,
      "author_name": "yxzf",
      "author_url": "",
      "post_date": "2018-05-12T11:05:04.480000",
      "content": "<p>Congrats! How to deal with \"counts of previous  days\" if it is first day?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326038,
      "author_name": "yyqing",
      "author_url": "",
      "post_date": "2018-05-09T07:19:04.090000",
      "content": "<p>Thank you for your sharing. I don't understand the last word of Restacking part, what does \"NNs were on par with LGB models\" mean ? Could you explain more ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 326554,
          "author_name": "Callum Gundlach",
          "author_url": "",
          "post_date": "2018-05-09T22:51:22.573000",
          "content": "<p>He means their neural net models were performing as well as the lgbm ones</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326563,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "2018-05-09T23:26:00.317000",
          "content": "<p>Got it. Thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325155,
      "author_name": "Meyk",
      "author_url": "",
      "post_date": "2018-05-08T07:16:48.887000",
      "content": "<p>Congratulations! Really nice job ;)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325078,
      "author_name": "Johnny Liu",
      "author_url": "",
      "post_date": "2018-05-08T05:14:28.050000",
      "content": "<p>hello, I have read the ppt in your tip 2, and I have a question. The metric '<strong>Information value</strong>'  show predictiveness of original feature or WOE?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 325335,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-05-08T10:12:43.660000",
          "content": "<p>Yes -it can be used as a measure to gauge the \"predictiveness\" of a feature.  You can see another exampl here : <a href=\"http://ucanalytics.com/blogs/information-value-and-weight-of-evidencebanking-case/\">http://ucanalytics.com/blogs/information-value-and-weight-of-evidencebanking-case/</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 325389,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-05-08T11:09:16.953000",
          "content": "<p>Thank you very much,your tip is so useful</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325069,
      "author_name": "Araks Stepanyan",
      "author_url": "",
      "post_date": "2018-05-08T04:55:05.907000",
      "content": "<p>Thanks for sharing. Lots of useful tips.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 324953,
      "author_name": "Sun ZhiHao",
      "author_url": "",
      "post_date": "2018-05-08T02:13:06.190000",
      "content": "<p>Thank you for sharing and congratulation!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 324931,
      "author_name": "YulinGUO",
      "author_url": "",
      "post_date": "2018-05-08T01:47:49.950000",
      "content": "<p>Thanks for sharing. May I ask one question, if you don't mind? How do you deal with tr-APPS?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 325397,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-05-08T11:30:24.930000",
          "content": "<p>Many app-focused features and made certain app was in restacking too (as categorical) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 324903,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-05-08T01:27:03.010000",
      "content": "<p>Thanks for sharing and congratulations on a strong finish.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325704,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T18:49:50.883000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325314,
      "author_name": "miewmiewman",
      "author_url": "",
      "post_date": "2018-05-08T09:46:49.400000",
      "content": "<p>as usual learn a lot from you.. thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325198,
      "author_name": "Feiyang Pan",
      "author_url": "",
      "post_date": "2018-05-08T08:03:45.357000",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 324922,
      "author_name": "Jeongwoo",
      "author_url": "",
      "post_date": "2018-05-08T01:40:28.360000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 324880,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T01:09:48.957000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 324866,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T00:55:50.637000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 324864,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T00:51:35.753000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325055,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T04:26:22.360000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "324861": "I would like to start via praising my teamates for an excellent effort and congratulate the winners for  an intense last-days' race. \n\nAlso thank you to all the people who shared code and ideas - they made it a great competition (with the exception of the latest high-scoring kernels) . Special thanks to [anttip][1] for his overall contribution with [wordbatch][2] and kernels in general .\n\nMy favourite positions in a kaggle  competition are 1st and 4th. 1st you get most points/money. 4th, you dont get as many points, but you dont have to reproduce your solution :)\n\nOur validation schema was as follows: \n\nWe used days 7,8 for training and we were making predictions for 9th day (hours [4,14])\nFor test predictions, we were training on all 7,8,9 and making predictions for the test day (10th)\n\nThings that work for us apart from what it is already in forums/public kernels:\n\n1) [Restacking][3] - When we first started doing Stacking, we could barely get 1,2 points our of it (like from 0.9818 to 0.9820). After adding ALL the features used in our standard modelling to the predictions of the ninth day - we got another +3 boost (to 0.9823). We ended up having around 50 models - mostly lightgbms, but also nns , FMs and some linear models. NNs were on par with LGB models. \n\n2) We got another +4 from creating WoE ([Weights of Evidence][4]) features for many combinations of all variables (ip,app,device,os)  . This is very strange , because we tried standard target encoding and **it did not work**. This is very strange, because the ordering of likelihood features and woe should be the same/similar (only the range in woe is more condensed) . We are not sure why this happens, maybe it has to do with the binning of Lightgbm?\n\n3)  We got a boost of +4 via sorting ties of time in the same groups of ip,app,device,os, making certain the is_attributed==1 **is always last** [as it was pointed out in the forums][5]. \n\nEach one of the team members had different features - mine were more related with time-series. Like counts of previous (and next ) days, hours, minutes and seconds of ips,apps,device (and their combinations) . \n\nOther than that features that measure time between next/previous clicks for various sortings (like ip and app or ip,app,device and os)  were also important. \n\nThe most important feature (importance-wise)  was by far the app (treated as categorical). You could see that certain apps had very different (HIGH) probabilities of **is_attributed**  - I called them for fun tr-APPS!\n\n  [1]: https://www.kaggle.com/anttip\n  [2]: https://github.com/anttttti/Wordbatch\n  [3]: https://github.com/kaz-Anova/StackNet#restacking-mode\n  [4]: https://github.com/h2oai/h2o-meetups/blob/master/2017_11_29_Feature_Engineering/Feature%20Engineering.pdf\n  [5]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677",
    "324869": "Thanks for sharing and I have one question. For your no.3 tips, you sort ties of time in the same group to make sure is_attributed = 1 is the last, but I think it is not doable in test data because we dont know which one has is_attributed = 1 for test data, right? Thank you again look forward to receiving your reply.",
    "325102": "@KazAnova, thanks for sharing and congrats for your 4th place. Your thoughts on WoE vs target encoding with LightGBM are very interesting indeed.",
    "325716": "Congrats! Thanks for sharing your experience!",
    "325636": "Congrats. <br> May I ask if you did feature engineering for each day separately or you concat train + test_supplement to do feature engineering? ",
    "325309": "Congrats @KazAnova and thanks for sharing. ",
    "325297": "Thanks for this nice infos! As you I learn in this competition wordbatch and I will definitely have a more thorough look at it now that the competition is over. This lib (wordbatch) seems very promising",
    "325283": "Congratz Marios and team ... especially on coming 4th so you don't need to reproduce the solution... Was that on purpose ? haha  :) \nInteresting to see FMs in here, we tried with them a lot and they did not help - but did not add to a stack. ",
    "325060": "Congratulations,I should learn carefully tomorrow!",
    "325051": "&gt; 3) We got a boost of +4 via sorting ties of time in the same groups of ip,app,device,os, making certain the is_attributed==1 is always last\n\nThanks for the quick and nice sharing. One more question here: For my model, I trusted the original ordering of clicks to work on the time_till_next_click feature, and that gave me reasonably good results. But as you've mentioned, have you reordered clicks in groupby (ip, app, device, os) for time-delta features in training data?",
    "325019": "Thanks for sharing KazAnova and congrats! I tried target encoding features and they didn't work too in LightGBM.\n\nYour participation and other top Kagglers' are the main reasons why I still want to participate in Kaggle, because you guys show that hard &amp; smart works always pay off at the end, not just by blending other public kernels/csvs !",
    "324917": "thank you for sharing. This is the moment, that people share top solutions, I waited for two month.",
    "324885": "Thanks for sharing! Could you share the hardware used for training + validation of your team's models? How much RAM did it take?",
    "324884": "Thanks for the quick share and the github links, good stuff to learn :)",
    "324878": "Спасибо",
    "324867": "Thanks for sharing so quickly, and congrats on the result!",
    "365184": "Sorry this seems late. May I ask when you tried target encoding / WOE, are the time series factor taken into account or did you simply just encode w.r.t the target. I.e we ignore the time component since encoding throughout will result in past / future data might mixed. ",
    "328910": "Nice solution, I didn't notice about duplicate problem at all. btw, I'm curious about how much the score will become better if we combine the best submission. could you upload your best submission if possible? 1st &amp; 5th &amp; 6th has already uploaded. https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423",
    "327747": "Congrats! How to deal with \"counts of previous  days\" if it is first day?",
    "326038": "Thank you for your sharing. I don't understand the last word of Restacking part, what does \"NNs were on par with LGB models\" mean ? Could you explain more ?",
    "325155": "Congratulations! Really nice job ;)",
    "325078": "hello, I have read the ppt in your tip 2, and I have a question. The metric '**Information value**'  show predictiveness of original feature or WOE?",
    "325069": "Thanks for sharing. Lots of useful tips.",
    "324953": "Thank you for sharing and congratulation!",
    "324931": "Thanks for sharing. May I ask one question, if you don't mind? How do you deal with tr-APPS?",
    "324903": "Thanks for sharing and congratulations on a strong finish.",
    "325704": "",
    "325314": "as usual learn a lot from you.. thanks",
    "325198": "Congrats and thanks for sharing!",
    "324922": "Thanks for sharing",
    "324880": "Thanks for sharing. Good Job.",
    "324866": "Thanks for sharing and congrats!",
    "324864": "thanks for your sharing ",
    "325055": "Thanks for sharing and congratulation."
  }
}