{
  "id": 56279,
  "title": "9th place",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/writeups/brute-force-attack-9th-place",
  "author_name": "",
  "post_date": "2018-05-08T11:41:50.033Z",
  "votes": 49,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Congrats to all the winners, and disappointed to see so many people affected by the late sharing of scripts - so sad to see a 0.9811 sub uploaded on last day and affect people's hard work.   </p>\n\n<p>Our solution was a couple of LGB's at public 0.9821 or 0.9820, and nnet at public 0.9816. We did a lot of feature engineering in <code>data.table</code> this is very nice for building features on such a dataset, especially, time based features. Entropy helped - <a href=\"https://github.com/owenzhang/kaggle-avito/blob/a7a2cc853b0ca86f07cdb9dd483779b2927b99ee/avito_utils.R\">linky</a>  ; particularly entropy over time based features like minute. Device <code>3032</code> was all, or nearly all, no downloads and was not in test, so we got circa 0.0004 from removing. We also got split second times and incorporated that into lead times - this was done by turning the ordering per second, and count of clicks per second, into sub-second time - if there are 100 clicks in a second <code>16:00:00</code>, first click gets <code>16:00:00.00</code>, second gets <code>16:00:00.01</code>, next <code>16:00:00.02</code> ... etc. This gave us good lift, around 0.001 early in the competition, but we did not test the lift later to see the drop by leaving it out. Stacking gave us, circa .0003. Counts of previous periods, eg. period this time yesterday, helped also. \nA lot of other things tried with no, or small lift - used about 30 features in all.  </p>",
  "messages": [
    {
      "id": "325205",
      "postDate": "05/08/2018 08:13:14",
      "content": "<p>Congrats to all the winners, and disappointed to see so many people affected by the late sharing of scripts - so sad to see a 0.9811 sub uploaded on last day and affect people's hard work.   </p>\n\n<p>Our solution was a couple of LGB's at public 0.9821 or 0.9820, and nnet at public 0.9816. We did a lot of feature engineering in <code>data.table</code> this is very nice for building features on such a dataset, especially, time based features. Entropy helped - <a href=\"https://github.com/owenzhang/kaggle-avito/blob/a7a2cc853b0ca86f07cdb9dd483779b2927b99ee/avito_utils.R\">linky</a>  ; particularly entropy over time based features like minute. Device <code>3032</code> was all, or nearly all, no downloads and was not in test, so we got circa 0.0004 from removing. We also got split second times and incorporated that into lead times - this was done by turning the ordering per second, and count of clicks per second, into sub-second time - if there are 100 clicks in a second <code>16:00:00</code>, first click gets <code>16:00:00.00</code>, second gets <code>16:00:00.01</code>, next <code>16:00:00.02</code> ... etc. This gave us good lift, around 0.001 early in the competition, but we did not test the lift later to see the drop by leaving it out. Stacking gave us, circa .0003. Counts of previous periods, eg. period this time yesterday, helped also. \nA lot of other things tried with no, or small lift - used about 30 features in all.  </p>",
      "rawMarkdown": "Congrats to all the winners, and disappointed to see so many people affected by the late sharing of scripts - so sad to see a 0.9811 sub uploaded on last day and affect people's hard work.   \n   \nOur solution was a couple of LGB's at public 0.9821 or 0.9820, and nnet at public 0.9816. We did a lot of feature engineering in `data.table` this is very nice for building features on such a dataset, especially, time based features. Entropy helped - [linky][1]  ; particularly entropy over time based features like minute. Device `3032` was all, or nearly all, no downloads and was not in test, so we got circa 0.0004 from removing. We also got split second times and incorporated that into lead times - this was done by turning the ordering per second, and count of clicks per second, into sub-second time - if there are 100 clicks in a second `16:00:00`, first click gets `16:00:00.00`, second gets `16:00:00.01`, next `16:00:00.02` ... etc. This gave us good lift, around 0.001 early in the competition, but we did not test the lift later to see the drop by leaving it out. Stacking gave us, circa .0003. Counts of previous periods, eg. period this time yesterday, helped also. \nA lot of other things tried with no, or small lift - used about 30 features in all.  \n\n\n  [1]: https://github.com/owenzhang/kaggle-avito/blob/a7a2cc853b0ca86f07cdb9dd483779b2927b99ee/avito_utils.R",
      "votes": null
    },
    {
      "id": "325208",
      "postDate": "05/08/2018 08:19:15",
      "content": "<p>Well done</p>",
      "rawMarkdown": "Well done",
      "votes": null
    },
    {
      "id": "325214",
      "postDate": "05/08/2018 08:25:07",
      "content": "<p>Great job~Congrats and thx for sharing~</p>",
      "rawMarkdown": "Great job~Congrats and thx for sharing~",
      "votes": null
    },
    {
      "id": "325219",
      "postDate": "05/08/2018 08:31:37",
      "content": "<p>Thanks for sharing, and congrats on your result!</p>",
      "rawMarkdown": "Thanks for sharing, and congrats on your result!",
      "votes": null
    },
    {
      "id": "325224",
      "postDate": "05/08/2018 08:36:50",
      "content": "<p>Thanks CPMP. Well done on your score and position, I'm interested to hear your validation strategy - mine had +/- .0002 variance which made it difficult to assess features - I used test hours from day 9 to validate, and full day 8 to train. </p>",
      "rawMarkdown": "Thanks CPMP. Well done on your score and position, I'm interested to hear your validation strategy - mine had +/- .0002 variance which made it difficult to assess features - I used test hours from day 9 to validate, and full day 8 to train.",
      "votes": null
    },
    {
      "id": "325237",
      "postDate": "05/08/2018 08:50:42",
      "content": "<p>Congrats and thx for sharing. In many competions , there is less R kernal. so  when i see data.table, i am very excited! \ni hope master can give us kernal , please!  </p>",
      "rawMarkdown": "Congrats and thx for sharing. In many competions , there is less R kernal. so  when i see data.table, i am very excited! \ni hope master can give us kernal , please!",
      "votes": null
    },
    {
      "id": "325239",
      "postDate": "05/08/2018 08:52:37",
      "content": "<p>I used two validation set, hour 4 of day 9, and hours 5, 9, 10, 13, 14.  Given lgb stops on the the first that does not improve it helps remove some of the variability.  I also watched the train auc in the early days, discarding features that improved validation but also increased the gap with train a lot.  I stopped watching train auc in the last week to speed up things, but last time I checked I had a quite small gap.</p>\n\n<p>I also used LB, i.e. only kept a feature if local auc and LB improved.  Yes, I know this can lead to overfit, but given I was filtering first on local validation I think I escaped it for the most part.</p>",
      "rawMarkdown": "I used two validation set, hour 4 of day 9, and hours 5, 9, 10, 13, 14.  Given lgb stops on the the first that does not improve it helps remove some of the variability.  I also watched the train auc in the early days, discarding features that improved validation but also increased the gap with train a lot.  I stopped watching train auc in the last week to speed up things, but last time I checked I had a quite small gap.\n\nI also used LB, i.e. only kept a feature if local auc and LB improved.  Yes, I know this can lead to overfit, but given I was filtering first on local validation I think I escaped it for the most part.",
      "votes": null
    },
    {
      "id": "325249",
      "postDate": "05/08/2018 08:59:48",
      "content": "<p>Thanks Alex, the repo is a bit messy currently but here is an example of making the split second lead time in data.table - calculating the feature runs on full data set in less than a minute - most of the time is data load.</p>\n\n<p>Repo link  -- <a href=\"https://github.com/darraghdog/talking-data-click-fraud/blob/1827a7a437dddce85642e01f561953c3f0066573/features/make/secdistR_2703.R\">secdistR_2703.R</a></p>",
      "rawMarkdown": "Thanks Alex, the repo is a bit messy currently but here is an example of making the split second lead time in data.table - calculating the feature runs on full data set in less than a minute - most of the time is data load.\n\nRepo link  -- [secdistR_2703.R][1]\n\n\n  [1]: https://github.com/darraghdog/talking-data-click-fraud/blob/1827a7a437dddce85642e01f561953c3f0066573/features/make/secdistR_2703.R",
      "votes": null
    },
    {
      "id": "325253",
      "postDate": "05/08/2018 09:01:13",
      "content": "<p>thanks. i am looking forward to yours</p>",
      "rawMarkdown": "thanks. i am looking forward to yours",
      "votes": null
    },
    {
      "id": "325270",
      "postDate": "05/08/2018 09:13:41",
      "content": "<p>Nice idea to monitor the variation between train and test...</p>",
      "rawMarkdown": "Nice idea to monitor the variation between train and test...",
      "votes": null
    },
    {
      "id": "325274",
      "postDate": "05/08/2018 09:15:37",
      "content": "<p>I find it to be the best way to detect overfiting.</p>",
      "rawMarkdown": "I find it to be the best way to detect overfiting.",
      "votes": null
    },
    {
      "id": "325279",
      "postDate": "05/08/2018 09:17:44",
      "content": "<p>Thanks Darragh for sharing and well done! I like that you relate feature engineering to entropy. Do you have any theoretical paper on this to share?</p>",
      "rawMarkdown": "Thanks Darragh for sharing and well done! I like that you relate feature engineering to entropy. Do you have any theoretical paper on this to share?",
      "votes": null
    },
    {
      "id": "325312",
      "postDate": "05/08/2018 09:45:37",
      "content": "<p>Sorry, I have no paper - but I think of it as a step up from ctuniq - where ctuniq counts the unique groups in a feature level; entropy tells how concentrated it is in certain groups. For example entropy over minutes; is there generally an even spread over minutes or do we have bursts of clicks at certain times.   </p>",
      "rawMarkdown": "Sorry, I have no paper - but I think of it as a step up from ctuniq - where ctuniq counts the unique groups in a feature level; entropy tells how concentrated it is in certain groups. For example entropy over minutes; is there generally an even spread over minutes or do we have bursts of clicks at certain times.",
      "votes": null
    },
    {
      "id": "325353",
      "postDate": "05/08/2018 10:26:31",
      "content": "<p>Congrats @Darragh and @Giba,</p>\n\n<p>And thanks for sharing the approach and repository link. This is very helpful for R users like me. </p>",
      "rawMarkdown": "Congrats @Darragh and @Giba,\n\nAnd thanks for sharing the approach and repository link. This is very helpful for R users like me.",
      "votes": null
    },
    {
      "id": "325356",
      "postDate": "05/08/2018 10:41:16",
      "content": "<p>It is really interesting that removing device  3032 gives you such boost ;) In my case particularly because I was thinking about it and then gave it up. More lessons learned.</p>\n\n<p>Congratulations! </p>",
      "rawMarkdown": "It is really interesting that removing device  3032 gives you such boost ;) In my case particularly because I was thinking about it and then gave it up. More lessons learned.\n\nCongratulations!",
      "votes": null
    },
    {
      "id": "325387",
      "postDate": "05/08/2018 11:06:04",
      "content": "<p>Thanks for sharing. amazing boost from removing the 3032 device</p>",
      "rawMarkdown": "Thanks for sharing. amazing boost from removing the 3032 device",
      "votes": null
    },
    {
      "id": "325393",
      "postDate": "05/08/2018 11:18:10",
      "content": "<p>Thanks you for sharing. It's nice to see <strong>data.table</strong>-based solution.</p>",
      "rawMarkdown": "Thanks you for sharing. It's nice to see **data.table**-based solution.",
      "votes": null
    },
    {
      "id": "325405",
      "postDate": "05/08/2018 11:39:50",
      "content": "<p>I find this to be a brilliant idea! Congrats!</p>",
      "rawMarkdown": "I find this to be a brilliant idea! Congrats!",
      "votes": null
    },
    {
      "id": "325807",
      "postDate": "05/08/2018 22:02:10",
      "content": "<p>Thanks for your sharing!</p>",
      "rawMarkdown": "Thanks for your sharing!",
      "votes": null
    },
    {
      "id": "326008",
      "postDate": "05/09/2018 06:51:00",
      "content": "<p>Thanks Darragh for the feedback.</p>",
      "rawMarkdown": "Thanks Darragh for the feedback.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 325208,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "05/08/2018 08:19:15",
      "content": "<p>Well done</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325214,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "05/08/2018 08:25:07",
      "content": "<p>Great job~Congrats and thx for sharing~</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325219,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/08/2018 08:31:37",
      "content": "<p>Thanks for sharing, and congrats on your result!</p>",
      "votes": null,
      "replies": [
        {
          "id": 325224,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "05/08/2018 08:36:50",
          "content": "<p>Thanks CPMP. Well done on your score and position, I'm interested to hear your validation strategy - mine had +/- .0002 variance which made it difficult to assess features - I used test hours from day 9 to validate, and full day 8 to train. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325239,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/08/2018 08:52:37",
          "content": "<p>I used two validation set, hour 4 of day 9, and hours 5, 9, 10, 13, 14.  Given lgb stops on the the first that does not improve it helps remove some of the variability.  I also watched the train auc in the early days, discarding features that improved validation but also increased the gap with train a lot.  I stopped watching train auc in the last week to speed up things, but last time I checked I had a quite small gap.</p>\n\n<p>I also used LB, i.e. only kept a feature if local auc and LB improved.  Yes, I know this can lead to overfit, but given I was filtering first on local validation I think I escaped it for the most part.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325270,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "05/08/2018 09:13:41",
          "content": "<p>Nice idea to monitor the variation between train and test...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325274,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/08/2018 09:15:37",
          "content": "<p>I find it to be the best way to detect overfiting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325237,
      "author_name": "alexyung757",
      "author_url": "",
      "post_date": "05/08/2018 08:50:42",
      "content": "<p>Congrats and thx for sharing. In many competions , there is less R kernal. so  when i see data.table, i am very excited! \ni hope master can give us kernal , please!  </p>",
      "votes": null,
      "replies": [
        {
          "id": 325249,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "05/08/2018 08:59:48",
          "content": "<p>Thanks Alex, the repo is a bit messy currently but here is an example of making the split second lead time in data.table - calculating the feature runs on full data set in less than a minute - most of the time is data load.</p>\n\n<p>Repo link  -- <a href=\"https://github.com/darraghdog/talking-data-click-fraud/blob/1827a7a437dddce85642e01f561953c3f0066573/features/make/secdistR_2703.R\">secdistR_2703.R</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325253,
          "author_name": "alexyung757",
          "author_url": "",
          "post_date": "05/08/2018 09:01:13",
          "content": "<p>thanks. i am looking forward to yours</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325279,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "05/08/2018 09:17:44",
      "content": "<p>Thanks Darragh for sharing and well done! I like that you relate feature engineering to entropy. Do you have any theoretical paper on this to share?</p>",
      "votes": null,
      "replies": [
        {
          "id": 325312,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "05/08/2018 09:45:37",
          "content": "<p>Sorry, I have no paper - but I think of it as a step up from ctuniq - where ctuniq counts the unique groups in a feature level; entropy tells how concentrated it is in certain groups. For example entropy over minutes; is there generally an even spread over minutes or do we have bursts of clicks at certain times.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326008,
          "author_name": "ericbenhamou",
          "author_url": "",
          "post_date": "05/09/2018 06:51:00",
          "content": "<p>Thanks Darragh for the feedback.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325353,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "05/08/2018 10:26:31",
      "content": "<p>Congrats @Darragh and @Giba,</p>\n\n<p>And thanks for sharing the approach and repository link. This is very helpful for R users like me. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325356,
      "author_name": "meykds",
      "author_url": "",
      "post_date": "05/08/2018 10:41:16",
      "content": "<p>It is really interesting that removing device  3032 gives you such boost ;) In my case particularly because I was thinking about it and then gave it up. More lessons learned.</p>\n\n<p>Congratulations! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325387,
      "author_name": "mrbeer",
      "author_url": "",
      "post_date": "05/08/2018 11:06:04",
      "content": "<p>Thanks for sharing. amazing boost from removing the 3032 device</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325393,
      "author_name": "kailex",
      "author_url": "",
      "post_date": "05/08/2018 11:18:10",
      "content": "<p>Thanks you for sharing. It's nice to see <strong>data.table</strong>-based solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325405,
      "author_name": "profetul",
      "author_url": "",
      "post_date": "05/08/2018 11:39:50",
      "content": "<p>I find this to be a brilliant idea! Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325807,
      "author_name": "bestfitting",
      "author_url": "",
      "post_date": "05/08/2018 22:02:10",
      "content": "<p>Thanks for your sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "325205": "Congrats to all the winners, and disappointed to see so many people affected by the late sharing of scripts - so sad to see a 0.9811 sub uploaded on last day and affect people's hard work.   \n   \nOur solution was a couple of LGB's at public 0.9821 or 0.9820, and nnet at public 0.9816. We did a lot of feature engineering in `data.table` this is very nice for building features on such a dataset, especially, time based features. Entropy helped - [linky][1]  ; particularly entropy over time based features like minute. Device `3032` was all, or nearly all, no downloads and was not in test, so we got circa 0.0004 from removing. We also got split second times and incorporated that into lead times - this was done by turning the ordering per second, and count of clicks per second, into sub-second time - if there are 100 clicks in a second `16:00:00`, first click gets `16:00:00.00`, second gets `16:00:00.01`, next `16:00:00.02` ... etc. This gave us good lift, around 0.001 early in the competition, but we did not test the lift later to see the drop by leaving it out. Stacking gave us, circa .0003. Counts of previous periods, eg. period this time yesterday, helped also. \nA lot of other things tried with no, or small lift - used about 30 features in all.  \n\n\n  [1]: https://github.com/owenzhang/kaggle-avito/blob/a7a2cc853b0ca86f07cdb9dd483779b2927b99ee/avito_utils.R",
    "325208": "Well done",
    "325214": "Great job~Congrats and thx for sharing~",
    "325219": "Thanks for sharing, and congrats on your result!",
    "325224": "Thanks CPMP. Well done on your score and position, I'm interested to hear your validation strategy - mine had +/- .0002 variance which made it difficult to assess features - I used test hours from day 9 to validate, and full day 8 to train.",
    "325237": "Congrats and thx for sharing. In many competions , there is less R kernal. so  when i see data.table, i am very excited! \ni hope master can give us kernal , please!",
    "325239": "I used two validation set, hour 4 of day 9, and hours 5, 9, 10, 13, 14.  Given lgb stops on the the first that does not improve it helps remove some of the variability.  I also watched the train auc in the early days, discarding features that improved validation but also increased the gap with train a lot.  I stopped watching train auc in the last week to speed up things, but last time I checked I had a quite small gap.\n\nI also used LB, i.e. only kept a feature if local auc and LB improved.  Yes, I know this can lead to overfit, but given I was filtering first on local validation I think I escaped it for the most part.",
    "325249": "Thanks Alex, the repo is a bit messy currently but here is an example of making the split second lead time in data.table - calculating the feature runs on full data set in less than a minute - most of the time is data load.\n\nRepo link  -- [secdistR_2703.R][1]\n\n\n  [1]: https://github.com/darraghdog/talking-data-click-fraud/blob/1827a7a437dddce85642e01f561953c3f0066573/features/make/secdistR_2703.R",
    "325253": "thanks. i am looking forward to yours",
    "325270": "Nice idea to monitor the variation between train and test...",
    "325274": "I find it to be the best way to detect overfiting.",
    "325279": "Thanks Darragh for sharing and well done! I like that you relate feature engineering to entropy. Do you have any theoretical paper on this to share?",
    "325312": "Sorry, I have no paper - but I think of it as a step up from ctuniq - where ctuniq counts the unique groups in a feature level; entropy tells how concentrated it is in certain groups. For example entropy over minutes; is there generally an even spread over minutes or do we have bursts of clicks at certain times.",
    "325353": "Congrats @Darragh and @Giba,\n\nAnd thanks for sharing the approach and repository link. This is very helpful for R users like me.",
    "325356": "It is really interesting that removing device  3032 gives you such boost ;) In my case particularly because I was thinking about it and then gave it up. More lessons learned.\n\nCongratulations!",
    "325387": "Thanks for sharing. amazing boost from removing the 3032 device",
    "325393": "Thanks you for sharing. It's nice to see **data.table**-based solution.",
    "325405": "I find this to be a brilliant idea! Congrats!",
    "325807": "Thanks for your sharing!",
    "326008": "Thanks Darragh for the feedback."
  },
  "source": "meta"
}