{
  "id": 56368,
  "title": "28th place, A 0.0006x boosting trend feature and non-overfitting target encoding (attributed rates)",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/writeups/agree-with-24th-trust-cv-28th-place-a-0-0006x-boos",
  "author_name": "",
  "post_date": "2018-05-09T12:59:09.390Z",
  "votes": 26,
  "comment_count": 21,
  "views": 0,
  "content": "<p>At first, I want to say Thanks to all you nice and friendly kagglers who've shared their enlightening ideas. This is my first competition on Kaggle and I find people here are really willing to share and help. Really learnt a lot here. I'm so touched (＊￣︶￣＊)</p>\n\n<p>Although 28th is not a very good place, I do have some useful features to share which I didn't use at last. Why didn't I use them? Because I am such a bonehead to trust the deceptive public LB rather than my reliable local CV. Dear CV, I'm so sorry that I got you wrong. ( ´･ω･)ﾉ(._.`)</p>\n\n<h1>The first one: <strong>About the overfitting target encoding</strong></h1>\n\n<p>First, my special thanks to @NanoMathias and his Notebook <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">https://www.kaggle.com/nanomathias/feature-engineering-importance-testing</a> which enlightened me a lot. As many have commented, I encounter the overfitting problem when I tried the target encoding (attributed rate) features.</p>\n\n<p>Intuitively, the historical attributed rate is helpful as it depicts the quality of an app or the habit of a user and the like in a way or another. But afterwards I find more than 25% (143786 records) downloads came from the users (ip-os-device) who clicked only once and downloaded only once. And much more users are there who clicked once but did not download. I suspect the overfitting might result from these, and here are some tricks I come up with to cope with it:</p>\n\n<ol>\n<li><p>Use the data from the previous day only. This cause data on day 6 to have all NaN but this helps any way. For instance, when you calculate the attributed rate for day 8, use data on day 7.</p></li>\n<li><p>Use all data before the current day. This results in a more precise historical attributed rate and is better than 1). Putting it clearly, use the data from day 6 to calculate the attributed rate for day 7 and use data from day 6 and day 7 to calculate the attributed rate for day 8 and day 6,7,8 for day 9 and day 6,7,8,9 for day 10.</p></li>\n<li><p>Use all data excluding the current day. Intuitively, this gives the closet results to the original one which use all data to calculate the \"historical\" attributed rate and this did boost most. You use day 7,8,9 to calculate for day 6 and day 6,7,9 for day 8 and so on.</p></li>\n</ol>\n\n<p>We are approaching the \"real\" attributed rates piecemeal. By now, the third way boosted me most and a further step will take me into overfitting. I don't know if there exist a not that radical step to improve over it with no overfitting based on 3).</p>\n\n<p>I use the attributed rate on the following 6 combinations: ip-app, ip-os, ip-os-device, ip-app-os-device, app-channel, app (Not a real combination). Besides, smoothing helps a little. One way is to use the \"log(count)\" given in the Notebook abovementioned. Another way is to use Laplacian smoothing. However, I didn't find a good parameter for Laplacian smoothing to beat the \"log(count)\"-fashion. One last interesting thing, Combining the unsmoothed with the smoothed ones improves a little further (0.0001 over using only the unsmoothed ones)</p>\n\n<p>Laplacian smoothing:</p>\n\n<p>(downloads + m * average_rate) / (counts + m)</p>\n\n<p>Need to pick a good \"m\". The average_rate is the mean over the attributed rates you got for a combination like the ones I mentioned above.</p>\n\n<p>Sorry I don't know how to insert a formula.</p>\n\n<h1>The second one: <strong>Trend feature is a good thing</strong></h1>\n\n<p>This is a serendipity I got when one day I happened to glimpse my cache files used to calculate daily attributed rates (Right, the intermediate results for the target encoding features). I find some apps' daily attributed rates are increasing or decreasing continuously. So I gave it a shot. I use the downloads (Note that it's the number of downloads not clicks nor attributed rates) on the previous day to divide that of the day before the previous day. For instance, </p>\n\n<pre><code>    trend_day_10 = downloads_on_day_9 / downloads_on_day_8\n</code></pre>\n\n<p>I calculated this on only 3 combinations: ip-app-os-device, ip-os-device, app. And it boosts me from 0.9821 to nearly 0.9828 on private LB, which is rather helpful than my hotchpotch-like blending of all my weird XGBs and LGBs. I didn't use it at last because it drops 0.0001 on public LB and I am such a bonehead to trust the deceptive public LB rather than my reliable local CV. Although this results in all NaN on day 6 and day 7 (as you don't have data from the two previous days), it did help a lot. Besides, I also come up with another way to represent the trend but had no time to try: 1. get the differences of the series of downloads on day[6,7,8,9]. 2. get the sign (representing the increase or decrease) of the differences. 3. Encode it. For instance:</p>\n\n<p>Downloads on day 6,7,8,9: [1, 2, 3, 2]</p>\n\n<p>After step 1: [1, 1, -1]</p>\n\n<p>After step 2: [+, +, -]</p>\n\n<p>And then you assign different integers to different patterns. I don't know if this work.</p>\n\n<p>Other interesting features (helps a little, about 0.0001):</p>\n\n<ol>\n<li><p>ip-app-os-device counts / ip-os-device counts, which describes the preference for apps of a user.</p></li>\n<li><p>ip-app / ip.</p></li>\n</ol>\n\n<p>At last, the competition gave me a really profound lesson: TRUST YOUR LOCAL CV :(</p>\n\n<p>English is not my native language so if I didn't put it clearly and make you mixed up, feel free to comment!! :)</p>\n\n<p>Thanks to all you nice kagglers!! :)</p>",
  "messages": [
    {
      "id": "325914",
      "postDate": "05/09/2018 03:54:21",
      "content": "<p>At first, I want to say Thanks to all you nice and friendly kagglers who've shared their enlightening ideas. This is my first competition on Kaggle and I find people here are really willing to share and help. Really learnt a lot here. I'm so touched (＊￣︶￣＊)</p>\n\n<p>Although 28th is not a very good place, I do have some useful features to share which I didn't use at last. Why didn't I use them? Because I am such a bonehead to trust the deceptive public LB rather than my reliable local CV. Dear CV, I'm so sorry that I got you wrong. ( ´･ω･)ﾉ(._.`)</p>\n\n<h1>The first one: <strong>About the overfitting target encoding</strong></h1>\n\n<p>First, my special thanks to @NanoMathias and his Notebook <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">https://www.kaggle.com/nanomathias/feature-engineering-importance-testing</a> which enlightened me a lot. As many have commented, I encounter the overfitting problem when I tried the target encoding (attributed rate) features.</p>\n\n<p>Intuitively, the historical attributed rate is helpful as it depicts the quality of an app or the habit of a user and the like in a way or another. But afterwards I find more than 25% (143786 records) downloads came from the users (ip-os-device) who clicked only once and downloaded only once. And much more users are there who clicked once but did not download. I suspect the overfitting might result from these, and here are some tricks I come up with to cope with it:</p>\n\n<ol>\n<li><p>Use the data from the previous day only. This cause data on day 6 to have all NaN but this helps any way. For instance, when you calculate the attributed rate for day 8, use data on day 7.</p></li>\n<li><p>Use all data before the current day. This results in a more precise historical attributed rate and is better than 1). Putting it clearly, use the data from day 6 to calculate the attributed rate for day 7 and use data from day 6 and day 7 to calculate the attributed rate for day 8 and day 6,7,8 for day 9 and day 6,7,8,9 for day 10.</p></li>\n<li><p>Use all data excluding the current day. Intuitively, this gives the closet results to the original one which use all data to calculate the \"historical\" attributed rate and this did boost most. You use day 7,8,9 to calculate for day 6 and day 6,7,9 for day 8 and so on.</p></li>\n</ol>\n\n<p>We are approaching the \"real\" attributed rates piecemeal. By now, the third way boosted me most and a further step will take me into overfitting. I don't know if there exist a not that radical step to improve over it with no overfitting based on 3).</p>\n\n<p>I use the attributed rate on the following 6 combinations: ip-app, ip-os, ip-os-device, ip-app-os-device, app-channel, app (Not a real combination). Besides, smoothing helps a little. One way is to use the \"log(count)\" given in the Notebook abovementioned. Another way is to use Laplacian smoothing. However, I didn't find a good parameter for Laplacian smoothing to beat the \"log(count)\"-fashion. One last interesting thing, Combining the unsmoothed with the smoothed ones improves a little further (0.0001 over using only the unsmoothed ones)</p>\n\n<p>Laplacian smoothing:</p>\n\n<p>(downloads + m * average_rate) / (counts + m)</p>\n\n<p>Need to pick a good \"m\". The average_rate is the mean over the attributed rates you got for a combination like the ones I mentioned above.</p>\n\n<p>Sorry I don't know how to insert a formula.</p>\n\n<h1>The second one: <strong>Trend feature is a good thing</strong></h1>\n\n<p>This is a serendipity I got when one day I happened to glimpse my cache files used to calculate daily attributed rates (Right, the intermediate results for the target encoding features). I find some apps' daily attributed rates are increasing or decreasing continuously. So I gave it a shot. I use the downloads (Note that it's the number of downloads not clicks nor attributed rates) on the previous day to divide that of the day before the previous day. For instance, </p>\n\n<pre><code>    trend_day_10 = downloads_on_day_9 / downloads_on_day_8\n</code></pre>\n\n<p>I calculated this on only 3 combinations: ip-app-os-device, ip-os-device, app. And it boosts me from 0.9821 to nearly 0.9828 on private LB, which is rather helpful than my hotchpotch-like blending of all my weird XGBs and LGBs. I didn't use it at last because it drops 0.0001 on public LB and I am such a bonehead to trust the deceptive public LB rather than my reliable local CV. Although this results in all NaN on day 6 and day 7 (as you don't have data from the two previous days), it did help a lot. Besides, I also come up with another way to represent the trend but had no time to try: 1. get the differences of the series of downloads on day[6,7,8,9]. 2. get the sign (representing the increase or decrease) of the differences. 3. Encode it. For instance:</p>\n\n<p>Downloads on day 6,7,8,9: [1, 2, 3, 2]</p>\n\n<p>After step 1: [1, 1, -1]</p>\n\n<p>After step 2: [+, +, -]</p>\n\n<p>And then you assign different integers to different patterns. I don't know if this work.</p>\n\n<p>Other interesting features (helps a little, about 0.0001):</p>\n\n<ol>\n<li><p>ip-app-os-device counts / ip-os-device counts, which describes the preference for apps of a user.</p></li>\n<li><p>ip-app / ip.</p></li>\n</ol>\n\n<p>At last, the competition gave me a really profound lesson: TRUST YOUR LOCAL CV :(</p>\n\n<p>English is not my native language so if I didn't put it clearly and make you mixed up, feel free to comment!! :)</p>\n\n<p>Thanks to all you nice kagglers!! :)</p>",
      "rawMarkdown": "At first, I want to say Thanks to all you nice and friendly kagglers who've shared their enlightening ideas. This is my first competition on Kaggle and I find people here are really willing to share and help. Really learnt a lot here. I'm so touched (＊￣︶￣＊)\n\nAlthough 28th is not a very good place, I do have some useful features to share which I didn't use at last. Why didn't I use them? Because I am such a bonehead to trust the deceptive public LB rather than my reliable local CV. Dear CV, I'm so sorry that I got you wrong. ( ´･ω･)ﾉ(._.`)\n\n# The first one: **About the overfitting target encoding**\n\nFirst, my special thanks to @NanoMathias and his Notebook https://www.kaggle.com/nanomathias/feature-engineering-importance-testing which enlightened me a lot. As many have commented, I encounter the overfitting problem when I tried the target encoding (attributed rate) features.\n \nIntuitively, the historical attributed rate is helpful as it depicts the quality of an app or the habit of a user and the like in a way or another. But afterwards I find more than 25% (143786 records) downloads came from the users (ip-os-device) who clicked only once and downloaded only once. And much more users are there who clicked once but did not download. I suspect the overfitting might result from these, and here are some tricks I come up with to cope with it:\n\n1. Use the data from the previous day only. This cause data on day 6 to have all NaN but this helps any way. For instance, when you calculate the attributed rate for day 8, use data on day 7.\n\n2. Use all data before the current day. This results in a more precise historical attributed rate and is better than 1). Putting it clearly, use the data from day 6 to calculate the attributed rate for day 7 and use data from day 6 and day 7 to calculate the attributed rate for day 8 and day 6,7,8 for day 9 and day 6,7,8,9 for day 10.\n\n3. Use all data excluding the current day. Intuitively, this gives the closet results to the original one which use all data to calculate the \"historical\" attributed rate and this did boost most. You use day 7,8,9 to calculate for day 6 and day 6,7,9 for day 8 and so on.\n\nWe are approaching the \"real\" attributed rates piecemeal. By now, the third way boosted me most and a further step will take me into overfitting. I don't know if there exist a not that radical step to improve over it with no overfitting based on 3).\n\nI use the attributed rate on the following 6 combinations: ip-app, ip-os, ip-os-device, ip-app-os-device, app-channel, app (Not a real combination). Besides, smoothing helps a little. One way is to use the \"log(count)\" given in the Notebook abovementioned. Another way is to use Laplacian smoothing. However, I didn't find a good parameter for Laplacian smoothing to beat the \"log(count)\"-fashion. One last interesting thing, Combining the unsmoothed with the smoothed ones improves a little further (0.0001 over using only the unsmoothed ones)\n\nLaplacian smoothing:\n\n(downloads + m * average_rate) / (counts + m)\n\nNeed to pick a good \"m\". The average_rate is the mean over the attributed rates you got for a combination like the ones I mentioned above.\n\nSorry I don't know how to insert a formula.\n\n# The second one: **Trend feature is a good thing**\n\nThis is a serendipity I got when one day I happened to glimpse my cache files used to calculate daily attributed rates (Right, the intermediate results for the target encoding features). I find some apps' daily attributed rates are increasing or decreasing continuously. So I gave it a shot. I use the downloads (Note that it's the number of downloads not clicks nor attributed rates) on the previous day to divide that of the day before the previous day. For instance, \n\n        trend_day_10 = downloads_on_day_9 / downloads_on_day_8\n\nI calculated this on only 3 combinations: ip-app-os-device, ip-os-device, app. And it boosts me from 0.9821 to nearly 0.9828 on private LB, which is rather helpful than my hotchpotch-like blending of all my weird XGBs and LGBs. I didn't use it at last because it drops 0.0001 on public LB and I am such a bonehead to trust the deceptive public LB rather than my reliable local CV. Although this results in all NaN on day 6 and day 7 (as you don't have data from the two previous days), it did help a lot. Besides, I also come up with another way to represent the trend but had no time to try: 1. get the differences of the series of downloads on day[6,7,8,9]. 2. get the sign (representing the increase or decrease) of the differences. 3. Encode it. For instance:\n\nDownloads on day 6,7,8,9: [1, 2, 3, 2]\n\nAfter step 1: [1, 1, -1]\n\nAfter step 2: [+, +, -]\n\nAnd then you assign different integers to different patterns. I don't know if this work.\n\nOther interesting features (helps a little, about 0.0001):\n\n1. ip-app-os-device counts / ip-os-device counts, which describes the preference for apps of a user.\n\n2. ip-app / ip.\n\nAt last, the competition gave me a really profound lesson: TRUST YOUR LOCAL CV :(\n\nEnglish is not my native language so if I didn't put it clearly and make you mixed up, feel free to comment!! :)\n\nThanks to all you nice kagglers!! :)",
      "votes": null
    },
    {
      "id": "325921",
      "postDate": "05/09/2018 04:28:36",
      "content": "<p>The trend stuff is really interesting and is easy to be ignored. As to trend features of the data in day 6, you just leave them with NaN?</p>",
      "rawMarkdown": "The trend stuff is really interesting and is easy to be ignored. As to trend features of the data in day 6, you just leave them with NaN?",
      "votes": null
    },
    {
      "id": "325924",
      "postDate": "05/09/2018 04:32:09",
      "content": "<p>Yes! NaN for day 6 and day 7. Filling with 0 or something else may help further. At least it helps for the attributed rate features.</p>",
      "rawMarkdown": "Yes! NaN for day 6 and day 7. Filling with 0 or something else may help further. At least it helps for the attributed rate features.",
      "votes": null
    },
    {
      "id": "325926",
      "postDate": "05/09/2018 04:37:23",
      "content": "<p>Wow, it's a little bit weird for me but it helps! Thanks for reply!</p>",
      "rawMarkdown": "Wow, it's a little bit weird for me but it helps! Thanks for reply!",
      "votes": null
    },
    {
      "id": "325927",
      "postDate": "05/09/2018 04:39:42",
      "content": "<p>It's also weird for me. Thanks also for your asking!</p>",
      "rawMarkdown": "It's also weird for me. Thanks also for your asking!",
      "votes": null
    },
    {
      "id": "325942",
      "postDate": "05/09/2018 05:11:06",
      "content": "<p>Nice work! Did you use log(count) to compute the attribute rate?</p>",
      "rawMarkdown": "Nice work! Did you use log(count) to compute the attribute rate?",
      "votes": null
    },
    {
      "id": "325944",
      "postDate": "05/09/2018 05:16:06",
      "content": "<p>Yes. At last I fed both the log-fashion-smoothed and unsmoothed rates to my model. Smoothing helps more when I use less data (data only on the previous days), but helps less when I use all the days excluding the current day.</p>",
      "rawMarkdown": "Yes. At last I fed both the log-fashion-smoothed and unsmoothed rates to my model. Smoothing helps more when I use less data (data only on the previous days), but helps less when I use all the days excluding the current day.",
      "votes": null
    },
    {
      "id": "325950",
      "postDate": "05/09/2018 05:30:00",
      "content": "<p>I, too, used previous day to generate current day target encoding and got 0.003 improvement which was nice.</p>",
      "rawMarkdown": "I, too, used previous day to generate current day target encoding and got 0.003 improvement which was nice.",
      "votes": null
    },
    {
      "id": "325952",
      "postDate": "05/09/2018 05:31:34",
      "content": "<p>Wow 0.003! That's a big improvement.</p>",
      "rawMarkdown": "Wow 0.003! That's a big improvement.",
      "votes": null
    },
    {
      "id": "325966",
      "postDate": "05/09/2018 06:07:00",
      "content": "<p>Thanks for your reply. but do you use log(count + 1) or log(count)? It seems that if the click count is 1, after log transforming, if will be 0.</p>",
      "rawMarkdown": "Thanks for your reply. but do you use log(count + 1) or log(count)? It seems that if the click count is 1, after log transforming, if will be 0.",
      "votes": null
    },
    {
      "id": "325970",
      "postDate": "05/09/2018 06:16:00",
      "content": "<p>You're right. I did miss this. Actually when I was extracting the features I got warnings about this, but I thought they are not so important and my unsmoothed ones can compensate for this so I just left it there. In a nutshell, I'm too lazy to change my code XD. I thought your +1 is more reasonable. Thanks for your pointing out.</p>",
      "rawMarkdown": "You're right. I did miss this. Actually when I was extracting the features I got warnings about this, but I thought they are not so important and my unsmoothed ones can compensate for this so I just left it there. In a nutshell, I'm too lazy to change my code XD. I thought your +1 is more reasonable. Thanks for your pointing out.",
      "votes": null
    },
    {
      "id": "325993",
      "postDate": "05/09/2018 06:43:26",
      "content": "<p>Thanks for sharing and congrats.  28th place is an extremely good place for a first competition, you should be proud of  it.</p>\n\n<p>Nice catch about trending! </p>",
      "rawMarkdown": "Thanks for sharing and congrats.  28th place is an extremely good place for a first competition, you should be proud of  it.\n\nNice catch about trending!",
      "votes": null
    },
    {
      "id": "326001",
      "postDate": "05/09/2018 06:46:13",
      "content": "<p>Thanks a lot for your encouraging! And I also learnt much from you. You're really awesome.</p>",
      "rawMarkdown": "Thanks a lot for your encouraging! And I also learnt much from you. You're really awesome.",
      "votes": null
    },
    {
      "id": "326003",
      "postDate": "05/09/2018 06:49:49",
      "content": "<p>Thanks for your reply.</p>",
      "rawMarkdown": "Thanks for your reply.",
      "votes": null
    },
    {
      "id": "326025",
      "postDate": "05/09/2018 07:04:32",
      "content": "<p>Thank you for sharing and well done!</p>",
      "rawMarkdown": "Thank you for sharing and well done!",
      "votes": null
    },
    {
      "id": "326029",
      "postDate": "05/09/2018 07:06:39",
      "content": "<p>Thanks! Your work is really good!</p>",
      "rawMarkdown": "Thanks! Your work is really good!",
      "votes": null
    },
    {
      "id": "326032",
      "postDate": "05/09/2018 07:11:20",
      "content": "<p>Wow! Feel excited to be 翻牌 by Da Lao.</p>",
      "rawMarkdown": "Wow! Feel excited to be 翻牌 by Da Lao.",
      "votes": null
    },
    {
      "id": "326206",
      "postDate": "05/09/2018 12:50:33",
      "content": "<p>都是🇨🇳人</p>",
      "rawMarkdown": "都是🇨🇳人",
      "votes": null
    },
    {
      "id": "326218",
      "postDate": "05/09/2018 13:04:40",
      "content": "<p>Trend feature is almost NULL in day 8 ，because day 6 has too little data，so this mean  this kinds of feature only have meaning in day 9 and day 10??</p>",
      "rawMarkdown": "Trend feature is almost NULL in day 8 ，because day 6 has too little data，so this mean  this kinds of feature only have meaning in day 9 and day 10??",
      "votes": null
    },
    {
      "id": "326221",
      "postDate": "05/09/2018 13:10:26",
      "content": "<p>For the trend features，I get one method:we can calculate some value(such as click count or download rate)of each hour of yesterday,and then we can get the delta of every adjacent two hours，and last we can encode this。I will try this</p>",
      "rawMarkdown": "For the trend features，I get one method:we can calculate some value(such as click count or download rate)of each hour of yesterday,and then we can get the delta of every adjacent two hours，and last we can encode this。I will try this",
      "votes": null
    },
    {
      "id": "326227",
      "postDate": "05/09/2018 13:15:37",
      "content": "<p>You are likely to be right. I didn't inspect the distribution of this feature carefully, rather I just extracted and fed them into my model, and hopefully it worked.</p>",
      "rawMarkdown": "You are likely to be right. I didn't inspect the distribution of this feature carefully, rather I just extracted and fed them into my model, and hopefully it worked.",
      "votes": null
    },
    {
      "id": "326234",
      "postDate": "05/09/2018 13:21:28",
      "content": "<p>I'm afraid the data within one hour is not sufficient to cover enough users and will cause sparsity, but it will work I presume. Looking forward to your discovery :)</p>",
      "rawMarkdown": "I'm afraid the data within one hour is not sufficient to cover enough users and will cause sparsity, but it will work I presume. Looking forward to your discovery :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 325921,
      "author_name": "fangao",
      "author_url": "",
      "post_date": "05/09/2018 04:28:36",
      "content": "<p>The trend stuff is really interesting and is easy to be ignored. As to trend features of the data in day 6, you just leave them with NaN?</p>",
      "votes": null,
      "replies": [
        {
          "id": 325924,
          "author_name": "rayarrow",
          "author_url": "",
          "post_date": "05/09/2018 04:32:09",
          "content": "<p>Yes! NaN for day 6 and day 7. Filling with 0 or something else may help further. At least it helps for the attributed rate features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325926,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/09/2018 04:37:23",
          "content": "<p>Wow, it's a little bit weird for me but it helps! Thanks for reply!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325927,
          "author_name": "rayarrow",
          "author_url": "",
          "post_date": "05/09/2018 04:39:42",
          "content": "<p>It's also weird for me. Thanks also for your asking!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325942,
      "author_name": "wuzuping",
      "author_url": "",
      "post_date": "05/09/2018 05:11:06",
      "content": "<p>Nice work! Did you use log(count) to compute the attribute rate?</p>",
      "votes": null,
      "replies": [
        {
          "id": 325944,
          "author_name": "rayarrow",
          "author_url": "",
          "post_date": "05/09/2018 05:16:06",
          "content": "<p>Yes. At last I fed both the log-fashion-smoothed and unsmoothed rates to my model. Smoothing helps more when I use less data (data only on the previous days), but helps less when I use all the days excluding the current day.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325966,
          "author_name": "wuzuping",
          "author_url": "",
          "post_date": "05/09/2018 06:07:00",
          "content": "<p>Thanks for your reply. but do you use log(count + 1) or log(count)? It seems that if the click count is 1, after log transforming, if will be 0.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325970,
          "author_name": "rayarrow",
          "author_url": "",
          "post_date": "05/09/2018 06:16:00",
          "content": "<p>You're right. I did miss this. Actually when I was extracting the features I got warnings about this, but I thought they are not so important and my unsmoothed ones can compensate for this so I just left it there. In a nutshell, I'm too lazy to change my code XD. I thought your +1 is more reasonable. Thanks for your pointing out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326003,
          "author_name": "wuzuping",
          "author_url": "",
          "post_date": "05/09/2018 06:49:49",
          "content": "<p>Thanks for your reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325950,
      "author_name": "konohayui",
      "author_url": "",
      "post_date": "05/09/2018 05:30:00",
      "content": "<p>I, too, used previous day to generate current day target encoding and got 0.003 improvement which was nice.</p>",
      "votes": null,
      "replies": [
        {
          "id": 325952,
          "author_name": "rayarrow",
          "author_url": "",
          "post_date": "05/09/2018 05:31:34",
          "content": "<p>Wow 0.003! That's a big improvement.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325993,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/09/2018 06:43:26",
      "content": "<p>Thanks for sharing and congrats.  28th place is an extremely good place for a first competition, you should be proud of  it.</p>\n\n<p>Nice catch about trending! </p>",
      "votes": null,
      "replies": [
        {
          "id": 326001,
          "author_name": "rayarrow",
          "author_url": "",
          "post_date": "05/09/2018 06:46:13",
          "content": "<p>Thanks a lot for your encouraging! And I also learnt much from you. You're really awesome.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326025,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "05/09/2018 07:04:32",
      "content": "<p>Thank you for sharing and well done!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326029,
      "author_name": "panfeiyang",
      "author_url": "",
      "post_date": "05/09/2018 07:06:39",
      "content": "<p>Thanks! Your work is really good!</p>",
      "votes": null,
      "replies": [
        {
          "id": 326032,
          "author_name": "rayarrow",
          "author_url": "",
          "post_date": "05/09/2018 07:11:20",
          "content": "<p>Wow! Feel excited to be 翻牌 by Da Lao.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326206,
          "author_name": "huayupeng",
          "author_url": "",
          "post_date": "05/09/2018 12:50:33",
          "content": "<p>都是🇨🇳人</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326218,
      "author_name": "huayupeng",
      "author_url": "",
      "post_date": "05/09/2018 13:04:40",
      "content": "<p>Trend feature is almost NULL in day 8 ，because day 6 has too little data，so this mean  this kinds of feature only have meaning in day 9 and day 10??</p>",
      "votes": null,
      "replies": [
        {
          "id": 326227,
          "author_name": "rayarrow",
          "author_url": "",
          "post_date": "05/09/2018 13:15:37",
          "content": "<p>You are likely to be right. I didn't inspect the distribution of this feature carefully, rather I just extracted and fed them into my model, and hopefully it worked.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326221,
      "author_name": "huayupeng",
      "author_url": "",
      "post_date": "05/09/2018 13:10:26",
      "content": "<p>For the trend features，I get one method:we can calculate some value(such as click count or download rate)of each hour of yesterday,and then we can get the delta of every adjacent two hours，and last we can encode this。I will try this</p>",
      "votes": null,
      "replies": [
        {
          "id": 326234,
          "author_name": "rayarrow",
          "author_url": "",
          "post_date": "05/09/2018 13:21:28",
          "content": "<p>I'm afraid the data within one hour is not sufficient to cover enough users and will cause sparsity, but it will work I presume. Looking forward to your discovery :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "325914": "At first, I want to say Thanks to all you nice and friendly kagglers who've shared their enlightening ideas. This is my first competition on Kaggle and I find people here are really willing to share and help. Really learnt a lot here. I'm so touched (＊￣︶￣＊)\n\nAlthough 28th is not a very good place, I do have some useful features to share which I didn't use at last. Why didn't I use them? Because I am such a bonehead to trust the deceptive public LB rather than my reliable local CV. Dear CV, I'm so sorry that I got you wrong. ( ´･ω･)ﾉ(._.`)\n\n# The first one: **About the overfitting target encoding**\n\nFirst, my special thanks to @NanoMathias and his Notebook https://www.kaggle.com/nanomathias/feature-engineering-importance-testing which enlightened me a lot. As many have commented, I encounter the overfitting problem when I tried the target encoding (attributed rate) features.\n \nIntuitively, the historical attributed rate is helpful as it depicts the quality of an app or the habit of a user and the like in a way or another. But afterwards I find more than 25% (143786 records) downloads came from the users (ip-os-device) who clicked only once and downloaded only once. And much more users are there who clicked once but did not download. I suspect the overfitting might result from these, and here are some tricks I come up with to cope with it:\n\n1. Use the data from the previous day only. This cause data on day 6 to have all NaN but this helps any way. For instance, when you calculate the attributed rate for day 8, use data on day 7.\n\n2. Use all data before the current day. This results in a more precise historical attributed rate and is better than 1). Putting it clearly, use the data from day 6 to calculate the attributed rate for day 7 and use data from day 6 and day 7 to calculate the attributed rate for day 8 and day 6,7,8 for day 9 and day 6,7,8,9 for day 10.\n\n3. Use all data excluding the current day. Intuitively, this gives the closet results to the original one which use all data to calculate the \"historical\" attributed rate and this did boost most. You use day 7,8,9 to calculate for day 6 and day 6,7,9 for day 8 and so on.\n\nWe are approaching the \"real\" attributed rates piecemeal. By now, the third way boosted me most and a further step will take me into overfitting. I don't know if there exist a not that radical step to improve over it with no overfitting based on 3).\n\nI use the attributed rate on the following 6 combinations: ip-app, ip-os, ip-os-device, ip-app-os-device, app-channel, app (Not a real combination). Besides, smoothing helps a little. One way is to use the \"log(count)\" given in the Notebook abovementioned. Another way is to use Laplacian smoothing. However, I didn't find a good parameter for Laplacian smoothing to beat the \"log(count)\"-fashion. One last interesting thing, Combining the unsmoothed with the smoothed ones improves a little further (0.0001 over using only the unsmoothed ones)\n\nLaplacian smoothing:\n\n(downloads + m * average_rate) / (counts + m)\n\nNeed to pick a good \"m\". The average_rate is the mean over the attributed rates you got for a combination like the ones I mentioned above.\n\nSorry I don't know how to insert a formula.\n\n# The second one: **Trend feature is a good thing**\n\nThis is a serendipity I got when one day I happened to glimpse my cache files used to calculate daily attributed rates (Right, the intermediate results for the target encoding features). I find some apps' daily attributed rates are increasing or decreasing continuously. So I gave it a shot. I use the downloads (Note that it's the number of downloads not clicks nor attributed rates) on the previous day to divide that of the day before the previous day. For instance, \n\n        trend_day_10 = downloads_on_day_9 / downloads_on_day_8\n\nI calculated this on only 3 combinations: ip-app-os-device, ip-os-device, app. And it boosts me from 0.9821 to nearly 0.9828 on private LB, which is rather helpful than my hotchpotch-like blending of all my weird XGBs and LGBs. I didn't use it at last because it drops 0.0001 on public LB and I am such a bonehead to trust the deceptive public LB rather than my reliable local CV. Although this results in all NaN on day 6 and day 7 (as you don't have data from the two previous days), it did help a lot. Besides, I also come up with another way to represent the trend but had no time to try: 1. get the differences of the series of downloads on day[6,7,8,9]. 2. get the sign (representing the increase or decrease) of the differences. 3. Encode it. For instance:\n\nDownloads on day 6,7,8,9: [1, 2, 3, 2]\n\nAfter step 1: [1, 1, -1]\n\nAfter step 2: [+, +, -]\n\nAnd then you assign different integers to different patterns. I don't know if this work.\n\nOther interesting features (helps a little, about 0.0001):\n\n1. ip-app-os-device counts / ip-os-device counts, which describes the preference for apps of a user.\n\n2. ip-app / ip.\n\nAt last, the competition gave me a really profound lesson: TRUST YOUR LOCAL CV :(\n\nEnglish is not my native language so if I didn't put it clearly and make you mixed up, feel free to comment!! :)\n\nThanks to all you nice kagglers!! :)",
    "325921": "The trend stuff is really interesting and is easy to be ignored. As to trend features of the data in day 6, you just leave them with NaN?",
    "325924": "Yes! NaN for day 6 and day 7. Filling with 0 or something else may help further. At least it helps for the attributed rate features.",
    "325926": "Wow, it's a little bit weird for me but it helps! Thanks for reply!",
    "325927": "It's also weird for me. Thanks also for your asking!",
    "325942": "Nice work! Did you use log(count) to compute the attribute rate?",
    "325944": "Yes. At last I fed both the log-fashion-smoothed and unsmoothed rates to my model. Smoothing helps more when I use less data (data only on the previous days), but helps less when I use all the days excluding the current day.",
    "325950": "I, too, used previous day to generate current day target encoding and got 0.003 improvement which was nice.",
    "325952": "Wow 0.003! That's a big improvement.",
    "325966": "Thanks for your reply. but do you use log(count + 1) or log(count)? It seems that if the click count is 1, after log transforming, if will be 0.",
    "325970": "You're right. I did miss this. Actually when I was extracting the features I got warnings about this, but I thought they are not so important and my unsmoothed ones can compensate for this so I just left it there. In a nutshell, I'm too lazy to change my code XD. I thought your +1 is more reasonable. Thanks for your pointing out.",
    "325993": "Thanks for sharing and congrats.  28th place is an extremely good place for a first competition, you should be proud of  it.\n\nNice catch about trending!",
    "326001": "Thanks a lot for your encouraging! And I also learnt much from you. You're really awesome.",
    "326003": "Thanks for your reply.",
    "326025": "Thank you for sharing and well done!",
    "326029": "Thanks! Your work is really good!",
    "326032": "Wow! Feel excited to be 翻牌 by Da Lao.",
    "326206": "都是🇨🇳人",
    "326218": "Trend feature is almost NULL in day 8 ，because day 6 has too little data，so this mean  this kinds of feature only have meaning in day 9 and day 10??",
    "326221": "For the trend features，I get one method:we can calculate some value(such as click count or download rate)of each hour of yesterday,and then we can get the delta of every adjacent two hours，and last we can encode this。I will try this",
    "326227": "You are likely to be right. I didn't inspect the distribution of this feature carefully, rather I just extracted and fed them into my model, and hopefully it worked.",
    "326234": "I'm afraid the data within one hour is not sufficient to cover enough users and will cause sparsity, but it will work I presume. Looking forward to your discovery :)"
  },
  "source": "meta"
}