{
  "id": 56304,
  "title": "34th place, thoughts on CV and Target Encoding in R",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/writeups/tenuki-34th-place-thoughts-on-cv-and-target-encodi",
  "author_name": "",
  "post_date": "2018-05-08T19:21:36.860Z",
  "votes": 23,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi!\nI wanted to share some of my findings during this competition. \nIt's my first ever competition or ML task, so much of what I tried was really bad, but it taught me good intuition making these horrible mistakes and have a better \"radar\" for possible leaks and over fitting.  </p>\n\n<p>I learned a lot on R's data.table, LGBM, XGBoost, I experimented with keras and H2O NNs (didn't get good results) and also experimented with AWS's VMs.</p>\n\n<p>My public lb is 65 and private is 34, i think this increase came from good generalized CV and features, after trying many things and participating in many discussions on the matter.</p>\n\n<p>I ended up using 2 LGBM models (second one uses the predictions of the first, details in item #3),\nwith main features including: Count clicks by groups, UniqueN by groups, Next click by groups, and the target encoding based features detailed in item #3)</p>\n\n<p>I won't be long on all the mistakes i did, i will focus on the things I did that gave the significant boosts along the way and seemed to have generalized well.\nI hope this will be of any value to any of you.</p>\n\n<ol>\n<li><p>Dropping any concerns for over fitting my model to day 9 hour 4, although the difference of CV and LB score was lowest, it obviously overfit.\nAfter some discussions on the matter, I ended up training on days 7 and 8, CV with day 9, and then retraining on same rounds on entire data. \n**I read @CPMP's post and learned that he used 1.2 multiplier on the Num_Rounds with the full training set, will definitely try it in the future!!</p></li>\n<li><p>I calculated all of my feature calculations on train + test supplement data.\nThis is after trying to run features only where hour exists in the test set, or only on the train\\ regular test set.</p></li>\n<li><p>My last boost up from 0.9806 to 0.9816 was using an interesting approach to target encoding i encountered somewhere in the discussions (I don't remember who suggested the approach so sorry for not giving the credit).\nThe method involved taking my best lgb model and predict on all train and test supplement data.\nThen using the predicitions to do target encoding on the entire data and using it to create features involving the \"is_attributed\" field.\nAfter some experimentation, i ended up with this:\nI took the prediction and converted it 3 times to a boolean 1/0 \"is_attributed\" value.</p>\n\n<ul><li>Once if pred &gt;= 0.99 then 1 otherwise 0</li>\n<li>then pred &gt;= 0.997 then 1 otherwise 0</li>\n<li>Then pred &gt;= 0.999 then 1 otherwise 0</li></ul></li>\n</ol>\n\n<p>I did this because i wanted better sensitivity for the model, i am ok missing some 1's but don't want many 0's to be mistakenly predicted as 1's.\nI then calculated 2 sets of features 3 times, once for each prediction result set.\nI had mean target features by groups (past is_attributed = 1 divided by count of clicks), and another set of features that included, amongst other features:</p>\n\n<ul>\n<li><p>how many times did user (ip, os, device) ever download any app</p></li>\n<li><p>how many times did he download the current app</p></li>\n<li><p>how many unique apps did he click on.</p></li>\n</ul>\n\n<p>After running some test runs of the new features, i found the highest increase in score with mean target when using the &gt;=0.997 prediction, and the the other set of features maxing with the &gt;=0.99 prediction.\nI then used both of the sets in a new lgb model that uses only the Quantity (non categorical) features from the first lgb model  + the new target encoded features.</p>\n\n<p>4, I used bayesian optimization to optimize hyper parameters for the 2 models separately.\nUnfortunately, even after testing for different methods of optimization (GP Upper bound, expected probability etc.) with different Kappa\\Epsilon results. I did not find it running differently or optimizing better in any way.\nIt was basically running like random search.\nI still need to learn more on it, but taking the best found hyper parameters did give me a decent boost in score.</p>\n\n<p>Thanks for reading and thanks for being such an awesome community!!\nAnd of course to all the awesome Kagglers that assisted and shared their wisdom during the competition.\nThis was quite an intense 2 months =)</p>",
  "messages": [
    {
      "id": "325392",
      "postDate": "05/08/2018 11:18:08",
      "content": "<p>Hi!\nI wanted to share some of my findings during this competition. \nIt's my first ever competition or ML task, so much of what I tried was really bad, but it taught me good intuition making these horrible mistakes and have a better \"radar\" for possible leaks and over fitting.  </p>\n\n<p>I learned a lot on R's data.table, LGBM, XGBoost, I experimented with keras and H2O NNs (didn't get good results) and also experimented with AWS's VMs.</p>\n\n<p>My public lb is 65 and private is 34, i think this increase came from good generalized CV and features, after trying many things and participating in many discussions on the matter.</p>\n\n<p>I ended up using 2 LGBM models (second one uses the predictions of the first, details in item #3),\nwith main features including: Count clicks by groups, UniqueN by groups, Next click by groups, and the target encoding based features detailed in item #3)</p>\n\n<p>I won't be long on all the mistakes i did, i will focus on the things I did that gave the significant boosts along the way and seemed to have generalized well.\nI hope this will be of any value to any of you.</p>\n\n<ol>\n<li><p>Dropping any concerns for over fitting my model to day 9 hour 4, although the difference of CV and LB score was lowest, it obviously overfit.\nAfter some discussions on the matter, I ended up training on days 7 and 8, CV with day 9, and then retraining on same rounds on entire data. \n**I read @CPMP's post and learned that he used 1.2 multiplier on the Num_Rounds with the full training set, will definitely try it in the future!!</p></li>\n<li><p>I calculated all of my feature calculations on train + test supplement data.\nThis is after trying to run features only where hour exists in the test set, or only on the train\\ regular test set.</p></li>\n<li><p>My last boost up from 0.9806 to 0.9816 was using an interesting approach to target encoding i encountered somewhere in the discussions (I don't remember who suggested the approach so sorry for not giving the credit).\nThe method involved taking my best lgb model and predict on all train and test supplement data.\nThen using the predicitions to do target encoding on the entire data and using it to create features involving the \"is_attributed\" field.\nAfter some experimentation, i ended up with this:\nI took the prediction and converted it 3 times to a boolean 1/0 \"is_attributed\" value.</p>\n\n<ul><li>Once if pred &gt;= 0.99 then 1 otherwise 0</li>\n<li>then pred &gt;= 0.997 then 1 otherwise 0</li>\n<li>Then pred &gt;= 0.999 then 1 otherwise 0</li></ul></li>\n</ol>\n\n<p>I did this because i wanted better sensitivity for the model, i am ok missing some 1's but don't want many 0's to be mistakenly predicted as 1's.\nI then calculated 2 sets of features 3 times, once for each prediction result set.\nI had mean target features by groups (past is_attributed = 1 divided by count of clicks), and another set of features that included, amongst other features:</p>\n\n<ul>\n<li><p>how many times did user (ip, os, device) ever download any app</p></li>\n<li><p>how many times did he download the current app</p></li>\n<li><p>how many unique apps did he click on.</p></li>\n</ul>\n\n<p>After running some test runs of the new features, i found the highest increase in score with mean target when using the &gt;=0.997 prediction, and the the other set of features maxing with the &gt;=0.99 prediction.\nI then used both of the sets in a new lgb model that uses only the Quantity (non categorical) features from the first lgb model  + the new target encoded features.</p>\n\n<p>4, I used bayesian optimization to optimize hyper parameters for the 2 models separately.\nUnfortunately, even after testing for different methods of optimization (GP Upper bound, expected probability etc.) with different Kappa\\Epsilon results. I did not find it running differently or optimizing better in any way.\nIt was basically running like random search.\nI still need to learn more on it, but taking the best found hyper parameters did give me a decent boost in score.</p>\n\n<p>Thanks for reading and thanks for being such an awesome community!!\nAnd of course to all the awesome Kagglers that assisted and shared their wisdom during the competition.\nThis was quite an intense 2 months =)</p>",
      "rawMarkdown": "Hi!\nI wanted to share some of my findings during this competition. \nIt's my first ever competition or ML task, so much of what I tried was really bad, but it taught me good intuition making these horrible mistakes and have a better \"radar\" for possible leaks and over fitting.  \n\nI learned a lot on R's data.table, LGBM, XGBoost, I experimented with keras and H2O NNs (didn't get good results) and also experimented with AWS's VMs.\n\nMy public lb is 65 and private is 34, i think this increase came from good generalized CV and features, after trying many things and participating in many discussions on the matter.\n\nI ended up using 2 LGBM models (second one uses the predictions of the first, details in item #3),\nwith main features including: Count clicks by groups, UniqueN by groups, Next click by groups, and the target encoding based features detailed in item #3)\n\nI won't be long on all the mistakes i did, i will focus on the things I did that gave the significant boosts along the way and seemed to have generalized well.\nI hope this will be of any value to any of you.\n\n1. Dropping any concerns for over fitting my model to day 9 hour 4, although the difference of CV and LB score was lowest, it obviously overfit.\nAfter some discussions on the matter, I ended up training on days 7 and 8, CV with day 9, and then retraining on same rounds on entire data. \n**I read @CPMP's post and learned that he used 1.2 multiplier on the Num_Rounds with the full training set, will definitely try it in the future!!\n\n2. I calculated all of my feature calculations on train + test supplement data.\nThis is after trying to run features only where hour exists in the test set, or only on the train\\ regular test set.\n\n3. My last boost up from 0.9806 to 0.9816 was using an interesting approach to target encoding i encountered somewhere in the discussions (I don't remember who suggested the approach so sorry for not giving the credit).\nThe method involved taking my best lgb model and predict on all train and test supplement data.\nThen using the predicitions to do target encoding on the entire data and using it to create features involving the \"is_attributed\" field.\nAfter some experimentation, i ended up with this:\nI took the prediction and converted it 3 times to a boolean 1/0 \"is_attributed\" value.\n- Once if pred &gt;= 0.99 then 1 otherwise 0\n- then pred &gt;= 0.997 then 1 otherwise 0\n- Then pred &gt;= 0.999 then 1 otherwise 0\n\nI did this because i wanted better sensitivity for the model, i am ok missing some 1's but don't want many 0's to be mistakenly predicted as 1's.\nI then calculated 2 sets of features 3 times, once for each prediction result set.\nI had mean target features by groups (past is_attributed = 1 divided by count of clicks), and another set of features that included, amongst other features:\n\n- how many times did user (ip, os, device) ever download any app\n\n- how many times did he download the current app\n\n- how many unique apps did he click on.\n\nAfter running some test runs of the new features, i found the highest increase in score with mean target when using the &gt;=0.997 prediction, and the the other set of features maxing with the &gt;=0.99 prediction.\nI then used both of the sets in a new lgb model that uses only the Quantity (non categorical) features from the first lgb model  + the new target encoded features.\n\n\n4, I used bayesian optimization to optimize hyper parameters for the 2 models separately.\nUnfortunately, even after testing for different methods of optimization (GP Upper bound, expected probability etc.) with different Kappa\\Epsilon results. I did not find it running differently or optimizing better in any way.\nIt was basically running like random search.\nI still need to learn more on it, but taking the best found hyper parameters did give me a decent boost in score.\n\nThanks for reading and thanks for being such an awesome community!!\nAnd of course to all the awesome Kagglers that assisted and shared their wisdom during the competition.\nThis was quite an intense 2 months =)",
      "votes": null
    },
    {
      "id": "325429",
      "postDate": "05/08/2018 12:02:56",
      "content": "<p>Thanks for sharing, congrats on the end result!  Your item 3 is very interesting.</p>",
      "rawMarkdown": "Thanks for sharing, congrats on the end result!  Your item 3 is very interesting.",
      "votes": null
    },
    {
      "id": "325495",
      "postDate": "05/08/2018 13:22:26",
      "content": "<p>Thanks @CPMP.\nYou were extremely helpful during the competition, any where i go, any kernel or discussion, you are always there with your insight and experience.\nThanks! </p>\n\n<p>Do you plan on taking 3 days break to recover your sleep and start the new Avito? =D</p>",
      "rawMarkdown": "Thanks @CPMP.\nYou were extremely helpful during the competition, any where i go, any kernel or discussion, you are always there with your insight and experience.\nThanks! \n\nDo you plan on taking 3 days break to recover your sleep and start the new Avito? =D",
      "votes": null
    },
    {
      "id": "325507",
      "postDate": "05/08/2018 13:38:32",
      "content": "<p>Thanks.  I think I'll go for TrackML rather.  I would love to do Avito too, but downloading all data for one of them will take me a week at least.  I'll use that week to recover, as you suggest!</p>",
      "rawMarkdown": "Thanks.  I think I'll go for TrackML rather.  I would love to do Avito too, but downloading all data for one of them will take me a week at least.  I'll use that week to recover, as you suggest!",
      "votes": null
    },
    {
      "id": "325617",
      "postDate": "05/08/2018 16:28:46",
      "content": "<p>Nice job and your no.3 tips is excellent! May I ask would you open source your solution? I'm really interested in that!</p>",
      "rawMarkdown": "Nice job and your no.3 tips is excellent! May I ask would you open source your solution? I'm really interested in that!",
      "votes": null
    },
    {
      "id": "325670",
      "postDate": "05/08/2018 17:39:24",
      "content": "<p>Thanks <a href=\"/snorlax\">@snorlax</a>.\nMy R code in this competition is a very big mess, nothing you can just run and expect a result.\nIf you're looking for a specific part i can \ngladly  copy paste it for you </p>",
      "rawMarkdown": "Thanks @snorlax.\nMy R code in this competition is a very big mess, nothing you can just run and expect a result.\nIf you're looking for a specific part i can \ngladly  copy paste it for you",
      "votes": null
    },
    {
      "id": "325688",
      "postDate": "05/08/2018 18:03:58",
      "content": "<p>Oh my goddddd! You are such a nice person! I'm specifically interested in the code of your no.3 tips(the encoding part), thank you! And I'm ok even you think your code is messy, for me, it is treasure.</p>",
      "rawMarkdown": "Oh my goddddd! You are such a nice person! I'm specifically interested in the code of your no.3 tips(the encoding part), thank you! And I'm ok even you think your code is messy, for me, it is treasure.",
      "votes": null
    },
    {
      "id": "325701",
      "postDate": "05/08/2018 18:48:39",
      "content": "<p>@AmirH, could you post it in Kernels? I'm very interested too. Thanks.</p>",
      "rawMarkdown": "AmirH, could you post it in Kernels? I'm very interested too. Thanks.",
      "votes": null
    },
    {
      "id": "326344",
      "postDate": "05/09/2018 15:36:13",
      "content": "<p>The different features  between the first LGB model and the second LGB model is just the \"category features\" and the target encoder??I mean do you just remove the category features from the first LGB model and add target encoder in the second LGB model?? thanks </p>",
      "rawMarkdown": "The different features  between the first LGB model and the second LGB model is just the \"category features\" and the target encoder??I mean do you just remove the category features from the first LGB model and add target encoder in the second LGB model?? thanks",
      "votes": null
    },
    {
      "id": "326636",
      "postDate": "05/10/2018 04:55:15",
      "content": "<p>Hi @ChinaBoy,\nNot quite,\nLGB2 contains all LGB1 features that are non categorical + features that are BASED ON the target encoding, not the target encoding itself.\nFeatures like Avg(is_attributed) by groups and those which are based on USER as i stated:\ne.g:\nhow many times did user (ip, os, device) ever download any app</p>\n\n<p>how many times did he download the current app</p>\n\n<p>how many unique apps did he click on.</p>",
      "rawMarkdown": "Hi @ChinaBoy,\nNot quite,\nLGB2 contains all LGB1 features that are non categorical + features that are BASED ON the target encoding, not the target encoding itself.\nFeatures like Avg(is_attributed) by groups and those which are based on USER as i stated:\ne.g:\nhow many times did user (ip, os, device) ever download any app\n\nhow many times did he download the current app\n\nhow many unique apps did he click on.",
      "votes": null
    },
    {
      "id": "327127",
      "postDate": "05/10/2018 21:26:02",
      "content": "<p>Hi @Snorlax &amp; @Terracota\nI hope this small snippet helps (it is in R)\nI couldn't run in in kernel so i couldn't share  </p>\n\n<pre><code>   ##Load my LGB model #1 AUC 0.906\n  model&lt;- lgb.load(\"e:/Project/Latest Files/Models/LGBM1.mod\") \n\n  #Import Training + Test Supp + Test data including all their features (after preprocessing)\n  TrainData2&lt;-fread(file =\"f:/AllDataTrainCVTest.csv\");\n  TrainData2[,click_time:=fastPOSIXct(click_time,tz=\"UTC\")]\n\n  #Which columns i need for LGB model # 1 to predict - keep only these columns\n  FeatureCols&lt;- c(\"app\", \"device\", \"os\", \"channel\", \"hour\",  \"ipOnly_dayCount\", \"ip_app_dayCount\", \"ipOnly_day_hourCount\",  \"ipOnly_Count\", \"ip_app_Count\",  \"User_UniqueApps\", \"User_Day_UniqueApps\", \"ip_UniqueUsers\", \"ip_UniqueApps\",  \"ip_Day_UniqueApps\", \"ip_UniqueDevices\",  \"ip_UniqueOs\",    \"App_UniqueChannel\",      \"User_app_NextClick\", \"User_app_NextClick2\",   \"User_Count\",  \"User_app_Count\", \"User_dayHourCount\"\n)   \n\ncols&lt;-FeatureCols\nTrainData2&lt;-TrainData2[,..cols]\n\n  #Convert to matrix and predict\n  TrainData2&lt;-as.matrix(TrainData2);gc()\n  LGBPrediction2&lt;-data.table(LGBPrediction2=predict(model, data =TrainData2) ) ;gc() \n  fwrite(LGBPrediction2,\"f:/LGBPrediction2.csv\")\n\n  #Add predictions as additional column to the data\n  TrainData2&lt;-cbind.data.frame(TrainData2,LGBPrediction2)\n\n  #Add columns that state prediction as 1 or 0 based on different thresholds of prediction.\n  #The threshold were chosen carefuly after examining confusion matrices and sensitivity values\n  #for different values, tried to minimize false positives\n  TrainData2[,LGBPred2IntOnlyModel99:=ifelse(LGBPrediction2==0.99,1,0)]\n  TrainData2[,LGBPred2IntOnlyModel997:=ifelse(LGBPrediction2==0.997,1,0)]\n  TrainData2[,LGBPred2IntOnlyModel999:=ifelse(LGBPrediction2==0.999,1,0)]\n\n  #Add different features a few times each, so can check for each one, which threshold worked best\n  TrainData2[,User_appDownloadsBasedOnModel99:=sum(LGBPred2IntOnlyModel99),by=.(ip,device,os,app)]\n  TrainData2[,User_appDownloadsBasedOnModel997:=sum(LGBPred2IntOnlyModel997),by=.(ip,device,os,app)]\n  TrainData2[,User_appDownloadsBasedOnModel999:=sum(LGBPred2IntOnlyModel999),by=.(ip,device,os,app)]\n\n  TrainData2[,User_appDownloadsBeforeBasedOnModel99:=cumsum(LGBPred2IntOnlyModel99),by=.(ip,device,os,app)]\n  TrainData2[,User_appDownloadsBeforeBasedOnModel997:=cumsum(LGBPred2IntOnlyModel997),by=.(ip,device,os,app)]\n  TrainData2[,User_appDownloadsBeforeBasedOnModel999:=cumsum(LGBPred2IntOnlyModel999),by=.(ip,device,os,app)]\n\n  TrainData2[,User_AppsDownloadedBasedOnModel99:=uniqueN(app),by=.(ip,device,os,LGBPred2IntOnlyModel99==1)]\n  TrainData2[,User_AppsDownloadedBasedOnModel997:=uniqueN(app),by=.(ip,device,os,LGBPred2IntOnlyModel997==1)]\n  TrainData2[,User_AppsDownloadedBasedOnModel999:=uniqueN(app),by=.(ip,device,os,LGBPred2IntOnlyModel999==1)]\n\n  TrainData2[,User_DivAppsDownloadedandClickedBasedOnModel99:=User_AppsDownloadedBasedOnModel99 / User_UniqueApps]\n  TrainData2[,User_DivAppsDownloadedandClickedBasedOnModel997:=User_AppsDownloadedBasedOnModel997 / User_UniqueApps]\n  TrainData2[,User_DivAppsDownloadedandClickedBasedOnModel999:=User_AppsDownloadedBasedOnModel999 / User_UniqueApps]\n\n\n#Run model a few times with different threshold features to test AUC\n  RunTest(c(\"User_appDownloadsBasedOnModel99\",   \"User_appDownloadsBeforeBasedOnModel99\",\"User_AppsDownloadedBasedOnModel99\",\"User_DivAppsDownloadedandClickedBasedOnModel99\"),\"LGBM2UserDownloads99\")\n  RunTest(c(\"User_appDownloadsBasedOnModel997\",   \"User_appDownloadsBeforeBasedOnModel997\",\"User_AppsDownloadedBasedOnModel997\",\"User_DivAppsDownloadedandClickedBasedOnModel997\"),\"LGBM2UserDownloads997\")\n  RunTest(c(\"User_appDownloadsBasedOnModel999\",   \"User_appDownloadsBeforeBasedOnModel999\",\"User_AppsDownloadedBasedOnModel999\",\"User_DivAppsDownloadedandClickedBasedOnModel999\"),\"LGBM2UserDownloads999\")\n</code></pre>",
      "rawMarkdown": "Hi @Snorlax &amp; @Terracota\nI hope this small snippet helps (it is in R)\nI couldn't run in in kernel so i couldn't share  \n\n       ##Load my LGB model #1 AUC 0.906\n      model&lt;- lgb.load(\"e:/Project/Latest Files/Models/LGBM1.mod\") \n    \n      #Import Training + Test Supp + Test data including all their features (after preprocessing)\n      TrainData2&lt;-fread(file =\"f:/AllDataTrainCVTest.csv\");\n      TrainData2[,click_time:=fastPOSIXct(click_time,tz=\"UTC\")]\n    \n      #Which columns i need for LGB model # 1 to predict - keep only these columns\n      FeatureCols&lt;- c(\"app\", \"device\", \"os\", \"channel\", \"hour\",  \"ipOnly_dayCount\", \"ip_app_dayCount\", \"ipOnly_day_hourCount\",  \"ipOnly_Count\", \"ip_app_Count\",  \"User_UniqueApps\", \"User_Day_UniqueApps\", \"ip_UniqueUsers\", \"ip_UniqueApps\",  \"ip_Day_UniqueApps\", \"ip_UniqueDevices\",  \"ip_UniqueOs\",    \"App_UniqueChannel\",      \"User_app_NextClick\", \"User_app_NextClick2\",   \"User_Count\",  \"User_app_Count\", \"User_dayHourCount\"\n    )   \n    \n    cols&lt;-FeatureCols\n    TrainData2&lt;-TrainData2[,..cols]\n    \n      #Convert to matrix and predict\n      TrainData2&lt;-as.matrix(TrainData2);gc()\n      LGBPrediction2&lt;-data.table(LGBPrediction2=predict(model, data =TrainData2) ) ;gc() \n      fwrite(LGBPrediction2,\"f:/LGBPrediction2.csv\")\n    \n      #Add predictions as additional column to the data\n      TrainData2&lt;-cbind.data.frame(TrainData2,LGBPrediction2)\n    \n      #Add columns that state prediction as 1 or 0 based on different thresholds of prediction.\n      #The threshold were chosen carefuly after examining confusion matrices and sensitivity values\n      #for different values, tried to minimize false positives\n      TrainData2[,LGBPred2IntOnlyModel99:=ifelse(LGBPrediction2==0.99,1,0)]\n      TrainData2[,LGBPred2IntOnlyModel997:=ifelse(LGBPrediction2==0.997,1,0)]\n      TrainData2[,LGBPred2IntOnlyModel999:=ifelse(LGBPrediction2==0.999,1,0)]\n    \n      #Add different features a few times each, so can check for each one, which threshold worked best\n      TrainData2[,User_appDownloadsBasedOnModel99:=sum(LGBPred2IntOnlyModel99),by=.(ip,device,os,app)]\n      TrainData2[,User_appDownloadsBasedOnModel997:=sum(LGBPred2IntOnlyModel997),by=.(ip,device,os,app)]\n      TrainData2[,User_appDownloadsBasedOnModel999:=sum(LGBPred2IntOnlyModel999),by=.(ip,device,os,app)]\n    \n      TrainData2[,User_appDownloadsBeforeBasedOnModel99:=cumsum(LGBPred2IntOnlyModel99),by=.(ip,device,os,app)]\n      TrainData2[,User_appDownloadsBeforeBasedOnModel997:=cumsum(LGBPred2IntOnlyModel997),by=.(ip,device,os,app)]\n      TrainData2[,User_appDownloadsBeforeBasedOnModel999:=cumsum(LGBPred2IntOnlyModel999),by=.(ip,device,os,app)]\n    \n      TrainData2[,User_AppsDownloadedBasedOnModel99:=uniqueN(app),by=.(ip,device,os,LGBPred2IntOnlyModel99==1)]\n      TrainData2[,User_AppsDownloadedBasedOnModel997:=uniqueN(app),by=.(ip,device,os,LGBPred2IntOnlyModel997==1)]\n      TrainData2[,User_AppsDownloadedBasedOnModel999:=uniqueN(app),by=.(ip,device,os,LGBPred2IntOnlyModel999==1)]\n    \n      TrainData2[,User_DivAppsDownloadedandClickedBasedOnModel99:=User_AppsDownloadedBasedOnModel99 / User_UniqueApps]\n      TrainData2[,User_DivAppsDownloadedandClickedBasedOnModel997:=User_AppsDownloadedBasedOnModel997 / User_UniqueApps]\n      TrainData2[,User_DivAppsDownloadedandClickedBasedOnModel999:=User_AppsDownloadedBasedOnModel999 / User_UniqueApps]\n    \n    \n    #Run model a few times with different threshold features to test AUC\n      RunTest(c(\"User_appDownloadsBasedOnModel99\",   \"User_appDownloadsBeforeBasedOnModel99\",\"User_AppsDownloadedBasedOnModel99\",\"User_DivAppsDownloadedandClickedBasedOnModel99\"),\"LGBM2UserDownloads99\")\n      RunTest(c(\"User_appDownloadsBasedOnModel997\",   \"User_appDownloadsBeforeBasedOnModel997\",\"User_AppsDownloadedBasedOnModel997\",\"User_DivAppsDownloadedandClickedBasedOnModel997\"),\"LGBM2UserDownloads997\")\n      RunTest(c(\"User_appDownloadsBasedOnModel999\",   \"User_appDownloadsBeforeBasedOnModel999\",\"User_AppsDownloadedBasedOnModel999\",\"User_DivAppsDownloadedandClickedBasedOnModel999\"),\"LGBM2UserDownloads999\")",
      "votes": null
    },
    {
      "id": "327152",
      "postDate": "05/10/2018 23:00:47",
      "content": "<p>@AmirH, Thanks. It's great.</p>",
      "rawMarkdown": "AmirH, Thanks. It's great.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 325429,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/08/2018 12:02:56",
      "content": "<p>Thanks for sharing, congrats on the end result!  Your item 3 is very interesting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 325495,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "05/08/2018 13:22:26",
          "content": "<p>Thanks @CPMP.\nYou were extremely helpful during the competition, any where i go, any kernel or discussion, you are always there with your insight and experience.\nThanks! </p>\n\n<p>Do you plan on taking 3 days break to recover your sleep and start the new Avito? =D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325507,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/08/2018 13:38:32",
          "content": "<p>Thanks.  I think I'll go for TrackML rather.  I would love to do Avito too, but downloading all data for one of them will take me a week at least.  I'll use that week to recover, as you suggest!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325617,
      "author_name": "wythhh",
      "author_url": "",
      "post_date": "05/08/2018 16:28:46",
      "content": "<p>Nice job and your no.3 tips is excellent! May I ask would you open source your solution? I'm really interested in that!</p>",
      "votes": null,
      "replies": [
        {
          "id": 325670,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "05/08/2018 17:39:24",
          "content": "<p>Thanks <a href=\"/snorlax\">@snorlax</a>.\nMy R code in this competition is a very big mess, nothing you can just run and expect a result.\nIf you're looking for a specific part i can \ngladly  copy paste it for you </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325688,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "05/08/2018 18:03:58",
          "content": "<p>Oh my goddddd! You are such a nice person! I'm specifically interested in the code of your no.3 tips(the encoding part), thank you! And I'm ok even you think your code is messy, for me, it is treasure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325701,
          "author_name": "ybwu01",
          "author_url": "",
          "post_date": "05/08/2018 18:48:39",
          "content": "<p>@AmirH, could you post it in Kernels? I'm very interested too. Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327127,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "05/10/2018 21:26:02",
          "content": "<p>Hi @Snorlax &amp; @Terracota\nI hope this small snippet helps (it is in R)\nI couldn't run in in kernel so i couldn't share  </p>\n\n<pre><code>   ##Load my LGB model #1 AUC 0.906\n  model&lt;- lgb.load(\"e:/Project/Latest Files/Models/LGBM1.mod\") \n\n  #Import Training + Test Supp + Test data including all their features (after preprocessing)\n  TrainData2&lt;-fread(file =\"f:/AllDataTrainCVTest.csv\");\n  TrainData2[,click_time:=fastPOSIXct(click_time,tz=\"UTC\")]\n\n  #Which columns i need for LGB model # 1 to predict - keep only these columns\n  FeatureCols&lt;- c(\"app\", \"device\", \"os\", \"channel\", \"hour\",  \"ipOnly_dayCount\", \"ip_app_dayCount\", \"ipOnly_day_hourCount\",  \"ipOnly_Count\", \"ip_app_Count\",  \"User_UniqueApps\", \"User_Day_UniqueApps\", \"ip_UniqueUsers\", \"ip_UniqueApps\",  \"ip_Day_UniqueApps\", \"ip_UniqueDevices\",  \"ip_UniqueOs\",    \"App_UniqueChannel\",      \"User_app_NextClick\", \"User_app_NextClick2\",   \"User_Count\",  \"User_app_Count\", \"User_dayHourCount\"\n)   \n\ncols&lt;-FeatureCols\nTrainData2&lt;-TrainData2[,..cols]\n\n  #Convert to matrix and predict\n  TrainData2&lt;-as.matrix(TrainData2);gc()\n  LGBPrediction2&lt;-data.table(LGBPrediction2=predict(model, data =TrainData2) ) ;gc() \n  fwrite(LGBPrediction2,\"f:/LGBPrediction2.csv\")\n\n  #Add predictions as additional column to the data\n  TrainData2&lt;-cbind.data.frame(TrainData2,LGBPrediction2)\n\n  #Add columns that state prediction as 1 or 0 based on different thresholds of prediction.\n  #The threshold were chosen carefuly after examining confusion matrices and sensitivity values\n  #for different values, tried to minimize false positives\n  TrainData2[,LGBPred2IntOnlyModel99:=ifelse(LGBPrediction2==0.99,1,0)]\n  TrainData2[,LGBPred2IntOnlyModel997:=ifelse(LGBPrediction2==0.997,1,0)]\n  TrainData2[,LGBPred2IntOnlyModel999:=ifelse(LGBPrediction2==0.999,1,0)]\n\n  #Add different features a few times each, so can check for each one, which threshold worked best\n  TrainData2[,User_appDownloadsBasedOnModel99:=sum(LGBPred2IntOnlyModel99),by=.(ip,device,os,app)]\n  TrainData2[,User_appDownloadsBasedOnModel997:=sum(LGBPred2IntOnlyModel997),by=.(ip,device,os,app)]\n  TrainData2[,User_appDownloadsBasedOnModel999:=sum(LGBPred2IntOnlyModel999),by=.(ip,device,os,app)]\n\n  TrainData2[,User_appDownloadsBeforeBasedOnModel99:=cumsum(LGBPred2IntOnlyModel99),by=.(ip,device,os,app)]\n  TrainData2[,User_appDownloadsBeforeBasedOnModel997:=cumsum(LGBPred2IntOnlyModel997),by=.(ip,device,os,app)]\n  TrainData2[,User_appDownloadsBeforeBasedOnModel999:=cumsum(LGBPred2IntOnlyModel999),by=.(ip,device,os,app)]\n\n  TrainData2[,User_AppsDownloadedBasedOnModel99:=uniqueN(app),by=.(ip,device,os,LGBPred2IntOnlyModel99==1)]\n  TrainData2[,User_AppsDownloadedBasedOnModel997:=uniqueN(app),by=.(ip,device,os,LGBPred2IntOnlyModel997==1)]\n  TrainData2[,User_AppsDownloadedBasedOnModel999:=uniqueN(app),by=.(ip,device,os,LGBPred2IntOnlyModel999==1)]\n\n  TrainData2[,User_DivAppsDownloadedandClickedBasedOnModel99:=User_AppsDownloadedBasedOnModel99 / User_UniqueApps]\n  TrainData2[,User_DivAppsDownloadedandClickedBasedOnModel997:=User_AppsDownloadedBasedOnModel997 / User_UniqueApps]\n  TrainData2[,User_DivAppsDownloadedandClickedBasedOnModel999:=User_AppsDownloadedBasedOnModel999 / User_UniqueApps]\n\n\n#Run model a few times with different threshold features to test AUC\n  RunTest(c(\"User_appDownloadsBasedOnModel99\",   \"User_appDownloadsBeforeBasedOnModel99\",\"User_AppsDownloadedBasedOnModel99\",\"User_DivAppsDownloadedandClickedBasedOnModel99\"),\"LGBM2UserDownloads99\")\n  RunTest(c(\"User_appDownloadsBasedOnModel997\",   \"User_appDownloadsBeforeBasedOnModel997\",\"User_AppsDownloadedBasedOnModel997\",\"User_DivAppsDownloadedandClickedBasedOnModel997\"),\"LGBM2UserDownloads997\")\n  RunTest(c(\"User_appDownloadsBasedOnModel999\",   \"User_appDownloadsBeforeBasedOnModel999\",\"User_AppsDownloadedBasedOnModel999\",\"User_DivAppsDownloadedandClickedBasedOnModel999\"),\"LGBM2UserDownloads999\")\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327152,
          "author_name": "ybwu01",
          "author_url": "",
          "post_date": "05/10/2018 23:00:47",
          "content": "<p>@AmirH, Thanks. It's great.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326344,
      "author_name": "huayupeng",
      "author_url": "",
      "post_date": "05/09/2018 15:36:13",
      "content": "<p>The different features  between the first LGB model and the second LGB model is just the \"category features\" and the target encoder??I mean do you just remove the category features from the first LGB model and add target encoder in the second LGB model?? thanks </p>",
      "votes": null,
      "replies": [
        {
          "id": 326636,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "05/10/2018 04:55:15",
          "content": "<p>Hi @ChinaBoy,\nNot quite,\nLGB2 contains all LGB1 features that are non categorical + features that are BASED ON the target encoding, not the target encoding itself.\nFeatures like Avg(is_attributed) by groups and those which are based on USER as i stated:\ne.g:\nhow many times did user (ip, os, device) ever download any app</p>\n\n<p>how many times did he download the current app</p>\n\n<p>how many unique apps did he click on.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "325392": "Hi!\nI wanted to share some of my findings during this competition. \nIt's my first ever competition or ML task, so much of what I tried was really bad, but it taught me good intuition making these horrible mistakes and have a better \"radar\" for possible leaks and over fitting.  \n\nI learned a lot on R's data.table, LGBM, XGBoost, I experimented with keras and H2O NNs (didn't get good results) and also experimented with AWS's VMs.\n\nMy public lb is 65 and private is 34, i think this increase came from good generalized CV and features, after trying many things and participating in many discussions on the matter.\n\nI ended up using 2 LGBM models (second one uses the predictions of the first, details in item #3),\nwith main features including: Count clicks by groups, UniqueN by groups, Next click by groups, and the target encoding based features detailed in item #3)\n\nI won't be long on all the mistakes i did, i will focus on the things I did that gave the significant boosts along the way and seemed to have generalized well.\nI hope this will be of any value to any of you.\n\n1. Dropping any concerns for over fitting my model to day 9 hour 4, although the difference of CV and LB score was lowest, it obviously overfit.\nAfter some discussions on the matter, I ended up training on days 7 and 8, CV with day 9, and then retraining on same rounds on entire data. \n**I read @CPMP's post and learned that he used 1.2 multiplier on the Num_Rounds with the full training set, will definitely try it in the future!!\n\n2. I calculated all of my feature calculations on train + test supplement data.\nThis is after trying to run features only where hour exists in the test set, or only on the train\\ regular test set.\n\n3. My last boost up from 0.9806 to 0.9816 was using an interesting approach to target encoding i encountered somewhere in the discussions (I don't remember who suggested the approach so sorry for not giving the credit).\nThe method involved taking my best lgb model and predict on all train and test supplement data.\nThen using the predicitions to do target encoding on the entire data and using it to create features involving the \"is_attributed\" field.\nAfter some experimentation, i ended up with this:\nI took the prediction and converted it 3 times to a boolean 1/0 \"is_attributed\" value.\n- Once if pred &gt;= 0.99 then 1 otherwise 0\n- then pred &gt;= 0.997 then 1 otherwise 0\n- Then pred &gt;= 0.999 then 1 otherwise 0\n\nI did this because i wanted better sensitivity for the model, i am ok missing some 1's but don't want many 0's to be mistakenly predicted as 1's.\nI then calculated 2 sets of features 3 times, once for each prediction result set.\nI had mean target features by groups (past is_attributed = 1 divided by count of clicks), and another set of features that included, amongst other features:\n\n- how many times did user (ip, os, device) ever download any app\n\n- how many times did he download the current app\n\n- how many unique apps did he click on.\n\nAfter running some test runs of the new features, i found the highest increase in score with mean target when using the &gt;=0.997 prediction, and the the other set of features maxing with the &gt;=0.99 prediction.\nI then used both of the sets in a new lgb model that uses only the Quantity (non categorical) features from the first lgb model  + the new target encoded features.\n\n\n4, I used bayesian optimization to optimize hyper parameters for the 2 models separately.\nUnfortunately, even after testing for different methods of optimization (GP Upper bound, expected probability etc.) with different Kappa\\Epsilon results. I did not find it running differently or optimizing better in any way.\nIt was basically running like random search.\nI still need to learn more on it, but taking the best found hyper parameters did give me a decent boost in score.\n\nThanks for reading and thanks for being such an awesome community!!\nAnd of course to all the awesome Kagglers that assisted and shared their wisdom during the competition.\nThis was quite an intense 2 months =)",
    "325429": "Thanks for sharing, congrats on the end result!  Your item 3 is very interesting.",
    "325495": "Thanks @CPMP.\nYou were extremely helpful during the competition, any where i go, any kernel or discussion, you are always there with your insight and experience.\nThanks! \n\nDo you plan on taking 3 days break to recover your sleep and start the new Avito? =D",
    "325507": "Thanks.  I think I'll go for TrackML rather.  I would love to do Avito too, but downloading all data for one of them will take me a week at least.  I'll use that week to recover, as you suggest!",
    "325617": "Nice job and your no.3 tips is excellent! May I ask would you open source your solution? I'm really interested in that!",
    "325670": "Thanks @snorlax.\nMy R code in this competition is a very big mess, nothing you can just run and expect a result.\nIf you're looking for a specific part i can \ngladly  copy paste it for you",
    "325688": "Oh my goddddd! You are such a nice person! I'm specifically interested in the code of your no.3 tips(the encoding part), thank you! And I'm ok even you think your code is messy, for me, it is treasure.",
    "325701": "AmirH, could you post it in Kernels? I'm very interested too. Thanks.",
    "326344": "The different features  between the first LGB model and the second LGB model is just the \"category features\" and the target encoder??I mean do you just remove the category features from the first LGB model and add target encoder in the second LGB model?? thanks",
    "326636": "Hi @ChinaBoy,\nNot quite,\nLGB2 contains all LGB1 features that are non categorical + features that are BASED ON the target encoding, not the target encoding itself.\nFeatures like Avg(is_attributed) by groups and those which are based on USER as i stated:\ne.g:\nhow many times did user (ip, os, device) ever download any app\n\nhow many times did he download the current app\n\nhow many unique apps did he click on.",
    "327127": "Hi @Snorlax &amp; @Terracota\nI hope this small snippet helps (it is in R)\nI couldn't run in in kernel so i couldn't share  \n\n       ##Load my LGB model #1 AUC 0.906\n      model&lt;- lgb.load(\"e:/Project/Latest Files/Models/LGBM1.mod\") \n    \n      #Import Training + Test Supp + Test data including all their features (after preprocessing)\n      TrainData2&lt;-fread(file =\"f:/AllDataTrainCVTest.csv\");\n      TrainData2[,click_time:=fastPOSIXct(click_time,tz=\"UTC\")]\n    \n      #Which columns i need for LGB model # 1 to predict - keep only these columns\n      FeatureCols&lt;- c(\"app\", \"device\", \"os\", \"channel\", \"hour\",  \"ipOnly_dayCount\", \"ip_app_dayCount\", \"ipOnly_day_hourCount\",  \"ipOnly_Count\", \"ip_app_Count\",  \"User_UniqueApps\", \"User_Day_UniqueApps\", \"ip_UniqueUsers\", \"ip_UniqueApps\",  \"ip_Day_UniqueApps\", \"ip_UniqueDevices\",  \"ip_UniqueOs\",    \"App_UniqueChannel\",      \"User_app_NextClick\", \"User_app_NextClick2\",   \"User_Count\",  \"User_app_Count\", \"User_dayHourCount\"\n    )   \n    \n    cols&lt;-FeatureCols\n    TrainData2&lt;-TrainData2[,..cols]\n    \n      #Convert to matrix and predict\n      TrainData2&lt;-as.matrix(TrainData2);gc()\n      LGBPrediction2&lt;-data.table(LGBPrediction2=predict(model, data =TrainData2) ) ;gc() \n      fwrite(LGBPrediction2,\"f:/LGBPrediction2.csv\")\n    \n      #Add predictions as additional column to the data\n      TrainData2&lt;-cbind.data.frame(TrainData2,LGBPrediction2)\n    \n      #Add columns that state prediction as 1 or 0 based on different thresholds of prediction.\n      #The threshold were chosen carefuly after examining confusion matrices and sensitivity values\n      #for different values, tried to minimize false positives\n      TrainData2[,LGBPred2IntOnlyModel99:=ifelse(LGBPrediction2==0.99,1,0)]\n      TrainData2[,LGBPred2IntOnlyModel997:=ifelse(LGBPrediction2==0.997,1,0)]\n      TrainData2[,LGBPred2IntOnlyModel999:=ifelse(LGBPrediction2==0.999,1,0)]\n    \n      #Add different features a few times each, so can check for each one, which threshold worked best\n      TrainData2[,User_appDownloadsBasedOnModel99:=sum(LGBPred2IntOnlyModel99),by=.(ip,device,os,app)]\n      TrainData2[,User_appDownloadsBasedOnModel997:=sum(LGBPred2IntOnlyModel997),by=.(ip,device,os,app)]\n      TrainData2[,User_appDownloadsBasedOnModel999:=sum(LGBPred2IntOnlyModel999),by=.(ip,device,os,app)]\n    \n      TrainData2[,User_appDownloadsBeforeBasedOnModel99:=cumsum(LGBPred2IntOnlyModel99),by=.(ip,device,os,app)]\n      TrainData2[,User_appDownloadsBeforeBasedOnModel997:=cumsum(LGBPred2IntOnlyModel997),by=.(ip,device,os,app)]\n      TrainData2[,User_appDownloadsBeforeBasedOnModel999:=cumsum(LGBPred2IntOnlyModel999),by=.(ip,device,os,app)]\n    \n      TrainData2[,User_AppsDownloadedBasedOnModel99:=uniqueN(app),by=.(ip,device,os,LGBPred2IntOnlyModel99==1)]\n      TrainData2[,User_AppsDownloadedBasedOnModel997:=uniqueN(app),by=.(ip,device,os,LGBPred2IntOnlyModel997==1)]\n      TrainData2[,User_AppsDownloadedBasedOnModel999:=uniqueN(app),by=.(ip,device,os,LGBPred2IntOnlyModel999==1)]\n    \n      TrainData2[,User_DivAppsDownloadedandClickedBasedOnModel99:=User_AppsDownloadedBasedOnModel99 / User_UniqueApps]\n      TrainData2[,User_DivAppsDownloadedandClickedBasedOnModel997:=User_AppsDownloadedBasedOnModel997 / User_UniqueApps]\n      TrainData2[,User_DivAppsDownloadedandClickedBasedOnModel999:=User_AppsDownloadedBasedOnModel999 / User_UniqueApps]\n    \n    \n    #Run model a few times with different threshold features to test AUC\n      RunTest(c(\"User_appDownloadsBasedOnModel99\",   \"User_appDownloadsBeforeBasedOnModel99\",\"User_AppsDownloadedBasedOnModel99\",\"User_DivAppsDownloadedandClickedBasedOnModel99\"),\"LGBM2UserDownloads99\")\n      RunTest(c(\"User_appDownloadsBasedOnModel997\",   \"User_appDownloadsBeforeBasedOnModel997\",\"User_AppsDownloadedBasedOnModel997\",\"User_DivAppsDownloadedandClickedBasedOnModel997\"),\"LGBM2UserDownloads997\")\n      RunTest(c(\"User_appDownloadsBasedOnModel999\",   \"User_appDownloadsBeforeBasedOnModel999\",\"User_AppsDownloadedBasedOnModel999\",\"User_DivAppsDownloadedandClickedBasedOnModel999\"),\"LGBM2UserDownloads999\")",
    "327152": "AmirH, Thanks. It's great."
  },
  "source": "meta"
}